The opportunity
Anthropic's Infrastructure organization is the engine that powers our mission. Every breakthrough in AI safety research and every interaction users have with Claude depends on the systems we build and operate: massive clusters for training frontier models, production…
What you'll do
Drive cross-functional programs to improve developer environments, CI/CD: infrastructure, and release processes that enable rapid innovation while maintaining high security standards
Coordinate large-scale migrations and platform modernization efforts across engineering teams
Partner with teams to measure and improve developer productivity metrics,: identifying bottlenecks and driving systematic improvements
Lead initiatives to integrate AI tools into development workflows, helping: Anthropic be at the forefront of AI-assisted research and engineering
Drive programs to establish and achieve reliability targets across training: infrastructure and production services
Coordinate incident response improvements, post-mortem processes, and on-call: rotations that help teams operate effectively
What they're looking for
- Establish metrics and dashboards to track infrastructure health, capacity: utilization, and operational excellence
- Serve as the critical bridge between infrastructure teams, research, and: product, translating technical complexities into clear updates for a variety of audiences
- Consult with stakeholders to deeply understand infrastructure, data, and: compute needs, identifying solutions to support frontier research and product development
- Drive alignment on priorities and timelines across teams with competing constraints