The opportunity
Anthropic's Infrastructure organization builds and operates the distributed systems that train, serve, and secure our AI models, including data pipelines, some of the largest Kubernetes clusters in the world, along with the databases, observability, and developer tooling that…
What you'll do
Team placement happens after the interview process, based on your interests: and experience alongside organizational needs. This lets us match you with the team where you'll have the most impact.
Independently scope and lead complex, multi-month infrastructure projects,: from an ambiguous starting point through to a production system
Make architectural decisions that shape the foundation of infrastructure other engineers and teams build on
Drive alignment on technical direction across multiple teams, working through ambiguous problem spaces
Partner with research and product teams to understand their infrastructure: and compute needs, and translate them into technical designs
Take ownership of the reliability, scalability, and security of the systems: you build as usage and complexity grow
What they're looking for
- + years of software engineering experience, not including internships
- Experience with machine learning infrastructure, such as GPUs, TPUs, or: Trainium, and associated networking infrastructure like NCCL
- Low-level systems experience, such as Linux kernel tuning or eBPF
- Background in security or privacy engineering best practices
- Prior experience as a technical lead or mentor for other engineers