The opportunity
Anthropic's Infrastructure organization builds and operates the distributed systems that train, serve, and secure our AI models. Every other team at Anthropic depends on these systems.
What you'll do
Independently scope and lead complex, multi-month infrastructure projects,: from an ambiguous starting point through to a production system
Build deep partnerships with researchers and Research teams to understand their needs and deliver for them
Mentor other engineers and help raise the technical bar for the team
Build alignment on technical direction across multiple teams, working through ambiguous problem spaces
Take ownership of the reliability, scalability, and security of the systems: you build as usage and complexity grow
Lead the improvement of operational processes across Infrastructure, such as: incident response, postmortems, and on-call rotations, so the team learns from every incident
What they're looking for
- + years of software engineering experience, not including internships
- Experience with machine learning infrastructure (e.g., GPUs, TPUs, Trainium): and associated networking infrastructure (e.g., NCCL)
- Low-level systems experience (e.g., Linux kernel tuning, eBPF)
- Experience applying security or privacy engineering best practices