The opportunity
Our Inference team is responsible for building and scaling the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry’s largest compute-agnostic inference deployments.
What you'll do
Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide
Develop resilient, flexible systems that adapt in real time to real world events
Develop intelligent request routing, load balancing, and traffic management: systems across thousands of accelerators and multiple cloud providers
Maximize compute efficiency and optimize cost across the fleet by autoscaling: and orchestrating production, research, and experimental workloads across multiple cloud providers
Build and operate production-grade deployment pipelines for releasing new models to users
Provide high-performance inference infrastructure that enables researchers to develop next-generation models
What they're looking for
- Experience with high-performance, large-scale distributed systems
- Experience implementing and deploying machine learning systems at scale
- Experience with load balancing, request routing, or traffic management systems
- Familiarity with LLM inference optimization, batching, and caching strategies
- Experience with Kubernetes and cloud infrastructure (AWS, GCP, Azure)
- Proficiency in Python or Rust