The opportunity
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments.
What you'll do
Design, build, and maintain the distributed systems that serve Claude to millions of users worldwide
Develop resilient, flexible systems that adapt in real time to real-world events
Develop intelligent request routing, load balancing, and traffic management: systems across thousands of accelerators
Maximize compute efficiency across the fleet by autoscaling and orchestrating: production, research, and experimental workloads
Build and operate production-grade deployment pipelines for releasing new models to users
Provide high-performance inference infrastructure that enables researchers to develop next-generation models
What they're looking for
- Significant experience with high-performance, large-scale distributed systems
- Experience implementing and deploying machine learning systems at scale
- Experience building load balancing, request routing, or traffic management systems
- Familiarity with LLM inference optimization, batching, and caching strategies
- Deep experience operating Kubernetes and cloud infrastructure at scale
- Experience with AI accelerator platforms (GPUs, TPUs, or emerging hardware)