The opportunity
Our Inference team is responsible for building and maintaining the critical systems that serve Claude to millions of users worldwide. We bring Claude to life by serving our models via the industry's largest compute-agnostic inference deployments.
What you'll do
High-performance, large-scale distributed systems
Implementing and deploying machine learning systems at scale
Load balancing, request routing, or traffic management systems
LLM inference optimization, batching, and caching strategies
Kubernetes and cloud infrastructure (AWS, GCP)
Python or Rust
What they're looking for
- Have significant software engineering experience, particularly with distributed systems
- Are results-oriented, with a bias towards flexibility and impact
- Pick up slack, even if it goes outside your job description
- Want to learn more about machine learning systems and infrastructure