The opportunity
The Cloud Inference team scales and optimizes Claude to serve the massive audiences of developers and enterprise companies across AWS, GCP, Azure, and future cloud service providers (CSPs). We own the end-to-end product of Claude on each cloud platform, from API integration and…
What you'll do
Design, build, and own backend services and infrastructure that serve Claude: across multiple CSPs, accounting for differences in compute hardware, networking, APIs, and operational models
Work cross-functionally with internal inference, product API, systems, and: security teams, among others, and with CSP partners to stand up the full serving stack on new cloud platforms, resolve operational issues, and influence provider roadmaps
Build and evolve CI/CD automation systems, including validation and: deployment pipelines, that reliably ship new model versions to millions of users across cloud platforms without regressions
Design interfaces and tooling abstractions across CSPs that enable: cost-effective inference management, scale across providers, and reduce per-platform complexity
Contribute to capacity planning, autoscaling, and workload routing strategies: that match supply with demand and direct requests to the most cost-effective accelerator and region
Analyze observability data across providers to identify performance: bottlenecks, cost anomalies, and regressions, and drive remediation based on real-world production workloads
What they're looking for
- Direct experience working with CSPs to scale infrastructure or products: across multiple platforms, navigating differences in networking, security, privacy, billing, and managed service offerings
- Hands-on experience with capacity management, cost optimization, or resource: planning at scale across heterogeneous environments
- Solid understanding of multi-region deployments, geographic routing, and global traffic management
- Proficiency in Python or Rust