The opportunity
We run one of the industry's largest AI compute fleets, spanning multiple cloud providers and datacenters, to train, research, and serve frontier AI models. Those fleets run on Kubernetes, and the Kubernetes Platform team owns the control plane that makes them work.
What you'll do
Own, operate, and extend the Kubernetes scheduler for Anthropic's accelerator: fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption
Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to: support clusters far beyond typical limits, and find the next bottleneck before it finds us
Design, build, and operate core cluster services such as service discovery: that every workload in the fleet depends on
Build and maintain custom controllers, operators, and CRDs
Partner with research, training, and inference to understand workload shapes: and turn their requirements into platform capabilities
Collaborate with cloud providers on required features and escalations
What they're looking for
- Experience with Kubernetes internals or contributions: kube-scheduler / scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar
- Experience building or operating cluster schedulers or batch systems (e.g.,: Kueue, Volcano, Slurm, or in-house equivalents)
- Background scaling control planes or coordination systems (etcd, ZooKeeper,: Consul, or large DNS/service-mesh deployments)
- Familiarity with ML infrastructure: GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; collective networking such as NCCL
- Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code
- Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF
- + years of relevant industry experience, including time leading large, ambiguous infrastructure projects