The opportunity
Anthropic's compute fleet is one of the largest and most varied in the world, and everything we do from training frontier models to serving Claude depends on getting the right work onto the right hardware at the right time. Our Scheduler team owns that problem.
What you'll do
Lead and grow a team of engineers building Anthropic's scheduling platform,: job-launch tooling, and fleet-efficiency systems, owning planning, execution, and delivery against key milestones
Set technical direction for scheduling, placement, queueing, and quota across Anthropic's compute fleet
Partner with capacity planning, research, inference, and product teams to: bring workloads onto the paved path and make efficient scheduling decisions
Drive the roadmap for scheduler capabilities, fleet utilization, and the: developer experience of launching and managing jobs
Define and track the metrics that measure fleet efficiency and scheduling: quality (utilization, queue wait, job-start latency, etc.) and hold the team accountable to them
Create clarity for the team and stakeholders in an ambiguous, fast-moving: environment where demand for compute routinely exceeds supply
What they're looking for
- + years of engineering management experience, including leading infrastructure, platform, or compute teams
- Experience owning a cluster scheduler, job orchestration system, or resource manager at scale
- Familiarity with scheduling ML workloads on accelerators and the tradeoffs: between utilization, fairness, and latency
- Experience building developer tooling that other engineers rely on every day
- A background in observability or incident response for control-plane systems,: and a track record of improving production reliability
- A track record of building a culture of belonging and of engineering excellence
- Low ego, high empathy, and a habit of leading by example