The opportunity
As demand for AI continues to accelerate, intelligent capacity management becomes one of the company's most strategic challenges. Every customer commitment, model launch, and infrastructure investment depends on making the right capacity decisions at the right time.
What you'll do
Run weekly capacity planning and daily capacity and deployment tracking with: Engineering, product and operations team. Own fleet utilization reporting and forecasting
Drive capacity planning for new customer deployments and major model launches
Drive continuous improvement and stakeholder adoption of new capacity management platform
Drive org level strategic initiatives related to capacity expansion,: improving fleet efficiency and maximizing effective utilization of available systems
Lead planning around major infrastructure events including but not limited to: new customer commits, new model releases, change to DC/cluster architecture, etc. that impacts capacity and fleet utilization. Update capacity plans and forecasts accordingly.
Maintain Jira EPICs and Confluence pages related to capacity planning,: reporting and change management to ensure execution transparency across teams
What they're looking for
- + years of TPM, technical program management, or product operations: experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning
- Experience leading large cross-functional programs involving Engineering, Product, and Operations
- Comfort with the inference serving stack: model replicas, batching, prefill/decode, KV cache, accelerator scheduling
- Strong data fluency: SQL, Grafana, basic Python or Flux to pull your own numbers without waiting for an analyst
- Track record of running a recurring cross-functional ritual involving senior engineers and LT
- Direct experience with AI accelerator fleet operations such as Habana, TPU pods, Inferentia, Trainium
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.