Technical Program Manager, Compute InfrastructureActive$257K–$335K

The opportunity

The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for…

What you'll do

  • Lead end-to-end delivery of both New Compute SKUs and large-scale GPU: clusters across an external partner ecosystem while supporting capacity planning for training and inference.

  • Ability to contextually drive multi-threaded bring-up programs spanning: hardware, networking, power, and cooling—owning plans, dependencies, and critical paths.

  • Interface with chip providers to derisk long-term onboarding to new hardware: platforms by working across kernels, comms, hardware, and scheduling engineering teams.

  • Build and operationalize program mechanisms (roadmaps, milestones, risk: registers, runbooks) that make delivery predictable at massive scale.

  • Partner with engineering to improve cluster turn-up reliability,: repeatability, and automation, reducing time-to-serve for new capacity.

  • Support network operations and end-to-end physical and logical bring-up of: OpenAI network Points-of-Presence (PoPs), including on-site deployment, rack cabling, and close collaboration with engineering teams.

What they're looking for

  • Coordinate cross-functional readiness (security, finance, operations,: product/research stakeholders) to ship production-ready compute.
  • Manage integration and handoffs across teams and partners—ensuring consistent: execution, clear communication, and fast issue resolution.
  • Identify bottlenecks and systemic gaps, then drive durable fixes across: tooling, process, and partner interfaces.
  • Provide crisp executive visibility on progress, tradeoffs, and risks across a: large portfolio of concurrent programs.