The opportunity
The compute infrastructure team runs the GPU fleet and large-scale compute clusters that serve the models backing ChatGPT and the API, while also supporting training workloads for our next generation models. We operate a large, modern GPU fleet and provide a unified platform for…
What you'll do
Lead end-to-end delivery of both New Compute SKUs and large-scale GPU: clusters across an external partner ecosystem while supporting capacity planning for training and inference.
Ability to contextually drive multi-threaded bring-up programs spanning: hardware, networking, power, and cooling—owning plans, dependencies, and critical paths.
Interface with chip providers to derisk long-term onboarding to new hardware: platforms by working across kernels, comms, hardware, and scheduling engineering teams.
Build and operationalize program mechanisms (roadmaps, milestones, risk: registers, runbooks) that make delivery predictable at massive scale.
Partner with engineering to improve cluster turn-up reliability,: repeatability, and automation, reducing time-to-serve for new capacity.
Support network operations and end-to-end physical and logical bring-up of: OpenAI network Points-of-Presence (PoPs), including on-site deployment, rack cabling, and close collaboration with engineering teams.
What they're looking for
- Coordinate cross-functional readiness (security, finance, operations,: product/research stakeholders) to ship production-ready compute.
- Manage integration and handoffs across teams and partners—ensuring consistent: execution, clear communication, and fast issue resolution.
- Identify bottlenecks and systemic gaps, then drive durable fixes across: tooling, process, and partner interfaces.
- Provide crisp executive visibility on progress, tradeoffs, and risks across a: large portfolio of concurrent programs.