The opportunity
OpenAI’s Infrastructure organization builds the systems that power frontier AI workloads at global scale. As compute demand accelerates, our ability to rapidly convert infrastructure investments into usable production capacity has become mission critical.
What you'll do
Lead end-to-end execution of CPU / GPU cluster activation programs across: OpenAI’s global infrastructure footprint
Drive readiness to convert contracted compute capacity into schedulable production clusters
Own deployment programs for new PoPs, backbone nodes, WAN expansion, and interconnection initiatives
Build integrated schedules spanning procurement, logistics, installation,: storage readiness, network turn-up, testing, and production handoff
Coordinate BOM readiness, server delivery, racks, optics, cabling, storage hardware, and vendor milestones
Partner with engineering teams to align compute, storage, and networking: dependencies before cluster activation
What they're looking for
- Manage deployment of storage systems supporting training and inference: workloads, including readiness, validation, performance checks, and scaling plans
- Coordinate backbone capacity expansion, cross-connects, inter-region pathing,: and cloud interconnect readiness with Azure and third-party providers
- Lead physical deployment execution including rack-and-stack, hardware: bring-up, L1 validation, and site acceptance criteria
- Build repeatable deployment playbooks, dashboards, governance cadences, and operating mechanisms for scale