The opportunity
The Industrial Compute team is responsible for building the physical infrastructure that powers OpenAI’s largest-scale AI systems. We design, deploy, and operate next-generation compute infrastructure across a rapidly expanding global footprint, combining OpenAI-owned…
What you'll do
Own operational engagement with third-party infrastructure providers,: ensuring consistent execution against operational commitments, service-level agreements (SLAs), and performance expectations.
Develop operational governance frameworks with strategic partners, including: business reviews, operational scorecards, escalation processes, executive reporting, and performance improvement plans.
Define, track, and continuously improve key operational metrics related to: infrastructure availability, deployment execution, incident response, operational health, service quality, and partner performance.
Build dashboards and reporting mechanisms that provide clear visibility into: partner operational performance, risks, trends, and areas requiring executive attention.
Drive cross-functional coordination between OpenAI teams and external: infrastructure providers to resolve operational issues, remove execution blockers, and improve delivery outcomes.
Lead operational escalations involving infrastructure availability,: deployment execution, hardware operations, capacity delivery, or service performance, ensuring timely resolution and clear executive communication.
What they're looking for
- Establish repeatable operating rhythms with external partners, including: weekly operational reviews, executive business reviews, service reviews, action tracking, and long-term improvement initiatives.
- Partner with Capacity Planning, Hardware Operations, Networking, Deployment,: Reliability Engineering, and Supply Chain teams to ensure external infrastructure providers remain aligned with OpenAI’s operational priorities.
- Identify systemic operational risks across partner organizations and: proactively drive corrective actions that improve long-term operational effectiveness.
- + years of experience in Technical Program Management, Infrastructure: Operations, Cloud Operations, Service Delivery, or Technical Account Management within large-scale infrastructure environments.