The opportunity
The Applied team safely brings OpenAI’s technology to the world, powering products like ChatGPT, and the APIs for GPT and more. Behind these products is a complex and rapidly evolving infrastructure platform that enables scale, performance, and safety.
What you'll do
Serve as the DRI for complex infrastructure programs spanning CPU planning,: orchestration, and other resource management domains (e.g. networking, storage).
Build and operationalize systems to capture demand signals, model future: capacity needs, and align infrastructure planning across internal teams and partners external to the company.
Partner closely with Infrastructure, Product and Finance teams to forecast: infrastructure usage patterns and ensure supply/demand alignment.
Lead cost attribution and quota enforcement programs to promote stability and: ensure equitable access to resources across teams.
Drive simplification and standardization of infrastructure tooling and: processes across Applied and Infra organizations.
Drive cross functional programs to evolve our infrastructure to support new growth and scale
What they're looking for
- Work with external vendors and partners (e.g., cloud providers) to manage: delivery schedules, capacity projections, and cost growth.
- Communicate project risks, progress, and key decisions clearly to technical and executive stakeholders.
- Have 7+ years of experience leading large-scale, technical infrastructure: programs, especially in compute resource planning, system scalability, or platform engineering.
- Understand how cloud infrastructure, distributed systems, and foundational: infra components (compute, storage, networking) come together to power complex products.