The opportunity
OpenAI’s Industrial Compute organization is responsible for ensuring our compute infrastructure scales efficiently to support millions of users and increasingly sophisticated AI models.
What you'll do
Build statistical and machine learning models to profile and improve GPU: utilization, latency, throughput, and overall fleet efficiency.
Develop forecasting models for inference demand across products, regions, and model families.
Analyze production workloads to identify latency bottlenecks and capacity: constraints, highlighting optimization opportunities.
Partner with Capacity Systems Engineering to inform infrastructure planning: and long-term GPU investment strategies.
Design experiments and simulations to evaluate scheduling policies, serving: strategies, and infrastructure tradeoffs.
Build dashboards and operational metrics that enable leadership to make data-driven capacity decisions.
What they're looking for
- Collaborate with Product, Research, Finance, and Infrastructure teams to: align compute planning with business growth and model roadmaps.
- Communicate technical findings clearly to both engineering teams and executive leadership.
- MS or PhD in Statistics, Computer Science, Operations Research, Applied: Mathematics, Economics, or related quantitative discipline (or equivalent industry experience).
- + years of experience working in the infrastructure data science space.