The opportunity
The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models.
What you'll do
Port and validate key inference and training workloads on new platforms/SKUs: as they arrive; drive correctness, performance, and stability to an internal readiness bar.
Build a suite of benchmarks and stress tests that capture real E2E behavior: of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts.
Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries)
Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects
Create repeatable test harnesses that run in CI / lab environments and: produce actionable outputs (pass/fail, performance score, regression detection).
Partner with systems + fleet bring-up engineers to ensure the platform is not: only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops).
What they're looking for
- Work cross-functionally with vendors and internal stakeholders by producing: clear bug reports, minimal repros, and prioritized issue lists.
- BS in CS/EE (or equivalent practical experience).
- + years in one or more of: ML systems, performance engineering, distributed systems, or HPC.
- Strong hands-on experience with: PyTorch and modern LLM training/inference stacks