Software Engineer, Workload EnablementActive$293K–$385K

The opportunity

The Scaling team is responsible for the architectural and engineering backbone of OpenAI’s infrastructure. We design and deliver advanced systems that support the deployment and operation of cutting-edge AI models.

What you'll do

  • Port and validate key inference and training workloads on new platforms/SKUs: as they arrive; drive correctness, performance, and stability to an internal readiness bar.

  • Build a suite of benchmarks and stress tests that capture real E2E behavior: of our workloads by exercising all aspects of a system, including CPU, GPU, memory subsystem, frontend, scale-up, and scale-out networking (including WAN traffic, NVlink and RDMA collectives), storage, thermals, and any other relevant parts.

  • Deep-dive performance on distributed training/inference: Collective performance and tuning (across NCCL/RCCL and internal libraries)

  • Overlap of compute/communication, kernel-level bottlenecks, memory bandwidth and scheduling effects

  • Create repeatable test harnesses that run in CI / lab environments and: produce actionable outputs (pass/fail, performance score, regression detection).

  • Partner with systems + fleet bring-up engineers to ensure the platform is not: only stable and performant, but also operationally usable and scalable (containerization, K8s integration, telemetry hooks, failure triage loops).

What they're looking for

  • Work cross-functionally with vendors and internal stakeholders by producing: clear bug reports, minimal repros, and prioritized issue lists.
  • BS in CS/EE (or equivalent practical experience).
  • + years in one or more of: ML systems, performance engineering, distributed systems, or HPC.
  • Strong hands-on experience with: PyTorch and modern LLM training/inference stacks