The opportunity
We bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API.
What you'll do
Analyze and optimize performance across application, middleware, runtime, and: infrastructure layers—networking, storage, Python runtime, GPU utilization, and beyond.
Develop tooling and metrics that provide deep observability into system performance.
Collaborate closely with infra, platform, training, and product teams to: identify key performance goals and drive systemic improvements.
Influence architecture and design decisions to prioritize latency, throughput, and efficiency at scale.
Lead investigations into high-impact performance regressions or scalability issues in production.
Drive performance testing strategies and help define SLAs/SLOs around latency: and throughput for critical systems.
What they're looking for
- Have 7+ years of experience in software engineering with a strong track: record in performance or reliability of high-scale distributed systems.
- Are deeply comfortable with performance profiling tools and tracing systems.
- Have experience optimizing performance across one or more layers of the stack: (e.g., database, networking, storage, application runtime, GC tuning, Python/Golang internals, GPU utilization).
- Have a strong understanding of OS internals, scheduling, memory management, and IO patterns.