The opportunity
ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond to operational issues, and use infrastructure effectively at scale.
What you'll do
Build and evolve the platform used to deploy, configure, and manage models across ChatGPT.
Develop systems for deployment orchestration, model rollouts, operational: visibility, and production readiness.
Create abstractions and tooling that simplify complex infrastructure and improve the developer experience.
Automate operational workflows, including incident detection, diagnosis, mitigation, and recovery.
Improve the reliability, scalability, and efficiency of model deployments and: the infrastructure that supports them.
Build systems that support capacity planning, resource allocation, and infrastructure utilization.
What they're looking for
- Five or more years of software engineering experience building production infrastructure.
- Strong programming skills in Go, Python, C++, Rust, or a comparable language.
- Experience designing or operating highly available distributed systems.
- Experience with platform engineering, infrastructure engineering, production: engineering, site reliability engineering, or similar disciplines.
- Strong systems design, debugging, and operational problem-solving skills.
- Strong communication skills and experience collaborating across engineering teams.