The opportunity
Join the engineering teams that bring OpenAI’s ideas safely to the world!
What you'll do
Design and implement solutions to ensure the scalability of our: infrastructure to meet rapidly increasing demands.
Build and maintain the load, chaos and synthetic-testing software leveraged: by development teams to make the systems they design and operate more reliable.
Build and maintain automation tools to streamline repetitive tasks and improve system reliability.
Build and maintain the platform for CPU, storage, GPU, and network lifecycle: management to drive efficiency, accountability and dynamic optimization of our resources.
Implement fault-tolerant and resilient design patterns to minimize service disruptions.
Develop and maintain service level objectives (SLOs) and service level: indicators (SLIs) to measure and ensure system reliability.
What they're looking for
- Partner with researchers, engineers, product managers, and designers to bring: new features and research capabilities to the world.
- Participate in an on-call rotation to respond to critical incidents and ensure 24/7 system availability.
- Have a track record of accelerating engineering reliability by building: world-class tooling and systems from 0→1.
- Have a humble attitude, an eagerness to help your colleagues, and a desire to: do whatever it takes to make the team succeed.