The opportunity
We’re hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas—including Core Distributed Systems, Reliability Engineering,…
What you'll do
Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI
Systems Engineering: Work across layers of the stack—debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability.
Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience.
Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale.
Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely.
Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads.
What they're looking for
- Databases: Building high performance, distributed database systems that power all of OpenAI's product stack.
- Design, build, and maintain reliable and performant systems used across: engineering. Work with your team to define technical strategy, architecture, and long-term goals.
- Collaborate with other engineers, product managers, and researchers to build: infrastructure that meets evolving needs.
- Improve internal tooling, automation, and developer experience.