The opportunity
This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly.
What you'll do
Build and operate reliable infrastructure for research workloads and research-facing services.
Support and improve systems across data infrastructure, processing, crawl and: ingest, caching, search, observability, and clusterwide services.
Improve cluster bootstrapping, provisioning, automation, and deployment workflows.
Debug issues across networking, compute, storage, orchestration, and service reliability layers.
Build software and automation that reduce manual operational work and improve system reliability.
Partner closely with researchers, infrastructure engineers, and service: owners to understand system needs and translate them into durable solutions.
What they're looking for
- Help evolve existing infrastructure toward more scalable, maintainable, and standard patterns.
- Take ownership of critical systems and drive work independently from problem definition through execution.
- Have strong systems fundamentals and understand how infrastructure scales in practice.
- Are comfortable with Linux, networking, Kubernetes, provisioning, and distributed systems operations.