The opportunity
Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products.
What you'll do
Design, build, and operate Kubernetes-based controllers and distributed: services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.
Define APIs and resource models that let clients request and track lifecycle: operations through consistent interfaces across hardware platforms and providers.
Build provisioning and configuration services that coordinate network boot,: hardware management interfaces, and the deployment of firmware, operating-system images, drivers, and host configuration.
Develop lifecycle management for discovery, allocation, provisioning,: upgrades, maintenance, recovery, and decommissioning, integrating with health and validation systems.
Design reliable reconciliation and recovery through concurrent changes,: interrupted operations, and partial failures, with staged rollouts that limit disruption across nodes, racks, and clusters.
Improve control-plane throughput, API latency, and the time infrastructure: takes to reach its desired state, while respecting the limits of site systems and provider APIs.
What they're looking for
- Build the software integrations that bring new sites and GPU hardware: generations into the platform, partnering with hardware, networking, data-center, and other infrastructure teams.
- Have strong software engineering fundamentals and experience designing,: implementing, and owning production distributed systems or infrastructure services.
- Have experience developing infrastructure systems that use Kubernetes APIs: and reconciliation to manage resources.
- Understand how a bare-metal node moves from power-on to a configured,: workload-ready system, with depth in one or more areas such as PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management.