The opportunity
The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI’s global compute fleet.
What you'll do
Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
Build and evolve health checks that detect, remediate, and verify failures at scale.
Ensure critical health checks execute with minimal latency to maximize workload uptime.
Investigate hardware failures and system-level issues across large-scale compute environments.
Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes.
Build automation and tooling that enables global cluster management with minimal manual intervention.
What they're looking for
- Partner with workload, reliability, and provider teams to integrate health: signals into training and inference systems.
- + years of industry experience in software or infrastructure engineering.
- Strong proficiency with Python and shell scripting.
- Experience building large-scale distributed systems or infrastructure platforms.