The opportunity
OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads.
What you'll do
Drive technical triage and resolution of complex hardware failures impacting production systems.
Partner with Fleet Health Engineering to investigate recurring hardware: issues, identify failure patterns, and improve fleet reliability.
Lead root cause analysis (RCA) efforts for critical hardware incidents and: develop corrective and preventive action plans.
Collaborate with Cloud Service Provider operations teams and OEM vendors to: coordinate repairs, replacements, upgrades, and hardware lifecycle activities.
Establish and continuously improve hardware maintenance procedures,: operational runbooks, and troubleshooting standards.
Analyze hardware failure trends and operational metrics to identify: reliability risks and improvement opportunities.
What they're looking for
- Experience supporting large-scale GPU clusters or AI/ML infrastructure environments.
- Familiarity with fleet health systems, telemetry platforms, and hardware monitoring tools.
- Ability to identify appropriate data and perform detailed analysis to support: all elements of this role, including dashboard development
- Experience with failure analysis methodologies such as FRACAS, RCCA, 5-Why, Fishbone, or FMEA.
- Knowledge of Linux system administration and hardware validation workflows.
- Experience supporting hyperscale datacenter operations or HPC environments.
- Familiarity with server manufacturing, rack integration, and NPI-to-sustaining transitions.
- Industry certifications such as CompTIA Server+, OEM hardware certifications, or equivalent experience.