Hardware Operations EngineerActive$157K–$302K

The opportunity

OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem. Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads.

What you'll do

  • Drive technical triage and resolution of complex hardware failures impacting production systems.

  • Partner with Fleet Health Engineering to investigate recurring hardware: issues, identify failure patterns, and improve fleet reliability.

  • Lead root cause analysis (RCA) efforts for critical hardware incidents and: develop corrective and preventive action plans.

  • Collaborate with Cloud Service Provider operations teams and OEM vendors to: coordinate repairs, replacements, upgrades, and hardware lifecycle activities.

  • Establish and continuously improve hardware maintenance procedures,: operational runbooks, and troubleshooting standards.

  • Analyze hardware failure trends and operational metrics to identify: reliability risks and improvement opportunities.

What they're looking for

  • Experience supporting large-scale GPU clusters or AI/ML infrastructure environments.
  • Familiarity with fleet health systems, telemetry platforms, and hardware monitoring tools.
  • Ability to identify appropriate data and perform detailed analysis to support: all elements of this role, including dashboard development
  • Experience with failure analysis methodologies such as FRACAS, RCCA, 5-Why, Fishbone, or FMEA.
  • Knowledge of Linux system administration and hardware validation workflows.
  • Experience supporting hyperscale datacenter operations or HPC environments.
  • Familiarity with server manufacturing, rack integration, and NPI-to-sustaining transitions.
  • Industry certifications such as CompTIA Server+, OEM hardware certifications, or equivalent experience.