The opportunity
OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers,…
What you'll do
Own the operational health, availability, and reliability of production AI: network infrastructure across Industrial Compute's data centers.
Monitor, troubleshoot, and resolve network incidents while meeting: service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).
Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.
Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.
Manage the hardware lifecycle, including switch and optics replacements, RMA: coordination, software upgrades, and preventive maintenance.
Support new AI cluster deployments, data center expansions, and: infrastructure migrations in partnership with deployment and engineering teams.
What they're looking for
- Partner with cloud service providers (CSPs), colocation providers, Smart: Hands teams, and hardware vendors to maintain production infrastructure.
- Perform root-cause analysis (RCA) for production incidents and drive: permanent corrective actions that eliminate recurring issues.
- Build and maintain monitoring, telemetry, dashboards, and alerting to improve: network observability and proactive issue detection.
- Develop and improve operational runbooks, playbooks, troubleshooting: documentation, and standard operating procedures.