Network Operations Engineer, AI NetworkingActive$157K–$302K

The opportunity

OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers,…

What you'll do

  • Own the operational health, availability, and reliability of production AI: network infrastructure across Industrial Compute's data centers.

  • Monitor, troubleshoot, and resolve network incidents while meeting: service-level objectives (SLOs), reducing Mean Time to Detect (MTTD), and minimizing Mean Time to Recovery (MTTR).

  • Operate and maintain large-scale Ethernet fabrics supporting GPU compute, storage, and management networks.

  • Execute production network changes, maintenance windows, and capacity expansions with minimal customer impact.

  • Manage the hardware lifecycle, including switch and optics replacements, RMA: coordination, software upgrades, and preventive maintenance.

  • Support new AI cluster deployments, data center expansions, and: infrastructure migrations in partnership with deployment and engineering teams.

What they're looking for

  • Partner with cloud service providers (CSPs), colocation providers, Smart: Hands teams, and hardware vendors to maintain production infrastructure.
  • Perform root-cause analysis (RCA) for production incidents and drive: permanent corrective actions that eliminate recurring issues.
  • Build and maintain monitoring, telemetry, dashboards, and alerting to improve: network observability and proactive issue detection.
  • Develop and improve operational runbooks, playbooks, troubleshooting: documentation, and standard operating procedures.