Senior Technical Program Manager – AI Infrastructure, Site OperationsActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Own end-to-end technical programs for data center and site operations

  • Act as single-threaded owner across: Hardware & Systems Engineering

  • AI Cloud Infrastructure & Operations

  • Network & Storage Engineering

  • Facilities, power, cooling, and colo partners

  • Drive site readiness for Cerebras Wafer-Scale Engine systems

What they're looking for

  • Partner on installation, commissioning, change management, and break/fix workflows
  • Lead incident reviews and postmortems; ensure corrective actions are closed
  • Define and own operational metrics and KPIs, including: Availability and reliability
  • Incident rate, severity, MTTR / MTTD