Principal SRE - AI InferenceActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Define and implement a robust strategy for delivering and running software: reliably and at scale across multiple datacenters and cloud-based solutions.

  • Architect self-service platforms and internal tooling that let product teams,: external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.

  • Define and evolve reliability practices for inference workloads, including: SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.

  • Mentor senior SREs, support critical incident escalations, and use production: pain points to prioritize the highest-leverage automation work.

  • Measure and drive impact through clear metrics, including toil reduction,: deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.

  • + years in SRE, infrastructure engineering, or platform engineering, with a: record of setting technical direction and delivering reliability improvements at large scale in FAANG, hyperscaler, frontier AI, or similarly demanding production environments.

What they're looking for

  • Deep experience with large-scale compute fleets, internal control planes,: schedulers, orchestration systems, capacity management, and reliability automation.
  • Experience defining and driving cross-team architecture for production: control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership.
  • Strong judgment in converging fragmented workflows, tools, and teams into: coherent architectures that improve reliability, efficiency, and operational leverage.
  • Ability to lead complex, ambiguous technical programs end to end; influence: senior cross-functional stakeholders; mentor senior engineers; and communicate technical strategy clearly.