Staff Site Reliability Engineer – Automation and PlatformActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Define and implement a robust strategy for delivering and running software: reliably and at scale across multiple datacenters and cloud-based solutions.

  • Architect self-service platforms and internal tooling that lets product: teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.

  • Define and evolve reliability practices for inference workloads, including: SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.

  • Mentor mid-level SREs, support critical incident escalations, and use: production pain points to prioritize the highest-leverage automation work.

  • Measure and drive impact through clear metrics, including toil reduction,: deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.

  • + years in SRE, infrastructure engineering, or platform engineering, with a: strong record of improving automation and reliability at large scale in FAANG, hyperscaler, or similarly demanding environments.

What they're looking for

  • Deep expertise operating large scale heterogenous clusters with a proprietary cloud control plane
  • Proven track record designing and delivering CI/CD or GitOps systems using: Argo CD or similar tools, with strong safety and observability built in.
  • Hands-on experience with observability systems such as Loki, Tempo, Mimir, and Prometheus
  • Ability to lead complex projects end to end, influence cross-functional: stakeholders, and communicate technical direction clearly.