The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Define and implement a robust strategy for delivering and running software: reliably and at scale across multiple datacenters and cloud-based solutions.
Architect self-service platforms and internal tooling that lets product: teams, external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.
Define and evolve reliability practices for inference workloads, including: SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.
Mentor mid-level SREs, support critical incident escalations, and use: production pain points to prioritize the highest-leverage automation work.
Measure and drive impact through clear metrics, including toil reduction,: deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
+ years in SRE, infrastructure engineering, or platform engineering, with a: strong record of improving automation and reliability at large scale in FAANG, hyperscaler, or similarly demanding environments.
What they're looking for
- Deep expertise operating large scale heterogenous clusters with a proprietary cloud control plane
- Proven track record designing and delivering CI/CD or GitOps systems using: Argo CD or similar tools, with strong safety and observability built in.
- Hands-on experience with observability systems such as Loki, Tempo, Mimir, and Prometheus
- Ability to lead complex projects end to end, influence cross-functional: stakeholders, and communicate technical direction clearly.