The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Define and implement a robust strategy for delivering and running software: reliably and at scale across multiple datacenters and cloud-based solutions.
Architect self-service platforms and internal tooling that let product teams,: external customers, and cluster operators safely trigger and observe critical workflows with minimal handoffs.
Define and evolve reliability practices for inference workloads, including: SLOs and SLIs for latency, throughput, and accuracy stability; error budgets; blameless postmortems; chaos testing; and capacity forecasting across multi-datacenter and on-prem environments.
Mentor senior SREs, support critical incident escalations, and use production: pain points to prioritize the highest-leverage automation work.
Measure and drive impact through clear metrics, including toil reduction,: deployment velocity, SLO compliance, MTTR, and adoption of self-service workflows.
+ years in SRE, infrastructure engineering, or platform engineering, with a: record of setting technical direction and delivering reliability improvements at large scale in FAANG, hyperscaler, frontier AI, or similarly demanding production environments.
What they're looking for
- Deep experience with large-scale compute fleets, internal control planes,: schedulers, orchestration systems, capacity management, and reliability automation.
- Experience defining and driving cross-team architecture for production: control planes, capacity orchestration, fleet management, or self-service infrastructure platforms with clear operational ownership.
- Strong judgment in converging fragmented workflows, tools, and teams into: coherent architectures that improve reliability, efficiency, and operational leverage.
- Ability to lead complex, ambiguous technical programs end to end; influence: senior cross-functional stakeholders; mentor senior engineers; and communicate technical strategy clearly.