The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Drive a no-downtime operating model across all facilities
Own uptime performance end-to-end including incident response and RCA
Build a culture where failures are analyzed and not repeated
Architect and run a high-discipline vendor operating model
Convert vendor relationships into measurable systems
Lead deep operational reviews and enforce accountability
What they're looking for
- Align SLAs, incentives, and performance
- Eliminate ambiguity in ownership and response
- Build a zero-incident safety culture
- Ensure rigor in high-risk procedures