The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Define and own the reliability strategy for internal IT systems, including: SLOs, error budgets, and operational health reporting.
Build and lead a team of IT SRE engineers focused on automation,: observability, and incident response for corporate systems.
Design and implement automation to eliminate manual IT work across: provisioning, access management, patching, and lifecycle operations.
Instrument internal services and SaaS integrations with monitoring, alerting, and on-call workflows.
Run incident response for IT outages, including root cause analysis and durable remediation.
Drive infrastructure-as-code and GitOps practices across IT-owned systems.
What they're looking for
- Partner with security and networking teams on identity, access, and network reliability.
- Minimum 8 years of experience in SRE, DevOps, or IT engineering roles, with: at least 2 years in a leadership capacity.
- Direct hands on experience with AI coding tools, building and deploying AI agents for triage and bug fixes.
- Strong software engineering background with hands-on experience in Python,: Go, or similar, and comfort writing production-grade automation.