IT SRE Team LeadActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Define and own the reliability strategy for internal IT systems, including: SLOs, error budgets, and operational health reporting.

  • Build and lead a team of IT SRE engineers focused on automation,: observability, and incident response for corporate systems.

  • Design and implement automation to eliminate manual IT work across: provisioning, access management, patching, and lifecycle operations.

  • Instrument internal services and SaaS integrations with monitoring, alerting, and on-call workflows.

  • Run incident response for IT outages, including root cause analysis and durable remediation.

  • Drive infrastructure-as-code and GitOps practices across IT-owned systems.

What they're looking for

  • Partner with security and networking teams on identity, access, and network reliability.
  • Minimum 8 years of experience in SRE, DevOps, or IT engineering roles, with: at least 2 years in a leadership capacity.
  • Direct hands on experience with AI coding tools, building and deploying AI agents for triage and bug fixes.
  • Strong software engineering background with hands-on experience in Python,: Go, or similar, and comfort writing production-grade automation.
IT SRE Team Lead at Cerebras | Role Match