Site Reliability Engineer - Ops & AutomationActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Remain hands-on with operational execution (releases, capacity changes,: cluster upgrades) over the next year as we build robust continuous delivery pipelines and self-service capabilities

  • Contribute to the development of self-service CD pipelines for key workflows: using our stack: Kubernetes, Bazel, Prometheus/Grafana/InfluxDB, Python, and Go.

  • Build reusable automation and internal developer tools that minimize operational toil and cross-team friction

  • Develop and extend telemetry, observability and alerting solutions to ensure operational reliability at scale

  • Collaborate with Cluster Ops and development teams to identify high-impact: automation opportunities and iterate quickly

  • Contribute to reliability practices (SLOs, post-mortems, capacity planning)

What they're looking for

  • + years in SRE with a strong operations or automation focus
  • Production Kubernetes experience
  • Solid Python or Go for building tools and automation
  • Proficiency with Prometheus, Grafana, and observability-driven workflows