The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Remain hands-on with operational execution (releases, capacity changes,: cluster upgrades) over the next year as we build robust continuous delivery pipelines and self-service capabilities
Contribute to the development of self-service CD pipelines for key workflows: using our stack: Kubernetes, Bazel, Prometheus/Grafana/InfluxDB, Python, and Go.
Build reusable automation and internal developer tools that minimize operational toil and cross-team friction
Develop and extend telemetry, observability and alerting solutions to ensure operational reliability at scale
Collaborate with Cluster Ops and development teams to identify high-impact: automation opportunities and iterate quickly
Contribute to reliability practices (SLOs, post-mortems, capacity planning)
What they're looking for
- + years in SRE with a strong operations or automation focus
- Production Kubernetes experience
- Solid Python or Go for building tools and automation
- Proficiency with Prometheus, Grafana, and observability-driven workflows