The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Design and implement observability instrumentation across services and platforms
Build and maintain telemetry pipelines for metrics, logs, and traces at scale
Develop internal observability platforms, libraries, and tooling
Define and operationalize SLIs, SLOs, and alerting strategies
Partner with engineers to make systems debuggable by design
Reduce MTTR by enabling fast root-cause analysis during incidents
What they're looking for
- Create clear, actionable dashboards and alerts that reflect real system health
- Balance telemetry signal vs cost, noise, and performance impact
- Improve the developer experience around observability and debugging
- Strong experience in backend or systems software engineering