The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Lead the design and implementation of system-level debugging, validation, and observability platforms.
Develop automated systems for collecting and analyzing numerical, and execution anomalies.
Create visualization and analysis tools to enable efficient root-cause investigation.
Build frameworks for failure classification, regression detection, and anomaly monitoring.
Extend compilers, runtimes, and programming interfaces to support advanced profiling and instrumentation.
Improve system bring-up, low-level debug, and validation workflows.
What they're looking for
- Partner cross-functionally with compiler, hardware, firmware, runtime, and infrastructure teams.
- Establish best practices for debuggability, reliability, and operational excellence.
- Lead high-impact initiatives.
- Support incident response and drive long-term corrective actions.