The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Participate in bring-up of next-generation AI hardware systems and supporting software infrastructure.
Debug complex system-level issues spanning hardware and software interactions.
Investigate failures occurring during system bring-up and identify root: causes using logs, telemetry, and diagnostic tools.
Build automation frameworks and internal tooling that improve system validation and debugging workflows.
Develop software used to test, validate, and stress distributed hardware: systems during development and production cycles.
Collaborate closely with hardware engineers to isolate and resolve system integration issues.
What they're looking for
- Improve system observability by building tools that surface failures quickly and accelerate debugging.
- Reproduce, triage, and diagnose difficult issues that arise during early hardware deployment.
- Support validation and qualification of new hardware generations as systems move toward production readiness.
- Continuously improve internal engineering workflows related to debugging, testing, and automation.