The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Contribute to the technical roadmap and execution for kernel-centric: reliability of our internal and customer-facing systems.
Partner with System and Cluster Operations teams to reduce system and service: downtime after failure through tooling, analysis, and hands-on debugging support.
Work with the Debug Team to enhance debug tools with the goal of speeding up failure analysis.
Collaborate with software teams to improve the software stack—including: kernels—to improve on-field debugging and failure analysis.
Work with ASIC and hardware architecture teams to co-design next-generation: architectures with reliability and ease of debug in mind.
Participate in incident response, root-cause analysis, and post-mortems;: drive follow-ups that measurably improve reliability over time.
What they're looking for
- We recognize great engineers come from different backgrounds. If you're: excited about the role, we encourage you to apply even if you don't meet every qualification.
- Required (or demonstrated through projects/internships/coursework): Strong programming skills in C/C++ and Python.
- Solid foundations in operating systems, computer architecture, and systems programming fundamentals.
- Ability to debug complex issues using logs, traces, and standard debugging: workflows; interest in root-cause analysis.