Software Engineer, Kernel ReliabilityActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Contribute to the technical roadmap and execution for kernel-centric: reliability of our internal and customer-facing systems.

  • Partner with System and Cluster Operations teams to reduce system and service: downtime after failure through tooling, analysis, and hands-on debugging support.

  • Work with the Debug Team to enhance debug tools with the goal of speeding up failure analysis.

  • Collaborate with software teams to improve the software stack—including: kernels—to improve on-field debugging and failure analysis.

  • Work with ASIC and hardware architecture teams to co-design next-generation: architectures with reliability and ease of debug in mind.

  • Participate in incident response, root-cause analysis, and post-mortems;: drive follow-ups that measurably improve reliability over time.

What they're looking for

  • We recognize great engineers come from different backgrounds. If you're: excited about the role, we encourage you to apply even if you don't meet every qualification.
  • Required (or demonstrated through projects/internships/coursework): Strong programming skills in C/C++ and Python.
  • Solid foundations in operating systems, computer architecture, and systems programming fundamentals.
  • Ability to debug complex issues using logs, traces, and standard debugging: workflows; interest in root-cause analysis.