CoDesign & NextGen Performance EngineerActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Bring up and optimize performance on new generations of the Cerebras WSE.

  • Build performance models (kernel-level, end-to-end) to estimate the: performance of state of the art and customer ML models.

  • Optimize and debug our kernel micro code and compiler algorithms to elevate: ML model inference speed, throughput and compute utilization on the Cerebras WSE.

  • Debug and understand runtime performance on the system and cluster.

  • Develop tools and infrastructure to help visualize performance data collected: from the Wafer Scale Engine and our compute cluster.

  • Bachelors / Masters / PhD in Electrical Engineering or Computer: Science. Strong background in computer architecture.

What they're looking for

  • Exposure to and understanding of low-level deep learning / LLM math.
  • Strong analytical and problem-solving mindset.
  • + years of experience in a relevant domain (Computer Architecture, CPU/GPU: Performance, Kernel Optimization, HPC).
  • Experience working on CPU/GPU simulators.
CoDesign & NextGen Performance Engineer at Cerebras | Role Match