Staff Kernel Optimzation EngineerActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Develop design specifications for new machine learning and linear algebra: kernels and mapping to the Cerebras WSE System using various parallel programming algorithms.

  • Develop and debug kernel library of highly optimized low level assembly: instruction and C-like domain specific language routines to implement algorithms targeting the Cerebras hardware system.

  • Develop and debug high-performance kernel routines in low-level assembly and: a custom C-like (CSL) language, implementing algorithms optimized for the Cerebras hardware system.

  • Using mathematical models and analysis to measure the software performance and inform design decisions.

  • Develop and integrate unit and system testing methodologies to verify correct: functionality and performance of kernel libraries.

  • Study emerging trends in Machine Learning applications and help evolve Kernel: library architecture to address computational challenges of the start-of-the-art Neural Networks.

What they're looking for

  • Interact with chip and system architects to optimize instruction sets,: microarchitecture, and IO of next generation systems.
  • Bachelor’s, Master’s, PhD or foreign equivalents in Computer Science,: Computer Engineering, Mathematics, or related fields.
  • Understanding of hardware architecture concepts: must be comfortable learning the details of a new hardware architecture.
  • Skilled in C++ and Python programming languages.