Lead Full Stack Machine Learning EngineerActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Contribute to the end-to-end bring up of frameworks for RL, inference: serving, ML models on Cerebras CSX systems.

  • Work across the stack: model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.

  • Debug performance and correctness issues spanning model code, compiler IRs,: runtime behavior, and hardware utilization.

  • Propose and prototype improvements across tools, APIs, or automation flows to accelerate future bring ups.

  • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or a related field with 10+ years’ experience.

  • Comfort navigating the full AI toolchain: Python modelling code, compiler IRs, performance profiling, etc.

What they're looking for

  • Strong debugging skills across performance, numerical accuracy, and runtime integration.
  • Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and: familiarity with model internals (e.g., attention, MoE, diffusion).
  • Proficiency in C/C++ programming and experience with low-level optimization.
  • Strong background in optimization techniques, particularly those involving NP-hard problems.