The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Contribute to the end-to-end bring up of ML models on Cerebras CSX systems.
Work across the stack: model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.
Debug performance and correctness issues spanning model code, compiler IRs,: runtime behavior, and hardware utilization.
Propose and prototype improvements across tools, APIs, or automation flows to accelerate future bring ups.
Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or a related field.
Comfort navigating the full AI toolchain: Python modeling code, compiler IRs, performance profiling, etc.
What they're looking for
- Strong debugging skills across performance, numerical accuracy, and runtime integration.
- Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and: familiarity with model internals (e.g., attention, MoE, diffusion).
- Proficiency in C/C++ programming and experience with low-level optimization.
- Proven experience in compiler development, particularly with LLVM and/or MLIR.