The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Contribute to the end-to-end bring up of frameworks for RL, inference: serving, ML models on Cerebras CSX systems.
Work across the stack: model architecture translation, graph lowering, compiler optimizations, runtime integration, and performance tuning.
Debug performance and correctness issues spanning model code, compiler IRs,: runtime behavior, and hardware utilization.
Propose and prototype improvements across tools, APIs, or automation flows to accelerate future bring ups.
Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or a related field with 10+ years’ experience.
Comfort navigating the full AI toolchain: Python modelling code, compiler IRs, performance profiling, etc.
What they're looking for
- Strong debugging skills across performance, numerical accuracy, and runtime integration.
- Experience with deep learning frameworks (e.g., PyTorch, TensorFlow) and: familiarity with model internals (e.g., attention, MoE, diffusion).
- Proficiency in C/C++ programming and experience with low-level optimization.
- Strong background in optimization techniques, particularly those involving NP-hard problems.