The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Direct experience with either AMD (ROCm / HIP) or NVIDIA software stacks.
Experience building workload replay tools, ML evaluation pipelines, or MLPerf Inference benchmark suites.
Familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or C++
Build a breakthrough AI platform beyond the constraints of the GPU.
Publish and open source their cutting-edge AI research.
Work on one of the fastest AI supercomputers in the world.
What they're looking for
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.