Staff GPU Inference SDETNew

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Direct experience with either AMD (ROCm / HIP) or NVIDIA software stacks.

  • Experience building workload replay tools, ML evaluation pipelines, or MLPerf Inference benchmark suites.

  • Familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or C++

  • Build a breakthrough AI platform beyond the constraints of the GPU.

  • Publish and open source their cutting-edge AI research.

  • Work on one of the fastest AI supercomputers in the world.

What they're looking for

  • Enjoy job stability with startup vitality.
  • Our simple, non-corporate work culture that respects individual beliefs.