Software Engineer, GPU InferenceActive

The opportunity

Cerebras is building a new generation of disaggregated AI inference systems that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine.

What you'll do

  • Productionize the GPU inference stack. Design, build, deploy, and maintain: the complete GPU prefill path, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.

  • Own GPU operational readiness. Establish deployment, upgrade, rollback,: health-checking, capacity-management, and failure-recovery practices for the AMD GPU fleet. Build automation that makes driver, firmware, runtime, model, and container compatibility explicit and reproducible.

  • Drive reliability in production. Define service-level indicators and: objectives for GPU-backed inference. Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation across the serving stack.

  • Improve inference performance. Profile and optimize time to first token,: request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity under representative production workloads.

  • Optimize model-serving behavior. Tune and improve scheduling, continuous: batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.

  • Debug across system layers. Diagnose complex failures and performance: regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.

What they're looking for

  • Experience with AMD Instinct accelerators and the ROCm ecosystem, including: HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, Composable Kernel, or related libraries and tools.
  • Deep CUDA experience that demonstrates an ability to transfer GPU systems: knowledge across accelerator platforms.
  • Experience modifying or contributing to vLLM, SGLang, PyTorch, Triton,: TensorRT-LLM, or another open-source ML systems project.
  • Experience optimizing prefill-heavy or disaggregated prefill/decode inference architectures.
  • Understanding of KV-cache transfer, prefix caching, continuous batching,: chunked prefill, request scheduling, and memory-aware admission control.
  • Experience with multi-GPU and multi-node inference, including tensor: parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling.
  • Experience optimizing Mixture-of-Experts or multimodal models.
  • Knowledge of GPU kernel optimization, operator fusion, graph capture,: attention kernels, GEMM tuning, and communication/computation overlap.
Software Engineer, GPU Inference at Cerebras | Role Match