The opportunity
Cerebras is building a new generation of disaggregated AI inference systems that combine GPU-accelerated prefill with ultra-fast decode on the Cerebras Wafer-Scale Engine.
What you'll do
Productionize the GPU inference stack. Design, build, deploy, and maintain: the complete GPU prefill path, spanning API services, model-serving workers, vLLM, PyTorch, ROCm, GPU nodes, networking, and rack-scale infrastructure.
Own GPU operational readiness. Establish deployment, upgrade, rollback,: health-checking, capacity-management, and failure-recovery practices for the AMD GPU fleet. Build automation that makes driver, firmware, runtime, model, and container compatibility explicit and reproducible.
Drive reliability in production. Define service-level indicators and: objectives for GPU-backed inference. Improve fault isolation, graceful degradation, automated recovery, incident response, and post-incident remediation across the serving stack.
Improve inference performance. Profile and optimize time to first token,: request throughput, tokens per second per GPU, tail latency, GPU utilization, memory efficiency, and rack-level capacity under representative production workloads.
Optimize model-serving behavior. Tune and improve scheduling, continuous: batching, prefix caching, KV-cache management, tensor and expert parallelism, request admission, quantization, graph execution, and distributed communication.
Debug across system layers. Diagnose complex failures and performance: regressions across application code, vLLM, PyTorch, ROCm/HIP, collective communication libraries, kernels, drivers, firmware, networking, and hardware.
What they're looking for
- Experience with AMD Instinct accelerators and the ROCm ecosystem, including: HIP, RCCL, rocprofiler, AMD SMI, AITER, hipBLASLt, Composable Kernel, or related libraries and tools.
- Deep CUDA experience that demonstrates an ability to transfer GPU systems: knowledge across accelerator platforms.
- Experience modifying or contributing to vLLM, SGLang, PyTorch, Triton,: TensorRT-LLM, or another open-source ML systems project.
- Experience optimizing prefill-heavy or disaggregated prefill/decode inference architectures.
- Understanding of KV-cache transfer, prefix caching, continuous batching,: chunked prefill, request scheduling, and memory-aware admission control.
- Experience with multi-GPU and multi-node inference, including tensor: parallelism, pipeline parallelism, expert parallelism, RDMA, collective communication, and failure handling.
- Experience optimizing Mixture-of-Experts or multimodal models.
- Knowledge of GPU kernel optimization, operator fusion, graph capture,: attention kernels, GEMM tuning, and communication/computation overlap.