The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Build production ML inference APIs. Design, implement, and maintain APIs for: chat completions, text generation, streaming, model configuration, tool calling, structured outputs, multimodal inputs, and other emerging inference capabilities.
Deliver a unified serving experience. Create consistent request and response: semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends.
Enable new models and capabilities. Integrate emerging foundation models,: tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features into the serving platform.
Own API compatibility and evolution. Maintain compatibility with widely: adopted inference interfaces while designing Cerebras-specific extensions. Establish clear versioning, deprecation, validation, and backward compatibility practices.
Integrate with model-serving runtimes. Extend and integrate custom inference: services with vLLM, PyTorch, Hugging Face libraries, the AMD ROCm stack, and Cerebras runtime components.
Support disaggregated inference. Build the control and data paths required to: coordinate GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management.
What they're looking for
- Improve serving performance. Optimize streaming behavior, time to first: token, request latency, throughput, batching, serialization, tokenization, scheduling, and communication between serving components.
- Ensure functional and numerical correctness. Build validation systems for: tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and compatibility across serving backends.
- Strengthen reliability and observability. Define end-to-end service: indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling for production inference traffic.
- Develop testing and qualification infrastructure. Create conformance tests,: workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates.