Staff Software Engineer, Inference APINew

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Build production ML inference APIs. Design, implement, and maintain APIs for: chat completions, text generation, streaming, model configuration, tool calling, structured outputs, multimodal inputs, and other emerging inference capabilities.

  • Deliver a unified serving experience. Create consistent request and response: semantics across GPU prefill, Cerebras decode, and other heterogeneous inference backends.

  • Enable new models and capabilities. Integrate emerging foundation models,: tokenizers, prompt formats, sampling methods, attention variants, multimodal inputs, and model-specific features into the serving platform.

  • Own API compatibility and evolution. Maintain compatibility with widely: adopted inference interfaces while designing Cerebras-specific extensions. Establish clear versioning, deprecation, validation, and backward compatibility practices.

  • Integrate with model-serving runtimes. Extend and integrate custom inference: services with vLLM, PyTorch, Hugging Face libraries, the AMD ROCm stack, and Cerebras runtime components.

  • Support disaggregated inference. Build the control and data paths required to: coordinate GPU prefill with Cerebras decode, including request routing, state transfer, error handling, retries, and lifecycle management.

What they're looking for

  • Improve serving performance. Optimize streaming behavior, time to first: token, request latency, throughput, batching, serialization, tokenization, scheduling, and communication between serving components.
  • Ensure functional and numerical correctness. Build validation systems for: tokenization, sampling, logits, generated outputs, precision changes, model upgrades, determinism, and compatibility across serving backends.
  • Strengthen reliability and observability. Define end-to-end service: indicators and build structured logging, tracing, metrics, dashboards, health checks, and diagnostic tooling for production inference traffic.
  • Develop testing and qualification infrastructure. Create conformance tests,: workload-replay tools, model-validation suites, performance benchmarks, integration tests, and release gates.