The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Design, build, and maintain test infrastructure and automation for deploying: and validating the Cerebras Inference Platform.
Validate the platform across environments: from cloud-managed Kubernetes to deployments running on Cerebras hardware.
Test and verify deployment infrastructure including Kubernetes workloads,: CI/CD pipelines, ingress and service discovery, NGINX, and load balancing.
Collaborate closely with the Inference Platform development team to ensure: new features and platform capabilities ship reliably.
Investigate and debug complex issues spanning networking, orchestration, deployment, and distributed services.
Develop and maintain testbeds used to validate platform performance, scalability, and reliability.
What they're looking for
- Identify failure points, bottlenecks, and edge cases that impact platform stability and inference performance.
- Contribute to test plans and validation strategies for new platform features and releases.
- Improve observability, diagnostics, and debugging workflows across the inference platform stack.
- Partner with engineering teams to ensure high-quality, production-ready: releases of the Cerebras Inference Platform. Minimum Skills & Qualifications