The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Design, develop, test, and maintain production software, with: responsibilities spanning testing, continuous development, observability, security, networking, debugging, and productionization.
Platform Direction. Help shape the technical direction for the Inference: Platform, Kubernetes custom resource definitions, failure domains, service boundaries, and system evolution over time, and own the roadmap for major technical areas.
Reliability & Performance. Architect active-active systems with rapid: failover, graceful degradation, and clear SLOs. Drive system-level improvements in latency, throughput, capacity efficiency, and resilience under unpredictable demand.
Execution on Critical Paths. Write and review production code in the most: important parts of the platform. Make high-consequence architectural decisions within your area and set the technical bar through design reviews, code reviews, and sound engineering judgment.
Production Leadership. Lead on the hardest production issues and cross-system: bottlenecks. Drive observability, incident response, capacity planning, and post-incident improvement with a high standard for operational rigor.
Technical Influence. Partner with ML, Product, Infrastructure, and Cloud: teams to translate product and business requirements into scalable system designs, and drive alignment on shared technical decisions within your domain and adjacent platform surfaces.
What they're looking for
- + years of experience in software engineering, with experience building and: operating large-scale distributed systems or cloud infrastructure.
- Experience in distributed systems, ideally with Kubernetes.
- Experience building highly available, latency-sensitive systems at scale.
- Experience with security (certificates, TLS, mTLS).