The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Platform Direction. Help shape the technical direction for the Inference: Cloud Platform, including multi-region topology, failure domains, service boundaries, and system evolution over time, and own the roadmap for major technical areas.
Core Cloud Systems. Design and build critical platform components such as: service discovery, request routing, load balancing, caching, batching, and traffic management for AI inference workloads.
Reliability & Performance. Architect active-active systems with rapid: failover, graceful degradation, and clear SLOs. Drive system-level improvements in latency, throughput, capacity efficiency, and resilience under unpredictable demand.
Traffic Control & Service Tiers. Define platform mechanisms for admission: control, quota management, rate limiting, and differentiated quality of service across workload types and customer tiers.
Execution on Critical Paths. Write and review production code in the most: important parts of the platform. Make high-consequence architectural decisions within your area and set the technical bar through design reviews, code reviews, and sound engineering judgment.
Production Leadership. Lead on the hardest production issues and cross-system: bottlenecks. Drive observability, incident response, capacity planning, and post-incident improvement with a high standard for operational rigor.
What they're looking for
- Technical Influence. Partner with ML, Product, Infrastructure, and Platform: teams to translate product and business requirements into scalable system designs, and drive alignment on shared technical decisions within your domain and adjacent platform surfaces.
- Mentorship. Raise the effectiveness of senior engineers through design: feedback, pairing, and clear technical standards.
- + years of experience in software engineering, with substantial individual: contributor experience building and operating large-scale distributed systems or cloud infrastructure.
- Deep expertise in distributed systems architecture in cloud environments,: including networking, compute orchestration, container platforms, and multi-region production services.