The opportunity
Cerebras operates AI clusters across a growing number of data centers.
What you'll do
Build and operate software for managing large fleets of AI clusters
Give operators clear, actionable views of cluster health, capacity, performance, and ongoing issues.
Develop services and integrations that bring together data and workflows from multiple infrastructure systems.
Automate repetitive work and improve the tools teams use to investigate incidents and restore service.
Design systems that remain reliable as the fleet grows and continue to: function through component and site failures.
Work closely with users of the platform to understand their needs and make: practical product and engineering decisions.
What they're looking for
- Lead projects from initial design to production, measure their impact, and: use operational feedback to guide improvements.