The opportunity
Cerebras Systems is a pioneer in large-scale AI Supercomputers. These multi-exaflop supercomputers are deployed in some of the biggest datacenters.
What you'll do
Automate bare-metal configuration of networking, OS, and application software: in large clusters of Cerebras WSE, servers, and switches.
Additional push button workflows for cluster upgrades, downgrades, and: security patching with key metrics to minimize downtime on clusters.
An orchestration and scheduler system for resource allocation, job submission: C placements for a multi-user environment on a cluster.
Seamless support for both on-premise and cloud mode deployment and operations.
A robust system for monitoring, detecting and handling failures for a variety: of resources on the clusters (including High Availability of clusters).
Broad cluster and job monitoring and visualization capabilities, along with alerting systems.
What they're looking for
- User facing tools to monitor the status of jobs and collect metrics.
- Administrator facing tools to manage and operate large clusters.