The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Develop and maintain automation for deployment workflows, including: provisioning, configuration, validation, and operational handoff.
Turn manual deployment steps into tested, repeatable pushbutton workflows.
Participate in hands-on cluster deployments to build practical debugging and operational expertise.
Troubleshoot issues across Linux systems, bare-metal servers, networking,: storage, Kubernetes, and connectivity.
Contribute to infrastructure-as-code and GitOps workflows using tools such as: Terraform, Ansible, pull requests, and code review.
Add health checks, observability, dashboards, and validation logic to improve deployment reliability.
What they're looking for
- Partner with networking, infrastructure, security, and operations teams to: deliver secure and reproducible data center deployments.
- + years of mid: to large-scale data center deployment
- Strong fundamentals in Python and Bash, with the ability to write scripts and small programs.
- Basic Linux experience, including command-line usage, processes, filesystems, and disk troubleshooting.