The opportunity
We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute clusters. These clusters would provide the candidate with an opportunity to work with the world's largest computer chip, the Wafer-Scale…
What you'll do
Deploy, configure, and debug container-based services using Docker.
Build and own software solutions that power cluster operations, including: monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling.
Collaborate with cross-functional teams to translate operational requirements: into scalable O&M products and platform capabilities.
Develop APIs, automation services, and integrations that improve operational: visibility, incident response, and fleet management across global AI infrastructure.
Manage and operate multiple advanced AI compute infrastructure clusters.
Monitor and oversee cluster health, proactively identifying and resolving potential issues.
What they're looking for
- Maximize compute capacity through optimization and efficient resource allocation.
- Provide 24/7 monitoring and support, leveraging automated tools and: performing hands-on troubleshooting as needed.
- Handle engineering escalations and collaborate with other teams to resolve complex technical challenges.
- Stay up-to-date with the latest advancements in AI compute infrastructure and related technologies.