Cluster Operations Software EngineerNew

The opportunity

We are seeking a highly skilled and experienced AI Cluster Operations Engineer to manage and operate our cutting-edge machine learning compute clusters. These clusters would provide the candidate with an opportunity to work with the world's largest computer chip, the Wafer-Scale…

What you'll do

  • Deploy, configure, and debug container-based services using Docker.

  • Build and own software solutions that power cluster operations, including: monitoring platforms, workflow automation systems, operational dashboards, and reliability tooling.

  • Collaborate with cross-functional teams to translate operational requirements: into scalable O&M products and platform capabilities.

  • Develop APIs, automation services, and integrations that improve operational: visibility, incident response, and fleet management across global AI infrastructure.

  • Manage and operate multiple advanced AI compute infrastructure clusters.

  • Monitor and oversee cluster health, proactively identifying and resolving potential issues.

What they're looking for

  • Maximize compute capacity through optimization and efficient resource allocation.
  • Provide 24/7 monitoring and support, leveraging automated tools and: performing hands-on troubleshooting as needed.
  • Handle engineering escalations and collaborate with other teams to resolve complex technical challenges.
  • Stay up-to-date with the latest advancements in AI compute infrastructure and related technologies.