Senior Site Reliability Engineer - FleetNew$240K–$356K

The opportunity

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you'll do

  • Build and operate monitoring and alerting for cluster health: fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively

  • Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible

  • Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand

  • Create runbooks and automated remediations for common cluster failure modes,: designed so Support and HPC Support can run them safely

  • Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL,: GPU-direct, fabric, switching, and power — working closely with on-site deployment teams

  • Participate in on-call rotations and lead incident response for cluster-level problems

What they're looking for

  • Contribute to and maintain Standard Operating Procedures, and feed clear: requirements back to other engineering teams on simplification, stability, and operational efficiency
  • + years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role
  • Have a strong understanding of modern AI infrastructure, from GPU: architectures to hardware performance optimization
  • Strong understanding of Linux-based systems in a distributed environment
Senior Site Reliability Engineer - Fleet at Lambda | Role Match