Senior Site Reliability Engineer - Core Cloud PlatformActive$240K–$356K

The opportunity

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you'll do

  • Operate and scale critical platform services across Lambda’s data centers.

  • Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems.

  • Build monitoring, alerting, and tracing for service health, provisioning: latency, and customer-impacting failures.

  • Define SLIs, SLOs, error budgets, and operational readiness standards.

  • Automate detection and remediation of configuration drift, failed workflows, and orphaned resources.

  • Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps.

What they're looking for

  • Design fault-isolation mechanisms that reduce blast radius and prevent cascading failures.
  • Lead production incident response, postmortems, and durable corrective actions.
  • Partner with Compute, Networking, Storage, Security, and Support teams.
  • Participate in on-call and improve its sustainability through automation and better tooling.