The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Operate and maintain bare-metal Kubernetes clusters , scaling up to thousands of nodes
Handle cluster degradation, recovery, resizing, and incident response using fleet management tools
Participate in a well-managed on-call rotation for critical incidents
Assist customers with Kubernetes questions, workload integration, storage, and authentication
Work closely with our HPC Ops and Datacenter Ops teams for low-level or cross-functional issues
Use Python and Golang to create tooling and automate the validation of platform quality.
What they're looking for
- Design, build, and maintain scalable control plane services, operators, and custom controllers for Kubernetes
- Develop automation for cluster lifecycle management: provisioning, upgrades, patching, and deletion.
- Define and implement SLOs and SLIs for Kubernetes services, workloads, and platform reliability.
- + years of experience in a SRE, operations engineer, or similar role, with a: deep knowledge of running Linux clusters and systems