The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Architect, deploy, and operate Kubernetes clusters across Lambda's bare-metal datacenters.
Build and maintain automation for cluster lifecycle management: provisioning,: upgrades, and scaling. Similar to SRE-level coding abilities (small scripts, k8s operators/CRDs a plus)
Own the reliability, performance, and security of Kubernetes workloads in production.
Implement observability, logging, and alerting for clusters and critical workloads.
Partner with product teams to design scalable, cloud-native services.
Set the standards for resource management, networking, and RBAC across the platform.
What they're looking for
- Lead incident response, root-cause analysis, and post-mortems for platform issues.
- Mentor junior engineers and raise the bar for platform engineering across the org.
- + years in Platform, Infrastructure, or SRE roles, including running Kubernetes in production at scale.
- Deep knowledge of Kubernetes internals and day-2 operations (upgrades, scaling, troubleshooting).