The opportunity
Lambda is building the AI Cloud of the future. We are seeking a Staff Engineer to help our development of our Managed Kubernetes platform.
What you'll do
Drive technical vision for Lambda's Managed Kubernetes bare-metal platform,: including control plane scalability, multi-tenancy, cluster lifecycle management, and high availability
Integrate and extend NVIDIA's open-source ecosystem: GPU Operator, Network Operator, DCGM, NCCL, and emerging projects like AICR and Topograph for topology-aware scheduling and placement
Design GPU-aware orchestration systems
Lead development of services that power our managed services
Inform on and help with networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect. You will work closely with our Network team to define and drive requirements
Inform and help with storage architecture requirements for AI workloads. You: will partner with Storage teams on what managed K8s, Slurm, and future services need
What they're looking for
- Experience building and operating managed Kubernetes services (GKE, EKS, AKS,: or similar) or working on Kubernetes control plane components
- Hands-on experience with NVIDIA's open-source ecosystem beyond GPU Operator:: Network Operator, NCCL tuning, Topograph, AICR, or similar emerging projects
- Familiarity with HPC and traditional job schedulers (Slurm) and: Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
- Background in confidential computing
- Experience migrating customers or workloads from legacy/bespoke infrastructure to standardized platforms
- Contributions to CNCF projects, Kubernetes SIGs, or NVIDIA open-source projects
- Familiarity with security and compliance in multi-tenant environments: RBAC, Pod Security Standards, network policies, workload isolation
- Background in ML infrastructure: training clusters, inference serving, simulation