The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Drive technical vision for Lambda's Managed Kubernetes bare-metal platform,: including control plane scalability, multi-tenancy, cluster lifecycle management, and high availability
Integrate and extend NVIDIA's open-source ecosystem: GPU Operator, Network Operator, DCGM, NCCL, and emerging projects like AICR and Topograph for topology-aware scheduling and placement
Design GPU-aware orchestration systems
Lead development of services that power our managed services
Inform on and help with networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect. You will work closely with our Network team to define and drive requirements
Inform and help with storage architecture requirements for AI workloads. You: will partner with Storage teams on what managed K8s, Slurm, and future services need
What they're looking for
- Build the foundation for Managed Slurm on Kubernetes, enabling traditional: HPC workloads to run seamlessly alongside Kubernetes workload
- Design higher-level platform services for inference, including model serving: infrastructure, autoscaling based on inference load, and multi-model deployment patterns
- Design self-healing systems and automation for incident response, root cause analysis, and platform resilience
- Lead chaos engineering efforts to validate system behavior under failure conditions at scale