Staff Software Engineer - Managed KubernetesActive$314K–$465K

The opportunity

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you'll do

  • Drive technical vision for Lambda's Managed Kubernetes bare-metal platform,: including control plane scalability, multi-tenancy, cluster lifecycle management, and high availability

  • Integrate and extend NVIDIA's open-source ecosystem: GPU Operator, Network Operator, DCGM, NCCL, and emerging projects like AICR and Topograph for topology-aware scheduling and placement

  • Design GPU-aware orchestration systems

  • Lead development of services that power our managed services

  • Inform on and help with networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect. You will work closely with our Network team to define and drive requirements

  • Inform and help with storage architecture requirements for AI workloads. You: will partner with Storage teams on what managed K8s, Slurm, and future services need

What they're looking for

  • Build the foundation for Managed Slurm on Kubernetes, enabling traditional: HPC workloads to run seamlessly alongside Kubernetes workload
  • Design higher-level platform services for inference, including model serving: infrastructure, autoscaling based on inference load, and multi-model deployment patterns
  • Design self-healing systems and automation for incident response, root cause analysis, and platform resilience
  • Lead chaos engineering efforts to validate system behavior under failure conditions at scale