The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Design, build, and maintain scalable control plane services, operators, and: custom Kubernetes controllers; develop automation in Go/Python for end-to-end cluster lifecycle management — provisioning, upgrades, patching, and deletion
Build GPU-aware orchestration systems, working within the platform: architecture to support GPU scheduling and resource allocation
Partner with the Network team on networking solutions for AI workloads: CNI integration (Cilium, Multus), high-performance fabrics (InfiniBand, RoCE), RDMA, and GPUDirect
Write resilient systems that handle failure gracefully: timeouts, retries, backoff, and degraded-mode operation — across large-scale distributed environments
Develop platform services for inference: model serving infrastructure, autoscaling based on inference load, and multi-model deployment patterns
Build internal tools and CLIs that let ML/AI teams deploy and monitor their own inference services
What they're looking for
- Support and debug production issues through on-call rotation
- Have 6+ years of experience in software engineering, with a track record of: owning significant technical scope within a team (e.g., driving a project from design through production, or acting as a de facto tech lead on a workstream)
- Deep understanding of Kubernetes internals: controllers, schedulers, operators, CRDs, CSI, CNI, and the extension patterns that make Kubernetes powerful
- Solid grasp of distributed systems fundamentals: fault tolerance, graceful degradation, and failure handling in large-scale environments