The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
High-Performance Distributed Storage Solutions and Protocols: We engineer the protocols and systems that serve massive datasets at the speeds demanded by modern clustered GPUs.
Dynamic Networking: We design advanced networks that provide multi-tenant security and intelligent routing without compromising performance, using the latest in AI networking hardware.
Compute Clustering and Virtualization: We enable cutting-edge virtualization and clustering that allows AI researchers and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs.
Design and build a vendor-agnostic control plane that provisions, scales,: heals, and meters storage across the platforms our customers actually demand, VAST Data, WEKA, DDN, Pure, NetApp, Ceph, MinIO, and the ones that don't exist yet.
Define the internal abstraction layer that hides vendor-specific APIs,: failure semantics, QoS knobs, and telemetry formats behind one declarative interface, so a new vendor integration is a driver, not a re-architecture.
Build reconciliation-loop and CRD-based orchestration (Kubernetes: controllers, operators, custom schedulers) that manages capacity, tenancy, encryption domains, and placement across data centers and availability zones.
What they're looking for
- Own multi-tenant isolation end to end: namespace and subsystem partitioning, per-tenant QoS and rate limiting, credential and key lifecycle, blast-radius containment, noisy-neighbor detection.
- Design the capacity and placement engine: PCIe-topology-aware, NUMA-aware, failure-domain-aware. On our platforms a drive behind the same PCIe switch as the GPU it serves beats a faster drive on a different root port, and the control plane needs to know that.
- Instrument everything: SLI/SLO definitions, fleet-wide performance regression detection, and the observability pipeline that makes a petabyte fleet debuggable at 3 a.m.
- Bachelor's or Master's degree in Computer Science or a related field.