Staff Software Engineer - Infrastructure StorageActive$314K–$465K

The opportunity

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you'll do

  • High-Performance Distributed Storage Solutions and Protocols: We engineer the protocols and systems that serve massive datasets at the speeds demanded by modern clustered GPUs.

  • Dynamic Networking: We design advanced networks that provide multi-tenant security and intelligent routing without compromising performance, using the latest in AI networking hardware.

  • Compute Virtualization: We enable cutting-edge virtualization and clustering that allows AI researchers and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs.

  • Technical Leadership: Set technical direction for storage software architecture across petabyte-scale deployments, authoring and reviewing design docs, mentoring senior engineers, and serving as the technical anchor for cross-functional initiatives spanning storage, networking, compute, and control plane teams. Represent the storage software team in architectural reviews, roadmap planning, and customer-facing technical discussions.

  • Execution: Design, develop, and maintain high-performance storage systems: software across file (NFS, SMB, Lustre), block (NVMe-oF, iSCSI), and object (S3) protocols. Build distributed systems for orchestrating storage resources, integrate with NVMe/GPU-direct/DPU-accelerated hardware, and troubleshoot complex production issues across performance, protocol, and hardware failure domains. Own the full lifecycle from requirements and design through deployment, monitoring, and maintenance, including benchmarking, profiling, and capacity planning tooling.

  • Collaboration: Partner closely with storage software, networking, control plane, Kubernetes, observability, compute, and fleet engineering teams to deliver cross-functional infrastructure initiatives, define and track storage SLOs/SLIs, and ensure reliable deployment and maintenance of distributed storage infrastructure.

What they're looking for

  • Innovate: Stay current with AI and HPC storage research, evaluate emerging: protocols and hardware (from open-source filesystems to vendor-specific accelerated storage), and optimize solutions for AI workloads including checkpoint I/O, high-throughput dataset serving, and latency-sensitive inference pipelines.
  • Experience: 10+ years in storage systems engineering, with 5+ years in a: technical lead or Staff+ IC role. Proven track record designing and operating multi-petabyte storage infrastructure in production data center or cloud environments. Background in HPC, AI/ML infrastructure, or large-scale cloud storage.
  • Systems-Level Programming: Strong proficiency in C, C++, Rust, or Go. Ability to write high-performance, concurrent, production-grade systems code. Familiarity with DPDK/SPDK and kernel-bypass data paths is a plus; kernel-level storage driver or storage daemon experience is even better.
  • Storage Protocol & API Expertise: Deep hands-on experience with two or more protocols, object (S3), block (iSCSI, NVMe-oF), or file (NFS, SMB, Lustre, DAOS), including implementing or maintaining protocol servers/clients in production, not just consuming them.