The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Architect high-performance networking solutions that power cloud platforms,: with a focus on ultra-low-latency and high-bandwidth connectivity.
Define the network topology and architectural patterns for large-scale GPU: clusters, storage backends, and multi-tenant environments.
Evaluate, benchmark, and select next-generation network technologies (e.g.,: InfiniBand NDR/XDR, RoCE, 400G-1.6T Ethernet, Ultra Ethernet, ESUN, OCI, etc) to meet AI workload requirements.
Develop and maintain network architecture standards, reference designs, and: scalability roadmaps for multi-site and hybrid environments.
Partner with compute and storage architects to ensure seamless end-to-end data flow and fault tolerance.
Guide network automation strategies and tooling to enable efficient: provisioning, telemetry, and operational visibility.
What they're looking for
- Mentor engineers and cross-functional teams on advanced network concepts,: troubleshooting, and architectural best practices.
- Proven experience (7+ years) architecting high-performance data center: networks, preferably for HPC, AI/ML, or large-scale cloud infrastructure.
- Deep expertise with InfiniBand (HDR/NDR) and advanced Ethernet fabrics, including RoCE and RDMA protocols.
- Strong understanding of data center switching architectures, congestion: control (PFC, ECN), QoS, and network virtualization technologies such as VXLAN and EVPN.