The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
High-Performance Distributed Storage Solutions and Protocols: We engineer the protocols and systems that serve massive datasets at the speeds demanded by modern clustered GPUs.
Dynamic Networking: We design advanced networks that provide multi-tenant security and intelligent routing without compromising performance, using the latest in AI networking hardware.
Compute Clustering and Virtualization: We enable cutting-edge virtualization and clustering that allows AI researchers and engineers to focus on AI workloads, not AI infrastructure, unleashing the full compute bandwidth of clustered GPUs.
Design, develop, and maintain software for storage systems, focusing on: performance, scalability, and reliability.
Implement and optimize storage protocol APIs for file (e.g., NFS, SMB), block: (e.g., iSCSI, Fibre Channel), and object (e.g., S3) access.
Develop distributed systems for managing and orchestrating storage resources: across multiple storage solutions and redundant arrays.
What they're looking for
- Collaborate with hardware and system architects to integrate software with: various storage solutions, including NVMe and GPU-direct storage.
- Troubleshoot and debug complex issues in a production data center environment.
- Contribute to the full software development lifecycle, from requirements: gathering and design to deployment and maintenance.
- Bachelor's or Master's degree in Computer Science or a related field.