The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Serve as the hands-on technical lead for integrating OEM and white-label HPC: AI/ML, general purpose compute, storage, and network hardware into Lambda’s HPC platform reference architectures.
Drive the end-to-end process of new product introduction (NPI) for hardware: systems, including system bring-up, documentation, vendor technical engagement, production readiness, and closure of hardware risks.
Identify, debug, and resolve hardware issues across different hardware: engineering domains during hardware NPI; support closure of critical fleet issues that require hardware design, vendor corrective action, or platform configuration changes.
Partner with HPC architects to translate platform blueprints into concrete: hardware selections and system configurations.
Partner with the supply chain team on new vendor evaluation and QBR/HBR feedback on established vendors.
Own the hardware platform through NPI, working with PMO to de-risk execution,: drive cross-functional closure of hardware readiness issues, and ensure platforms reach production on schedule.
What they're looking for
- Collaborate with the quality team and fleet reliability team during hardware: NPI and after production to continuously improve product quality and reliability at scale.
- Work cross-functionally with fleet engineering, deployment, operation and: datacenter engineering teams to ensure on-time delivery and deployment, quality, compatibility, performance, and scalability of new systems.
- Serve as the technical lead to evaluate, enable, and prototype new hardware in labs.
- Review BOMs to ensure configuration accuracy, component compatibility, and: alignment of key commodities and components to Lambda platform requirements.