The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Design and architect advanced HPC systems optimized for large-scale: computational workloads and AI applications.
Collaborate with internal teams and stakeholders to define system requirements and performance goals.
Develop comprehensive testing frameworks to rigorously assess system: performance, scalability, and reliability.
Evaluate emerging technologies and architectural approaches to continuously: enhance infrastructure capabilities.
Create detailed architectural plans, documentation, and blueprints to guide implementation teams.
Provide technical leadership and mentoring to engineering teams, fostering best practices in HPC architecture.
What they're looking for
- + years of experience designing and architecting large-scale HPC and distributed computing systems.
- Expert-level knowledge of HPC hardware including GPU clusters, compute nodes,: high-speed networking (InfiniBand, Ethernet), and distributed storage.
- Hands-on experience with direct-to-chip liquid cooling systems.
- Proven expertise in creating robust performance benchmarks, capacity planning, and system validation.