The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Architect and define scalable compute platforms optimized for AI/ML,: simulation, and high-throughput workloads.
Develop compute system standards and design patterns to ensure consistency,: performance, and maintainability across infrastructure.
Evaluate emerging CPU, GPU, and accelerator technologies, owning: architectural tradeoff decisions that impact compute density, power, cooling, and total cost.
Collaborate with product and engineering teams to map workload requirements: to compute platform capabilities across bare metal and cloud deployments.
Experience converting ambiguous business or customer needs into measurable: platform requirements, technical specifications, acceptance criteria, and architecture decisions.
Define compute platform roadmaps and architectural reference designs that: guide hardware selection, firmware baselines, rack-level, and cluster design.
What they're looking for
- Act as a technical lead during new platform introductions, guiding validation: and performance characterization efforts.
- Mentor systems engineers and cross-functional stakeholders on compute: performance tuning, sizing, and architectural decisions.
- Proven experience (7+ years) architecting large-scale 10k-100k+ GPU HPC or cloud compute platforms.
- Deep knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies.