The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Designing and implementing a highly available and reliable GPU and CPU “host: and instance lifecycle” control plane.
Guide technical decisions involving semiconductor architecture, BIOS/Firmware: settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
Guide design of compute platform multi-tenant security model
Provide technical leadership and mentorship for senior engineers across: several teams to execute on complex infrastructure roadmaps and technical strategy.
Collaborate with product and data center organizations to translate customer: requirements into scalable infrastructure capabilities.
Work with customers on translating vague customer technical requirements into: concrete engineering deliverables.
What they're looking for
- Set engineering standards and lead design reviews for mission-critical cloud software at scale.
- + years of experience working on compute control plane distributed systems: used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
- Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
- Basic knowledge of software defined networking fundamentals that informs: secure, multi-tenant distributed systems.