The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Develop and Maintain Production Systems: Design, implement, and improve the software that powers fleet lifecycle management, machine configuration, and cluster state at scale.
Automate Provisioning and Deployment: Build and enhance automation that takes clusters from logical design and racking through OS provisioning, configuration, validation, and customer hand-off.
Support New Hardware and Site Bring-Up: Enable bring-up, validation, and production readiness for new server, accelerator, and network platforms, as well as new datacenter sites.
Improve Machine Lifecycle Workflows: Refine bare metal provisioning, firmware and DPU updates, imaging, and system health monitoring across the fleet.
Keep Fleet State Consistent and Healthy: Build systems that reconcile intended against actual configuration, catch drift before it causes deployment failures, and maintain production SLAs.
Debug Hardware and Firmware Issues: Investigate failures across BIOS, BMC, firmware, DPUs, networking, storage, and boot flows.
What they're looking for
- Collaborate Across Teams: Work closely with datacenter and deployment operations, networking, architecture, security, and product engineering teams to build scalable, maintainable solutions.
- Have 5+ years of engineering experience
- Are fluent in Python, Go, or similar, and comfortable with APIs, distributed systems, and automation pipelines
- Work confidently in Linux environments and can debug across the OS, hardware, and networking layers