The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Build and operate core cloud platform services for compute lifecycle, bare: metal hosts, capacity, placement, and maintenance workflows.
Design reliable APIs, backend services, state machines, and orchestration: systems that power Lambda’s GPU cloud.
Work on bare metal lifecycle systems including launch, terminate,: restart/reboot, host reclaim, validation, quarantine, and return-to-pool workflows.
Improve deployment, observability, testing, alerting, runbooks, and: operational readiness for business-critical control-plane services.
Debug complex production issues across distributed services, infrastructure: dependencies, networking, and cloud workflows.
Partner with infrastructure, networking, fleet, security, support, and: product teams to define cross-system contracts and deliver end-to-end cloud capabilities.
What they're looking for
- Contribute to architecture, design docs, code reviews, incident follow-through, and mentoring across the team.
- Bachelor's degree or equivalent working experience.
- Have 6+ years of professional software engineering experience building: production backend or distributed systems.
- Are strong in Python, Go, or a similar backend/system language.