The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Design, develop and maintain software for GPU/CPU compute infrastructure with: focus on performance, scalability, and reliability.
Implement and develop services for baremetal and VM instancing.
Develop distributed systems for managing and orchestrating compute resources across various SKU’s.
Troubleshoot and debug complex issues in a production and development environment.
On-call and incident ownership
Collaboration across multiple teams and drive ambiguity in requirements or solutions on RFC’s.
What they're looking for
- + years of experience working with Go (Golang) or Python in production environments.
- + years of experience with bare metal & virtualization hardware management and configuration.
- Are comfortable working in Linux environments and debugging issues at the OS, hardware, and networking layers.
- Can independently troubleshoot complex systems and communicate effectively: across software, infrastructure, and vendor teams.