The opportunity
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.
What you'll do
Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure
Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs
Develop tooling and automation to reduce operational toil and improve reliability
Collaborate with software, platform, and networking teams to improve service: reliability and deployment workflows
Deploy and maintain network monitoring, observability, and management tools
Improve deployment safety through CI/CD pipelines, GitOps workflows, testing, and progressive rollouts
What they're looking for
- Drive operational excellence through observability, incident management,: capacity planning, postmortems, and participation in the on-call rotation
- Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role
- Have experience operating and supporting large-scale distributed systems in production
- Have experience with Kubernetes application lifecycle management, upgrades,: troubleshooting, and production operations