Senior Site Reliability Engineer - SDNActive$240K–$312K

The opportunity

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you'll do

  • Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure

  • Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs

  • Develop tooling and automation to reduce operational toil and improve reliability

  • Collaborate with software, platform, and networking teams to improve service: reliability and deployment workflows

  • Deploy and maintain network monitoring, observability, and management tools

  • Improve deployment safety through CI/CD pipelines, GitOps workflows, testing, and progressive rollouts

What they're looking for

  • Drive operational excellence through observability, incident management,: capacity planning, postmortems, and participation in the on-call rotation
  • Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role
  • Have experience operating and supporting large-scale distributed systems in production
  • Have experience with Kubernetes application lifecycle management, upgrades,: troubleshooting, and production operations