AI Operations Engineer - IT/Internal InfrastructureNew$206K–$275K

The opportunity

Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers.

What you'll do

  • Design, write, and deliver software and services to improve the availability,: scalability, reliability, and efficiency of Lambda’s internal IT systems and platforms.

  • Solve problems relating to mission critical services and build automation to: prevent problem recurrence with the goal of automating response to all non-exceptional events.

  • Work with Lambda Engineering and internal teams to Influence and create new: designs, architectures, standards, and methods for large-scale distributed systems.

  • Engage in service capacity planning and demand forecasting, software performance analysis, and system tuning.

  • Be an excellent communicator, producing documentation and related artifacts: for the systems you are responsible for.

  • Have a keen interest in system design, architecting for performance,: scalability, and experience with multiple cloud infrastructure platforms (AWS, GCP, Azure, etc.).

What they're looking for

  • Think carefully about systems: edge cases, failure modes, behaviors, and specific implementations.
  • Know and prefer configuration management systems and toolchains (Chef,: Ansible, Terraform, GitHub Actions, etc.)
  • Have solid programming skills: Python, Go, etc.
  • Have an urge to collaborate and communicate asynchronously, combined with a: desire to record and document issues and solutions.