The opportunity
We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB.
What you'll do
Own the reliability architecture of the platform across regions and cloud providers
Collaborate with the teams building the platform, providing internal support: and guidance on operability, capacity, and best practices
Set operational standards for the team: on-call quality, incident response, SLO discipline
Mentor and technically develop the SRE team
Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure
+ years of experience working on software and operating distributed systems,: with deep Kubernetes expertise, including designing or evolving multi-cluster platforms
What they're looking for
- + years of experience working on software and operating distributed systems,: with deep Kubernetes expertise, including designing or evolving multi-cluster platforms
- Proficiency in Python, Go, or a similar programming language
- Understand workload isolation at the systems level: containers, virtual machines, and the trade-offs between them for running untrusted code
- Possess a customer-focused mindset
- Value efficiency in processes and operations, and display a strong preference: for automation over manual processes
- Be intimately familiar with the infrastructure primitives of at least one of: AWS, GCP, or Azure, and comfortable reasoning about differences between them
- Have a track record of driving infrastructure architecture across teams and mentoring engineers