Senior Site Reliability EngineerActive

The opportunity

We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using MongoDB.

What you'll do

  • Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads

  • Build for reliability, making services and infrastructure available,: resilient, fault-tolerant, and self-healing

  • Identify and configure key metrics to detect incidents and quantify service: health, availability, and performance

  • Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure

  • Mentor early-career SREs and contribute to the team’s operational practices as it grows

  • Strong background in software development and operating distributed systems

What they're looking for

  • Strong background in software development and operating distributed systems
  • + years of experience building and operating distributed systems, with: proficiency in Python, Go, or a similar programming language
  • Experience operating Kubernetes in production and debugging below the: abstraction layer, including scheduling, cluster networking, and node-level issues
  • Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure
  • Strong understanding of Linux operating system internals and networking: concepts such as TCP/IP, DNS, TLS, and routing
  • Customer-focused mindset and strong verbal and written technical: communication skills, with a desire to collaborate with colleagues
  • Strong bias for efficient processes, operational simplicity, and automation over manual work
  • Bonus points for experience with Kubernetes networking, such as Istio or: Cilium, service mesh or edge load balancing in production, secure multi-tenant runtime environments at scale, multi-cloud infrastructure management, and virtualization or workload isolation technologies