Sr Platform Monitoring EngineerPosted today$145K

The opportunity

At Databricks, we are passionate about empowering data teams to tackle the world’s most challenging problems — from bringing the next mode of transportation to reality to accelerating the development of medical breakthroughs. We achieve this by building and operating the world’s…

What you'll do

  • Lead platform incident investigation, coordinating cross-functional teams: through rapid detection, mitigation, and resolution to minimize customer impact.

  • Conduct thorough post-incident root cause analysis across infrastructure,: services, and cloud providers to identify systemic patterns and prevent future occurrences.

  • Design and implement customer-focused alerting pipelines and end-to-end: observability workflows to enhance detection coverage and reduce mean time to detection.

  • Build automation tools, establish reusable monitoring patterns, and resolve: reliability gaps that directly impact customer experience.

  • Provide mentorship to junior engineers on observability patterns, alert design, and service health metrics.

  • Participate in on-call rotation

What they're looking for

  • Minimum of 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or similar role.
  • Production-level experience with at least one major cloud provider (AWS,: Azure, GCP) and proficiency in container and orchestration technologies (Docker, Kubernetes).
  • Hands-on experience with monitoring, logging, and alerting tools such as ELK,: Prometheus, Grafana, PagerDuty, etc. Ability to architect monitoring solutions that correlate metrics, logs, and traces.
  • Strong proficiency in Python or similar languages with the ability to build: production-quality automation tools.