Sr Platform Monitoring EngineerPosted today$145K
The opportunity
At Databricks, we are passionate about empowering data teams to tackle the world’s most challenging problems — from bringing the next mode of transportation to reality to accelerating the development of medical breakthroughs. We achieve this by building and operating the world’s…
What you'll do
Lead platform incident investigation, coordinating cross-functional teams: through rapid detection, mitigation, and resolution to minimize customer impact.
Conduct thorough post-incident root cause analysis across infrastructure,: services, and cloud providers to identify systemic patterns and prevent future occurrences.
Design and implement customer-focused alerting pipelines and end-to-end: observability workflows to enhance detection coverage and reduce mean time to detection.
Build automation tools, establish reusable monitoring patterns, and resolve: reliability gaps that directly impact customer experience.
Provide mentorship to junior engineers on observability patterns, alert design, and service health metrics.
Participate in on-call rotation
What they're looking for
- Minimum of 6 years of experience as an SRE, DevOps Engineer, Production Engineer, or similar role.
- Production-level experience with at least one major cloud provider (AWS,: Azure, GCP) and proficiency in container and orchestration technologies (Docker, Kubernetes).
- Hands-on experience with monitoring, logging, and alerting tools such as ELK,: Prometheus, Grafana, PagerDuty, etc. Ability to architect monitoring solutions that correlate metrics, logs, and traces.
- Strong proficiency in Python or similar languages with the ability to build: production-quality automation tools.