Incident ManagerActive$104K

The opportunity

At Databricks, we are passionate about empowering data teams to tackle the world’s most challenging problems — from bringing the next mode of transportation to reality to accelerating the development of medical breakthroughs. We achieve this by building and operating the world’s…

What you'll do

  • Lead critical incidents: coordinate multi-disciplinary response efforts across Databricks’ cloud-based services to rapidly mitigate impact and restore operations.

  • Drive technical root cause analysis and reliability improvements:

  • collaborate with engineering teams to trace and document underlying causes: across distributed systems, services, and data stores.

  • Summarize key learnings, clearly communicate action items, and ensure that: technical and procedural improvements are followed through.

  • Own communications during incidents: deliver frequent, high-quality updates to internal stakeholders (executives, engineering leadership, support) and compose and publish customer-facing notifications that are accurate, timely, and empathetic.

  • Mentor and train peers in both incident communication and technical response: disciplines to raise the overall quality of Databricks’ incident response.

What they're looking for

  • + years of experience in incident management, site reliability engineering,: or production operations supporting large-scale, cloud-native systems.
  • Proven ability to lead and coordinate high-severity incidents, including: identifying impact, isolating fault domains, and managing multi-team response efforts.
  • Strong understanding of cloud infrastructure (AWS, Azure, or GCP): including compute, networking, storage, and observability components.
  • Deep expertise in log analysis and debugging: