Sr. Staff Production Engineer - Data PlatformActive$229K

The opportunity

At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI…

What you'll do

  • Architecting Agentic Reliability: Define and drive the design of future "self-healing" infrastructure at scale where AI agents proactively detect, diagnose, and remediate production incidents before they impact customers.

  • Data Platform Optimization: Own the operational integrity of the Data Platform that powers our internal AI models, ensuring 99.99% availability for the compute, storage, and control plane services used by thousands of Databricks engineers.

  • High-Scale Operational Excellence: Establish the next generation of "Change Safety" protocols, utilizing automation and agentic guardrails to manage complex deployments across 100+ global regions.

  • Leadership in Chaos & Scale: Serve as a technical bar-raiser for the team, evangelizing modern SRE practices (including Chaos Engineering) to navigate the structural transformation of the industry toward agentic, autonomous systems.

  • BS/MS/PhD in Computer Science, or a related field

  • Technical Depth: 10+ years of production-level experience as a Software Engineer or SRE in highly distributed, multi-cloud environments.

What they're looking for

  • Engineering Persona: You write code to solve operational problems. You are not a traditional sys-admin; you build frameworks, automation, and tooling (Scala, Java, Go, or Python) to eliminate toil.
  • Platform & AI Mindset: Deep understanding of distributed data platforms and a passion for leveraging AI/ML to revolutionize infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks is a significant plus.
  • Operational Grit: Proven ability to remain calm and decisive under pressure. You have navigated large-scale distributed systems through hyper-growth and have a track record of driving incident-to-roadmap loops.
  • Strategic Influence: Experience building long-range technical roadmaps and driving cross-functional alignment. You are comfortable challenging senior leadership with data-driven insights.