Sr. Staff Production Engineer - Data PlatformActive$229K
The opportunity
At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI…
What you'll do
Architecting Agentic Reliability: Define and drive the design of future "self-healing" infrastructure at scale where AI agents proactively detect, diagnose, and remediate production incidents before they impact customers.
Data Platform Optimization: Own the operational integrity of the Data Platform that powers our internal AI models, ensuring 99.99% availability for the compute, storage, and control plane services used by thousands of Databricks engineers.
High-Scale Operational Excellence: Establish the next generation of "Change Safety" protocols, utilizing automation and agentic guardrails to manage complex deployments across 100+ global regions.
Leadership in Chaos & Scale: Serve as a technical bar-raiser for the team, evangelizing modern SRE practices (including Chaos Engineering) to navigate the structural transformation of the industry toward agentic, autonomous systems.
BS/MS/PhD in Computer Science, or a related field
Technical Depth: 10+ years of production-level experience as a Software Engineer or SRE in highly distributed, multi-cloud environments.
What they're looking for
- Engineering Persona: You write code to solve operational problems. You are not a traditional sys-admin; you build frameworks, automation, and tooling (Scala, Java, Go, or Python) to eliminate toil.
- Platform & AI Mindset: Deep understanding of distributed data platforms and a passion for leveraging AI/ML to revolutionize infrastructure management. Familiarity with LLM infrastructure, training/inference pipelines, or agentic frameworks is a significant plus.
- Operational Grit: Proven ability to remain calm and decisive under pressure. You have navigated large-scale distributed systems through hyper-growth and have a track record of driving incident-to-roadmap loops.
- Strategic Influence: Experience building long-range technical roadmaps and driving cross-functional alignment. You are comfortable challenging senior leadership with data-driven insights.