Staff Software Engineer - Multi Cloud EfficiencyActive

The opportunity

At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI…

What you'll do

  • Architect at Scale: Lead the design and implementation of systems to optimize cloud spend, driving impact in the magnitude of hundreds of millions of dollars. You will architect solutions for complex problems like reserved-instance automation, tag-based cost attribution at scale, automated waste-cleanup policy engines, and cross-cloud resource governance.

  • Drive Technical Strategy: Identify and deliver high-impact, well-adopted projects that solve complex efficiency problems. You will define the roadmap for our efficiency pillar, conducting thorough design reviews and risk assessments for systems supporting the company's core infrastructure.

  • Engineer for Efficiency: Design scalable systems for regression detection and automated resource rightsizing. You will build the frameworks that allow product teams to measure their "efficiency footprint” — the cost and resource impact of their services and hold them accountable through automated release gates that block changes that regress agreed efficiency targets (SLOs).

  • Operational Excellence: Serve as a technical authority for incident management, addressing complex, system-wide performance issues. You will improve the stability and reliability of our internal platform tools, ensuring they are robust enough to manage our growing global scale.

  • Technical Leadership: Act as a force multiplier by mentoring junior and senior engineers, raising the bar for engineering quality, and driving engineering culture across the organization.

  • Build self-driving efficiency systems: Design the automation and agentic layers — AI-driven detection, decisioning, and autonomous remediation — that let cost governance run continuously at scale, shifting the team from manual cleanup to systems that keep themselves efficient.

What they're looking for

  • + years of experience building production-grade, distributed systems in Java, Scala, C++, or Go.
  • Proven track record of architecting solutions for large-scale infrastructure,: cloud computing (AWS/Azure/GCP), and container orchestration (Kubernetes/Docker).
  • Deep expertise in identifying and solving performance bottlenecks in: high-throughput, distributed environments.
  • Demonstrated ability to drive technical requirements for ambiguous problems: where the path forward is not obvious.