Principal Engineer, Compute Fleet ManagementActive$264K
The opportunity
At Databricks, we are passionate about enabling data teams to solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI…
What you'll do
Pioneering Fleet Optimization: Provisioning and pooling of O(Billion)s of cloud resources to achieve peak workload performance, industry-leading efficiency, and robust resource isolation.
Delivering Hyper-Scale Resilience: Build the architecture that guarantees horizontal scaling and resilience against zonal or even cloud account-level failures, ensuring Databricks is always on.
Owning the Critical Path: Lead the development of the lowest-dependency systems required to bootstrap and manage our massive compute platform.
High Availability: Achieve and maintain 99.99% availability for all batch and serving workloads.
Stellar Efficiency: Drive utilization to 60% or higher—a crucial metric that requires balancing high efficiency with unwavering tolerance for cloud failures.
Best-in-Class Isolation: Architect and enforce strong security and performance isolation across a diverse range of customer workloads.
What they're looking for
- Leading Transformative Projects: Taking ownership of complex, cross-team, cross-layer, and multi-quarter strategic engineering initiatives from concept to execution.
- Distributed Systems Mastery: Deep, hands-on experience developing and operating high-scale distributed systems on at least one major public cloud.
- Influence Without Authority: Proven ability to drive consensus, establish technical direction, and lead large technical efforts across organizational boundaries.
- Execution Discipline: Exceptional strength in planning, tracking project progress, and managing complex cross-organizational dependencies.