Engineering Manager, Serverless Compute PlatformActive$181K
The opportunity
At Databricks, we are passionate about helping data teams solve the world's toughest problems — from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best AI and data…
What you'll do
You will inherit a team of strong senior ICs who have already delivered an: initial preview. Your job is to build out the full vision, guide evolution, and scale the team.
You will ensure strong execution health and that the service launches with: production-grade reliability spanning a range of use cases, e.g. GPU onboarding, UDF generalization, and managed REPL.
Own a 0→1 service with platform-wide blast radius. Architect and launch the: Execution Sandbox Service from inception to production scale. This greenfield provisioning layer will power all non-Spark compute workloads on Serverless (Notebooks, AI Agents, Remote UDFs).
Unify a fragmented compute surface. Converge disparate CPU and GPU cluster: management paths into a single provisioning service, eliminating parity bugs and enabling consistent product experiences.
Collaborate across 5+ partner organizations. Drive alignment on API contracts: and shared milestones across Serverless Platform, AI Runtime, Lakeguard, and product teams.
Shape product strategy through deep technical understanding. Partner with: Product Management to leverage this new sandbox primitive for future offerings like serverless command execution APIs and FaaS-style workloads.
What they're looking for
- + years managing engineers building and operating distributed systems in: production, ideally control-plane or orchestration services
- BS or higher in Computer Science or a related field. Equivalent practical experience is equally valued.
- Deep technical fluency in infrastructure systems. Ability to deeply review: architecture docs, challenge design tradeoffs (e.g., state machine design, API boundaries), and coach senior ICs.
- Experience with multi-cloud or multi-region service deployment (AWS, Azure, GCP).