Sr. Engineering Manager, AI RuntimeActive$229K

The opportunity

At Databricks, we are passionate about enabling data teams to solve the world's toughest problems, from making the next mode of transportation a reality to accelerating the development of medical breakthroughs. We do this by building and running the world's best data and AI…

What you'll do

  • Lead, mentor, and grow a high-performing engineering team responsible for the: Custom Training product and its foundational infrastructure, including distributed training orchestration, cluster lifecycle, fault tolerance, and training efficiency.

  • Define and own the product and technical roadmap for AIR, balancing customer: experience, functionality, and foundational investments.

  • Collaborate closely with product, research, platform, infrastructure teams,: and customers to drive end-to-end delivery, from ideation and prioritization to launch and operation.

  • Drive architectural decisions and product design for managed GPU training at scale.

  • Advocate for customer needs through direct engagement, ensuring engineering: decisions translate to clear product impact.

  • Build observability and reliability practices for long-running, multi-node: training jobs, including checkpoint strategies, failure recovery, and operational runbooks.

What they're looking for

  • Partner with recruiting to attract, hire, and develop top-tier engineering talent.
  • + years of software engineering experience, with 3+ years in engineering management.
  • Track record building and operating managed GPU training infrastructure at scale (100s/1000s GPUs).
  • Deep familiarity with distributed training frameworks (PyTorch, DeepSpeed,: Composer, Megatron-LM) and parallelism strategies (FSDP, tensor/pipeline parallelism).