The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Collaborate with ML practitioners on models for perception, behavior: prediction, and planning, to understand their models and accelerate them onboard through custom NVIDIA GPU kernel development.
Deep dive into the NVIDIA ML software and runtime stack, from custom CUDA ops: to the XLA:GPU compiler and low-level libraries. Analyze numeric behaviors, debug complex compiler issues, and ensure inference results are stable and consistent. Develop tools/system software for optimal resource usage, hardware efficiency, and platform reliability in an ML serving system.
Analyze ML workload performance at the hardware level; apply manual and AI: agent-assisted techniques and develop highly optimized, custom CUDA/Triton operator libraries tailored to Waymo's specific architectures.
Build tools to benchmark, profile GPU execution, and productize deep learning: models for a streamlined and robust onboard and offboard deployment.
B.S. or M.S. in CS, EE, Deep Learning or a related field
+ years of industry experience on system performance, hardware-level GPU optimization, or ML compilers
What they're looking for
- Strong C++ and CUDA programming skills
- Extensive experience in NVIDIA GPU Kernel development to accelerate deep learning models
- Proven debugging and optimization experience on the XLA:GPU compiler, as well as the NVIDIA runtime stack
- Passion for developing and optimizing ML software stacks for modern ML: accelerator architectures (framework, runtime library, ML compiler, efficient deep learning etc.)