The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Architect and develop an efficient, high-performance ML runtime and serving: system tailored for both onboard autonomous vehicle compute and large-scale, offboard data center environments.
Lead the integration and feature development for ML inference runtimes across: both domains, balancing the strict real-time latency and memory constraints of onboard systems with the high-throughput, highly concurrent demands of offboard serving fleets.
Drive the strategic migration of ML workloads toward a JAX-native runtime: architecture, which includes extending and modifying underlying ML compilers and runtimes (e.g., OpenXLA/PjRT, TensorRT).
Collaborate with world-class Waymo ML practitioners across perception,: planner, and research to analyze system-level ML workloads and apply hardware-aware compute optimizations.
Design and build robust tooling for profiling, benchmarking, and identifying: system-level bottlenecks across the end-to-end ML software stack.
B.S. or M.S. in CS, EE, Deep Learning or a related field
What they're looking for
- + years of professional software engineering experience focused on building,: scaling, or maintaining ML systems and infrastructure.
- + years production programming in C++.
- + years of production experience in Python and major deep learning frameworks (e.g., PyTorch, JAX).
- Experience optimizing ML software for hardware accelerators (e.g., GPUs, TPUs, custom silicon).