The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Drive the design and development of instrumentation and onboard architecture: to proactively prevent, detect, and debug reliability issues across current and future Waymo compute platforms.
Design and implement long-term strategies to systematically mitigate compute: hardware reliability and numerical stability issues across heterogeneous platforms (CPU, TPU, GPU), creating scalable tools to automate triage, debugging, and resolution.
Lead cross-functional initiatives across Waymo, Alphabet, and external: partners to architect proactive solutions that prevent issues originating from compilers (LLVM, XLA, JAX), optimization (AutoFDO), or model changes.
Architect frameworks and diagnostics to proactively identify and eliminate: complex software and firmware faults, including deadlocks, memory corruption, race conditions, memory leaks, or tail latency spikes.
Influence the design and tooling of next-generation hardware platforms,: serving as a technical authority to ensure a coherent, forward-looking reliability architecture.
BS/MS in Comp Sci, EE, Robotics, Physics, Math, or related field (or equivalent experience).
What they're looking for
- Experience in C++.
- Experience with CUDA/GPU/TPU acceleration and model reliability or optimization.
- Experience with writing GPU kernels and evaluation/debugging of GPU workloads.
- MS or PhD in Computer Science, Robotics, similar technical field of study, or equivalent practical experience.