Staff Software Engineer, Compute ReliabilityPosted today$310K

The opportunity

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…

What you'll do

  • Drive the design and development of instrumentation and onboard architecture: to proactively prevent, detect, and debug reliability issues across current and future Waymo compute platforms.

  • Design and implement long-term strategies to systematically mitigate compute: hardware reliability and numerical stability issues across heterogeneous platforms (CPU, TPU, GPU), creating scalable tools to automate triage, debugging, and resolution.

  • Lead cross-functional initiatives across Waymo, Alphabet, and external: partners to architect proactive solutions that prevent issues originating from compilers (LLVM, XLA, JAX), optimization (AutoFDO), or model changes.

  • Architect frameworks and diagnostics to proactively identify and eliminate: complex software and firmware faults, including deadlocks, memory corruption, race conditions, memory leaks, or tail latency spikes.

  • Influence the design and tooling of next-generation hardware platforms,: serving as a technical authority to ensure a coherent, forward-looking reliability architecture.

  • BS/MS in Comp Sci, EE, Robotics, Physics, Math, or related field (or equivalent experience).

What they're looking for

  • Experience in C++.
  • Experience with CUDA/GPU/TPU acceleration and model reliability or optimization.
  • Experience with writing GPU kernels and evaluation/debugging of GPU workloads.
  • MS or PhD in Computer Science, Robotics, similar technical field of study, or equivalent practical experience.