Staff ML Engineer, Generative Model Performance & EfficiencyPosted today$310K

The opportunity

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…

What you'll do

  • Analyze model architectures and identify bottlenecks in training and: inference performance (e.g., memory bandwidth, compute, communication).

  • Apply and develop techniques such as quantization (e.g., FP8, INT4), pruning,: knowledge distillation, and efficient attention mechanisms.

  • Optimize model code for specific hardware accelerators (TPUs, GPUs),: leveraging compiler features and low-level libraries (e.g., XLA).

  • Experiment with different model partitioning and sharding strategies (e.g.,: data, tensor, pipeline parallelism, expert parallelism) to improve scalability and efficiency.

  • Design and implement low-latency, high-throughput serving solutions for: generative models and optimize training pipelines to reduce training time.

  • Build and maintain tools for performance analysis, profiling (e.g., xprof), and debugging of ML models.

What they're looking for

  • MS or PhD in Computer Science, Machine Learning, Robotics, or a related field.
  • + years of experience with deep learning architectures (especially: Transformers, Diffusion Models, MoEs), algorithms, and optimization techniques.
  • Proficiency in JAX, Flax, and potentially TensorFlow/PyTorch.
  • Expertise in using profiling tools (e.g., XProf, Perfetto, NVIDIA Nsight) to: diagnose performance issues in ML workloads.