The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Analyze model architectures and identify bottlenecks in training and: inference performance (e.g., memory bandwidth, compute, communication).
Apply and develop techniques such as quantization (e.g., FP8, INT4), pruning,: knowledge distillation, and efficient attention mechanisms.
Optimize model code for specific hardware accelerators (TPUs, GPUs),: leveraging compiler features and low-level libraries (e.g., XLA).
Experiment with different model partitioning and sharding strategies (e.g.,: data, tensor, pipeline parallelism, expert parallelism) to improve scalability and efficiency.
Design and implement low-latency, high-throughput serving solutions for: generative models and optimize training pipelines to reduce training time.
Build and maintain tools for performance analysis, profiling (e.g., xprof), and debugging of ML models.
What they're looking for
- MS or PhD in Computer Science, Machine Learning, Robotics, or a related field.
- + years of experience with deep learning architectures (especially: Transformers, Diffusion Models, MoEs), algorithms, and optimization techniques.
- Proficiency in JAX, Flax, and potentially TensorFlow/PyTorch.
- Expertise in using profiling tools (e.g., XProf, Perfetto, NVIDIA Nsight) to: diagnose performance issues in ML workloads.