The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Design, build, and optimize realistic simulation environments and business: logic running on TPUs using JAX and TensorFlow. Implement and optimize large-scale model and data parallelism strategies for training and running foundation models on TPU hardware.
Collaborate closely with modeling teams to integrate foundation models into simulation pipelines.
Drive technical architectures and system designs from data engineering: through simulation execution to meet business and performance objectives.
Profile systems, identify performance bottlenecks across ML accelerators, and: optimize end-to-end execution speed.
Translate product and business goals into concrete technical requirements and system deliverables.
+ years of professional software engineering experience, with at least 4: years focused on machine learning infrastructure (scaling, training, optimizing, and deploying large-scale ML systems).
What they're looking for
- Direct ML programming experience on TPU and GPU hardware using frameworks such as JAX, PyTorch, or TensorFlow.
- Proven hands-on experience scaling large models using model parallelism, data: parallelism, or distributed training techniques.
- Strong understanding of state-of-the-art ML models (e.g., autoregressive: transformers) and hands-on proficiency with ML accelerator profiling tools to diagnose bottlenecks.
- Demonstrated ability to independently lead ambiguous technical initiatives: end-to-end and build robust libraries, pipelines, and developer tooling.