The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Build scalable systems for training and fine-tuning large-scale models to: evaluate interesting driving behaviors.
Work at the intersection of data engineering, model development, and: simulation Provide guidance on architectural decisions and technical directions. Own large, complex systems, driving architectures that meet technical and business objectives.
Contribute to the production and optimization of machine learning models: aiming to assess Waymo’s expansive fleet of vehicles that cumulatively travel millions of miles.
Design and scale large distributed systems covering the ML lifecycle,: supporting planet-scale dataset generation, model training, and evaluation.
Collaborate cross-functionally to derive performance and system-level: requirements for large ML systems. Translate product/business goals into measurable technical deliverables, ensuring system component alignment.
M.S. or Ph.D. degree Computer Science, Machine Learning, Artificial: Intelligence, or a related technical field, or equivalent practical experience.
What they're looking for
- + years in machine learning infrastructure such as developing, designing,: scaling, training, deploying, and optimizing large-scale machine learning systems from data to model.
- A history of contributions to machine learning tooling and frameworks e.g.: PyTorch, Jax, Tensorflow, Ray, or similar. The candidate should understand both the user facing API and the internal workings.
- Strong expertise in distributed training techniques, including gradient: sharding and optimization strategies for scaling large models across ML accelerator profiling tools to uncover performance bottlenecks.
- + years in machine learning infrastructure such as developing, designing,: scaling, training, deploying, and optimizing large-scale machine learning systems from data to model.