Staff Machine Learning Engineer, InfrastructureActive$310K

The opportunity

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…

What you'll do

  • Design, build, and optimize realistic simulation environments and business: logic running on TPUs using JAX and TensorFlow. Implement and optimize large-scale model and data parallelism strategies for training and running foundation models on TPU hardware.

  • Collaborate closely with modeling teams to integrate foundation models into simulation pipelines.

  • Drive technical architectures and system designs from data engineering: through simulation execution to meet business and performance objectives.

  • Profile systems, identify performance bottlenecks across ML accelerators, and: optimize end-to-end execution speed.

  • Translate product and business goals into concrete technical requirements and system deliverables.

  • + years of professional software engineering experience, with at least 4: years focused on machine learning infrastructure (scaling, training, optimizing, and deploying large-scale ML systems).

What they're looking for

  • Direct ML programming experience on TPU and GPU hardware using frameworks such as JAX, PyTorch, or TensorFlow.
  • Proven hands-on experience scaling large models using model parallelism, data: parallelism, or distributed training techniques.
  • Strong understanding of state-of-the-art ML models (e.g., autoregressive: transformers) and hands-on proficiency with ML accelerator profiling tools to diagnose bottlenecks.
  • Demonstrated ability to independently lead ambiguous technical initiatives: end-to-end and build robust libraries, pipelines, and developer tooling.