Senior Machine Learning Infrastructure Engineer, SimulationPosted today$263K

The opportunity

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…

What you'll do

  • Be part of a world-class, high-performing research engineering team to: advance the state of the art of ultra realistic multi-agent simulations using foundation models.

  • Collaborate closely with the core Google DeepMind and Waymo Realism Modeling: teams in London, and Waymo Oxford to use the large models to improve sim realism.

  • Provide deep technical leadership on large-scale ML model architectures,: especially for autonomous vehicle models. Work at the intersection of data engineering, model development, and deployment, and provide guidance on architectural decisions and technical directions. Own large, complex systems, driving architectures that meet technical and business objectives.

  • Design and scale large distributed systems covering the ML lifecycle,: supporting planet-scale dataset generation and model training.

  • Collaborate cross-functionally to derive performance and system-level: requirements for large ML systems. Translate product/business goals into measurable technical deliverables, ensuring system component alignment.

  • Mentor junior engineers, growing their expertise and fostering a collaborative culture.

What they're looking for

  • BS in Computer Science, Robotics, similar technical field of study, or equivalent practical experience
  • + years of professional software engineering experience, with at least 3: years in machine learning infrastructure such as developing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model.
  • + years of professional software engineering experience, with at least 5: years in machine learning infrastructure such as developing, designing, scaling, training, deploying, and optimizing large-scale machine learning systems from data to model.
  • Solid experience in the development and optimization of machine learning: infrastructure tools like DeepSpeed, PyTorch, TensorFlow, or similar frameworks.