The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Design, build, and maintain scalable data pipelines to process many petabytes: of complex sensor data, making it ready for efficient model training and evaluation.
Develop infrastructure to produce reliable, high-quality datasets for a wide: range of ML models, from real-time on-car models to large-scale offboard foundation models.
Build towards an automated, unified data flywheel -: a datagen and ingestion solution that seamlessly connects data curation to model training.
Develop infrastructure for Perception-wide model training and release-ready: packaging, ensuring the model development lifecycle is robust, efficient, and reproducible.
Maintain and support critical data generation infrastructure and data refreshes for the Perception team.
Automate data quality and validation checks to ensure the integrity,: consistency, and trustworthiness of our datasets as we scale to new cities and vehicle platforms
What they're looking for
- Collaborate closely with ML engineers, research scientists, and core: infrastructure teams to understand user needs and deliver impactful ML workflows.
- Outstanding programming skills in C++ or Python
- Experience in ML data engineering, including data pipelines, data curation, data balancing, etc.
- Experience with the ML development lifecycle, including data engineering,: model training, model evaluation, and model deployment.