The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Design, and improve distributed input data pipelines for large-scale ML training workloads.
Collaborate with researchers and ML engineers to resolve bottlenecks in data pipeline performance.
Improve runtime goodput of ML training workload, including optimizing input: data processing systems, ensuring scalability and reliability across distributed environments.
Implement and maintain advanced ML infrastructure tools, including ML Pathways, Grain, JAX, and TensorFlow.
Evaluate and integrate modern technologies to enhance the performance and scalability of ML systems.
Promote best practices for distributed systems architecture and contribute to: technical leadership within the team.
What they're looking for
- B.S. in Computer Science, Math, or 5+ years equivalent real-world experience.
- Proficient in distributed systems design with an understanding of ML data pipeline optimization.
- Experience with ML frameworks, including TensorFlow and JAX.
- Hands-on experience libraries like Grain or tf.data service.