The opportunity
Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…
What you'll do
Develop and train state-of-the-art computer vision / multimodal models (e.g.,: Gemini) to extract the rich semantic information (e.g., object attributes, scene properties, interaction dynamics) required by the AI agent.
Design and implement a scalable AI agent framework that integrates large: foundation models (e.g., Gemini) with the outputs of our perception models and internal knowledge bases.
Develop and apply Fine-tuning and Reinforcement Learning (RL) techniques to: create a "data flywheel," continuously improving the system's captioning and reasoning abilities through automated feedback.
Develop and prototype novel prompting strategies for Vision-Language Models: (VLMs) to elicit complex, causal reasoning about driving scenarios.
Collaborate closely with the ML Infra, Perception, Behavior, and AI: Foundation teams to define data requirements and integrate the captioning system into the broader ML development lifecycle.
Own the full system lifecycle, from advanced model development and: prototyping to production deployment and scaling for massive data generation
What they're looking for
- Master’s degree in Computer Science, or a related technical field.
- + years of hands-on experience training and shipping deep learning models for: computer vision tasks (e.g., detection, segmentation, video understanding) using Python and frameworks like PyTorch, JAX, or TensorFlow.
- + years of demonstrated experience working with large language models (LLMs): or vision-language models (VLMs) in areas such as fine-tuning, prompting, or Retrieval-Augmented Generation (RAG).
- Strong software engineering fundamentals, including designing scalable and reliable systems.