Site Reliability Engineer, LeadActive$368K

The opportunity

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on building the Waymo Driver—The World's Most Experienced Driver™—to improve access to…

What you'll do

  • Drive incident response processes, retrospectives, and continuous improvement: to prevent recurring issues. Organize and participate in on call rotations for release oriented outages, along with disaster recovery planning and business continuity

  • Harden software, develop playbooks, improve oncall metrics and: post-mitigation workstreams, and improve robustness of the software system running our most critical release processes

  • Lead the engineering analysis and response to novel, rare or unexpected: events encountered during our development analysis cycle.

  • Champion a culture of reliability engineering across the organization,: including partnering with organization leads and others across the Onboard, Eval and Simulator organization to educate and improve large-scale software engineering best practices.

  • Contribute to system software architecture to help improve its robustness and debuggability

  • Master's degree or PhD in Computer Science, Engineering, or a related technical field

What they're looking for

  • + years of experience working on large-scale production software systems, especially machine learned systems
  • + years on hands on coding experience with C++
  • Experience organizing and running oncall rotations
  • Excellent influence without authority skills