Engineering Manager, Site ReliabilityNew

The opportunity

Instacart is transforming the grocery industry by building technology that connects customers, shoppers, retailers, and brands through a reliable, convenient online marketplace. The Site Reliability Engineering team helps ensure that this experience remains resilient, scalable, and dependable as Instacart grows.

What you'll do

  • Lead, mentor, and develop a team of Site Reliability Engineers, establishing: clear goals, providing actionable feedback, and supporting career growth and professional development.

  • Set the team’s technical direction and priorities for improving the: reliability, scalability, availability, performance, and operational readiness of Instacart’s systems.

  • Partner with engineering, product, security, infrastructure, and other: cross-functional teams to define reliability standards, influence system design, and deliver initiatives that improve the customer and developer experience.

  • Drive incident management and operational excellence, including incident: response, post-incident learning, service-level objectives, capacity planning, observability, and continuous risk reduction.

  • Promote automation and self-service tooling that reduce operational toil,: improve deployment confidence, and enable engineering teams to own and operate their services effectively.

  • Balance near-term operational needs with long-term investments, making: thoughtful tradeoffs in a high-growth environment where priorities and requirements can change quickly.

What they're looking for

  • Experience leading Site Reliability Engineering, platform engineering,: infrastructure engineering, or developer productivity teams.
  • Experience operating highly available services at significant scale and: improving service-level objectives, observability, capacity, or disaster recovery capabilities.
  • Experience with infrastructure as code, continuous delivery, monitoring,: logging, tracing, and automated remediation.
  • Experience building or evolving reliability programs across multiple: engineering teams, including shared standards, operational reviews, and service ownership practices.
  • Demonstrated ability to create alignment across teams, navigate ambiguity,: and turn complex technical challenges into clear, achievable plans.
  • A leadership approach grounded in empathy, transparency, direct: communication, collaboration, and a commitment to inclusive team development.