The opportunity
Instacart is transforming the grocery industry by building technology that connects customers, shoppers, retailers, and brands through a reliable, convenient online marketplace. The Site Reliability Engineering team helps ensure that this experience remains resilient, scalable, and dependable as Instacart grows.
What you'll do
Lead, mentor, and develop a team of Site Reliability Engineers, establishing: clear goals, providing actionable feedback, and supporting career growth and professional development.
Set the team’s technical direction and priorities for improving the: reliability, scalability, availability, performance, and operational readiness of Instacart’s systems.
Partner with engineering, product, security, infrastructure, and other: cross-functional teams to define reliability standards, influence system design, and deliver initiatives that improve the customer and developer experience.
Drive incident management and operational excellence, including incident: response, post-incident learning, service-level objectives, capacity planning, observability, and continuous risk reduction.
Promote automation and self-service tooling that reduce operational toil,: improve deployment confidence, and enable engineering teams to own and operate their services effectively.
Balance near-term operational needs with long-term investments, making: thoughtful tradeoffs in a high-growth environment where priorities and requirements can change quickly.
What they're looking for
- Experience leading Site Reliability Engineering, platform engineering,: infrastructure engineering, or developer productivity teams.
- Experience operating highly available services at significant scale and: improving service-level objectives, observability, capacity, or disaster recovery capabilities.
- Experience with infrastructure as code, continuous delivery, monitoring,: logging, tracing, and automated remediation.
- Experience building or evolving reliability programs across multiple: engineering teams, including shared standards, operational reviews, and service ownership practices.
- Demonstrated ability to create alignment across teams, navigate ambiguity,: and turn complex technical challenges into clear, achievable plans.
- A leadership approach grounded in empathy, transparency, direct: communication, collaboration, and a commitment to inclusive team development.