The opportunity
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences…
What you'll do
Develop a roadmap with a longer-term vision for Reliability and serve as a: strategic thought partner within the organization.
Design, implement and influence company-wide SRE architecture, innovation, engineering, and standards.
Create incident management processes that can scale with the organization as: it continues its rapid growth. Assess how the organization manages incidents and responds to them; reduce operational toil stemming from incident management.
Foster the SRE/Reliability model that takes into consideration the nuances of: an engineering culture that has a great sense of ownership over their services.
Bring a strong customer focus to the Reliability function, centered on: optimizing the infrastructure and platform, and ensuring systems are highly available and performant.
Develop Production Readiness standards to ensure service reliability.: Automate as much as possible and always configure as code. Predict future failures and work proactively to mitigate them. Advocate and implement reliable design patterns (circuit breakers, graceful degradation, etc.)
What they're looking for
- Culture, Influence and Team Leadership
- Create a culture where Reliability is a state of mind, instilling a proactive: approach to seeing patterns and opportunities to increase leverage and tooling.
- Build deep partnerships with engineering leaders. Work closely with product: engineering teams on design and implementation choices of large-scale distributed systems.
- Partner with the broader organization to learn from incidents through a blameless post mortem process.