The opportunity
Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades.
What you'll do
Serve as a senior technical leader driving the long-term reliability and: observability strategy across Robinhood’s infrastructure
Partner closely across many different types of engineers to raise the bar for: operational excellence and incident response
Lead incident mitigation efforts by coordinating service owners, facilitating: time-sensitive decisions like rollbacks, traffic shifts, and maintaining a clear source of truth during active incidents
Develop and maintain incident management processes and procedures to ensure: timely resolution and minimize customer impact
Own incident discovery at the company level by defining and maintaining: global dashboards and alerts tied to critical user journeys (CUJs), availability, and business-impact metrics
Own and evolve incident response tooling and processes, including education,: adoption, and measurement of MTTD/MTTR improvements
What they're looking for
- Drive post-incident governance and learning, defining standards for: postmortems, SEV reviews, and follow-up tracking to ensure durable reliability improvements
- Design and implement next-generation failure mitigation strategies that avoid: full-region or full-datacenter failovers
- Define and build frameworks to improve monitoring, alerting, and: observability across hundreds of services and systems
- Define and own the roadmap of bringing observability to critical user journeys for Robinhood’s products