The opportunity
At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences.
What you'll do
Own the reliability posture of production services in your area: availability, latency, capacity, efficiency, performance, and the monitoring and alerting that makes them visible
Define, instrument, and operate against SLIs and SLOs, and use error budgets to drive engineering priorities
Identify trends and problem areas that threaten stability, and provide a path: forward to mitigate risk before it reaches customers
Drive down repair items and prevent classes of incidents rather than resolving them one at a time
Improve detection, response, and recovery: reducing time to acknowledge, engage, mitigate, and restore, with fewer people pulled in
Design for failure: strengthen failure domains, validate recovery paths, and make production changes safer to ship and safer to roll back
What they're looking for
- Participate in on-call for the services you support, and lead the response when production is degraded
- Write post-mortems that identify true root causes, and drive the follow-up work to completion
- Oversee efforts to identify, diagnose, report, and document production: problems across all reliability dimensions
- Write, configure, and deploy code that measurably improves service: reliability — maintainable, reviewed, documented, and well tested