The opportunity
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.
What you'll do
Defining and enforcing reliability standards across engineering
Designing incident response processes and improving recovery times
Building observability systems and reliability tooling
Driving SLO adoption and production readiness reviews
Reducing operational toil through automation
Increase platform uptime and reduce incident frequency and duration
What they're looking for
- Establish and operationalize SLIs/SLOs across services
- Improve MTTR through better tooling, automation, and runbooks
- Strengthen production readiness standards
- Drive long-term systemic reliability improvements