Site Reliability EngineerActive$150K–$200K

The opportunity

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.

What you'll do

  • Defining and enforcing reliability standards across engineering

  • Designing incident response processes and improving recovery times

  • Building observability systems and reliability tooling

  • Driving SLO adoption and production readiness reviews

  • Reducing operational toil through automation

  • Increase platform uptime and reduce incident frequency and duration

What they're looking for

  • Establish and operationalize SLIs/SLOs across services
  • Improve MTTR through better tooling, automation, and runbooks
  • Strengthen production readiness standards
  • Drive long-term systemic reliability improvements