Site Reliability EngineerActive

The opportunity

Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform.

What you'll do

  • Partner with service teams to define meaningful SLIs and SLOs grounded in: customer experience, and build the error budget policies that turn them into engineering decisions

  • Own and evolve the Operational Readiness Review (ORR) process: conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation

  • Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes

  • Act as the reliability expert teams pull in for architecture reviews, failure: mode analysis, dependency mapping, and resilience design

  • Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it

  • Help teams design sustainable on-call practices: alert quality, escalation paths, runbook coverage, and noise reduction

What they're looking for

  • Track and report on org-wide operational maturity, surfacing systemic gaps and driving remediation
  • Have 7+ years of experience in SRE, production engineering, or: reliability-focused roles, including experience shaping SRE practices and driving adoption across engineering teams
  • Have a software engineering mindset: you write code and build tools, not just configure them
  • Have hands-on experience defining and operationalizing SLOs/SLIs at scale,: including error budget policies that actually influenced engineering decisions