The opportunity
Supabase manages millions of Postgres instances and is growing. We have strong teams across observability, release engineering, and incident management — and we're concentrating our reliability efforts into a dedicated SRE practice that ties the discipline together across the platform.
What you'll do
Partner with service teams to define meaningful SLIs and SLOs grounded in: customer experience, and build the error budget policies that turn them into engineering decisions
Own and evolve the Operational Readiness Review (ORR) process: conducting reviews for new services and major changes across observability, alerting, runbooks, capacity, and graceful degradation
Strengthen the incident-to-improvement pipeline: connecting postmortem findings to operational readiness gaps, identifying repeat failure patterns, and driving systemic fixes
Act as the reliability expert teams pull in for architecture reviews, failure: mode analysis, dependency mapping, and resilience design
Identify and quantify operational toil across the org, and build or advocate for automation that eliminates it
Help teams design sustainable on-call practices: alert quality, escalation paths, runbook coverage, and noise reduction
What they're looking for
- Track and report on org-wide operational maturity, surfacing systemic gaps and driving remediation
- Have 7+ years of experience in SRE, production engineering, or: reliability-focused roles, including experience shaping SRE practices and driving adoption across engineering teams
- Have a software engineering mindset: you write code and build tools, not just configure them
- Have hands-on experience defining and operationalizing SLOs/SLIs at scale,: including error budget policies that actually influenced engineering decisions