The opportunity
We are hiring a AI Platform Engineer to build the execution layer for Supabase's internal AI systems.
What you'll do
Ship the agent platform to production. An event-triggered queue, a headless: model-agnostic runtime (choosing the runtime is an open decision you will close), durable state that survives a failed run, a human review gate, atomic rollback, and complete run logging in the warehouse.
Own the evaluation layer, and switch on the gate that depends on it. Golden: suites with behavioral assertions rather than intuition, judge criteria with a written rubric, safety cases that must pass on every run, and a CI gate that blocks a regression from merging. This is the precondition for every agent that does anything beyond read internal data.
Build and register the agent portfolio. Reporting, drafting, linting, triage: and question-answering agents across the executive, team-lead and individual-contributor layers, plus a meta layer that observes the platform and improves it.
Enforce governance in code. Risk tiers the pipeline actually enforces,: least-privilege credentials per agent, tool-permission gates, a decision audit log, and autonomy classes where the dangerous class has no code path rather than a warning label. Nothing runs without a registered owner, tier, tool grant and human gate.
Design how the system contacts people. A hard interruption budget per person,: message bundling instead of a stream of pings, and a structure that gives something useful before it asks for anything. Adoption depends on this more than on any other single design choice.
Own the platform tooling. The compiler and validator, inventory integrity,: and the paths that distribute context and capabilities into the repositories and chat surfaces where work happens.
What they're looking for
- Compute the operating measures from production data. A pipeline from raw: system activity through to a computed maturity grade per team, defensible enough that a team can dispute the result and be answered with the query rather than an opinion.
- Instrument the platform's own return. A ledger that logs the work each agent: absorbs and computes the monthly figure, so the value of the system is a measurement rather than a claim.
- Have shipped production LLM agent systems that other people depended on. Not: demos, not internal showcases. Systems with operational history, real users and at least one incident you can talk through. Prompt engineering on its own does not clear this bar.
- Design evaluations, not spot checks. You build golden sets, write behavioral: assertions, define judge rubrics, set pass thresholds and gate CI on the result. You can explain why "we reviewed a bunch of outputs and they looked good" is not evaluation.