AI Platform EngineerActive

The opportunity

We are hiring a AI Platform Engineer to build the execution layer for Supabase's internal AI systems.

What you'll do

  • Ship the agent platform to production. An event-triggered queue, a headless: model-agnostic runtime (choosing the runtime is an open decision you will close), durable state that survives a failed run, a human review gate, atomic rollback, and complete run logging in the warehouse.

  • Own the evaluation layer, and switch on the gate that depends on it. Golden: suites with behavioral assertions rather than intuition, judge criteria with a written rubric, safety cases that must pass on every run, and a CI gate that blocks a regression from merging. This is the precondition for every agent that does anything beyond read internal data.

  • Build and register the agent portfolio. Reporting, drafting, linting, triage: and question-answering agents across the executive, team-lead and individual-contributor layers, plus a meta layer that observes the platform and improves it.

  • Enforce governance in code. Risk tiers the pipeline actually enforces,: least-privilege credentials per agent, tool-permission gates, a decision audit log, and autonomy classes where the dangerous class has no code path rather than a warning label. Nothing runs without a registered owner, tier, tool grant and human gate.

  • Design how the system contacts people. A hard interruption budget per person,: message bundling instead of a stream of pings, and a structure that gives something useful before it asks for anything. Adoption depends on this more than on any other single design choice.

  • Own the platform tooling. The compiler and validator, inventory integrity,: and the paths that distribute context and capabilities into the repositories and chat surfaces where work happens.

What they're looking for

  • Compute the operating measures from production data. A pipeline from raw: system activity through to a computed maturity grade per team, defensible enough that a team can dispute the result and be answered with the query rather than an opinion.
  • Instrument the platform's own return. A ledger that logs the work each agent: absorbs and computes the monthly figure, so the value of the system is a measurement rather than a claim.
  • Have shipped production LLM agent systems that other people depended on. Not: demos, not internal showcases. Systems with operational history, real users and at least one incident you can talk through. Prompt engineering on its own does not clear this bar.
  • Design evaluations, not spot checks. You build golden sets, write behavioral: assertions, define judge rubrics, set pass thresholds and gate CI on the result. You can explain why "we reviewed a bunch of outputs and they looked good" is not evaluation.