Agent Post-Training, API & Power UsersActive$295K–$445K

The opportunity

The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and…

What you'll do

  • Design and run experiments that improve model behavior in API and power-user: workflows: function calling, tool use, coding, planning, long-horizon execution, factuality, instruction following, error recovery, and calibrated reasoning.

  • Build evals, graders, and environments from real developer and power-user: workflows, then turn observed failures into training data, model-behavior hypotheses, and shipped improvements.

  • Partner with API and power-users to identify high-leverage behavior gaps and: convert product signals into post-training interventions.

  • Improve how models behave when composed into systems: using tools reliably, respecting developer intent, handling partial failures, asking for clarification when appropriate, and maintaining coherence across multi-step tasks.

  • Own end-to-end model behavior projects, from qualitative failure analysis: through data generation, training experiments, eval design, integration into major runs, and launch readiness.

  • Develop feedback loops that use power-user traces, API usage patterns, and: production-like environments to discover the next frontier of agentic model failures and gaps.

What they're looking for

  • Help decide which agentic capabilities, behavioral fixes, and partner-team: integrations are ready for inclusion in major model runs.
  • Debug hard failures in shipped or near-shipped models by moving between: traces, evals, training data, model outputs, and product context.
  • Work on early-training and alignment interventions, including data mixtures,: objectives, synthetic data, and eval loops that shape downstream agent behavior.
  • Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
Agent Post-Training, API & Power Users at OpenAI | Role Match