The opportunity
The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and…
What you'll do
Design and run experiments that improve agentic model behavior across coding,: tool use, function calling, computer use, multi-agent collaboration, long-horizon tasks, factuality, instruction following, and calibrated reasoning.
Own end-to-end improvements to the post-training stack, including RL, data: pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis.
Build evals and environments that expose the next set of model failures, then: turn those failures into training data, product fixes, or new research directions.
Partner with Codex, API/platform, and ChatGPT product teams to understand: what users need and translate product signal into model improvements.
Work on early-training and alignment interventions, including data mixtures,: objectives, synthetic data, and eval loops that shape downstream agent behavior.
Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs.
What they're looking for
- Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness.
- Take on cross-functional projects that touch model training, product: infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments.
- Debug hard failures in shipped or near-shipped models and turn messy: qualitative behavior into concrete hypotheses, experiments, and fixes.
- Have strong technical fundamentals in machine learning, software engineering,: systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before.