The opportunity
The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and…
What you'll do
Develop a rigorous understanding of what makes an agent a great collaborator: across professional, creative, technical, and everyday work.
Turn qualitative judgments about model behavior into concrete hypotheses,: evals, graders, and training interventions.
Study explicit and implicit user signals to understand which behaviors create: trust, satisfaction, continued use, and successful outcomes.
Work with human experts and trainers to produce high-quality, tasteful: rollouts and preference data that capture excellent collaborative behavior.
Improve reward models and RL objectives for model behaviors.
Work with pretraining and early-training teams on data mixtures, objectives,: synthetic data, and other upstream choices that shape downstream personality.
What they're looking for
- Build sustainable pipelines for updating older training data as our: understanding of excellent model behavior evolves.
- Partner closely with ChatGPT, Codex, and other product teams to turn consumer: insight into model improvements and validate them in real workflows.
- Own projects end to end, from observing a subtle behavioral failure through: experimentation, training, evaluation, and launch.
- Think instinctively from the user’s perspective and care deeply about how: models feel to work with, not only how they perform on benchmarks.