Agent Post-Training, Frontier Evals and Environments ResearchActive$295K–$445K

San FranciscoOperations

The opportunity

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate…

What they're looking for

  • Have strong technical fundamentals in machine learning, software engineering,: systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before.
  • Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals,: graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
  • Are excited by open-ended problems where the path is unclear, the signal is: noisy, and the right answer requires both research taste and engineering execution.
  • Care about product impact and model behavior, not just benchmark movement.: You have opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with.