Research Engineer, Frontier Evals & EnvironmentsActive$205K–$380K

The opportunity

About the team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate…

What you'll do

  • Create ambitious RL environments to push our models to their limits, and: measure frontier model capabilities, skills, and behaviors

  • Develop new methodologies for automatically exploring the behavior of these models

  • Dive deep into the science of measurement, including understanding: scalability, reliability, and variance of our evaluation methodology

  • Help steer training for our largest training runs, and see the future first

  • Design scalable systems and processes to support continuous evaluation

  • Build self-improvement loops to automate model understanding

What they're looking for

  • Have strong technical fundamentals in machine learning, software engineering,: systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before.
  • Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals,: graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
  • Are excited by open-ended problems where the path is unclear, the signal is: noisy, and the right answer requires both research taste and engineering execution.
  • Care about product impact and model behavior, not just benchmark movement.: You have opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with.