Software Engineer, RL Training InfraActive$295K–$445K

Hybrid · San FranciscoTechnology

The opportunity

About the Team The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next). We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning.

What they're looking for

  • Become useful quickly in messy, ambiguous areas where ownership matters more than a perfectly scoped project.
  • Debug hard failures in shipped or near-shipped models and turn messy: qualitative behavior into concrete hypotheses, experiments, and fixes.
  • Are driven by having a large impact on the world and want to train and ship: the best model in the world to our users.
  • Are a strong generalist engineer with experience in some layer of ML infrastructure.
Software Engineer, RL Training Infra at OpenAI | Role Match