The opportunity
We’re looking for Research Scientists who can drive effective RL or mid-training research in a small-team setting. You’ll own ambiguous, hard research problems end-to-end: forming hypotheses, designing experiments, building the training/eval/data needed to test them, and pushing results into the next model.
What you'll do
Improve our understanding of RL, what it takes to handle longer horizon tasks, and train with less compute
Train graders to improve performance on coding tasks with non-verifiable reward
Improve the quality and difficulty of datapoints we use for training our models
Realtime RL for coding agents