The opportunity
As a Software Engineer on the RL Data team at Cursor, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.
What you'll do
Designing a task set that teaches a specific agent capability, then iterating: on it from traces and evals until the model actually gets better.
Reading a pile of agent traces, finding a failure mode or a surprising: behavior, and building a system that surfaces more of the same.
Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
Partnering with research on whether a dataset is actually teaching the thing we think it is.