Applied AI Researcher, Agent Systems & EvaluationNew$194K–$352K

The opportunity

Nuro is a self-driving technology company on a mission to make autonomy accessible to all. Founded in 2016, Nuro is building the world’s most scalable driver, combining cutting-edge AI with automotive-grade hardware.

What you'll do

  • Eval data collection. Where ground truth comes from. Mining production traces: for labeled outcomes, capturing human accept/reject/edit signal as it happens, building task sets that reflect the real distribution of work rather than the tasks that are easy to score, and knowing when a model-based judge is trustworthy and when it is laundering an assumption.

  • Eval loop construction. Turning a fuzzy objective into a measurement that: runs on every change. Noise floors, statistical standards for acceptance, task suites that resist gaming, and experiments designed for production settings where clean randomization is not always available.

  • Automated hill climbing. The payoff. Once a task has a trustworthy loop,: improvement can be searched rather than hand-crafted — prompts, context strategies, tool sets, routing, reasoning budgets, eventually model choice. This only works if the first two stages are sound; done wrong, it optimizes hard against a metric that means nothing. Post-train models on data nobody else has. This role includes hands-on model work: supervised fine-tuning and RL on open-source vision-language models, using the proprietary driving data Nuro has collected across years of real-world autonomous operation. You would have a labeling workforce available to you, which means you can specify the data you need rather than making do with what exists. Very few researchers get to run this loop, form a hypothesis about model behavior, commission the exact data to test it, post-train, and evaluate against real driving performance. The fungibility of frontier models is precisely why this matters: the weights are rentable, the data and the labeling capacity behind them are not.

  • Establish the evaluation foundation for the agent fleet we already run: eval data sources, task suites, noise floors, and the statistical standard the team uses to accept or reject a change.

  • Take one high-volume workflow from unmeasured to automatically hill-climbing,: end to end, as the template the rest of the system follows.

  • Run a first post-training experiment on an open-weight VLM against our: driving data, and establish whether the result justifies the pipeline.

What they're looking for

  • Put a defensible number on what the platform is worth: which workflows improved, by how much, with what confidence.
  • Graduate degree in CS, ML, statistics, or a related field, or equivalent: research experience. We care about demonstrated research judgment, not credentials.
  • Fluent in the current literature and able to judge it. You read papers: continuously, can tell a real result from a well-marketed one, and have opinions about which recent directions are overrated.
  • Deep understanding of how LLMs work: pretraining through the post-training stack, and what actually happens at inference. You reason from mechanism, not just from published numbers.