The opportunity
Nuro is a self-driving technology company on a mission to make autonomy accessible to all. Founded in 2016, Nuro is building the world’s most scalable driver, combining cutting-edge AI with automotive-grade hardware.
What you'll do
Eval data collection. Where ground truth comes from. Mining production traces: for labeled outcomes, capturing human accept/reject/edit signal as it happens, building task sets that reflect the real distribution of work rather than the tasks that are easy to score, and knowing when a model-based judge is trustworthy and when it is laundering an assumption.
Eval loop construction. Turning a fuzzy objective into a measurement that: runs on every change. Noise floors, statistical standards for acceptance, task suites that resist gaming, and experiments designed for production settings where clean randomization is not always available.
Automated hill climbing. The payoff. Once a task has a trustworthy loop,: improvement can be searched rather than hand-crafted — prompts, context strategies, tool sets, routing, reasoning budgets, eventually model choice. This only works if the first two stages are sound; done wrong, it optimizes hard against a metric that means nothing. Post-train models on data nobody else has. This role includes hands-on model work: supervised fine-tuning and RL on open-source vision-language models, using the proprietary driving data Nuro has collected across years of real-world autonomous operation. You would have a labeling workforce available to you, which means you can specify the data you need rather than making do with what exists. Very few researchers get to run this loop, form a hypothesis about model behavior, commission the exact data to test it, post-train, and evaluate against real driving performance. The fungibility of frontier models is precisely why this matters: the weights are rentable, the data and the labeling capacity behind them are not.
Establish the evaluation foundation for the agent fleet we already run: eval data sources, task suites, noise floors, and the statistical standard the team uses to accept or reject a change.
Take one high-volume workflow from unmeasured to automatically hill-climbing,: end to end, as the template the rest of the system follows.
Run a first post-training experiment on an open-weight VLM against our: driving data, and establish whether the result justifies the pipeline.
What they're looking for
- Put a defensible number on what the platform is worth: which workflows improved, by how much, with what confidence.
- Graduate degree in CS, ML, statistics, or a related field, or equivalent: research experience. We care about demonstrated research judgment, not credentials.
- Fluent in the current literature and able to judge it. You read papers: continuously, can tell a real result from a well-marketed one, and have opinions about which recent directions are overrated.
- Deep understanding of how LLMs work: pretraining through the post-training stack, and what actually happens at inference. You reason from mechanism, not just from published numbers.