The opportunity
Evaluation Infrastructure plays a critical role at Nuro, directly enabling L4 driverless deployment. The team supports two demanding workloads: day-to-day Autonomy Evaluation that powers rapid software iteration, and large-scale Driverless Safety Validation that produces the…
What you'll do
Build and own a unified metrics, evaluation, and validation platform: pipelines, introspection tooling, and analysis products that turn on-road and simulation logs into high-fidelity signals for autonomy iteration and driverless safety validation
Drive the technical bar for metric quality across both heuristic and ML-based: approaches; invest in the scale, reliability, and CI/CD of the evaluation stack to shorten time-to-signal for evaluation and time-to-confidence for validation, and to meet high SLAs for downstream stakeholders
Mentor and grow the Evaluation Infrastructure team, and champion AI-native: engineering practices that compound team velocity and code quality
Partner with Product, Autonomy, Systems & Safety, and Simulation teams to: define and execute the vision and strategy for evaluation at Nuro
You have a degree in B.Sc or M.Sc., plus 4 years of relevant work experience
Domain experience: Strong fluency in distributed systems, large-scale data and ML evaluation pipelines, metrics frameworks (heuristic and/or ML-based), and analytics platforms
What they're looking for
- Engineering leadership: Experience setting technical vision, roadmap, and prioritization for a team operating at the intersection of autonomy, safety, and data infrastructure; a clear, concise communicator who partners effectively with PMs, engineers, and cross-functional stakeholders across Autonomy, Systems & Safety, and Simulation
- Technical excellence: Ability and willingness to deep-dive into implementation; sets the technical bar for metric quality, pipeline rigor, and safety-critical engineering practice across the broader software organization; strong proficiency in Python, C++, or similar languages
- AI-native mindset: Daily user of modern AI coding assistants and agentic tools (Claude Code, Cursor, and similar), with strong intuition for where they accelerate engineering work and where they don't; eager to apply LLMs and ML systems to evaluation problems, from automated triage and metric generation to natural-language analysis of fleet behavior; raises the team's productivity, code quality, and signal density through thoughtful AI integration
- Knowledge of data engineering, and its tooling and best practices