Senior/Staff Software Engineer, AI Agent InfrastructureActive$194K–$352K

The opportunity

Nuro believes self-driving vehicles are the most immediate and profound opportunity for AI to drive positive change in the physical world. Safer streets, more time for what matters, and easier access to the world around us, that’s why we’re building a universal autonomy…

What you'll do

  • Build the closed-loop measurement layer that tells us, per workflow, whether: agent output is accepted, reverted, or overridden — and use it to decide where autonomy expands and where it gets pulled back.

  • Take the autoresearch loop from assisted to unattended for a bounded class of: experiments, including the eval and confidence machinery required to run it without a human in the loop.

  • Design the isolation and permissioning model that lets agents act on: production repositories and infrastructure with an auditable record of what they did and why.

  • + years of software engineering experience (or 4+ with a Master's) in: computer science, engineering, or equivalent practical experience. Staff-level candidates should bring correspondingly deeper scope and ownership.

  • Deep, current taste in LLM research. You understand how a model is trained: from scratch — data, tokenization, architecture, pretraining dynamics, the full post-training stack of supervised fine-tuning, preference optimization, and RL — and you can reason about what a training decision does to model behavior. You follow the literature because you want to, not because it is on a roadmap.

  • You know what happens under the hood at inference. Attention and KV-cache: behavior, batching and scheduling, quantization, speculative decoding, prefix caching, context handling, and how each trades off latency, throughput, and cost.

What they're looking for

  • You have built and operated LLM-based agent systems in production: tool use, orchestration, sandboxing, retrieval, memory — and you know where they break.
  • Strong backend and distributed systems background at scale: cloud infrastructure, service design, storage, queuing, and the judgment to build things that stay up.
  • Strong programming skills in Python.
  • You are opinionated about evaluation. You have argued with someone about: whether a benchmark measured anything real, and you were right.