Software Engineer, Agent Evaluation and QualityActive

The opportunity

As a Software Engineer on the Agent Quality team at SpaceXAI, you’ll build the measurement, evaluation, and feedback-loop infrastructure that makes the Cursor core agent reliably better over time.

What you'll do

  • Designing and building best-in-class AI evaluation system: curated datasets, offline replay, scorers / judges, regression alerts, and dashboards.

  • Designing feedback loops from real usage: collecting, cleaning, and interpreting user signals to inform model and harness changes.

  • Developing analysis tooling and workflows for debugging agent behavior: deep dives on failure modes, clustering themes, and surfacing actionable insights.

  • Improving reliability and guardrails by making quality measurable and: operational: defining “good/bad/degraded” sessions, alerting, and triage primitives.

  • You’ve built and operated evaluation or measurement systems, such as AI: evals, experimentation, ranking/relevance, or search quality. You can turn ambiguous “quality” questions into concrete metrics, pipelines, and decisions.

  • You have strong data acumen, and can collaborate effectively with data scientists and researchers.

What they're looking for

  • You have taste and strong opinions on model and agent behaviors. You stay: up-to-date and informed on emerging research and industry trends.
  • You have strong software engineering fundamentals and enjoy shipping production systems.
Software Engineer, Agent Evaluation and Quality at Cursor | Role Match