Director, Research - AI EvalsActive$348K

The opportunity

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI.

What you'll do

  • Own AI evaluation methods and operations for Figma's AI-powered experiences —: define quality dimensions, design how we measure them, and turn results into decision-ready signal

  • Build and maintain evaluation frameworks, rubrics, golden datasets, and: quality bars, combining human evaluation with automated/model-based approaches (e.g., LLM-as-judge) where appropriate

  • Partner with engineering to stand up repeatable, reproducible evaluation: pipelines and regression testing, so evaluation is a routine part of how AI features are built and shipped

  • Produce clear readouts and dashboards that let stakeholders confidently make: go/no-go and prioritization decisions

  • Socialize a shared definition of quality so evaluation standards are adopted: across teams rather than re-invented — and advocate for evaluation as a strategic partner in the product process

  • Manage a small team to execute our AI evals in partnership with contractors, internal staff, and/or LLMs

What they're looking for

  • + years of experience in product, research, applied research, or a closely: related field, including 2+ years of management experience
  • Direct, hands-on experience owning the evaluation of AI/LLM-powered products
  • Expertise designing and running AI evaluation: human evaluation programs, rubric and benchmark/golden-dataset construction, inter-rater reliability — and sound judgment about when and how to apply automated/model-based approaches (e.g., LLM-as-judge), including their limitations
  • Strength across both qualitative and quantitative methods, comfort with data: and metrics, and the ability to reason about model behavior