The opportunity
Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI.
What you'll do
Own AI evaluation methods and operations for Figma's AI-powered experiences —: define quality dimensions, design how we measure them, and turn results into decision-ready signal
Build and maintain evaluation frameworks, rubrics, golden datasets, and: quality bars, combining human evaluation with automated/model-based approaches (e.g., LLM-as-judge) where appropriate
Partner with engineering to stand up repeatable, reproducible evaluation: pipelines and regression testing, so evaluation is a routine part of how AI features are built and shipped
Produce clear readouts and dashboards that let stakeholders confidently make: go/no-go and prioritization decisions
Socialize a shared definition of quality so evaluation standards are adopted: across teams rather than re-invented — and advocate for evaluation as a strategic partner in the product process
Manage a small team to execute our AI evals in partnership with contractors, internal staff, and/or LLMs
What they're looking for
- + years of experience in product, research, applied research, or a closely: related field, including 2+ years of management experience
- Direct, hands-on experience owning the evaluation of AI/LLM-powered products
- Expertise designing and running AI evaluation: human evaluation programs, rubric and benchmark/golden-dataset construction, inter-rater reliability — and sound judgment about when and how to apply automated/model-based approaches (e.g., LLM-as-judge), including their limitations
- Strength across both qualitative and quantitative methods, comfort with data: and metrics, and the ability to reason about model behavior