The opportunity
Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences…
What you'll do
Define evaluation strategy and success metrics for GenAI systems, aligning: offline evaluation with online business and customer experience outcomes.
Build and scale evaluation frameworks (golden sets, synthetic data, automated: regressions, rubric-based grading, LLM-as-judge where appropriate) with strong controls for bias, drift, and reliability.
Design the data flywheel: instrumentation, feedback collection, data quality checks, labeling strategy, dataset versioning, and governance to support continuous improvement.
Lead cross-functional quality initiatives across product, ops, and: engineering, driving clarity on what “good” looks like and how teams act on evaluation results.
Develop and productionize pipelines for dataset creation, model monitoring,: evaluation-at-scale, and continuous testing (pre-deploy and post-deploy).
Drive technical decisions and architecture for evaluation and data: infrastructure, balancing speed, rigor, cost, and safety.
What they're looking for
- Educational Background: PhD in Computer Science, Mathematics, Statistics, or related technical field (or equivalent practical experience).
- Industry Experience: 10+ years building, testing, and shipping ML/AI systems end-to-end; including 2+ years of experience with GenAI/LLM systems in production.
- Leadership Experience: 5+ years leading large, ambiguous technical initiatives as a senior IC, influencing roadmap and engineering/science direction across teams.
- Technical Proficiency: Deep expertise in evaluation methodology (offline/online alignment, metric design, human-in-the-loop evaluation, A/B testing, power analysis, regression testing).