Senior Staff Machine Learning Engineer, Data & EvalActive$244K

The opportunity

Airbnb was born in 2007 when two hosts welcomed three guests to their San Francisco home, and has since grown to over 5 million hosts who have welcomed over 2 billion guest arrivals in almost every country across the globe. Every day, hosts offer unique stays and experiences…

What you'll do

  • Define evaluation strategy and success metrics for GenAI systems, aligning: offline evaluation with online business and customer experience outcomes.

  • Build and scale evaluation frameworks (golden sets, synthetic data, automated: regressions, rubric-based grading, LLM-as-judge where appropriate) with strong controls for bias, drift, and reliability.

  • Design the data flywheel: instrumentation, feedback collection, data quality checks, labeling strategy, dataset versioning, and governance to support continuous improvement.

  • Lead cross-functional quality initiatives across product, ops, and: engineering, driving clarity on what “good” looks like and how teams act on evaluation results.

  • Develop and productionize pipelines for dataset creation, model monitoring,: evaluation-at-scale, and continuous testing (pre-deploy and post-deploy).

  • Drive technical decisions and architecture for evaluation and data: infrastructure, balancing speed, rigor, cost, and safety.

What they're looking for

  • Educational Background: PhD in Computer Science, Mathematics, Statistics, or related technical field (or equivalent practical experience).
  • Industry Experience: 10+ years building, testing, and shipping ML/AI systems end-to-end; including 2+ years of experience with GenAI/LLM systems in production.
  • Leadership Experience: 5+ years leading large, ambiguous technical initiatives as a senior IC, influencing roadmap and engineering/science direction across teams.
  • Technical Proficiency: Deep expertise in evaluation methodology (offline/online alignment, metric design, human-in-the-loop evaluation, A/B testing, power analysis, regression testing).