The opportunity
Scale has been the leading AI data foundry, helping fuel the most exciting advancements in AI, including frontier model training, enterprise adoption, defense applications, and autonomous vehicles. Our mission is to develop reliable AI systems for the world's most important decisions.
What you'll do
Senior AI PM, Coding: Own Scale's Coding portfolio, built on SWE-Bench Pro (which cut frontier scores from 70%+ to roughly 23%), SWE Atlas, and our FrontierBench contributions. You'll decide what comes next as agents saturate current tasks, and turn that benchmark credibility into revenue from SFT and preference data, RL environments, and agentic coding evals. You'll also own the coding infrastructure roadmap and our expert network of professional software engineers. This role requires hands-on software engineering depth.
Senior AI PM, Leaderboards: Own and scale the SEAL Leaderboard portfolio across all of Scale's domains, turning cutting-edge evals into trusted industry benchmarks. You'll run the Leaderboard Steering Committee, decide which leaderboards get built, and set governance standards for evaluation integrity and release cadence. This is a horizontal role, so breadth across evaluation categories and a strong public-facing presence matter more than depth in any one domain.
Own the roadmap and strategy for your portfolio, setting priorities across: product development, launches, infrastructure investments, and expansion.
Drive alignment across AI-PM, ML, Engineering, Operations, and GTM stakeholders.
Evaluate, prioritize, and operationalize new product proposals, making sure: they align with customer demand, model capability frontiers, and company strategy.
Define and manage the end-to-end product lifecycle, from ideation and design: to launch, scaling, maintenance, and sunset decisions.
What they're looking for
- Partner with ML researchers and domain experts to develop trustworthy: evaluation methodologies, task specifications, scoring and grader design, and quality bars.
- Drive the roadmap for infrastructure, automation, and operational tooling to: improve scalability, reduce manual effort, and accelerate delivery.
- Establish governance processes for quality, evaluation integrity,: contamination prevention, reproducibility, auditability, and release management.
- Work directly with frontier AI labs and enterprise customers to understand: where their models fail, gather feedback, and shape future investments.