Research Scientist, Safety Post TrainingActive$46K–$183K

The opportunity

As the leading data and evaluation partner for frontier AI companies, Scale plays an integral role in understanding the capabilities and safeguarding AI models and systems. Building on this expertise, Scale Labs has launched a new team focused on policy research, to bridge the…

What you'll do

  • Design and run post-training pipelines to study how training choices affect: model safety, robustness, and alignment properties;

  • Develop interpretability-informed evaluations that reveal how and why models: produce unsafe, deceptive, or otherwise undesirable behaviors, and use those insights to guide targeted mitigations;

  • Collaborate with policymakers, engineers, and other researchers to translate: post-training and interpretability findings into actionable safety standards, evaluation benchmarks, and best practices.

  • Commitment to our mission of promoting safe, secure, and trustworthy AI: deployments in the industry as frontier AI capabilities continue to advance.

  • Experience with post-training and RL techniques such as RLHF, DPO, GRPO, and similar approaches.

  • A track record of published research in machine learning, particularly in generative AI.

What they're looking for

  • At least three years of experience addressing sophisticated ML problems,: whether in a research setting or in product development.
  • Strong written and verbal communication skills to operate in a cross-functional team.
  • Experience with mechanistic interpretability, probing, or other techniques for understanding model internals.
  • Familiarity with red-teaming or adversarial evaluation of post-trained models.
Research Scientist, Safety Post Training at Scale AI | Role Match