Senior Evaluation Algorithm EngineerPosted today

The opportunity

Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed,…

What you'll do

  • Design end-to-end LLM evaluation plans for business scenarios such as: dialogue and financial trading. Build evaluation metric systems and rubrics, transforming subjective model performance judgments into quantifiable, reproducible, and explainable evaluation conclusions.

  • Lead the design and construction of evaluation datasets. Define evaluation: dimensions and scenario coverage, establish high-quality data annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs and have discriminative power.

  • Analyze model capability boundaries and failure modes based on evaluation: results. Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration, making evaluation a critical component of the R&D loop.

  • Drive the automation and scaling of evaluation workflows. Build sustainable: evaluation platforms and toolchains to support high-frequency, stable evaluation needs during rapid model iteration.

  • Collaborate with algorithm, product, and data teams to translate business and: model objectives into clear evaluation standards, and turn evaluation findings into concrete R&D directions and drive their implementation.

What they're looking for

  • Master's degree or above in Computer Science, Artificial Intelligence,: Mathematics, Statistics, or related fields, with a solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes.
  • Hands-on LLM evaluation experience at a large tech company, with: participation in commercial deployment evaluation (not purely academic or offline benchmarking). Familiar with the full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D.
  • Familiar with mainstream evaluation methods (human evaluation, model-based: automatic evaluation / LLM-as-a-judge, metric computation) and their applicable boundaries. Able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, and discriminative rubrics.
  • Systematic control over evaluation data representativeness, annotation: consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions.
  • Proficient in Python, with experience in evaluation workflow automation,: benchmark construction, or evaluation platform development. Able to independently handle data processing, evaluation script writing, and result analysis.
  • Strong business understanding and communication skills, able to translate: evaluation findings into clear improvement directions and effectively drive cross-team collaboration.