Software Engineer, PretrainingNew

The opportunity

We’re looking for Software Engineers to build the data systems behind our frontier coding models’ initial training. You’ll work on large-scale crawling, data platform, and pipeline infrastructure, turning raw dumps into the datasets our models train on, and making iteration with researchers fast and reliable.

What you'll do

  • on the Data Quality Team: Build and own high-throughput, fully telemetered data pipelines that process frontier-scale data with end-to-end traceability. If something breaks or drifts, your systems will tell us before the training run does.

  • Train and ship models that classify, rank, filter, clean, and identify data: at extreme throughput. These models have to be both accurate and fast enough to sit in the critical path without becoming the bottleneck.

  • Design and run scaling-ladder experiments on data-mixture, repeatability, and: quality depth that turn “this dataset feels good” into hard evidence the training team can trust.

  • Partner tightly with Data Acquisition to hunt down missing or low-quality: sources, and with the training teams to close the loop on what actually moves loss and downstream evals.

  • Treat data quality as a systems problem and a research problem. You will: write performance-critical code one week and design careful experiments the next.

  • on the Data Platform Team: Build the platform that turns raw web, code, multimodal, and acquired data into training-ready datasets for frontier pretraining runs.

What they're looking for

  • Own the pipelines, orchestration, and tooling that make pretraining data: iteration fast, reliable, observable, and reproducible at scale.
  • Create clear signals for data quality, lineage, freshness, and pipeline: health so researchers can trust what goes into each run.
  • Partner with initial training, crawling, data quality, and acquisition teams: to turn new data ideas into measurable improvements in loss, evals, and model capability.
  • on the Crawling Team: Build and scale the web crawling systems that discover, schedule, fetch, and parse high-quality documents across the open web for initial training.