Training Performance EngineerActive$250K–$445K

The opportunity

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified,…

What you'll do

  • Profile end-to-end training runs to identify performance bottlenecks across: compute, communication, and storage.

  • Optimize GPU utilization and throughput for large-scale distributed model training.

  • Collaborate with runtime and systems engineers to improve kernel efficiency,: scheduling, and collective communication performance.

  • Implement model graph transforms to improve end to end throughput.

  • Build tooling to monitor and visualize MFU, throughput, and uptime across clusters.

  • Partner with researchers to ensure new model architectures scale efficiently during pre-training.

What they're looking for

  • Contribute to infrastructure decisions that improve reliability and efficiency of large training jobs.
  • Love optimizing performance and digging into systems to understand how every layer interacts.
  • Have strong programming skills in Python and C++ (Rust or CUDA a plus).
  • Have experience running distributed training jobs on multi-GPU systems or HPC clusters.