Training Performance EngineerActive$250K–$445K

Hybrid · San FranciscoTechnology

The opportunity

About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified,…

What they're looking for

  • Contribute to infrastructure decisions that improve reliability and efficiency of large training jobs.
  • Love optimizing performance and digging into systems to understand how every layer interacts.
  • Have strong programming skills in Python and C++ (Rust or CUDA a plus).
  • Have experience running distributed training jobs on multi-GPU systems or HPC clusters.