The opportunity
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified,…
What you'll do
Profile end-to-end training runs to identify performance bottlenecks across: compute, communication, and storage.
Optimize GPU utilization and throughput for large-scale distributed model training.
Collaborate with runtime and systems engineers to improve kernel efficiency,: scheduling, and collective communication performance.
Implement model graph transforms to improve end to end throughput.
Build tooling to monitor and visualize MFU, throughput, and uptime across clusters.
Partner with researchers to ensure new model architectures scale efficiently during pre-training.
What they're looking for
- Contribute to infrastructure decisions that improve reliability and efficiency of large training jobs.
- Love optimizing performance and digging into systems to understand how every layer interacts.
- Have strong programming skills in Python and C++ (Rust or CUDA a plus).
- Have experience running distributed training jobs on multi-GPU systems or HPC clusters.