The opportunity
About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs. With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified,…
What they're looking for
- Contribute to infrastructure decisions that improve reliability and efficiency of large training jobs.
- Love optimizing performance and digging into systems to understand how every layer interacts.
- Have strong programming skills in Python and C++ (Rust or CUDA a plus).
- Have experience running distributed training jobs on multi-GPU systems or HPC clusters.