The opportunity
Training Runtime designs the core distributed runtime that powers everything from early research experiments to frontier-scale model runs. We build robust, scalable, high performance components to support our distributed training workloads.
What you'll do
Work across our Python and Rust stack
Design, build, and maintain software to orchestrate and monitor machine: learning workloads on our largest supercomputers
Profile and optimize our software stack to support computation orchestration at frontier scale
Improve reliability, observability, and fault tolerance for long-running jobs
Debug complex distributed systems issues across large clusters
Respond to the changing shapes and needs of the ML systems to enable our researchers
What they're looking for
- Have experience developing distributed systems
- Enjoy understanding how large systems behave and fail at scale
- Love being both a developer and an operator
- Care deeply about performance, correctness, and reliability