The opportunity
AI is becoming vitally important in every function of our society. At Scale, our mission is to accelerate the development of AI applications.
What you'll do
Build, profile and optimize our training and inference framework.
Post-train state of the art models, developed both internally and from the: community, to define stable post-training recipes for our enterprise engagements.
Collaborate with ML teams to accelerate their research and development, and: enable them to develop the next generation of models and data curation..
Create a next-gen agent training algorithm for multi-agent/multi-tool rollouts.
At least 1-3 years of LLM training in a production environment
Passionate about system optimization
What they're looking for
- Experience with post-training methods like RLHF/RLVR and related algorithms like PPO/GRPO etc.
- Ability to demonstrate know-how on how to operate the architecture of the modern GPU cluster
- Experience with multi-node LLM training and inference
- Strong software engineering skills, proficient in frameworks and tools such: as CUDA, Pytorch, transformers, flash attention, etc.