The opportunity
Our team is organized around the north star goal of building an AI scientist – a system capable of solving the long term reasoning challenges and basic capabilities necessary to push the scientific frontier.
What you'll do
Design and implement large-scale infrastructure systems to support AI: scientist training, evaluation, and deployment across distributed environments
Identify and resolve infrastructure bottlenecks impeding progress toward scientific capabilities
Develop robust and reliable evaluation frameworks for measuring progress towards scientific AGI.
Build scalable and performant VM/sandboxing/container architectures to safely: execute long-horizon AI tasks and scientific workflows
Collaborate to translate experimental requirements into production-ready infrastructure
Develop large scale data pipelines to handle advanced language model training requirements
What they're looking for
- Optimize large scale training and inference pipelines for stable and efficient reinforcement learning
- Have 6+ years of highly-relevant experience in infrastructure engineering: with demonstrated expertise in large-scale distributed systems
- Are a strong communicator and enjoy working collaboratively
- Possess deep knowledge of performance optimization techniques and system: architectures for high-throughput ML workloads