The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Apply post-training techniques (e.g. RLVR, RLHF, GRPO etc.) techniques to improve model performance.
Build and maintain evaluation pipelines to measure model performance across tasks and domains.
Debug issues across the ML stack, including data pipelines, training jobs,: model outputs and mixed or lower precision computation.
Collaborate with researchers to translate ML ideas into efficient, scalable implementation.
Design, implement, and scale ML pipelines across all stages of LLM: development (pretraining, fine-tuning, alignment).
Work with large datasets, including dataset generation, filtering, and synthetic data approaches.
What they're looking for
- Optimize training and inference workflows for performance, efficiency, and reliability.
- Contribute high-quality, maintainable code to shared ML infrastructure.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- + years of experience (including internships, research, or industry: experience) working with machine learning systems; we are hiring multiple positions for various levels.