The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Implement and adapt transformer-based models (NLP and/or vision) to run on Cerebras hardware
Assist in optimizing models for inference performance (latency, throughput)
Run experiments, analyze results, and support model improvements
Help bring up and validate models on the Cerebras system
Debug and troubleshoot model or system issues with guidance from senior team members
Support profiling and performance analysis using internal tools
What they're looking for
- Collaborate with cross-functional teams (ML, software, hardware) on model integration
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field
- –3 years of experience in software engineering or machine learning in a similar capacity (internships count)
- Experience with Python and at least one ML framework (e.g., PyTorch, Transformers, vLLM or SGLang)