The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Design, implement, and optimize state-of-the-art transformer architectures: for NLP and computer vision on Cerebras hardware.
Research and prototype novel inference algorithms and model architectures: that exploit the unique capabilities of Cerebras hardware, with emphasis on speculative decoding, pruning/compression, sparse attention, and sparsity .
Train models to convergence, perform hyperparameter sweeps, and analyze results to inform next steps.
Bring up new models on the Cerebras system, validate functional correctness,: and troubleshoot any integration issues.
Profile and optimize model code using Cerebras tools to maximize throughput and minimize latency.
Develop diagnostic tooling or scripts to surface performance bottlenecks and: guide optimization strategies for inference workloads.
What they're looking for
- Collaborate across teams, including software, hardware, and product, to drive: projects from inception through delivery.
- One of the following education and experience combinations: Bachelor’s degree in Computer Science, Software Engineering, Computer Engineering, Electrical Engineering, or a related technical field AND 7+ years of ML software development experience, OR
- Master’s degree in Computer Science or related technical field AND 4+ years: of software development experience, OR
- PhD in Computer Science or related technical field with 2+ years of: relevant research or industry experience, OR