The opportunity
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
What you'll do
Build analytical and empirical performance models for state-of-the-art ML training and inference algorithms.
Characterize asymptotic behavior and identify how algorithmic trade-offs: change with model size, sequence length, batch size, parallelism, and hardware scale.
Construct Pareto frontiers across model quality, latency, throughput, memory: footprint, communication, and compute cost.
Develop prototype implementations and benchmarks for the Cerebras WSE and relevant GPU or software baselines.
Analyze system behavior to identify kernel, compiler, runtime, communication, and algorithmic bottlenecks.
Evaluate emerging techniques in areas such as parallel token generation,: diffusion and speculative decoding, attention, sparsity, mixture-of-experts, low-precision computation, and distributed training.
What they're looking for
- Partner with researchers and kernel, compiler, runtime, inference, and: architecture teams to recommend high-value implementation and co-design directions.
- Develop tools and visualizations that make performance projections,: measurements, and design trade-offs understandable across engineering and research teams.
- Clearly communicate conclusions, assumptions, limitations, and: recommendations through technical reports, presentations, and design reviews.
- Bachelor’s, Master’s, PhD, or equivalent practical experience in Computer: Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field.