ML Algorithm Mapping and Performance Engineer, Core MLNew

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Build analytical and empirical performance models for state-of-the-art ML training and inference algorithms.

  • Characterize asymptotic behavior and identify how algorithmic trade-offs: change with model size, sequence length, batch size, parallelism, and hardware scale.

  • Construct Pareto frontiers across model quality, latency, throughput, memory: footprint, communication, and compute cost.

  • Develop prototype implementations and benchmarks for the Cerebras WSE and relevant GPU or software baselines.

  • Analyze system behavior to identify kernel, compiler, runtime, communication, and algorithmic bottlenecks.

  • Evaluate emerging techniques in areas such as parallel token generation,: diffusion and speculative decoding, attention, sparsity, mixture-of-experts, low-precision computation, and distributed training.

What they're looking for

  • Partner with researchers and kernel, compiler, runtime, inference, and: architecture teams to recommend high-value implementation and co-design directions.
  • Develop tools and visualizations that make performance projections,: measurements, and design trade-offs understandable across engineering and research teams.
  • Clearly communicate conclusions, assumptions, limitations, and: recommendations through technical reports, presentations, and design reviews.
  • Bachelor’s, Master’s, PhD, or equivalent practical experience in Computer: Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field.