Senior Research Engineer - Inference MLNew

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Design, implement, and optimize state-of-the-art transformer architectures: for NLP and computer vision on Cerebras hardware.

  • Research and prototype novel  inference algorithms  and model architectures: that exploit the unique capabilities of Cerebras hardware, with emphasis on  speculative decoding, pruning/compression, sparse attention, and sparsity .

  • Train models to convergence, perform hyperparameter sweeps, and analyze results to inform next steps.

  • Bring up new models on the Cerebras system, validate functional correctness,: and troubleshoot any integration issues.

  • Profile and optimize model code using Cerebras tools to maximize throughput and minimize latency.

  • Develop diagnostic tooling or scripts to surface performance bottlenecks and: guide optimization strategies for inference workloads.

What they're looking for

  • Collaborate across teams, including software, hardware, and product, to drive: projects from inception through delivery.
  • One of the following education and experience combinations: Bachelor’s degree in Computer Science, Software Engineering, Computer Engineering, Electrical Engineering, or a related technical field  AND 7+ years  of ML software development experience,  OR
  • Master’s degree in Computer Science or related technical field  AND 4+ years: of software development experience,  OR
  • PhD in Computer Science or related technical field with  2+ years  of: relevant research or industry experience,  OR