ML Research Engineer (Inference)Active

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Implement and adapt transformer-based models (NLP and/or vision) to run on Cerebras hardware

  • Assist in optimizing models for inference performance (latency, throughput)

  • Run experiments, analyze results, and support model improvements

  • Help bring up and validate models on the Cerebras system

  • Debug and troubleshoot model or system issues with guidance from senior team members

  • Support profiling and performance analysis using internal tools

What they're looking for

  • Collaborate with cross-functional teams (ML, software, hardware) on model integration
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field
  • –3 years of experience in software engineering or machine learning in a similar capacity (internships count)
  • Experience with Python and at least one ML framework (e.g., PyTorch, Transformers, vLLM or SGLang)