Staff Python / PyTorch Developer — Frontend Inference Compiler – DubaiActive

The opportunity

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.

What you'll do

  • Drive and provide technical guidance to a team of software engineers working: on complex machine learning integration projects.

  • Design and implement ML features (e.g., structured outputs, biased sampling,: predicted outputs) that improve performance of generative AI models at inference time.

  • Design and implement high-throughput, low-latency multimodal inference models: that support delivery of image, audio, and video inputs and outputs.

  • Maintain our scalable serving backend for handling many concurrent requests per minute.

  • Scale our inference service by implementing detailed observability throughout the entire stack.

  • Analyze and improve latency, throughput, memory usage, and compute efficiency: on the service and the implementation of various features.

What they're looking for

  • Optimize software to accelerate generative LLM inference by achieving high throughput and low latency.
  • Stay up-to-date with advancements in machine learning and deep learning, and: apply state-of-the-art techniques to enhance our solutions.
  • Evaluate trade-offs between different approaches, clearly articulate design: choices, and develop detailed proposals for implementing new features.
  • Uncover, scope, and prioritize significant areas of technical debt across the: software stack to ensure continued high quality of the inference service.