Distributed LLM Inference EngineerActive$170K–$245K

The opportunity

At  Anyscale , we're on a mission to democratize distributed computing and make it accessible to software developers of all skill levels. We’re commercializing  Ray , a popular open-source project that's creating an ecosystem of libraries for scalable machine learning.

What you'll do

  • Iterate very quickly with product teams to ship the end to end solutions for: Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale

  • Work across the stack integrating Ray Data and LLM engine providing: optimizations achieving low cost solutions for large scale ML inference

  • Integrate with Open source software like vLLM, work closely with the: community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source

  • Follow the latest state-of-the-art in the open source and the research: community, implementing and extending best practices

  • Familiarity with running ML inference at large scale with high throughput and low latency

  • Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)

What they're looking for

  • Solid understanding of distributed systems, ML inference challenges
  • ML Systems knowledge
  • Experience using Ray
  • Work closely with community on LLM engines like vLLM, TensorRT-LLM