Member of Technical Staff - Research, InferenceActive$150K–$350K

The opportunity

AI needs a new infrastructure layer. We're building it at Modal.

What you'll do

  • Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic, and whatever else the research agenda calls for.

  • Train custom speculators against real production traffic and feed what you: learn back into target models -- acceptance length is the metric that decides the win.

  • Work directly with customers alongside our Forward Deployed Engineers to: deploy and tune models, and bring what you learn back into the research.

  • Carry and expand collaborations with outside research labs, for example: our work with ZLab on DFlash , a speculator design built on KV injection and blockwise parallel drafting

  • our work with SGLang on specdec and multimodal inference performance

  • our work on Flash Attention 4 kernels

What they're looking for

  • Work with engineering to turn frontier serving techniques into products: primitives for disaggregation, fast weight refresh for models that keep training after deployment, observability for quality and latency in production, or even a next-generation inference engine.
  • Help shape the research agenda. None of the above is prescriptive; your work will help guide our future.
  • A research-leaning or systems background in LLM inference, with work you can point to.
  • Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling.