The opportunity
AI needs a new infrastructure layer. We're building it at Modal.
What you'll do
Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic, and whatever else the research agenda calls for.
Train custom speculators against real production traffic and feed what you: learn back into target models -- acceptance length is the metric that decides the win.
Work directly with customers alongside our Forward Deployed Engineers to: deploy and tune models, and bring what you learn back into the research.
Carry and expand collaborations with outside research labs, for example: our work with ZLab on DFlash , a speculator design built on KV injection and blockwise parallel drafting
our work with SGLang on specdec and multimodal inference performance
our work on Flash Attention 4 kernels
What they're looking for
- Work with engineering to turn frontier serving techniques into products: primitives for disaggregation, fast weight refresh for models that keep training after deployment, observability for quality and latency in production, or even a next-generation inference engine.
- Help shape the research agenda. None of the above is prescriptive; your work will help guide our future.
- A research-leaning or systems background in LLM inference, with work you can point to.
- Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling.