The opportunity
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.
What you'll do
Define how we measure inference performance, including throughput, time to: first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.
Profile and diagnose performance problems across the serving stack, from: scheduling and memory management down to kernels and interconnect.
Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.
Turn what you learn into production-ready runtimes, configurations, and: defaults that customers benefit from automatically.
Work closely with product and infrastructure teams to shape how inference is offered on Runpod.
Keep up with the fast-moving inference ecosystem, including the open-source: community, and decide what's worth adopting, what's worth building, and what's worth contributing back.
What they're looking for
- + years of professional system engineering experience.
- Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine): in production or at serious benchmark scale.
- Strong software engineering skills in Python . You're comfortable working in: large, performance-critical codebases.
- A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.
- Experience with modern inference optimization techniques such as: quantization, speculative decoding, or distributed serving.
- Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.
- The ability to explain your results clearly in writing and turn them into decisions.