Senior ML Systems Engineer, InferencePosted today$150K–$220K

The opportunity

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.

What you'll do

  • Define how we measure inference performance, including throughput, time to: first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.

  • Profile and diagnose performance problems across the serving stack, from: scheduling and memory management down to kernels and interconnect.

  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.

  • Turn what you learn into production-ready runtimes, configurations, and: defaults that customers benefit from automatically.

  • Work closely with product and infrastructure teams to shape how inference is offered on Runpod.

  • Keep up with the fast-moving inference ecosystem, including the open-source: community, and decide what's worth adopting, what's worth building, and what's worth contributing back.

What they're looking for

  • + years of professional system engineering experience.
  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine): in production or at serious benchmark scale.
  • Strong software engineering skills in Python . You're comfortable working in: large, performance-critical codebases.
  • A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.
  • Experience with modern inference optimization techniques such as: quantization, speculative decoding, or distributed serving.
  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.
  • The ability to explain your results clearly in writing and turn them into decisions.