Research Engineer - InferencePosted today
The opportunity
We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production. The quality of our models only matters if they can be served fast, reliably, and at scale.
What you'll do
ElevenAgents enables businesses to deliver seamless and intelligent customer: experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.
ElevenCreative empowers creators and marketers to generate and edit speech,: music, image, and video across 70+ languages.
ElevenAPI gives developers access to our leading AI audio foundational models.
High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.
Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.
AI first: We use AI to move faster with higher-quality results. We do this: across the whole company—from engineering to growth to operations.
What they're looking for
- Experience deploying and serving ML models in production, ideally for: latency-sensitive or real-time applications.
- Strong engineering skills in GPU programming and inference optimization: (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).
- The capacity to autonomously profile, diagnose, and eliminate bottlenecks: across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it.