HPC Infrastructure Engineer - GPU ClustersActive
The opportunity
Every model we train runs on infrastructure this role owns. We operate NVIDIA GPU clusters across bare metal and rented capacity, and we're looking for an engineer to join our small research infrastructure team and make that compute fast, reliable, and boring - in the best sense.
What you'll do
ElevenAgents enables businesses to deliver seamless and intelligent customer: experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.
ElevenCreative empowers creators and marketers to generate and edit speech,: music, image, and video across 70+ languages.
ElevenAPI gives developers access to our leading AI audio foundational models.
High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.
Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.
AI first: We use AI to move faster with higher-quality results. We do this: across the whole company—from engineering to growth to operations.
What they're looking for
- Have run large-scale Linux server or GPU environments in production and enjoy both building and operating
- Know the NVIDIA stack well: drivers, CUDA, NCCL, DCGM — or have deep systems experience and learn hardware stacks fast
- Are comfortable with bare-metal environments, server hardware, and high-speed networking
- Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform
- Are happy digging into noisy data (metrics, logs, PromQL) to find what's actually wrong
- Like owning real scope end to end and being the person others rely on
- Don't consider any task above or beneath you: datacenter trips included