The opportunity
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.
What you'll do
Hardware Validation & Benchmarking: Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads.
Uptime & SLA Enforcement: Monitor fleet health to identify performance degradation. You will help audit downtime and provide the technical data needed to protect customer SLAs.
AI-Driven Operations: We operate with an AI-first mindset, powering our operations with the technology we host. You will work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet.
Incident Support: Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions.
Partner Technical Support: Support the growth of our infrastructure partners
Professional Background: 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
What they're looking for
- Datacenter Networking: Strong proficiency in standard datacenter networking and performance troubleshooting. Exposure to RDMA, InfiniBand, or RoCE is highly preferred.
- GPU & AI Stack: Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning.
- Systems & Diagnostics: Solid Linux system administration skills and experience with containerization (Docker). You are comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
- Effective Communication: Clear written and verbal communication skills. You can explain hardware or networking issues to both technical partners and internal leadership.