The opportunity
Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform.
What you'll do
Own capacity, durability, availability, and performance characteristics of: network volumes, local NVMe, and S3-compatible object storage.
Tune the full I/O path: device and filesystem configuration, caching and read-ahead strategies, replication and erasure coding trade-offs, and client-side mount behavior.
Diagnose hard performance problems end to end
Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption.
Work with Runpod and partner networking teams to design and tune the network: paths storage depends on: high-throughput east-west fabric, MTU and jumbo frames, congestion and flow control, multipath, and NIC/offload configuration.
Understand and optimize RDMA/RoCE and high-speed IB/Ethernet fabrics as they apply to storage traffic.
What they're looking for
- Work closely with network engineering on topology decisions, oversubscription: ratios, and cross-region data movement.
- Write production code (Go, Python, or similar) for storage control-plane: services, provisioning workflows, data movement pipelines, and monitoring
- Build against and extend APIs: our own control plane, S3-compatible interfaces, CSI drivers, Kubernetes APIs, vendor and cloud provider APIs.
- Automate the operations you'd otherwise do by hand. Manual runbooks are a starting point, not a destination.