Machine Learning Infrastructure EngineerActive$150K–$350K
The opportunity
We’re looking for seasoned ML Infrastructure engineers with experience designing, building and maintaining training and serving infrastructure for ML research.
What you'll do
Provide infrastructure support to our ML research and product
Build tooling to diagnose cluster issues and hardware failures
Monitor deployments, manage experiments, and generally support our research
Maximize GPU allocation and utilization for both serving and training
+ years of experience supporting the infrastructure within an ML environment
Experience in developing tools used to diagnose ML infrastructure problems and failures
What they're looking for
- Experience with cloud platforms (e.g., Compute Engine, Kubernetes, Cloud Storage)
- Experience working with GPUs
- Experience with large GPU clusters and high-performance computing/networking
- Experience with supporting large language model training