The opportunity
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.
What you'll do
Manage GPU/CPU infrastructure deployments to Top Secret datacenters
Manage and provide support for GPU as a service for external customers on: bare metal hardware and virtualized platforms
Design, validate, and productize solutions for AI clusters (100k+ GPU scale)
Develop automation to deploy and manage on-premise Kubernetes\AI clusters, and operating systems
Deploy and manage core infrastructure such as databases, monitoring and distributed storage
Closely collaborate with AI engineers to create highly scalable, operable, and maintainable products
What they're looking for
- Engage in and improve the whole lifecycle of services -: from inception and design, through deployment, operation and refinement
- Monitoring and alerting supporting systems to have high availability
- Identify areas for improvement and create innovative solutions that enable high system availability
- Bachelor’s degree in computer science, information systems/IT, or an: engineering discipline and 1+ years of professional experience in site reliability engineering or DevOps; OR 3+ years of professional experience in site reliability engineering or DevOps in lieu of a degree