Site Reliability Engineer — HPC & Automation (Silicon Engineering)Active$150K

The opportunity

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

What you'll do

  • Deploy, upgrade, operate, maintain, and scale our suite of clusters and services

  • Collaborate with engineers to develop automated, full turnkey solutions for: silicon simulation workflows to speed up project timelines

  • Manage our underlying infrastructure as code and use modern observability: tools to provide a complete picture of cluster and infrastructure health

  • Operate the continuous integration pipeline, build and release systems, and: version control across the environment

  • Identify and eliminate performance bottlenecks using measurement and creative engineering

  • Bachelor’s degree in computer science, information systems, or an engineering: discipline; OR 2+ years of professional experience in system administration, high performance computing, or site reliability engineering

What they're looking for

  • + years of development experience with Bash, Python, and/or other programming languages
  • + years of experience with Linux operating systems
  • Familiarity with containerization technologies (i.e. Docker, Kubernetes)
  • Knowledge in computer system concepts (computer architecture, computer: organization, operating systems and concurrency)