Site Reliability Engineer (Manufacturing Infrastructure)Active

The opportunity

SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.

What you'll do

  • Deploy, upgrade, operate, maintain, and scale compute, storage, and: networking for manufacturing systems across Starship, Starlink, Starshield, and Terafab

  • Manage infrastructure as code and use observability to provide a complete picture of platform health

  • Design for reliability, stability, and scale; find and remove bottlenecks with measurement and engineering

  • Practice proactive maintenance: capacity planning, lifecycle management, and reducing toil before it becomes an incident

  • Partner with software engineers, manufacturing stakeholders, and site teams: to build operable, maintainable systems

  • Improve the full lifecycle—from design through deployment, operation, and continuous refinement

What they're looking for

  • Practice sustainable incident response and blameless postmortems
  • Provide high-quality support to manufacturing and engineering users
  • Communicate clearly with stakeholders and teammates
  • Participate in on-call and travel to sites as needed for deployments, incidents, and cross-site reliability