Technical Site Reliability EngineerPosted today

Abu Dhabi, United Arab Emirates; London, England, United KingdomManufacturing

The opportunity

Anduril Industries is a defense technology company with a mission to transform U. S.

What you'll do

  • Maintain the simulation software stack: installation, configuration, updates, version management, and day-to-day functionality across the Simulation Center's tools and environments.

  • Own the underlying infrastructure: compute, networking, storage, and environment configuration that the simulation depends on; keep it provisioned, patched, and performant.

  • Build and maintain a post-release test suite: design, automate, and continually extend a regression and smoke-test process that runs after every software release or configuration change, so integration issues surface immediately rather than mid-exercise.

  • Forecast, diagnose, and eliminate failure modes: root-cause errors and bugs in the system, drive them to permanent resolution, and implement the guardrails, monitoring, or process changes that prevent recurrence.

  • Partner with development teams and stakeholders: review upcoming changes for reliability risk, surface concerns early, and implement mitigation strategies before releases land in the simulation environment.

  • Monitor overall system health: instrument and watch the environment, triage issues within your scope, and escalate clearly and quickly with the context needed for others to act when an issue exceeds your ability to resolve it.

What they're looking for

  • Document what you learn: runbooks, known issues, environment configuration, and release validation results, so the Simulation Center's operational knowledge isn't held in one person's head.
  • Proficiency in Python for automation, tooling, and test development.
  • Working knowledge of C++; enough to read, debug, build, and trace issues in the simulation codebase.
  • Solid general networking fundamentals: TCP/IP, UDP, multicast, DNS, routing, firewalls, and the ability to diagnose latency, packet loss, and connectivity problems across distributed systems.