Software Engineer - ML InfrastructureActive$220K

The opportunity

Anduril’s core Lattice platform integrates heterogeneous sensors across large fleets of assets to build a single common operating picture, providing mission-critical situational awareness and context to our users. As a Software Engineer on the ML Infrastructure team, you’ll be…

What you'll do

  • Build & operate our core model training & dataset management infrastructure :: Perception engineers rely on reliable, high quality data to continually retrain and improve our detection & tracing systems. You’ll own and operate our core training systems, as well as our dataset management infrastructure.

  • Improve perception engineer-facing tooling & SDKs: Perception engineers across Anduril rely on our core model training and evaluation framework. You’ll partner closely with these engineers to understand their pain-points & feature requests, and iterate with them directly to deliver them.

  • Integrate frontier models for data augmentation, curation, and synthetic data: generation: While this team is not ultimately responsible for delivering perception systems for particular products, you will be responsible for integrating frontier models at all stages of our dataset curation/annotation pipeline, as well as enabling semantic search across our broad data holdings.

  • Enable observability & traceability of the full ML lifecycle: You’ll build and develop systems that enable real-time monitoring of training jobs, tools to store/curate model evaluation reports, and our model registry. This integrates with our Fleet Management systems to continuously deliver model improvements to our edge systems.

  • Eagerly engage with our customers: We partner closely with business line engineering teams to understand their specific data requirements and iteratively deliver solutions that meet their particular needs and constraints.

  • Strong proficiency in modern Python & exposure to Rust/the Rust ↔ Python toolchain (PyO3, Maturin, etc).

What they're looking for

  • Strong proficiency in modern Python & exposure to Rust/the Rust ↔ Python toolchain (PyO3, Maturin, etc).
  • Experience with modern deep learning frameworks such as Tensorflow or: PyTorch, as well as distributed training paradigms
  • Experience leveraging modern MLOps tooling on the orchestration side (Modal,: Flyte, Airflow, Dagster, etc.), as well as common ML observability platforms (Weights & Biases, MLFlow, etc)
  • Familiarity with building and deploying backend services into a Kubernetes environment
  • U.S. Person status is required as this position needs to access export: controlled data. Eligibility to obtain/maintain a US Top Secret clearance is also desirable.