AI Infrastructure Engineer, pAGIActive$266K–$500K

The opportunity

pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling.

What you'll do

  • Build and operate infrastructure for large-scale training and evaluation,: improving reliability, throughput, and resource efficiency.

  • Develop shared inference and grading platforms with automated capacity: management, health monitoring, and visibility into performance.

  • Improve compute scheduling and resource allocation to reduce idle GPU time: and help workloads recover quickly from failures.

  • Diagnose bottlenecks across training, inference, and orchestration, and work: across teams to improve end-to-end performance.

  • Build self-service tools, automated validation, and observability that help: researchers launch experiments, diagnose issues, and compare results with less manual intervention.

  • Are excited about the potential of personal AGI and want to build the infrastructure that enables it.

What they're looking for

  • Have strong software engineering fundamentals and experience building or: operating large-scale distributed systems.
  • Have experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
  • Are highly self-motivated and comfortable taking ownership of open-ended problems.
  • Enjoy debugging across system boundaries and using measurements to guide: improvements in performance and reliability.