The opportunity
You will lead a team of engineers building the infrastructure used to train, test, and evaluate our models. This is one of the few places at SpaceXAI where infrastructure and model behavior meet directly: when something breaks, it's rarely obvious whether it's a systems bug or…
What you'll do
Building the rollout infrastructure that lets researchers run RL experiments: at scale without fighting the plumbing.
Designing eval pipelines that catch regressions before they ship, and give: researchers fast, trustworthy signal on whether a change actually helped.
Owning the environments in which models are trained and tested: sandboxed, reproducible, and fast enough that iteration speed isn't the bottleneck.
Bringing rigor to how the team measures quality and progress, in places where: "did it ship" isn't the same as "did it work?"
Partnering with research to translate model-level tradeoffs (latency,: quality, cost) into concrete infrastructure decisions.
Hiring and growing the team: sourcing, interviewing, and closing exceptional infrastructure engineers, while developing your engineers through coaching, mentorship, and high-leverage project assignments.
What they're looking for
- You've led engineering teams building infrastructure that trains, evaluates,: or serves ML models in production.
- You have strong infrastructure and distributed systems fundamentals: you know what reliability and performance look like under real load, not just in a design doc.
- You genuinely want to stay technical: you're comfortable writing code, reviewing PRs with depth, and using tools like Cursor itself to move fast.
- You’re comfortable operating in ambiguity: you ask the right questions, make sound decisions with incomplete information, and help the team find a path forward.