The opportunity
We're looking for an Engineering Manager to lead the Agent Context team in NYC. This team owns how Asana's AI systems search, retrieve, and reason over the work graph the search infrastructure, dense embedding pipelines, ranking systems, and evaluation frameworks that determine…
What you'll do
Own the technical direction and delivery of Asana's retrieval stack end to: end: lexical and semantic search, dense embedding generation and backfill at scale, chunking and ranking strategies, and RAG comprehensiveness across the work graph.
Build and operate the evaluation infrastructure that makes retrieval quality: measurable recall/precision benchmarks, offline and online evals, and comparative testing across retrieval backends - so quality decisions are made with data, not vibes.
Drive the cost, performance, and quality tradeoffs that define this space:: when semantic search earns its infrastructure cost over lexical, how to hit latency targets without sacrificing recall, and how retrieval improvements compound into cheaper, faster downstream LLM calls.
Set and enforce the bar for how other teams at Asana integrate with: retrieval: clear ownership of embedding decisions, rollout guidance, metrics to watch, and a platform posture that says no to unjustified infrastructure spend.
Hire, grow, and retain a team of strong senior engineers in NYC, and lead: effectively across three time zones with deliberate async communication practices.
Partner with your PM counterpart to translate a multi-year platform thesis: into a sequenced roadmap, and represent the team's technical strategy to engineering and product leadership.
What they're looking for
- + years of software engineering experience with 3+ years managing engineers,: including senior engineers, on infrastructure or ML systems teams. You've hired, coached, grown, and when necessary exited engineers and your former reports would work for you again.
- You have shipped and operated production search, retrieval, or ML-serving: systems at meaningful scale. You can speak concretely about systems you've run: the index architecture, the embedding models, the latency budgets, the incidents, and what you'd do differently.
- Deep working knowledge of the modern retrieval stack inverted indexes and: BM25, vector search and embedding models, hybrid retrieval, chunking strategies, re-ranking and strong opinions about when each is worth its cost. You should be able to argue both sides of "semantic search everywhere" and tell us where you actually land.
- You've built or heavily used evaluation systems for ML/AI quality: golden datasets, recall/precision metrics, LLM-as-judge, online experimentation. You believe unmeasured quality claims are noise.