The opportunity
Applied AI is moving faster than ever. New foundation models, reasoning techniques, agent architectures, and research papers emerge every week.
What you'll do
Design and deploy production AI agents that leverage the latest advances in: large language models, reasoning, retrieval, memory, and tool use.
Architect intelligent systems that combine LLMs, traditional machine: learning, structured knowledge, enterprise data, and deterministic software into reliable production workflows.
Engineer customer intelligence layers, retrieval pipelines, memory systems,: and knowledge representations that allow agents to reason over large, heterogeneous enterprise data.
Develop multi-agent systems that coordinate reasoning, planning, tool execution, and human oversight.
Translate frontier AI research into production systems by rapidly evaluating: new models, prompting techniques, reasoning paradigms, and agent architectures.
Own the full experimentation lifecycle, from hypothesis generation to production rollout.
What they're looking for
- Design rigorous evaluation frameworks using offline benchmarks, online A/B: experiments, golden datasets, regression suites, LLM-as-a-Judge, and human evaluation.
- Run controlled experiments and ablation studies to understand the: contribution of different models, prompts, retrieval strategies, reasoning techniques, memory systems, and agent architectures.
- Continuously evaluate newly released frontier models and determine where they: meaningfully improve quality, latency, reliability, or cost.
- Develop confidence estimation, reflection, and continuous learning systems: that improve agents over time using real-world feedback.