Staff Applied AI ScientistActive
The opportunity
We’re big believers in the power of IRL, so for most roles we ask Campers to work from their local Culture Amp office an average of 2 days a week to unlock connection, pace and culture together.
What you'll do
Own the end-to-end feedback loop: establish a rigorous cycle of prompt engineering, evaluation at scale, and continuous improvement. You will build LLM-powered analysis tools that diagnose performance shifts, provide deep-dive insights, and automate recommendations for prompt or system-level enhancements.
Contribute to Context engineering: design and optimise what actually enters the model's context: retrieval, memory across sessions, context assembly and compression, and managing context budget in long or multi-turn agentic flows. Validate each change against eval rather than opinion or adhoc testing.
Design and run evals: sampling, LLM-as-a-judge, and labelling systems over de-identified production traces (for example, with Langfuse) to build longitudinal evaluation monitoring and alerting.
Eval-driven agentic orchestration: contribute to the agent architecture (planning, tool use, routing, decomposition, verification/critique steps) and let eval findings drive structural changes — e.g. when a failure mode surfaces, add a self-check step, change tool selection, or re-route.
Model and provider selection: make and own model/routing decisions against quality, latency and cost trade-offs, including when to prompt vs fine-tune vs swap models.
Create and Monitor guardrails and safety in production: given sensitive coaching and people data, design input/output guardrails, PII handling, content-safety and jailbreak resistance as part of the system.
What they're looking for
- Enable others: through reusable frameworks, tooling and documentation so product and engineering teams run their own evaluations. Lead from the front, then hand over.
- Partner closely: with the AI Coach team, product, data science and people science so measured quality maps to real customer value.
- Stay current: with the latest evaluation, observability and LLMOps research and provider offerings.
- Experience building and turning production agentic systems, including context: engineering, RAG, memory, cost, model selection and performance.