The opportunity
At Datadog, AI agents are becoming first-class consumers of observability, security, and software delivery data — from third-party coding agents like Claude Code, Cursor, and Copilot, to our own Bits SRE, Bits Assistant, and Bits Dev Agent. The Agentic Interfaces team owns the…
What you'll do
Own the evaluation strategy for Datadog's AI agent integrations. Define the: metrics — offline and online, quality and cost, single-turn and trajectory-level — that the team and the broader organization optimize against.
Build the eval datasets, golden traces, and regression harnesses that catch: quality changes before they hit customers, and make those assets reusable by every team contributing tools to the platform.
Drive measurable improvements to retrieval relevance, tool-selection: accuracy, and context efficiency, partnering closely with the AI engineers on the team who build the underlying platform.
Run applied research on the open problems in agent–data interaction: tool selection under large catalogs, multi-turn agent evaluation, grounding and hallucination control on live telemetry, cost/quality tradeoffs at scale.
Partner with the Bits SRE, Bits Assistant, and Bits Dev Agent teams so: first-party agents benefit from the same measurement substrate as third-party integrations, and so learnings move freely in both directions.
Provide technical leadership across the Agentic Interfaces team and the: broader organization through design reviews, working groups, and mentorship, and represent the team externally through talks, blog posts, and contributions to the open agent ecosystem.
What they're looking for
- You have a BS/MS/PhD in a scientific field, or equivalent experience.
- + years of relevant engineering or applied science experience, including time as a technical lead.
- Proven track record of leading ML or GenAI initiatives in a product-driven: environment, from research through production.
- Significant experience with evaluation, experimentation, or measurement of ML systems at scale.