Staff Applied Scientist - Agentic InterfacesNew$354K

The opportunity

At Datadog, AI agents are becoming first-class consumers of observability, security, and software delivery data — from third-party coding agents like Claude Code, Cursor, and Copilot, to our own Bits SRE, Bits Assistant, and Bits Dev Agent. The Agentic Interfaces team owns the…

What you'll do

  • Own the evaluation strategy for Datadog's AI agent integrations. Define the: metrics — offline and online, quality and cost, single-turn and trajectory-level — that the team and the broader organization optimize against.

  • Build the eval datasets, golden traces, and regression harnesses that catch: quality changes before they hit customers, and make those assets reusable by every team contributing tools to the platform.

  • Drive measurable improvements to retrieval relevance, tool-selection: accuracy, and context efficiency, partnering closely with the AI engineers on the team who build the underlying platform.

  • Run applied research on the open problems in agent–data interaction: tool selection under large catalogs, multi-turn agent evaluation, grounding and hallucination control on live telemetry, cost/quality tradeoffs at scale.

  • Partner with the Bits SRE, Bits Assistant, and Bits Dev Agent teams so: first-party agents benefit from the same measurement substrate as third-party integrations, and so learnings move freely in both directions.

  • Provide technical leadership across the Agentic Interfaces team and the: broader organization through design reviews, working groups, and mentorship, and represent the team externally through talks, blog posts, and contributions to the open agent ecosystem.

What they're looking for

  • You have a BS/MS/PhD in a scientific field, or equivalent experience.
  • + years of relevant engineering or applied science experience, including time as a technical lead.
  • Proven track record of leading ML or GenAI initiatives in a product-driven: environment, from research through production.
  • Significant experience with evaluation, experimentation, or measurement of ML systems at scale.