Staff Research Scientist - Post-Training for AgentsPosted today$46K–$183K

The opportunity

As a Staff Research Scientist within Datadog AI Research (DAIR), you will drive research in post-training and autonomous agents as a hands-on individual contributor. You will advance reinforcement learning, simulated environments, synthetic data generation, and evaluation…

What you'll do

  • Drive research in agent post-training and reinforcement learning, shaping the: technical direction of ambitious research programs grounded in observability and security

  • Own research problems end to end, from framing the question through: experimentation, model development, and evaluation

  • Design and advance reinforcement learning approaches, post-training methods,: and training loops for agents operating in complex environments

  • Build simulated environments, synthetic data generation approaches, and: evaluation frameworks that enable scalable agent training and rigorous measurement

  • Raise the technical bar across the team by reviewing research directions,: mentoring researchers and research engineers, and setting standards for experimental rigor

  • Collaborate with cross-functional teams across Research, Product, and: Engineering to translate advances in autonomous agents into scalable Datadog capabilities

What they're looking for

  • Contribute to research publications, present at top-tier conferences such as: NeurIPS, ICLR, and ICML, and help open-source key model artifacts and benchmarks
  • You hold a PhD in Computer Science, Machine Learning, or a related field, or: have equivalent experience, with deep expertise in areas such as reinforcement learning, AI agents, post-training, or generative modeling
  • You have driven technically ambitious research at meaningful scale as an: individual contributor, whether in an industry research lab, startup, academic environment, or another research setting
  • You have extensive hands-on experience designing and implementing: reinforcement learning systems, agent training methods, simulated environments, synthetic data pipelines, or evaluation frameworks