The opportunity
AI Platform builds the foundations of Datadog's AI efforts. The org is 70+ people organised in three pillars: training and serving (GPU clusters, distributed training, low-level infrastructure), agents (agent harnesses, memory systems, the internal AI gateway that routes every…
What you'll do
Own the applied science direction for GenSim: set the methodology and the forward-looking technical calls on how simulated environments and post-training data should be built, on a team where that decision-making does not exist yet.
Define, measure and raise the quality of post-training data: basic correctness, representativeness against the real distribution of customer systems and production telemetry, and difficulty — and make those measures something the team can act on release over release.
Close the realism gap. Simulated environments today are too clean and the: injected problems are not yet hard enough; you'll drive the research and the engineering that make them look like real, imperfect production systems.
Build scalable, production-grade systems rather than research scripts. The: output is not just a dataset — it is a system of synthetic environments that must be reliable and invokable inside a training loop.
Determine how this data is best applied, in LLM post-training and in: evaluation, and own the agent and LLM application evaluation approaches for these environments.
Work cross-functionally with the engineers and applied scientists on adjacent: teams — Bits AI SRE, the model training effort, and the wider evaluation and experimentation pillar — so that what you learn moves freely in both directions.
What they're looking for
- You have a PhD, MS or equivalent research experience in a scientific field,: with strong applied mathematics grounding.
- + years of relevant applied science or ML engineering experience, including: setting technical direction for others.
- You have hands-on experience with LLM and agent post-training data: how it is created, managed, and how training-data quality is controlled. This is the requirement that matters most.
- You have real domain expertise in LLMs and agentic applications: not classical ML fine-tuning. Fine-tuning classifiers or traditional models is a different problem from the one this team is solving.