The opportunity
Anthropic's Safeguards team builds the systems that detect and mitigate misuse of our AI models, from individual policy violations to sophisticated, coordinated attacks. A growing part of that work depends on lightweight detection methods trained on model internals, which let us…
What you'll do
Build and scale the infrastructure and data pipelines behind Safeguards machine learning research
Own the training, evaluation, and scoring workflows researchers use, with a: focus on cutting the time between an idea and a result
Design tooling and interfaces, including libraries and command line tools,: that researchers can use directly without needing to understand the systems underneath
Build correctness and sanity checking into the stack, so results stay: trustworthy as models and workloads evolve
Take the highest-value research workflows from experiments to reliable, production-grade jobs
Improve the throughput, cost, and reliability of large-scale inference and scoring workloads
What they're looking for
- Experience with high-performance, large-scale machine learning systems
- Familiarity with language modeling and transformers, including working with model internals
- Experience with machine learning framework internals, GPU or accelerator: programming, or inference optimization
- Experience building experiment tracking, caching layers, or evaluation harnesses for research teams
- Experience with probes, interpretability, or classifier development
- Interest in the misuse risks of AI systems and a desire to work on mitigating them