The opportunity
Anthropic's Safeguards team is seeking a Red Team Engineer to help ensure the safety of our deployed AI systems and products. In this role, you'll take an adversarial approach to uncover vulnerabilities across our product ecosystem before they can be exploited by malicious actors.
What you'll do
Conduct comprehensive adversarial testing across Anthropic's product: surfaces, developing creative attack scenarios that combine multiple exploitation techniques
Research and implement novel testing approaches for emerging capabilities,: including agent systems, tool use, and new interaction paradigms
Design and execute "full kill chain" attacks that emulate real-world threat: actors attempting to achieve specific malicious objectives
Build and maintain systematic testing methodologies that evaluate every aspect of our systems
Develop automated testing frameworks to enable continuous assessment at scale
Collaborate with Product, Engineering, and Policy teams to translate findings into concrete improvements
What they're looking for
- Experience with AI/ML security or adversarial machine learning
- Understanding of AI safety considerations beyond traditional security,: including modern guardrails against jailbreaks
- Experience testing API security and rate-limiting systems
- Background in testing business logic vulnerabilities and authorization bypass techniques
- Background in anti-fraud, trust & safety, or abuse prevention systems
- Familiarity with distributed systems and infrastructure security
- Familiarity with abuse detection mechanisms and the ability to engineer novel bypasses
- Adaptability to understand and build engagements around emerging threats outside your direct area of expertise