The opportunity
As a Safeguards Analyst on the User Well-being team, you will be focused on supporting the design and deployment of mental health guardrails – iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing ones. Interventions range…
What you'll do
Support the design and execution of interventions, defining key metrics, and curating evaluation datasets
Partner with Engineering and Data Science teams to build, tune, and validate: detection models for automated intervention systems, including threshold-setting and precision and recall tradeoffs
Monitor how interventions and detection systems perform over time
Review flagged content to drive enforcement and policy improvements
Support the development of in-product features that connect users to crisis: resources, working with Product, Legal, and external partners on referral pathways and user-facing content
Support the Safeguards Policy Design team by providing detailed feedback on: policy gaps based on real scenarios
What they're looking for
- Subject matter expertise in mental health, whether developed in academia,: clinical practice, crisis intervention, trust & safety, or other related settings
- Experience building or evaluating LLM-based classification systems
- Experience using agentic tools (e.g. Claude Code) to scale analysis or automate recurring work
- Experience working within crisis support