Research Engineer, InterpretabilityActive$315K

Hybrid · San Francisco, CATechnology

The opportunity

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole.

What they're looking for

  • Pretraining: Training dictionary learning models looks a lot like model: pretraining - creating stable, performant training jobs for massively parameterized models across thousands of chips
  • Inference: Interp runs a customized inference stack. Day-to-day analysis: requires services that allow editing a model's internal activations mid-forward-pass - for example, adding a "steering vector"
  • Performance: Like all LLM work, we push up against the limits of hardware and: software. Rather than squeezing the last 0.1%, we are focused on finding bottlenecks, fixing them and moving ahead given rapidly evolving research and safety mission
  • Build and maintain the specialized inference and training infrastructure that: powers interpretability research - including instrumented forward/backward passes, activation extraction, and steering vector application