Product Designer, Evals & PromptsActive$305K

The opportunity

Prompts spec out what Claude does. Evals measure whether it did.

What you'll do

  • Write and revise the prompts behind Claude's tools, features, and behaviors: on a product surface; test the surface, turn findings into prompt fixes, ship them, and confirm the prompt users get is the one intended

  • Build the graders that prove a prompt fix and rerun on the next model; turn: designers' hand-run rubrics into automated evals, then read transcripts for what the eval missed

  • Build visual, low-code eval tools designers can use without an engineer: assemble a comparison set from real transcripts, turn a plain-English rubric into a grader, compare prompt variants across models side by side, and read results in the tool rather than a notebook

  • Watch designers use those tools and make them simpler

  • Support model releases: test each surface against the new model, write prompt fixes and migrations, and write prompts for features launching with it, so the surface owner's call has numbers behind it

  • Stand up and scale the eval harness: build the test environment that exercises our 50 to 100 tools with trustworthy settings, keep evals green across models, and call whether a regression is the harness or the model

What they're looking for

  • Has worked inside a model-launch cycle
  • A/B testing experience and the ability to connect offline evals to online outcomes
  • Front-end or notebook-to-app experience, and opinions about what makes an eval result legible at a glance
  • Has turned product rubrics into training signal: graders, human-feedback questions, or preference pairs
  • Cares how Claude behaves for the people using it, not only whether the metric moved