Member of Technical Staff (Answer Quality & Evals)Active$200K–$350K
The opportunity
Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and specialized data sources. The Answer Quality team ensures that our prompts, tools, search systems, datasets, and models work together to create the best possible experience for our users.
What you'll do
Build shared evaluation infrastructure that helps teams run reliable evals,: analyze results, and make product and model decisions
Develop the platform for replaying and analyzing agent traces to reproduce: production behavior and diagnose failures
Build and operate scalable systems for processing, storing, and monitoring: interaction, trace, and evaluation data
Partner with data scientists, engineers, and product teams to turn: answer-quality problems into evaluations, analyses, and product improvements
Operate in a small, high-impact team where your work directly shapes how: Perplexity measures and improves Answer Quality
What they're looking for
- + years of software, data, or machine learning engineering experience: shipping and operating production systems
- Strong proficiency in Python and SQL, with solid fundamentals in system: design, data modeling, and distributed systems
- Experience building big-data systems, including distributed compute,: large-scale storage, and high-volume pipelines
- Demonstrated ownership of ambiguous technical projects from initial design through production operation
- Ability to work effectively with data scientists, engineers, and product partners
- Experience building evaluation, experimentation, observability, or machine learning infrastructure
- Familiarity with LLM and agent systems, including tool use, execution traces, replay, and simulation
- Experience building on top of large-scale data processing platforms such as: Databricks, Snowflake, or ClickHouse