Member of Technical Staff (Search Crawler Analyst)Posted today
The opportunity
The internet is vast, containing trillions of URLs. Perplexity’s crawling and storage system is complex and has multiple stages (URL discovery, crawling, parsing, indexing).
What you'll do
Find and diagnose quality issues in our crawling pipeline
Train small models that optimize particular aspects of the pipeline (e.g. parsing quality)
Build datasets for model training, including LLM-as-a-judge labeling pipelines
Improve page selection algorithms for indexing
Design and analyze experiments to validate improvements
+ years of experience as a data analyst, ML engineer, or in a related role
What they're looking for
- Strong coding skills: you should be able to write production-grade code at the level of a mid-level backend engineer
- Experience designing metrics from scratch
- Experience training ML models that shipped to production with measurable metric improvements
- Direct experience working on web crawling or indexing pipelines