| Location | San Francisco, CA |
Perplexity serves tens of millions of users daily with reliable, high-quality answers grounded in an LLM-first search engine and our specialized data sources. We aim to use the latest models as they are released, but the intelligence frontier is a jagged one, and popular benchmarks do not effectively cover our use cases.
In this role, you will build specialized evals to improve answer quality across Perplexity, covering search-based LLM answers and other scenarios popular with our users.
Responsibilities ----------------
Qualifications --------------
PhD or MS in a technical field or equivalent experience 4+ years of experience in data science or machine learning Strong proficiency in Python and SQL (expected to write production-grade code) Experience building within a modern cloud data stack, specifically AWS and Databricks Comfortable with agentic coding workflows and using AI-assisted development tools to iterate faster
Preferred Qualifications ----------------------
1+ years of experience working with LLMs at scale, specifically with LLM-as-a-judge setups Prior experience working on customer-facing web products or consumer apps, with real user traffic at scale A strong research background, with experience applying research methods to real-world ML problems Experience defining evaluation metrics (e.g., factual consistency, hallucination rate, retrieval precision) and building ground truth datasets