Role: NLP Scientist — Claim Accuracy and ComplianceLocation: South San Francisco, CA (3 days onsite/week)Duration: 6+ monthsFocus: Evidence-Grounded Retrieval, Entailment and Claim Verification Overview: The goal is to build the capability that checks generated claims against approved evidence before they reach a human reviewer. Given a statement and a corpus of approved claims, product labeling, clinical study results and references, the system retrieves the relevant evidence, decomposes compound statements into checkable assertions, tests whether the evidence supports each one, and returns a decision with citations a reviewer can follow. A large part of the value is knowing when to refuse: the system must distinguish a directly supported claim from one supported only with a qualifier, one the evidence contradicts and one where evidence is insufficient — and abstain rather than guess. This capability assists Medical, Legal and Regulatory review; it does not replace that review or approve content.
Must Have: Evidence-Grounded NLP, Hybrid Retrieval & RAG, NLI/Entailment & Claim Verification, Claim Decomposition & Evidence Attribution, Evaluation & Human-in-the-Loop, Python & NLP Frameworks
Minimum Capabilities:- Strong Python production engineering with modern natural language processing frameworks
- Demonstrated work in evidence-grounded NLP: hybrid retrieval, natural-language inference and entailment, claim decomposition and evidence attribution
- Has measured whether answers were genuinely supported by their cited source — not only that a retrieval pipeline returned something
- Experience with scientific, technical or regulatory source material — studies, specifications, publications, labeling or contracts
- Experience building expert-labeled evaluation datasets, including annotation guidelines and inter-annotator agreement
- Reports error rates by direction, not only aggregate accuracy — false approval and false rejection carry very different costs
- Experience with human-in-the-loop design: confidence thresholds, abstention, escalation rules and safe failure behavior
- Experience building traceable systems where a past decision can be reconstructed from its model version, evidence set and reviewer action
- Experience with vector and lexical retrieval, model APIs and production evaluation infrastructure
- Ability to work with legal and regulatory stakeholders and treat process constraints as design requirements
Preferred:- Experience combining deterministic rules with model judgment in a single decision system
- Experience with knowledge graphs linking claims, evidence, references, products and indications
- Experience in any regulated or high-stakes review environment — legal, financial compliance, scientific publishing or fact-checking
- Familiarity with study design, statistical evidence and citation practice
- Pharmaceutical or life-sciences experience is welcome but not required — domain context and review workflow will be provided