NLP Scientist — Claim Accuracy and Compliance
Location: South San Francisco, CA (3 days onsite/week) | Duration: 6+ months
About the role
The goal is to build the capability that checks generated claims against approved evidence before they reach a human reviewer. Given a statement and a corpus of approved claims, product labeling, clinical study results and references, the system retrieves the relevant evidence, decomposes compound statements into checkable assertions, tests whether the evidence supports each one, and returns a decision with citations a reviewer can follow. A large part of the value is knowing when to refuse: the system must distinguish a directly supported claim from one supported only with a qualifier, one the evidence contradicts and one where evidence is insufficient — and abstain rather than guess. This capability assists Medical, Legal and Regulatory review; it does not replace that review or approve content.
Responsibilities
- Build evidence-grounded NLP systems for claim verification using hybrid retrieval, natural-language inference, and entailment
- Decompose compound statements into checkable assertions and attribute evidence to support or refute claims
- Design and implement evaluation frameworks, including expert-labeled datasets, annotation guidelines, and inter-annotator agreement metrics
- Measure error rates by direction (false approval vs. false rejection) to account for differing costs
- Develop human-in-the-loop systems with confidence thresholds, abstention rules, and safe failure behavior
- Ensure traceability of decisions by maintaining records of model versions, evidence sets, and reviewer actions
- Collaborate with legal and regulatory stakeholders, treating process constraints as design requirements
Requirements
- Strong Python production engineering with modern natural language processing frameworks
- Demonstrated experience in evidence-grounded NLP, including hybrid retrieval, natural-language inference, entailment, claim decomposition, and evidence attribution
- Experience measuring whether answers are genuinely supported by cited sources, not just retrieval performance
- Familiarity with scientific, technical, or regulatory source material (e.g., studies, specifications, publications, labeling, or contracts)
- Experience with vector and lexical retrieval, model APIs, and production evaluation infrastructure
Preferred Qualifications
- Experience combining deterministic rules with model judgment in a single decision system
- Experience with knowledge graphs linking claims, evidence, references, products, and indications
- Background in regulated or high-stakes review environments (e.g., legal, financial compliance, scientific publishing, or fact-checking)
- Familiarity with study design, statistical evidence, and citation practices
- Pharmaceutical or life-sciences experience (welcome but not required; domain context and workflow will be provided)