Postdoctoral Scholar, AI Evaluation & Standards
At J&J we are developing Generative AI solutions to support pharmaceutical R&D, including literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support. These systems need to be tested before teams use them in scientific workflows. In pharmaceutical R&D, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error. We are looking for a postdoctoral researcher to help design methods that test whether GenAI tools produce answers that are accurate, evidence-grounded, traceable, usable, and appropriate for the intended task. The role reports to the Associate Director, Generative AI Evaluation & Quality Standards. The team defines how J&J Innovative Medicine evaluates GenAI systems before use and helps determine when they are ready for release, expansion, or improvement.
Responsibilities
- Design evaluation frameworks, rubrics, and criteria for GenAI tools used across pharmaceutical R&D.
- Develop therapeutic-area-specific criteria with business and scientific teams to reflect domain and use-case quality needs.
- Build benchmark datasets, reference answer sets, annotation guides, and evaluation datasets.
- Run expert reviews with scientific, clinical, regulatory, medical, data science, and engineering teams.
- Test LLM, RAG, and agent performance, including accuracy, source grounding, retrieval quality, citation fidelity, task completion, robustness, safety, and usability.
- Analyze failure patterns such as unsupported claims, incorrect reasoning, poor evidence use, missing uncertainty, weak traceability, or failure to follow instructions.
- Translate evaluation findings into improvements in prompts, retrieval methods, agent workflows, tools, and user experience.
- Help define release criteria for systems moving from prototype to limited release, expanded use, or product support.
- Review emerging evaluation methods and adapt useful approaches for pharmaceutical R&D.
- Document methods, findings, and recommendations so teams can apply consistent evaluation practices.
- Design and develop agentic judge methods to evaluate GenAI outputs against defined criteria, flag evidence gaps or unsupported claims, and support expert review workflows.
Requirements
- PhD or equivalent research experience in biomedical science, computational biology, bioinformatics, AI/ML, data science, clinical research, regulatory science, biostatistics, pharmaceutical sciences, or a related field.
- Understanding of biomedical science, pharmaceutical R&D, therapeutic area science, translational science, clinical development, regulatory science, biomedical informatics, data science, or related areas.
- Experience translating expert judgment into criteria, rubrics, datasets, protocols, or measurable outcomes.
- Experience designing or applying evaluation methods, benchmark datasets, annotation protocols, validation studies, quality reviews, or assessment frameworks.
- Interest in testing GenAI systems, including LLMs, RAG, and AI agents.
- Proficiency in Python and common data science or machine learning tools.
- Ability to analyze model outputs, compare performance, identify failure patterns, and recommend improvements.
- Clear written and verbal communication skills.
Skills
- Required Skills: GenAI evaluation, Python, data science, benchmarking, rubric design, biomedical research, technical communication.
- Preferred Skills: RAG evaluation, agent evaluation, biomedical informatics, expert review, annotation protocols, therapeutic area expertise, responsible AI.
Qualifications
- Experience with LLM APIs, embeddings, vector databases, prompt engineering, agent frameworks, or AI evaluation tools.
- Experience evaluating retrieval quality, generated answers, multi-step workflows, tool use, scientific reasoning, citation quality, or evidence-grounded outputs.
- Experience designing expert review workflows, annotation instructions, adjudication processes, or inter-rater reliability analyses.
- Domain knowledge in one or more biomedical or therapeutic areas.
- Familiarity with biomedical data standards, structured scientific or clinical data, ontologies, knowledge graphs, CDISC, FHIR, or related frameworks.
- Publications or applied research in AI evaluation, NLP, biomedical informatics, machine learning, data science, computational biology, bioinformatics, or a related field.
Pay
The anticipated base pay range for this position is €43,600.00 - €70,150.00.
Benefits
- Annual bonus with set target (% of pay) depending on pay grade/location, based on employee and company performance.
- Sales commissions (where applicable).
- Vacation days.
- Parental leave for a minimum of 12 weeks.
- Bereavement leave, caregiver leave, and volunteer leave.
- Well-being reimbursement.
- Programs for financial, physical, and mental health.
- Service anniversary and recognition awards.
- Insurance plans for employees and, in some locations, eligible dependents (subject to plan terms).
This is for informative purposes only. Amounts and actual benefits may vary by location and are subject to change.