Postdoctoral Scholar, AI Evaluation & Standards
At J&J we are developing Generative AI solutions to support pharmaceutical R&D, including literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support. These systems need to be tested before teams use them in scientific workflows. In pharmaceutical R&D, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error. We are looking for a postdoctoral researcher to help design methods that test whether GenAI tools produce answers that are accurate, evidence-grounded, traceable, usable, and appropriate for the intended task. The role reports to the Associate Director, Generative AI Evaluation & Quality Standards. The team defines how J&J Innovative Medicine evaluates GenAI systems before use and helps determine when they are ready for release, expansion, or improvement.
All Job Posting Locations: Barcelona, Spain; Beerse, Antwerp, Belgium; Madrid, Spain; Raritan, New Jersey, United States of America; Titusville, New Jersey, United States of America
Responsibilities
- Design evaluation frameworks, rubrics, and criteria for GenAI tools used across pharmaceutical R&D.
- Develop therapeutic-area-specific criteria with business and scientific teams to reflect domain and use-case quality needs.
- Build benchmark datasets, reference answer sets, annotation guides, and evaluation datasets.
- Run expert reviews with scientific, clinical, regulatory, medical, data science, and engineering teams.
- Test LLM, RAG, and agent performance, including accuracy, source grounding, retrieval quality, citation fidelity, task completion, robustness, safety, and usability.
- Analyze failure patterns such as unsupported claims, incorrect reasoning, poor evidence use, missing uncertainty, weak traceability, or failure to follow instructions.
- Translate evaluation findings into improvements in prompts, retrieval methods, agent workflows, tools, and user experience.
- Help define release criteria for systems moving from prototype to limited release, expanded use, or product support.
- Review emerging evaluation methods and adapt useful approaches for pharmaceutical R&D.
- Document methods, findings, and recommendations so teams can apply consistent evaluation practices.
- Design and develop agentic judge methods to evaluate GenAI outputs against defined criteria, flag evidence gaps or unsupported claims, and support expert review workflows.
Requirements
- PhD or equivalent research experience in biomedical science, computational biology, bioinformatics, AI/ML, data science, clinical research, regulatory science, biostatistics, pharmaceutical sciences, or a related field.
- Understanding of biomedical science, pharmaceutical R&D, therapeutic area science, translational science, clinical development, regulatory science, biomedical informatics, data science, or related areas.
- Experience translating expert judgment into criteria, rubrics, datasets, protocols, or measurable outcomes.
- Experience designing or applying evaluation methods, benchmark datasets, annotation protocols, validation studies, quality reviews, or assessment frameworks.
- Interest in testing GenAI systems, including LLMs, RAG, and AI agents.
- Proficiency in Python and common data science or machine learning tools.
- Ability to analyze model outputs, compare performance, identify failure patterns, and recommend improvements.
- Clear written and verbal communication skills.
Qualifications
- Experience with LLM APIs, embeddings, vector databases, prompt engineering, agent frameworks, or AI evaluation tools.
- Experience evaluating retrieval quality, generated answers, multi-step workflows, tool use, scientific reasoning, citation quality, or evidence-grounded outputs.
- Experience designing expert review workflows, annotation instructions, adjudication processes, or inter-rater reliability analyses.
- Domain knowledge in one or more biomedical or therapeutic areas.
- Familiarity with biomedical data standards, structured scientific or clinical data, ontologies, knowledge graphs, CDISC, FHIR, or related frameworks.
- Publications or applied research in AI evaluation, NLP, biomedical informatics, machine learning, data science, computational biology, bioinformatics, or a related field.
Skills
- Required Skills: GenAI evaluation, Python, data science, benchmarking, rubric design, biomedical research, technical communication.
- Preferred Skills: RAG evaluation, agent evaluation, biomedical informatics, expert review, annotation protocols, therapeutic area expertise, responsible AI.
Pay
The anticipated base pay range for this position is €43,600.00 - €70,150.00.
Benefits
- Annual bonus with set target (% of pay) depending on pay grade/location, based on employee and company performance.
- Sales commissions (where applicable).
- Vacation days.
- Parental leave for a minimum of 12 weeks.
- Bereavement leave.
- Caregiver leave.
- Volunteer leave.
- Well-being reimbursement.
- Programs for financial, physical, and mental health.
- Service anniversary and recognition awards.
- Insurance plans for employees and, in some locations, eligible dependents (subject to plan terms).
For more information, visit Employee benefits | Supporting well-being & career growth | Johnson & Johnson Careers. This is for informative purposes only. Amounts and actual benefits may vary by location and are subject to change.