Jobs · New Jersey

Postdoctoral Scholar, AI Evaluation & Standards

Johnson & Johnson Innovative Medicine · Titusville, NJ · Yesterday
HybridFull-time

At J&J we are developing Generative AI solutions to support pharmaceutical R&D, including literature review, evidence synthesis, document Q&A, therapeutic area knowledge search, translational science workflows, and R&D decision support. These systems need to be tested before teams use them in scientific workflows. In pharmaceutical R&D, a useful AI response depends on the question, user, source material, therapeutic area, and risk of error. We are looking for a postdoctoral researcher to help design methods that test whether GenAI tools produce answers that are accurate, evidence-grounded, traceable, usable, and appropriate for the intended task. The role reports to the Associate Director, Generative AI Evaluation & Quality Standards. The team defines how J&J Innovative Medicine evaluates GenAI systems before use and helps determine when they are ready for release, expansion, or improvement.

All Job Posting Locations: Barcelona, Spain; Beerse, Antwerp, Belgium; Madrid, Spain; Raritan, New Jersey, United States of America; Titusville, New Jersey, United States of America

Responsibilities

  • Design evaluation frameworks, rubrics, and criteria for GenAI tools used across pharmaceutical R&D.
  • Develop therapeutic-area-specific criteria with business and scientific teams to reflect domain and use-case quality needs.
  • Build benchmark datasets, reference answer sets, annotation guides, and evaluation datasets.
  • Run expert reviews with scientific, clinical, regulatory, medical, data science, and engineering teams.
  • Test LLM, RAG, and agent performance, including accuracy, source grounding, retrieval quality, citation fidelity, task completion, robustness, safety, and usability.
  • Analyze failure patterns such as unsupported claims, incorrect reasoning, poor evidence use, missing uncertainty, weak traceability, or failure to follow instructions.
  • Translate evaluation findings into improvements in prompts, retrieval methods, agent workflows, tools, and user experience.
  • Help define release criteria for systems moving from prototype to limited release, expanded use, or product support.
  • Review emerging evaluation methods and adapt useful approaches for pharmaceutical R&D.
  • Document methods, findings, and recommendations so teams can apply consistent evaluation practices.
  • Design and develop agentic judge methods to evaluate GenAI outputs against defined criteria, flag evidence gaps or unsupported claims, and support expert review workflows.

Requirements

  • PhD or equivalent research experience in biomedical science, computational biology, bioinformatics, AI/ML, data science, clinical research, regulatory science, biostatistics, pharmaceutical sciences, or a related field.
  • Understanding of biomedical science, pharmaceutical R&D, therapeutic area science, translational science, clinical development, regulatory science, biomedical informatics, data science, or related areas.
  • Experience translating expert judgment into criteria, rubrics, datasets, protocols, or measurable outcomes.
  • Experience designing or applying evaluation methods, benchmark datasets, annotation protocols, validation studies, quality reviews, or assessment frameworks.
  • Interest in testing GenAI systems, including LLMs, RAG, and AI agents.
  • Proficiency in Python and common data science or machine learning tools.
  • Ability to analyze model outputs, compare performance, identify failure patterns, and recommend improvements.
  • Clear written and verbal communication skills.

Qualifications

  • Experience with LLM APIs, embeddings, vector databases, prompt engineering, agent frameworks, or AI evaluation tools.
  • Experience evaluating retrieval quality, generated answers, multi-step workflows, tool use, scientific reasoning, citation quality, or evidence-grounded outputs.
  • Experience designing expert review workflows, annotation instructions, adjudication processes, or inter-rater reliability analyses.
  • Domain knowledge in one or more biomedical or therapeutic areas.
  • Familiarity with biomedical data standards, structured scientific or clinical data, ontologies, knowledge graphs, CDISC, FHIR, or related frameworks.
  • Publications or applied research in AI evaluation, NLP, biomedical informatics, machine learning, data science, computational biology, bioinformatics, or a related field.

Skills

  • Required Skills: GenAI evaluation, Python, data science, benchmarking, rubric design, biomedical research, technical communication.
  • Preferred Skills: RAG evaluation, agent evaluation, biomedical informatics, expert review, annotation protocols, therapeutic area expertise, responsible AI.

Pay

The anticipated base pay range for this position is €43,600.00 - €70,150.00.

Benefits

  • Annual bonus with set target (% of pay) depending on pay grade/location, based on employee and company performance.
  • Sales commissions (where applicable).
  • Vacation days.
  • Parental leave for a minimum of 12 weeks.
  • Bereavement leave.
  • Caregiver leave.
  • Volunteer leave.
  • Well-being reimbursement.
  • Programs for financial, physical, and mental health.
  • Service anniversary and recognition awards.
  • Insurance plans for employees and, in some locations, eligible dependents (subject to plan terms).

For more information, visit Employee benefits | Supporting well-being & career growth | Johnson & Johnson Careers. This is for informative purposes only. Amounts and actual benefits may vary by location and are subject to change.

Similar jobs