Jobs · Analyst

Postdoctoral Fellow, Applied AI

Ancestry · Lehi, UT · 2 wk ago
RemoteRemoteAnalystTemporary

About the Company

Ancestry® is the global leader in family history, helping millions of people discover, preserve, and share their unique family stories. With over 65 billion records, 3.5 million subscribers, and 27 million people in our DNA network, we connect customers to their past through unparalleled historical and genealogical data. We foster a human-centered, inclusive, and diverse work environment where every idea is valued. Ancestry offers location flexibility, allowing employees to work from an office, home, or hybrid (subject to role and location restrictions).

About the Role

Ancestry is seeking a highly motivated Postdoctoral Research Fellow to join the AI Applied Science Content team. This fellowship is designed for a researcher at the intersection of industry-scale data and academic rigor. You will lead research in Document Understanding, designing and implementing AI-Native agentic systems to unlock insights from billions of historical and genealogical records. Your work will advance autonomous multi-agent workflows and transform unstructured historical records into structured, searchable knowledge.

Responsibilities

  • Innovate with state-of-the-art AI for Document Understanding tasks, including OCR/HTR, transcription, Named Entity Recognition (NER), Relation Extraction (RE), Coreference Resolution, Summarization, and Knowledge Graphs. Work with diverse genealogical collections such as newspapers, city directories, family history books, and vital records (birth, marriage, death).
  • Architect agentic systems using frameworks like LangChain, LangGraph, CrewAI, AutoGen, AgentCore, Strands, Google ADK, and A2A to automate complex multi-step reasoning tasks in historical document analysis.
  • Analyze and optimize multi-modal models (e.g., GPT, Gemini, Claude, Llama, Qwen) for zero-shot and few-shot scenarios in document understanding.
  • Apply expertise in Natural Language Processing (NLP) for NER, Relation Extraction, Coreference Resolution, Entity Resolution, and Knowledge Graphs (Neo4j) using tools like spaCy, NLTK, and BERT.
  • Utilize Computer Vision (CV) models (e.g., YOLO, Nougat, DONUT, OpenCV) for layout analysis, identifying text blocks, headers, tables, and nested lists.
  • Establish ensemble models and "LLM-as-a-Judge" frameworks using tools like Arize Phoenix, DeepEval, or RAGAS to monitor hallucination, drift, and bias.
  • Leverage AI coding assistants (e.g., Amazon Q, Cursor, Claude Code, Kiro) to accelerate development cycles.
  • Collaborate with ML Ops to deploy datasets, models, and pipelines in cloud environments like AWS (S3, SageMaker, Bedrock, ECS, EKS) and GCP (Vertex AI, Gemini API).

Requirements

  • Ph.D. (recent graduate or near completion) in Computer Science, Data Science, Statistics, Linguistics, Engineering, or a related quantitative field with a strong research focus.
  • Strong record of academic publications in NLP, CV, or Agentic AI (preferred).
  • Specialization in AI & LLMs, including familiarity with foundational models such as GPT, Gemini, Qwen, Llama, and Claude.
  • Research experience in inference efficiency and optimization (e.g., vLLM, LoRA, QLoRA, quantization).
  • Familiarity with embeddings, vector databases, and transformer models.
  • Strong proficiency in Python and relevant libraries for transformer and multi-modal models.
  • Familiarity with cloud platforms (AWS, GCP) and AI/ML services (e.g., Google Vertex AI, AWS SageMaker, Bedrock) is a plus.
  • Ability to present complex technical solutions to both technical and non-technical stakeholders.

Similar jobs