Jobs · Information Technology · California

Associate Director, Principal Data Engineer

Vir Biotechnology, Inc. · San Francisco, CA · 1 wk ago
On-siteInformation Technology$177k–$247k/yrFull-time

Vir Biotechnology, Inc. is a clinical-stage biopharmaceutical company focused on powering the immune system to transform lives by discovering and developing medicines for serious infectious diseases and cancer. Its clinical-stage portfolio includes programs for chronic hepatitis delta and multiple PRO-XTEN® dual-masked T-cell engagers across validated targets in solid tumor indications. Vir Biotechnology also has a preclinical portfolio of programs across a range of infectious diseases and oncologic malignancies.

About the role

Vir Biotechnology is seeking an Associate Director, Principal Data Engineer to lead scalable data, cloud, software, and AI-enabled research platforms that support scientific discovery, clinical data analysis, and program decision-making. This role will build the foundation for AI-powered research workflows, including AI agents for research, QC/QA, and scientific/regulatory reporting, while partnering across Research, Clinical, Data Science, TechOps, Competitive Intelligence, and IT.

The ideal candidate is a forward-thinking technical leader who combines expertise in data engineering, scientific software, AWS cloud platforms, and practical AI implementation to transform complex scientific data into connected, searchable, and actionable insights. They excel at partnering with scientists, data scientists, and engineers to build trusted, scalable solutions that accelerate research while maintaining the highest standards of data quality, governance, validation, and reliability.

Reporting to the Senior Director, Research, this role is based in San Francisco with an expectation of three days per week in the office.

Responsibilities

  • Build AI-enabled research workflows and agents that help scientists search, analyze, summarize, and connect Vir-generated data.
  • Develop AI agents for TCE research insights, experimental QC/QA review, and scientific/regulatory reporting.
  • Design trusted LLM, RAG, knowledge-search, and agentic AI workflows with traceability, validation, governance, and human oversight.
  • Lead development and integration of scientific platforms, including SeqAssembler, HPD/dAIsY, SASTRY, OPAL Miner, and others.
  • Build scalable research and clinical data pipelines for bioinformatics, genomics, machine learning, and scientific decision-making.
  • Architect and operate AWS cloud infrastructure for AI/ML, bioinformatics, clinical analysis, and large-scale scientific data.
  • Strengthen data quality, metadata, lineage, observability, security, governance, and platform reliability.
  • Partner cross-functionally to align data, cloud, application, and AI solutions with program needs.
  • Provide senior technical leadership across architecture, engineering practices, mentoring, and roadmap development.

Requirements

  • BS/MS in Computer Science, Data Engineering, Bioinformatics, Computational Biology, Engineering, or a related technical field.
  • 12+ years of experience in data engineering, software engineering, scientific computing, cloud architecture, or scientific application development.
  • Strong hands-on experience with Python and/or Java, including scalable data pipelines, APIs, services, and production-grade data systems.
  • Experience building AI/ML-enabled scientific applications using leading LLMs such as GPT, Claude, Gemini, or similar models, including RAG, knowledge search, agentic workflows, and governance.
  • Experience integrating LIMS, ELN, scientific data platforms, or experiment-management systems such as Experimenta to support data capture, analysis, quality review, and scientific reporting.
  • Strong AWS cloud infrastructure experience, including compute, storage, networking, security, data processing, monitoring, automation, and operational best practices.
  • Experience with modern cloud-based data platforms and workflow technologies such as Databricks, Snowflake, Redshift, Spark, Airflow, Nextflow, or similar.
  • Experience supporting bioinformatics, genomics, machine learning, clinical, healthcare, or other scientific data domains.
  • Strong knowledge of data modeling, metadata management, data quality, observability, governance, lineage, and platform engineering.
  • Demonstrated technical leadership, including architecture ownership, roadmap development, mentoring, and cross-functional collaboration.
  • Excellent communication and collaboration skills across Research, Clinical, Data Science, IT, and business teams.

Benefits

  • Competitive compensation package including base salary, bonus, and equity.
  • Health and welfare benefit plans.
  • Non-accrual paid time off and company shutdown for holidays.
  • Commuter benefits and 401K match.
  • Lunch provided daily in the office.

Pay

The expected salary range for this position is $176,500 to $246,500 per year. Actual pay will be determined based on experience, qualifications, geographic location, and other job-related factors.

Schedule

This role is based in San Francisco with an expectation of three days per week in the office.

Similar jobs