Senior Data Scientist - Multiomics and Population Cohorts
Michigan Medicine · Ann Arbor, MI · Yesterday
EngineeringFull-time
About the role
The Division of Cardiovascular Medicine at the University of Michigan is expanding a nationally and internationally funded research program at the cutting edge of cardiovascular medicine. Led by Drs. Venkatesh Murthy and Sascha Goonewardena, whose work has appeared in leading journals including NEJM AI, JAMA, and Circulation, the program fuses artificial intelligence, advanced cardiac imaging, multiomics (proteomics, metabolomics, and genomics), and cardiometabolic disease biology to develop precision diagnostics and identify novel therapeutic approaches, with a particular focus on coronary microvascular disease and other cardiovascular conditions that disproportionately affect women and remain poorly served by existing diagnostic tools.
Responsibilities
- Build, implement, and maintain proteomics and multi-omic analysis pipelines integrating high-throughput proteomic data with metabolomic, genomic, and clinical datasets from large prospective cohort studies
- Deploy AI models to biobank datasets
- Build and maintain cross-cohort data harmonization and proteomic cross-walk tools across derivation and validation datasets
- Perform causal and genetic inference analyses (including Mendelian randomization and mediation analysis) to identify disease mechanisms
- Evaluate pipeline and signature performance; prepare written and code-based analytical reports (Python or R)
- Contribute to manuscripts, grant applications, and presentations
- Coordinate data sharing and computational workflows with collaborative network partners and other research partners
Requirements
- Masters or doctoral degree in bioinformatics, computational biology, biostatistics, systems biology, or a closely related field
- Demonstrated experience analyzing large-scale proteomic datasets from high-throughput platforms (Olink or SomaScan)
- Proficiency in R and/or Python for statistical analysis and pipeline development
- Experience with multi-omic data integration combining at least two of: proteomics, metabolomics, genomics, transcriptomics
- Familiarity with statistical methods for high-dimensional biological data (dimensionality reduction, regularized regression, survival analysis)
- Strong organizational skills and attention to detail
- Ability to prepare and present written and code-based (Python or R) analytical reports
- Strong scientific writing skills; ability to contribute to manuscripts and grant applications
Qualifications
- Experience with large prospective cohort or biobank datasets (e.g., CARDIA, MESA, Framingham Heart Study, UK Biobank, or similar)
- Experience with mediation analysis, causal inference, or Mendelian randomization methods in an omics context
- Background in cardiovascular biology, vascular biology, or cardiometabolic disease
- Experience with cloud or HPC computing environments, including Slurm-based job scheduling
- Familiarity with SQL or PostgreSQL for data querying and management
- Track record of peer-reviewed publications as a computational contributor to biomedical research
- Experience with endothelial biology, inflammation, or microvascular disease
- Experience with Git and reproducible research practices
- Experience building reproducible pipelines using workflow managers (Snakemake, Nextflow, or equivalent)