Jobs · Engineering · California

Principal Data Scientist

Genoa Ventures · South San Francisco, CA · 2 days ago
Engineering$196k/yrFull-time

Responsibilities

  • Design, prototype, and rigorously evaluate novel classifier architectures for clinical diagnostics across oncology indications
  • Lead exploratory research into new quantification, normalization, and feature engineering methods for high-dimensional glycoproteomic data
  • Bring a diverse modeling toolkit — classical statistical methods, tree-based ensembles, deep learning, probabilistic and Bayesian approaches, foundation models, graph neural networks, and generative AI — and choose the right tool for the problem based on evidence rather than habit or hype
  • Develop cross-validation, calibration, and uncertainty-quantification strategies that hold up to the realities of small clinical cohorts and high feature counts
  • Investigate and mitigate batch, cohort, and site effects so that models generalize from discovery to bridging to locked panels
  • Drive cross-indication synthesis — separate shared disease biology from indication-conditioned signal, and from nonspecific inflammatory or acute-phase axes
  • Build multimodal models that combine glycan/motif information, proteomic grounding, and clinical covariates rather than relying on protein-quantity signal alone
  • Translate emerging techniques from the ML, AI, and computational-biology literature into production-ready methods
  • Mentor junior data scientists and raise the methodological bar across the team

Qualifications

  • Ph.D. in Statistics, Computer Science, Computational Biology, Bioinformatics, or a related quantitative field, plus 6+ years of experience building predictive models on biological data in industry or academia; alternatively, an MS in a similar field with 8+ years of relevant experience
  • Demonstrated track record of methodological innovation — first-author publications, novel methods deployed in production, open-source contributions, or comparable evidence of original work
  • Deep proficiency in Python and/or R, including the modern ML stack (scikit-learn, PyTorch or TensorFlow, XGBoost/LightGBM, and similar)
  • Methodological breadth across paradigms — comfortable moving between classical statistics, tree-based ML, deep learning, and modern AI (transformers, graph neural networks, foundation models, generative methods) — and the judgment to argue rigorously for one approach over another
  • Strong statistical foundation: cross-validation strategy, regularization, calibration, uncertainty quantification, and handling of confounders and class imbalance
  • Hands-on experience building and validating classifiers on high-dimensional, low-sample-size biological data (proteomics, glycoproteomics, transcriptomics, or genomics)
  • Experience with batch-effect correction and normalization techniques, and a healthy skepticism about how those choices propagate into downstream performance estimates
  • Prior experience in the clinical diagnostics industry with a solid understanding of analytical and clinical validation, locking classifiers, and bridging studies
  • Excellent written and verbal communication: able to explain novel methods clearly to wet-lab scientists, clinicians, and fellow statisticians alike
  • A genuine desire to impact patient lives and contribute to the broader scientific community

Similar jobs

Data Scientist

Molina HealthcareUnited States· 1 mo ago
RemoteInformation Technology$80k–$172k/yr

Data Scientist

ABSC (Absolute Business Solutions Corp.)Herndon, VA· 2 mo ago
Engineering$75k–$150k/yrapply on absc-us.clearcompany.com

Data Scientist

ManulifeBoston, MA· 2 wk ago
Engineering$90k/yrapply on careers.manulife.com

Data Scientist

Juul LabsUnited States· 2 wk ago
RemoteEngineering$165k/yrapply on grnh.se

Data Scientist

WaystarLehi, UT· 2 wk ago
Engineeringapply on rr.jobsyn.org

Data Scientist

KoahUnited States· 6 mo ago
RemoteEngineering$180k–$250k/yrapply on jobs.ashbyhq.com