Jobs · Engineering · Massachusetts

Scientist II / Senior ML Scientist, Data-Efficient Learning for Drug Discovery

Lila Sciences · Cambridge, MA · 3 wk ago
On-siteEngineering$228k/yrFull-time

Lila Sciences is seeking a Machine Learning Scientist to build models and learning strategies for settings where data is scarce, expensive, and intentionally generated. This role focuses on training useful models from low-quantity but high-quality datasets and deciding what data should be acquired next in tightly focused areas of chemical space.

About the role

This is an applied scientific ML role in a frontier research area, connecting model training with scientific decision-making. You will develop approaches across active learning, meta-learning, fine-tuning, uncertainty estimation, experimental design, and multimodal modeling to build closed-loop systems that learn efficiently from targeted data acquisition.

Responsibilities

  • Build ML models that perform well in low-data regimes for drug discovery and molecular optimization.
  • Design data acquisition strategies that identify which compounds, assays, DEL selections, simulations, structural predictions, or experiments should be run next to maximize learning.
  • Develop active learning, meta-learning, fine-tuning, transfer learning, and uncertainty-aware modeling approaches for focused chemical spaces.
  • Train models on low-quantity, high-quality datasets generated by Lila's experimental, computational, and agentic discovery systems.
  • Build multimodal models that integrate DEL data, simulation outputs, assay data, protein and structural information, chemical features, literature or text-derived signals, images, and experimental metadata.
  • Partner with experimental, computational, and drug discovery teams to ensure data acquisition plans are scientifically meaningful and operationally feasible.
  • Evaluate models through learning curves, prospective validation, retrospective benchmarks, uncertainty calibration, and decision-focused metrics.
  • Develop closed-loop learning workflows that continuously update models as new data arrives from experiments, simulations, and automated systems.
  • Translate model predictions and uncertainty into practical recommendations for compound selection, assay selection, batch design, or next experiments.
  • Work with platform and agent teams to expose model-driven recommendations as tools for scientists and AI agents.

Requirements

  • PhD or equivalent experience in machine learning, computational chemistry, computational biology, statistics, computer science, bioengineering, or a related field.
  • Strong experience training ML models in low-data regimes.
  • Experience with active learning, Bayesian optimization, experimental design, meta-learning, fine-tuning, transfer learning, uncertainty estimation, or related data-efficient learning methods.
  • Experience building ML models for scientific, molecular, biological, chemical, pharmacological, biochemical, or other high-dimensional experimental datasets.
  • Experience with multimodal learning or methods that combine heterogeneous data sources.
  • Ability to reason about data acquisition strategy, not only model fitting.
  • Strong scientific judgment and ability to connect model behavior to experimental decisions.
  • Practical experience with PyTorch, JAX, scikit-learn, or equivalent ML tools.
  • Ability to collaborate across ML, data, computational science, experimental, and drug discovery teams.

Qualifications (Bonus)

  • Drug discovery experience, especially in molecular optimization, screening, or design-make-test-learn workflows.
  • General understanding of pharmacology, biochemistry, or mechanisms of molecular activity.
  • Experience with DEL, high-throughput screening, medicinal chemistry, assay data, simulation-derived features, protein or structure-based features, text or literature features, or scientific images.
  • Experience with closed-loop experimentation, autonomous labs, or agent-driven scientific workflows.
  • Experience with generative molecular design, candidate prioritization, or batch selection workflows.
  • Familiarity with causal inference, optimal experimental design, decision theory, or Bayesian methods.
  • Comfort working with frontier ML techniques where standard out-of-the-box approaches are insufficient.

Benefits

  • U.S. Benefits: Full-time U.S. employees receive medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company-wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office-based employees; and a company-subsidized lunch program.
  • International Benefits: Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region.

Pay

Expected base salary range: $228,000 USD - $358,000 USD (U.S.-based positions; international salaries are set to local market). Competitive base compensation with bonus potential and generous early-stage equity.

About LILA

Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We combine advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery across medicine, materials, and energy. Guided by core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance.

Similar jobs