Data Scientist
Summary
The Client is seeking a Data Scientist to join the Enabling Technologies & Automation Group within Drug Substance Technologies - Synthetics. This team supports data-driven process development through high-throughput reaction screening, process optimization, crystallization development, kinetic profiling, and biocatalytic reaction development. The role sits at the intersection of process chemistry, data science, cheminformatics, and data engineering, with responsibility for developing and maintaining robust, user-focused scientific data tools and predictive modeling solutions. The successful candidate will enable data-driven workflows from experimental design through actionable predictions, while supporting the integration of process chemistry data into enterprise data systems.
Responsibilities
- Develop and implement regression and classification models for chemical reactivity using process parameters and molecular descriptors.
- Build predictive models that evaluate outcomes such as yield, conversion, selectivity, impurity formation, reaction success or failure, reaction rate robustness, and optimal operating conditions.
- Design models with a strong emphasis on statistical performance, interpretability, chemical relevance, and practical application by experimental teams.
- Maintain molecular and physicochemical descriptor databases used in modeling and process development workflows.
- Develop active learning and Bayesian optimization workflows to support data-driven experiment selection.
- Support iterative design-make-test-analyze cycles by incorporating newly generated experimental data into predictive models to improve coverage and accuracy.
- Collaborate with process chemists, analytical scientists, automation scientists, engineers, and other subject matter experts to collect, curate, and structure datasets for modeling and decision-making purposes.
- Provide recommendations on experimental design practices that improve data quality, model performance, usability, and interpretability.
- Translate analytical and modeling results into actionable recommendations for chemistry and process development teams.
- Communicate technical concepts, methodologies, and results effectively to both technical and non-technical audiences.
- Develop, maintain, and support scientific data tools that address real-world process chemistry challenges.
- Design and support workflows for ingesting, transforming, validating, and structuring process chemistry data.
- Support the integration of curated process chemistry datasets into enterprise data platforms and data infrastructure.
- Reimport and harmonize historical and legacy datasets from varying sources and formats.
- Capture and manage critical metadata to support data traceability, accessibility, and reuse.
- Ensure scientific datasets are suitable for both human review and machine learning applications.
- Contribute to sustainable software and data solutions that support process development activities throughout all stages of a lifecycle program.
Education Requirements
Degree in relevant scientific, engineering, data science, computational science, chemistry, or related discipline.
Experience Requirements
- Experience working in cross-functional teams consisting of chemists, engineers, and data scientists in support of data-driven chemical development.
- Experience developing predictive modeling solutions using experimental and process chemistry data.
- Experience applying cheminformatics methods and molecular descriptors to chemical modeling workflows.
- Experience supporting experimental design, active learning, optimization, and data-driven decision-making initiatives.
- Experience working with scientific datasets, including data ingestion, transformation, validation, standardization, and integration activities.
- Experience collaborating with subject matter experts to develop practical software and data solutions.
- Experience communicating technical findings and recommendations to diverse stakeholder groups.
Required Skills
- Strong knowledge of Python (version 3.10 or higher).
- Proficiency with data science and scientific computing libraries, including: NumPy, SciPy, Pandas, Scikit-learn, RDKit.
- Experience applying Density Functional Theory (DFT) for molecular property and descriptor calculations.
- Knowledge of regression and classification modeling techniques.
- Knowledge of molecular descriptors and physicochemical property analysis.
- Experience with data engineering and scientific data workflow development.
- Ability to develop robust, maintainable, and user-focused scientific software solutions.
- Strong analytical, problem-solving, and communication skills.
Preferred Skills
- Experience developing active learning workflows.
- Experience with Bayesian optimization techniques.
- Experience supporting high-throughput experimentation and process development environments.
- Familiarity with process chemistry, reaction optimization, crystallization development, kinetic profiling, or biocatalytic reaction development.
- Experience working with enterprise-scale scientific data platforms and reusable data assets.
Pay
$41.00 - $43.00 per hour. Non-exempt positions are eligible for overtime at a rate of 1.5 times the base hourly rate for all hours worked in excess of 40 in a work week, or as required by state or local law. Final offer amounts, within the base pay set forth above, are determined by factors including your relevant skills, education and experience.