Scientific Data Engineer
About the role
The Computational Biosciences Group at Lawrence Berkeley National Laboratory seeks a Scientific Data Engineer to develop new methods and software tools that enable scientific knowledge discovery using modern data management and machine learning technologies. This role focuses on multi-modal data modeling and analysis with applications in omics/structural biology and neurophysiology data. You will work as part of a multidisciplinary team of computer scientists, data scientists, and bioinformaticians, advancing the state-of-the-art in data-intensive analysis under the FAIR data science, AI, and modern data understanding paradigms.
Responsibilities
- Design and develop user-friendly software packages for scientific data management and analysis.
- Work with domain experts to develop FAIR data models and management solutions for bioscience applications.
- Support machine learning and AI use of biological data by making it well-structured, documented, and efficiently accessible.
- Collaborate with the Neurodata Without Borders and LinkML open source data ecosystems, as well as the Joint Genome Institute.
- Maintain and manage open source software products, including development priorities, software releases, continuous integration, and testing.
- Design, implement, and maintain high-performance computing and cloud solutions for visualization and analysis of complex biological data.
- Develop machine learning and AI solutions for biological data analysis in close collaboration with diverse teams of scientists.
- Train scientists and research software engineers in the use of developed software products at workshops and conferences.
- Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
- Network with senior internal and external personnel in your area of expertise.
Requirements
- Minimum of 5 years of related experience with a Bachelor’s degree in computer science, data science, machine learning, bioinformatics, or equivalent; or 3 years and a Master’s degree; or equivalent work experience designing and developing software for data modeling or analysis; or a PhD in a relevant STEM field.
- Demonstrated experience developing software in a scientific or research context, such as in a research group, scientific user facility, or on a scientific software project.
- Hands-on experience in a production environment, developing scientific software, scientific data models, or scientific data pipelines.
- Strong programming experience in Python. Working proficiency in C++ or JavaScript is a plus.
- Experience testing large code bases.
- Experience contributing to community-driven open source software.
- Demonstrated experience in one or more of the following areas: data management, scientific data analysis, machine learning.
- Works well in a collaborative team environment.
- Demonstrated capability with Git version control and continuous integration systems, such as GitHub or GitLab.
- Ability to work effectively with domain scientists whose expertise is outside computing, and to translate their requirements into technical designs.
- Excellent oral and written communication skills.
- Demonstrated ability to work effectively as part of a cross-disciplinary team.
Desired Skills
- Master’s or PhD in Computer Science or related field, with 5 or more years of professional experience designing and developing scientific data modeling or analysis software.
- Experience working with modern scientific data formats and database systems, such as HDF5, Zarr, MongoDB, PostgreSQL, MySQL, and Redis.
- Experience with Neurodata Without Borders, LinkML, or similar software ecosystems.
- Experience working with large biological data, such as in neurophysiology, microbiology, genomics, or protein design.
- Experience designing or working with structured data models, schemas, ontologies, or data standards.
- Familiarity with FAIR data principles, persistent identifiers, provenance, and controlled vocabularies and ontologies.
- Experience preparing scientific datasets for use by machine learning pipelines or LLM-based agents.
- Experience working with cloud object storage, cloud computing, High-Performance Computing, data lakehouse architecture, or containerization.
- Experience developing web-based graphical user interfaces (GUIs) or application programming interfaces (APIs) for scientific data analysis and management.
Benefits
- Exceptional health and retirement benefits, including pension or 401K-style plans.
- Winter Holiday Shutdown every year in addition to accrued vacation and sick time.
- Parental bonding leave (for both mothers and fathers).
- Pet insurance.
Pay
The expected salary for this position is $131,760 - $161,064, with a full salary range of $117,132 - $197,676 depending on skills, knowledge, and abilities, including education, certifications, and years of experience.
Schedule
- Full-time, 2-year term appointment with the possibility of extension or conversion to a Career appointment based on performance, funding, and operational needs.
- Work may be performed on-site or hybrid at Lawrence Berkeley National Lab, 1 Cyclotron Road, Berkeley, CA. Work must be performed within the United States.