Scientific Data Engineer (Institutional Informatics Team - Joint Genome Institute)
About the Role
The Scientific Data Engineer will provide technical expertise supporting raw scientific data generation and compute systems that enable laboratory operations, data analysis, and project management activities. This role involves implementing complex scientific and operational functional specifications for automated systems, supporting core systems such as laboratory workflow orchestration, genomic data generation, metadata management, status tracking, the Laboratory Information Management System (LIMS), and the Data Warehouse/Data Lakehouse.
The position requires technical prowess, sound judgment in selecting methods and techniques, and collaboration with the team to resolve technical issues creatively. You will proactively identify technical and system needs, refine implementation approaches, and deliver operationally and scientifically valuable solutions. This role may be hired at the Staff or Senior level, with an anticipated start date of October 1, 2026.
Responsibilities
- Staff Scientific Data Engineer:
- Translate functional specifications and system designs into implementation, system enhancements, integrations, and the evolution of shared data platforms.
- Develop, test, deploy, maintain, and support core automated systems, services, APIs, and workflows—including the Laboratory Management System (LMS), Data Warehouse/Data Lakehouse, and Proposal/Project Management systems—to enable genomic data generation, metadata management, and JGI operations.
- Resolve technical issues, integration gaps, and optimize system architecture and operational efficiency.
- Drive implementation efforts to ensure shared production systems meet high standards for reliability, scalability, interoperability, and performance.
- Participate in peer technical review, documentation, and continuous optimization of development workflows and service operations.
- Senior Scientific Data Engineer (in addition to the above):
- Translate complex scientific, operational, and user requirements into functional specifications, system designs, and implementation plans.
- Design, develop, test, deploy, maintain, and support core automated systems, services, APIs, and workflows.
- Proactively identify and resolve complex technical issues, integration gaps, and opportunities to optimize system architecture and operational efficiency.
- Establish and champion engineering best practices and provide technical guidance to team members.
Requirements
- Staff Scientific Data Engineer:
- A Bachelor’s Degree (or equivalent knowledge/training) in Computer Science or a related field and a minimum of 5 years of related professional work experience developing, integrating, deploying, and supporting production software applications and data systems for metadata management, workflow orchestration, data lifecycle operations, and broad user data access to scientific and operational data.
- Experience with various database and data storage technologies, including relational databases, object storage platforms, and systems supporting structured, semi-structured, and large-scale datasets.
- Experience with data engineering and event-driven technologies such as Airflow, Kafka, or related tools.
- Experience using AI-assisted development tools, with demonstrated sound judgment in evaluating and validating generated code for production suitability.
- Strong working knowledge of software and data engineering fundamentals for large-scale production systems, including system design, APIs, testing methodologies, concurrency, reliability, scalability, interoperability, and performance optimization.
- Proficiency in Python and experience with one or more additional programming languages.
- Excellent communication skills, including organizing and presenting complex technical information to internal teams and stakeholders.
- Demonstrated experience collaborating with stakeholders to translate complex scientific, operational, and user requirements into automated systems, technical specifications, and implementation plans.
- Senior Scientific Data Engineer (in addition to the above):
- A minimum of 8 years of related professional work experience.
- Demonstrated experience providing technical leadership in shaping system architecture and technical direction across cross-functional engineering groups.
- Demonstrated experience providing technical guidance and mentorship to team members while promoting and championing engineering best practices, documentation, and continuous optimization of development workflows and operational processes.
Benefits
This is a full-time, exempt (monthly paid), 2-year term appointment with the possibility of extension or conversion to a career appointment based on performance, funding, and operational needs. The position is eligible for relocation assistance and a hybrid work schedule, combining telework and on-site work at Lawrence Berkeley National Lab in Berkeley, CA. Individuals working a hybrid schedule must reside within 150 miles of Berkeley Lab.
Pay
- Staff Scientific Data Engineer (Job Code C71.2): $117,132 - $146,400 annually.
- Senior Scientific Data Engineer (Job Code C71.3): $139,440 - $174,312 annually.
It is not typical for an individual to be offered a salary at or near the top of the range. Salary will be commensurate with the final candidate’s qualifications and experience, including skills, knowledge, relevant education, certifications, and aligned with the internal peer group.