Scientific Software & DevOps Engineer
Jefferson Lab · Newport News, VA · 2 days ago
Engineering$92k–$145k/yrFull-time
About the Role
We are seeking a Scientific DevOps Engineer to bridge the gap between experimental nuclear physics workloads, distributed data management, and modern containerized deployment practices. The role is responsible for automating complex software stacks, operationalizing petabyte-scale data pipelines, and deploying local AI/LLM infrastructure on platforms like Kubernetes and OpenShift. Additionally, this position will support developer-facing application layers, custom front-end data portals, and core experimental metadata frameworks.
Responsibilities
- Take the lead on complex system configuration and software deployments in collaboration with other system administrators, developers, lab staff, and users.
- Deploy, scale, and manage critical services using platforms including Kubernetes (k8s), OKD, OpenShift/OCP, and Talos.
- Maintain, optimize, and automate infrastructure configuration pipelines using tools like Ansible and Puppet.
- Serve as expert resource to other members of the division in several areas of system administration and development.
- Remain up-to-date with emerging technology by researching and creating testbed configurations for next-generation technologies of interest including storage, compute scheduling, and configuration control.
Additional Responsibilities
- Design, configure, and manage distributed workflow systems and high-throughput data transfer pipelines using tools such as Rucio, FTS (File Transfer Service), InvenioRDM, Globus, Slurm, and the Open Science Grid (OSG).
- Deploy, operationalize, and maintain local and commercial Artificial Intelligence / Large Language Model (AI/LLM) environments, specifically managing gateway tools like liteLLM to serve developer and experimental research workloads.
- Provide secondary DBA support to optimize, query, and maintain relational database schemas (PostgreSQL, MySQL/MariaDB) underpinning the scientific application stack.
Qualifications
- 3 or more years related experience exercising the Knowledge, Skills, and Abilities noted below.
- Experience managing production-grade Kubernetes, OpenShift, or similar container platforms.
Education
- Required: Bachelor's Degree in Computer Science, software engineering, or closely related technical discipline.
- Preferred: Master's Degree in Computer Science or Software Engineering.
- Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.
Knowledge, Skills, and Abilities
- Solid foundations in Linux system administration, network troubleshooting, and security best practices.
- Strong scripting and automation skills (e.g., Python, Bash, or Perl) paired with experience using Configuration Management tools (Ansible/Puppet/Satellite/Foreman).
- RedHat ecosystem experience favored.
- Knowledgeable of IP networking, relational database principles, and computer architecture.
- Ability to clearly communicate and report on progress on tasks and projects.
- Strong interpersonal skills ability to effectively work on project teams.
Benefits
- Medical, Dental, and Vision Care Plans
- Flexible Spending Accounts
- Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and sick leave)
- 401(k) Plan – 9% Lab Contribution; 100% vested
- Flexible Work Arrangements (Remote & Alternate Work Schedules available)
- Tuition Assistance, Training and Professional Development Programs
Pay
$91,800 - $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.