Linux Systems Engineer 2
At PNNL, our core capabilities are divided among major departments called Directorates, each focused on a specific area of scientific research or function. The Integrated Discovery Sciences Directorate (IDSD) leads fundamental research across biology, chemistry, Earth and environmental sciences, materials science, advanced computing, artificial intelligence, quantum information science, mathematics, autonomy, and DOE national user facilities. Our vision is to accelerate scientific discovery by observing, understanding, simulating, predicting, and controlling dynamic biotic-abiotic processes and interactions within complex systems.
The Environmental Molecular Sciences Division, part of IDSD, comprises 18 interdisciplinary research teams focused on deciphering molecular-level interactions driving biological and environmental processes. The division also manages the Environmental Molecular Sciences Laboratory (EMSL), a Department of Energy, Office of Science user facility providing world-class expertise, instrumentation, and computational resources to scientists globally.
Responsibilities
- Install, configure, maintain, and troubleshoot Linux servers (RHEL/Rocky and derivatives) across CDO's HPC clusters and EMSL support infrastructure, including bare-metal provisioning.
- Administer HPC cluster operations on Tahoma and Boreal, including Slurm scheduler configuration, node provisioning with Warewulf, InfiniBand networking, and hardware diagnostics.
- Develop and maintain automation using Ansible, including playbooks, roles, and configuration pipelines.
- Support and manage large-scale storage systems (Ceph, VastData, BeeGFS, Lustre, and the Aurora/HPSS archive) underlying HPC compute and EMSL instrument data.
- Work with DevOps tooling and workflows to ensure reliable software and infrastructure deployments (GitLab CI/CD, Git-based workflows).
- Participate in on-call or rotational support for mission-critical systems.
- Monitor compute, storage, and network health using tools such as Prometheus, Grafana, Nagios, and ELK.
- Assist with containerized workloads (Docker, Kubernetes) where applicable to HPC operational workflows.
- Perform hands-on hardware work on-site—racking, cabling, component replacement, and diagnostics—as a regular part of data center operations.
- Document processes, procedures, and checklists to ensure operational consistency across environments.
- Collaborate closely with development, research, and infrastructure teams to improve cluster performance, reliability, and automation.
This position requires onsite work a minimum of three days per week. Regular hands-on hardware work in the data center is a core and ongoing part of this role. No security clearance is required.
Requirements
- Minimum Qualifications:
- BS/BA and 2 years of relevant experience, or
- MS/MA, or
- PhD
- Preferred Qualifications:
- Degree in Computer Science, Computer Information Systems, or a related field.
- Strong Linux administration background including installation, patching, tuning, and troubleshooting.
- Hands-on automation experience with Ansible (YAML, roles, playbooks) or other system configuration tools.
- Experience with virtualization technologies or cloud platforms (VirtualBox, Proxmox, AWS, Azure), or hybrid deployments.
- Experience with databases such as MySQL, MariaDB, Postgres, or MongoDB.
- Scripting skills (bash, Python).
- Comfort and physical ability to perform hands-on server hardware work—diagnostics, firmware updates, racking, and cabling—in a data center environment.
- Familiarity with GitLab CI or GitHub Actions and Git-centric workflows.
- Basic knowledge of containerization (Docker/Kubernetes) used in DevOps and ML/HPC contexts.
- Experience deploying or supporting HPC systems or large Linux installations.
- Experience administering an HPC job scheduler (Slurm preferred).
- Experience with bare-metal provisioning tools such as xCAT or Warewulf.
- Experience using Linux system packaging tools to deploy, remove, and package software.
- Experience with large-scale storage systems such as Ceph, VastData, Lustre, BeeGFS, or HPSS.
- Experience building dashboards, alerting, or automation on top of monitoring stacks such as Prometheus, Grafana, or ELK.
- Experience collaborating with AI on scripting, software development, and/or system administration work.
Benefits
- Medical, dental, and vision insurance, including robust telehealth care options and mental health benefits.
- Health savings account and flexible spending accounts.
- Basic life insurance and disability insurance.
- Employee assistance program and business travel insurance.
- Tuition assistance, relocation, backup childcare, legal benefits, and fertility support.
- Supplemental parental bonding leave, surrogacy and adoption assistance.
- Company-funded pension plan and 401(k) savings plan with company match.
- Up to 120 vacation hours per year and ten paid holidays per year (Research Associates excluded).
All benefits are dependent upon eligibility.
About PNNL
Pacific Northwest National Laboratory (PNNL) is a world-class research institution powered by a highly educated, diverse workforce committed to the values of Integrity, Creativity, Collaboration, Impact, and Courage. PNNL is located in eastern Washington State, known for its stellar outdoor recreation and affordable cost of living. The Lab’s campus is a 45-minute flight (or ~3-hour drive) from Seattle or Portland and is serviced by the convenient PSC airport, connected to 8 major hubs.