High-Performance Computing (HPC) Systems Engineer
About the role
We are seeking a High-Performance Computing (HPC) Systems Engineer to architect, deploy, and maintain the large-scale physical hardware, distributed filesystems, and low-latency networking infrastructure powering our scientific computing ecosystem. This role focuses on the bare-metal and system-level foundations of a petabyte-scale environment, ensuring high availability, peak storage performance, and reliable data movement for experimental nuclear physics workloads. The ideal candidate will blend deep Linux systems administration expertise with modern infrastructure-as-code automation to support state-of-the-art research computing clusters.
Responsibilities
- Collaboratively design and implement software infrastructure supporting High Performance and High Throughput computing using best-in-class containerization and orchestration tools.
- Deploy, maintain, and operate physical infrastructure for scientific computing services.
- Create highly available services through consideration of the entire stack from hardware and networking through the user application.
- Engage with users to meet service requirements for performance and availability.
- Develop software to address gaps in existing tools.
- Audit and recommend architectural changes to improve performance, security, and availability.
- Contribute to upstream development efforts of community tools.
- Work with students and interns on research projects.
Experience Required
- 3 or more years performing enterprise Linux system administration tasks including installation, configuration, and support of COTS/GOTS/FOSS software, file, network, and large-scale storage systems.
- Experience operating in a production environment with high availability requirements.
Preferred Experience
- Experience with automating test procedures.
- Experience with NP/HEP HPC infrastructure like Slurm, Rucio, Globus, XrootD.
Education
- Required: Bachelor’s Degree in Computer Science, Computer Engineering, or related degree with significant Computer Science coursework.
- Preferred: Master’s Degree in Computer Science, Data Science, or related discipline.
Knowledge, Skills, and Abilities
- Ability to architect, provision, tune, and maintain petabyte-scale, high-performance parallel and distributed storage systems, including Lustre, Ceph, CephFS, and NFS.
- Expert Kubernetes, Docker/Podman, and containerization administration.
- Expert knowledge of Linux system administration including installation and configuration management (e.g., Ansible/Puppet/Foreman).
- Ability to develop, debug, and test applications based upon design and performance requirements.
- Ability to communicate clearly in writing (e.g., email, presentations, drawings) and explain work to their supervisor and others in the group.
- Ability to work effectively with peers and participate in troubleshooting, including off-hours during outages and as part of an occasional on-call rotation.
Pay
The good-faith pay range for this role is $91,800 – $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate’s skills, experience, education, certifications, and work location.