Jobs · Information Technology · Virginia

High-Performance Computing (HPC) Systems Engineer

Jefferson Lab · Newport News, VA · Yesterday
Information Technology$92k–$145k/yrFull-time

About the role

We are seeking a High-Performance Computing (HPC) Systems Engineer to architect, deploy, and maintain the large-scale physical hardware, distributed filesystems, and low-latency networking infrastructure powering our scientific computing ecosystem. This role focuses on the bare-metal and system-level foundations of a petabyte-scale environment, ensuring high availability, peak storage performance, and reliable data movement for experimental nuclear physics workloads. The ideal candidate will blend deep Linux systems administration expertise with modern infrastructure-as-code automation to support state-of-the-art research computing clusters.

Responsibilities

  • Collaboratively design and implement software infrastructure supporting High Performance and High Throughput computing using best-in-class containerization and orchestration tools.
  • Deploy, maintain, and operate physical infrastructure for scientific computing services.
  • Create highly available services through consideration of the entire stack from hardware and networking through the user application.
  • Engage with users to meet service requirements for performance and availability.
  • Develop software to address gaps in existing tools.
  • Audit and recommend architectural changes to improve performance, security, and availability.
  • Contribute to upstream development efforts of community tools.
  • Work with students and interns on research projects.

Experience Required

  • 3 or more years performing enterprise Linux system administration tasks including installation, configuration, and support of COTS/GOTS/FOSS software, file, network, and large-scale storage systems.
  • Experience operating in a production environment with high availability requirements.

Preferred Experience

  • Experience with automating test procedures.
  • Experience with NP/HEP HPC infrastructure like Slurm, Rucio, Globus, XrootD.

Education

  • Required: Bachelor’s Degree in Computer Science, Computer Engineering, or related degree with significant Computer Science coursework.
  • Preferred: Master’s Degree in Computer Science, Data Science, or related discipline.

Knowledge, Skills, and Abilities

  • Ability to architect, provision, tune, and maintain petabyte-scale, high-performance parallel and distributed storage systems, including Lustre, Ceph, CephFS, and NFS.
  • Expert Kubernetes, Docker/Podman, and containerization administration.
  • Expert knowledge of Linux system administration including installation and configuration management (e.g., Ansible/Puppet/Foreman).
  • Ability to develop, debug, and test applications based upon design and performance requirements.
  • Ability to communicate clearly in writing (e.g., email, presentations, drawings) and explain work to their supervisor and others in the group.
  • Ability to work effectively with peers and participate in troubleshooting, including off-hours during outages and as part of an occasional on-call rotation.

Pay

The good-faith pay range for this role is $91,800 – $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate’s skills, experience, education, certifications, and work location.

Similar jobs