Research Computing Infrastructure Engineer
Employee Type: exempt full-time | Division: Enterprise Information Technology | Location: Frederick, MD, USA
About the Role
The Frederick National Laboratory, operated by Leidos Biomedical Research, Inc., addresses urgent and intractable problems in biomedical sciences, including cancer, AIDS, drug development, nanotechnology in medicine, and emerging infectious diseases. The Enterprise Information Technology (EIT) mission is to develop an enterprise-level IT infrastructure supporting basic, translational, and clinical research. The IT Operations Group (ITOG) focuses on computational servers, storage, virtual machine infrastructure, and network services, implementing enterprise IT best practices.
The Research Computing Infrastructure Engineer provides technical leadership for the design, standards, and implementation of the SOM virtualization and container platform and its supporting infrastructure. This role evaluates and implements high-impact solutions, serves as a technical mentor, and supports research computing programs within a federal compliance boundary.
Responsibilities
- Design, deploy, and operate the SOM virtualization and container platform as secure, compliant infrastructure within a federal boundary.
- Operate the virtualization substrate: cluster lifecycle, node management, upgrades, and underlying storage and networking.
- Manage downstream Kubernetes cluster provisioning, RBAC, and multi-tenant access for research groups.
- Own storage integration across the platform (software-defined storage plus POSIX and object storage such as VAST) for VM and container workloads.
- Build and maintain the security and compliance posture of the platform: SSP development, ATO support, continuous monitoring, and remediation under NIST 800-53/FISMA.
- Automate provisioning, configuration, and scaling through infrastructure-as-code and CI/CD practices (Ansible, Terraform, Packer, GitHub Actions, or equivalent).
- Partner with embedded bioinformatics staff to provide the platform and guardrails for user-facing VM provisioning, container workflows, and job templates.
Requirements
- Bachelor’s degree from an accredited college/university (or four years of relevant experience in lieu of a degree). Foreign degrees must be evaluated for U.S. equivalency.
- Minimum of eight years of related experience with strong Linux systems engineering and administration.
- Hands-on experience operating a virtualization platform in production (KVM/libvirt, VMware/vSphere, OpenStack, Harvester, or equivalent), including host lifecycle, live migration, and storage backends.
- Hands-on scripting/programming proficiency (Python, Bash, or Go) for automation and operations.
- Experience with infrastructure-as-code and automation tooling (Ansible, Terraform, Packer, or equivalent).
- Hands-on experience standing up and maintaining Kubernetes clusters, including upgrades, RBAC, networking, and storage.
- Experience working within a federal security and compliance framework: SSP, ATO, continuous monitoring, or NIST 800-53/FISMA controls.
- Comfortable with small-team environments and taking end-to-end ownership of compute infrastructure.
- Ability to obtain and maintain a security clearance.
Preferred Qualifications
- Direct experience with Harvester and/or Rancher in production, or demonstrated ability to rapidly own a new HCI/Kubernetes platform.
- Experience with KubeVirt or other VM-on-Kubernetes patterns.
- Experience with Longhorn or comparable software-defined storage.
- Understanding of storage integration with high-performance clusters (POSIX + object storage, VAST or similar).
- Familiarity with cloud GPU environments (AWS, GCP, Azure) and hybrid workflows.
- Familiarity with HPC scheduling (Slurm, batch workloads) and container workflow/pipeline tooling (Argo, Kubeflow, Ray, Prefect, Airflow).
- Good communication and documentation skills, with the ability to make complex infrastructure understandable to researchers and engineers.
Expected Competencies
- Operational command of Kubernetes and at least one virtualization platform, with the ability to own a virtualization stack end-to-end.
- Deep Linux systems administration, performance tuning, and automation, with sound security judgment inside a federal compliance boundary.
Pay
123,800.00 - 207,125.00 USD. The posted pay range is a general guideline and not a guarantee of compensation or salary. Additional factors considered include job responsibilities, education, experience, knowledge, skills, abilities, internal equity, and market data.
Benefits
Competitive compensation, Health and Wellness programs, Income Protection, Paid Leave, and Retirement.