Senior HPC & Infrastructure Engineer
Soni · Camden, NJ · Today
Information Technology$135k–$155k/yrFull-time
Location: Cherry Hill, NJ – 3-4 days/week on-site Position Overview Our client is seeking a highly skilled Senior HPC & Infrastructure Engineer to support and evolve enterprise technology platforms spanning both cloud and on-premises environments. This role will serve as a technical lead for Linux-based infrastructure and high-performance computing (HPC) resources that enable advanced engineering, scientific, and data-intensive workloads. The ideal candidate will bring deep expertise in Linux administration, cloud technologies, automation, security, and performance optimization while partnering closely with technical stakeholders to deliver scalable and reliable computing solutions. Key Responsibilities Lead the design, deployment, and ongoing support of enterprise infrastructure solutions across hybrid cloud and on-premises environments. Administer large-scale Linux platforms, including clustered computing environments used for modeling, simulation, analytics, and research activities. Manage software ecosystems by installing, configuring, and maintaining development tools, scientific applications, compilers, libraries, and supporting dependencies. Analyze system performance and implement improvements to maximize efficiency, throughput, and resource utilization. Configure and support resilient Linux environments, incorporating fault tolerance, redundancy, and failover capabilities for networking and storage systems. Oversee HPC resource management platforms, including job scheduling, queue administration, workload prioritization, and policy governance. Collaborate with engineers, researchers, and business users to understand compute, storage, and workflow requirements and translate them into effective infrastructure solutions. Maintain systems supporting data preparation, processing pipelines, and workflow orchestration across compute environments. Design, implement, and manage Microsoft Azure environments, including government cloud deployments, with a focus on security, governance, and compliance. Support infrastructure required for AI and machine learning initiatives, ensuring performance, scalability, and operational reliability. Establish and maintain security controls for critical systems, including identity management, auditing, data protection, encryption, and regulatory compliance practices. Investigate and resolve issues affecting compute nodes, storage systems, software environments, scheduling platforms, and user workloads. Develop automation solutions to streamline platform administration, deployments, monitoring, and operational consistency. Produce and maintain technical documentation, system standards, operational procedures, and knowledge base materials. Required Qualifications Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline, or equivalent professional experience. 10+ years of experience managing Linux-based enterprise infrastructure, including high-performance computing environments. Advanced expertise in Linux administration, preferably within Red Hat Enterprise Linux (RHEL) or Rocky Linux ecosystems. Hands-on experience with Microsoft technologies, including Azure, Entra ID, Active Directory, and Windows Server platforms. Strong knowledge of HPC scheduling and workload management solutions such as Altair PBS, Slurm, or IBM Spectrum LSF. Experience deploying and maintaining complex software stacks, development environments, and containerized applications. Demonstrated proficiency in automation and scripting using Bash, Python, or comparable technologies. Experience supporting virtualized environments, workstations, and GPU-enabled computing platforms. Strong understanding of infrastructure security principles, access controls, and best practices for protecting sensitive or regulated workloads. Ability to partner effectively with engineering, research, analytics, and technical operations teams. Excellent troubleshooting, communication, and documentation skills. Preferred Qualifications Experience with Infrastructure-as-Code technologies such as Terraform or Azure Bicep. Familiarity with configuration management platforms including Ansible, Puppet, or similar automation tools. Background supporting scientific computing, parallel processing applications, or computational engineering environments. Understanding of cybersecurity and compliance frameworks such as NIST, CIS Controls, or equivalent standards. Experience supporting cloud-native AI, advanced analytics, or data science platforms. Compensation: $135,000 to $155,000 annually Compensation is based on a range of factors that include relevant experience, knowledge, skills, other job-related qualifications.