Jobs · OTHR · Georgia

High Performance Computing Administrator III

Gulfstream Aerospace · Savannah, GA · 4 wk ago
OTHRFull-time

Requisition Number: 235640 | Full-time | First shift | Up to 25% travel

About the role

Contribute to the strategic objectives of the High Performance Computing environment in the Advanced Computing Technologies department. Develop operational plans, goals, and strategies to best serve the Computational Fluid Dynamics, Simulation, and Modeling Engineering business units. Provide technical oversight and operational support to ensure continued sustained functionality of the HPC environment. Responsible for the optimum integration of scientific applications to HPC technology and the exploration of new HPC technology to better meet business needs.

Responsibilities

  • Assume responsibility for the day-to-day operations of Gulfstream's production HPC cluster.
  • Assist end users running applications on the HPC cluster.
  • Provide third-level support for end users experiencing problems on engineering workstations and remote visualization systems.
  • Manage, maintain, monitor, and control interactive and batch processes, both scheduled and unscheduled (including on-request processing).
  • Ensure engineering-defined batch processing and backups are completed in the correct sequence and within established time periods.
  • Suggest improvements to processing capabilities and efficiencies through system tuning and other hardware/software optimizations.
  • Perform regular monitoring of utilization needs and efficiencies; report on tuning initiatives.
  • Perform proactive failure trend analysis and root cause analysis for all system failures.
  • Produce trend reports to highlight production issues and follow predetermined action/escalation procedures.
  • Monitor, verify, and adjust support for proper application executions.
  • Provide technical solutions that meet performance and processing objectives of business areas.
  • Perform upgrades compliant with corporate policies and industry best practices.
  • Lead HPC Administrators during system upgrades and outages; create thorough upgrade plans.
  • Assist in introducing new technologies to improve capabilities, productivity, and reduce total cost of ownership.
  • Participate in the design of HPC technical solutions.
  • Continuously evaluate efficiency, existing technology effectiveness, and interoperability; suggest areas for improvement.
  • Maintain technical relationships with multiple hardware and software vendors.
  • Work multiple operational windows as required; provide 24x7 on-call support.
  • Assist in development and implementation of technical, hardware, and software standards.
  • Perform other duties as assigned.

Requirements

  • Bachelor's Degree in Engineering, Computer Science, Information Technology, or related curriculum, or equivalent combination of education and experience. A Master's degree may offset one year of experience; a PhD may offset two years.
  • 7 years of experience in an HPC or scientific computing environment, including installation, configuration, and maintenance of Red Hat Package Manager (RPM)-based Linux distributions (RedHat, SuSE).
  • Experience managing Infiniband-based Linux HPC clusters, high-performance parallel storage, and cluster scheduling software.
  • Experience with low-latency, high-bandwidth HPC interconnects.
  • Experience supporting Linux-based scientific workstations running visualization applications.
  • Ability to read, write, speak, and understand English.

Skills

Essential Skills

  • Linux Competency: Installation, configuration, command-line proficiency, and package management.
  • Systems Management: Mastery of cluster deployment and configuration management tools (Satellite/Foreman, Warewulf, Ansible, Puppet, Terraform, etc.).
  • Scripting Experience: Ability to script administrative tasks in Bash, Python, Perl, or similar languages.
  • Hardware Support: Ability to troubleshoot advanced hardware failures and apply repairs.
  • HPC System Support: Scheduler management, complex code compilation, and multi-node application support.
  • HPC System Engineering: Ability to design, build, and configure large-scale production-level systems with minimal supervision.
  • HPC Networking: Experience designing, deploying, and administering HPC networking platforms (InfiniBand, RoCE, etc.).
  • Infrastructure Essentials: Experience with enterprise infrastructure tools (DNS, DHCP, PXE, identity management, etc.).
  • Productivity Tools: Proficiency with Excel, Word, PowerPoint, Teams, or similar applications.

Preferred Skills

  • Storage: Experience designing, deploying, and administering enterprise or HPC storage systems (Lustre, VAST, FlashArray, etc.).
  • Virtualization: Experience designing, deploying, and administering virtualization platforms (VMWare, HyperV, ProxMox, etc.).
  • Security: Experience with vulnerability scanning and compliance standards (CIS, STIG, etc.).

Similar jobs