High Performance Compute Systems Site Lead (Onsite - LANL)
This role has been designated as primarily work-from-home, with daily onsite work required in Los Alamos, New Mexico.
About the Company
Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today’s complex world. Our culture thrives on finding new and better ways to accelerate what’s next. We value varied backgrounds and embrace flexibility to manage work and personal needs. We make bold moves together and are a force for good.
About the Role
Join Hewlett Packard Enterprise as a Technical Site Lead supporting some of the world’s most advanced high-performance computing (HPC) and AI systems at Los Alamos National Laboratory (LANL), including large-scale Linux and HPE Cray environments. This senior, hands-on individual contributor position provides onsite technical and service-delivery leadership for HPE hardware engineers, Linux system administrators, software analysts, inventory specialists, and remote engineering resources supporting mission-critical HPE Cray and related platforms that enable scientific discovery and national security.
This is a technical leadership role with no direct reports. The Technical Site Lead combines hands-on Linux systems administration and HPC support experience with enterprise server hardware capability, structured troubleshooting, customer-facing leadership, incident coordination, and disciplined service delivery. The role is accountable for coordinating the onsite team's technical execution, maintaining operational readiness and service quality, managing escalations, planning maintenance, supporting major incidents, and preparing the site for next-generation HPC and AI platforms.
US Citizenship and the ability to obtain and maintain a DOE Q Clearance are required. This is not a remote or hybrid position.
Responsibilities
- Provide day-to-day technical leadership and guidance to onsite HPE hardware engineers, Linux system administrators, and software analysts, while coordinating work across inventory specialists and remote engineering resources supporting large-scale HPE Cray EX and related HPC and AI systems.
- Set and communicate daily and weekly technical priorities based on system health, open support cases, scheduled maintenance, customer priorities, operational risk, service commitments, and available staffing.
- Serve as the primary onsite technical focal point for day-to-day system support, technical escalations, maintenance activities, and service-delivery risks, partnering with the DSM on customer governance, executive escalations, and broader service-delivery matters.
- Maintain a current view of system health, support-case status, technical risks, maintenance actions, ownership, and outstanding commitments. Prepare and lead routine onsite operational reviews with the customer and onsite team in coordination with the DSM and HPC leadership.
- Ensure support cases contain accurate technical details, diagnostic evidence, business impact, troubleshooting history, current ownership, and clearly defined next actions.
- Drive timely escalation through established HPE processes and help team members engage next-tier support, engineering, product teams, and other resources needed to advance diagnosis and resolution.
- Plan and coordinate system upgrades, maintenance windows, installations, acceptance activities, and other production-impacting work in accordance with customer change-control requirements and HPE support processes. Review prerequisites, risks, execution steps, and expected outcomes with the onsite team and customer before work begins.
- Confirm required staffing, parts, tools, test equipment, documentation, communications, escalation contacts, rollback plans, and contingency coverage are in place for planned work.
- Lead the HPE onsite response during major incidents by aligning technical priorities with the customer-designated incident lead and DSM, organizing HPE resources, maintaining clear communications, tracking actions and decisions, and ensuring appropriate escalation.
- Coordinate HPE root-cause analysis, corrective actions, lessons learned, and documentation updates after significant incidents or recurring issues involving HPE-supported hardware, firmware, and infrastructure.
- Identify risks to system availability, service-level commitments, maintenance schedules, operational readiness, or customer satisfaction and escalate them promptly to the DSM.
- Coordinate onsite parts inventory, repair materials, tools, test equipment, and other HPE-owned resources using approved HPE business systems and controls.
- Build and maintain effective, professional working relationships with onsite team members, remote engineering organizations, HPE leadership, customer technical staff, and customer management.
- Facilitate regular team coordination discussions to organize work, confirm ownership, review progress, surface technical blockers, and ensure commitments are completed or appropriately escalated.
- Provide technical mentoring, coaching, and practical guidance to team members without assuming formal people-management authority.
- Promote a culture of accountability, disciplined troubleshooting, accurate documentation, safe work practices, knowledge sharing, and professional customer engagement.
- Track completion of required HPE, customer, security, safety, and technical training; notify team members of approaching deadlines and escalate overdue or at-risk requirements to the DSM.
- Coordinate onsite coverage using approved team schedules and planned leave information. Identify and escalate potential gaps in business-hours, maintenance, or on-call coverage before they affect service delivery.
- Provide the DSM with fact-based observations regarding technical performance, development needs, recognition opportunities, and issues requiring formal management attention.
- Coordinate site-specific onboarding and operational readiness for new team members, including required accounts, badges, site access, training, workspace, equipment, and introductions to key stakeholders.
- Maintain accurate site procedures, contact lists, escalation paths, team schedules, operational references, and other information required for effective day-to-day support.
- Maintain a clean, safe, secure, and organized working environment in accordance with HPE and customer requirements.
- Participate in the on-call rotation and provide additional onsite support when required for 24x7 operations, planned maintenance, system outages, and major incidents.
- Use Linux command-line and diagnostic tools in Red Hat Enterprise Linux (RHEL), SUSE Linux Enterprise Server (SLES), or comparable environments to analyze processes, filesystems, services, permissions, network state, system logs, hardware telemetry, and overall system health.
- Lead and participate in monitoring, diagnosis, maintenance, and restoration of large-scale HPC compute, high-speed interconnect, storage, and management infrastructure, along with HPE-supported power and cooling components.
- Apply a structured, evidence-based troubleshooting methodology that correlates Linux logs, hardware telemetry, network state, firmware information, and prior case history to isolate faults, test hypotheses, document findings, and determine the appropriate corrective action or escalation path.
- Diagnose and support repair of enterprise server hardware, compute nodes, management components, HPE-supported interconnect and network components, storage components, power systems, cabling, and other integrated HPC equipment.
- Perform or assist with hardware component replacement, cable and fiber management, rack-level work, equipment installation, electrostatic-discharge controls, and other hands-on data center activities in accordance with approved service procedures, safety requirements, and change controls.
- Use out-of-band management controllers and interfaces, including BMCs, Redfish, and IPMI, to assess hardware health, validate firmware inventory, review environmental conditions and event logs, manage power state, and verify component status.
- Read and interpret system documentation, hardware diagrams, rack elevations, network diagrams, cable maps, schematics, and service procedures to locate, identify, and diagnose system components.
- Support new-system installation, expansion, integration, acceptance testing, hardware and firmware upgrades, and transition-to-operations activities.
- Create and maintain site documentation, troubleshooting procedures, maintenance plans, workflows, technical checklists, incident records, and knowledge articles.
- Use Bash, Python, Git, and other approved scripting, version-control, and collaboration tools to collect information, automate repeatable tasks, analyze system data, and maintain operational documentation and configuration references.
Requirements
- US Citizenship and the ability to obtain and maintain a DOE Q Clearance.
- Must work onsite Monday through Friday in Los Alamos, New Mexico, with additional onsite work as required for planned maintenance, major incidents, and on-call support. This is not a remote or hybrid position.
- High school diploma or equivalent with at least 7 years of relevant technical experience, or an associate or bachelor’s degree in a technical field with at least 5 years of relevant technical experience.
- 5+ years of hands-on experience supporting complex electronic systems, enterprise server hardware, integrated data center infrastructure, or comparable production technology environments. This experience must include diagnosing, repairing, or maintaining enterprise server components, including processors, memory, storage devices, power supplies, BMCs, network adapters, optical connectivity, and copper or fiber cabling.
- 3+ years of hands-on experience supporting HPC systems, supercomputing environments, large-scale Linux clusters, or similarly complex Linux-based compute infrastructure. This experience must include troubleshooting interactions among compute nodes, management systems, high-speed interconnects, storage platforms, operating systems, workload managers or schedulers, power, cooling, and supporting infrastructure.
- 3+ years of experience providing technical leadership, mentoring, work coordination, or task direction for a multidisciplinary technical team. Direct people-management experience is not required.
- 3+ years of hands-on Linux system administration, production support, or troubleshooting experience with Red Hat Enterprise Linux (RHEL), SUSE Linux Enterprise Server (SLES), or a comparable enterprise Linux distribution. Candidates must be able to independently use Linux command-line tools to navigate filesystems, inspect and manage processes and services, analyze system logs, review permissions, and perform package management.
Schedule
Standard work hours are Monday through Friday, either 8:00 a.m. to 5:00 p.m. or 7:00 a.m. to 4:00 p.m., with additional onsite support required for planned maintenance, major incidents, and on-call lead responsibilities.