Systems Engineer
About the role
The Lead Systems Engineer for High Performance Computing (HPC) and Artificial Intelligence (AI) works as part of the Advanced Systems team within Research Computing. This role supports the hardware and system-level software on the University's centralized high-performance computing and other research systems. The Lead Systems Engineer engages with faculty, researchers, vendors, and other IT staff to specify, design, install, and administer research computing systems while providing insight into trends and technologies supporting AI research. The role also involves evaluating, piloting, and implementing systems that advance Princeton's HPC and AI technologies, enhancing Research Computing services.
Responsibilities
- Design, maintain, troubleshoot, and refine advanced HPC/AI cluster infrastructure including high-performance interconnects, cluster schedulers, and configuration management across research systems.
- Partner with colleagues in Advanced Data and Storage Management to align designs for scratch filesystems and data management with cluster designs.
- Develop data-transfer pathways and networks to support AI-driven computing workloads.
- Establish and maintain best practices for cluster management and usage to support AI-driven workloads.
- Develop documentation for users and technical staff that can be used by the larger community.
- Develop, enhance, and expand monitoring infrastructure and related protocols for research computing systems.
- Plan and implement scheduled maintenance of operations, including during off hours.
- Define and drive the institutional technical strategy for advanced AI and data-intensive HPC.
- Bring creativity, foresight, and mature professional judgment in anticipating and solving novel and complex problems, determining project objectives and requirements, and developing standards and governance for all research computing platforms.
- Leverage expertise in AI technologies to identify, evaluate, and pilot researcher-facing systems that enable the acceleration of research using AI.
- Lead the implementation and expand adoption of modern, automation-driven infrastructure and cluster management practices.
- Promote institution-wide collaboration as the community expert advising and working with faculty, researchers, and vendors on emerging trends and challenges in AI-enabled research computing.
- Cultivate a collaborative, knowledge-sharing environment by providing technical mentorship to systems specialists and analysts by sharing designs and operational expertise across data systems and HPC/AI infrastructure.
- Contribute to the strategic vision for HPC/AI systems; advise senior leadership and stakeholders on strategic investments, risks, and opportunities related to research infrastructure.
- Monitor HPC clusters, networks, and storage systems for abnormalities, and resolve issues.
- Analyze and solve problems in Linux and HPC/AI computing environments with software, data, and job submissions.
- Use scripting and programming tools to troubleshoot issues.
- Perform other tasks as assigned.
Requirements
- 10+ years of strong experience managing advanced research computing systems.
- Strong expertise with Linux system administration, installation, and troubleshooting.
- Advanced experience writing scripts in languages such as bash, Python, and/or Perl.
- Proficient in managing networking in HPC environments.
- Strong experience managing software in an advanced research computing environment.
- Experience supporting scheduling and managing jobs (SLURM) in large-scale computing environments.
- Strong oral and written communication skills, with the ability to proactively engage peers and communicate effectively across a diverse stakeholder community.
- Strong ability to solve complex and system infrastructure problems, and share expertise with colleagues at all levels.
- Demonstrated ability to collaborate across teams to solve systems and infrastructure challenges, aligning day-to-day operational needs with longer-term technical and organizational goals as technologies evolve.
- When provided access to personal, proprietary, and/or otherwise confidential data, maintain such data in the strictest confidence and follow procedures to ensure the privacy, security, and proper use of data.
Qualifications
Bachelor's degree in a related field or equivalent experience.
- Experience working in academic and research settings.
- Experience supporting AI-driven research in open and secure computing environments.
- Familiarity using and administering data-transfer technologies such as Globus that facilitate the transfer of large datasets.
- Experience using and supporting parallel file systems that are commonly used in HPC/AI systems.
- Experience supporting unstructured data in HPC/AI environments.
Schedule
On-call rotation is a mandatory facet of this role, requiring infrequent off-hour and weekend duty.