Systems Engineer
Haystack · Princeton, NJ · 3 wk ago
HybridFull-time
About the role
Design, maintain, troubleshoot, and refine advanced HPC/AI cluster infrastructure, including high-performance interconnects and cluster schedulers. Partner with colleagues to align designs for scratch filesystems and data management with cluster designs. Develop data-transfer pathways and networks to support AI-driven computing workloads. Define and drive the institutional technical strategy for advanced AI and data-intensive HPC. Promote institution-wide collaboration, advising faculty, researchers, and vendors on emerging trends in AI-enabled research computing. Monitor HPC clusters, networks, and storage systems for abnormalities, and resolve issues.
Requirements
- 10+ years of strong experience managing advanced research computing systems
- Strong expertise with Linux system administration, installation, and troubleshooting
- Advanced experience writing scripts in languages such as bash, Python, and/or Perl
- Proficient in managing networking in HPC environments
- Experience supporting scheduling and managing jobs (SLURM) in large-scale computing environments
- Strong oral and written communication skills with the ability to engage peers and diverse stakeholders effectively
Benefits
- Opportunity to work at a leading research institution
- Engage with cutting-edge HPC and AI technologies
- Collaborative and knowledge-sharing environment
- Comprehensive benefits program