Research Computing Engineer
About the Role
As the strategic anchor for the SCU High-Performance Computing (HPC) environment, the Research Computing Engineer serves as the primary technical partner between Santa Clara University's research community and its computational infrastructure. This role focuses on driving the "outer relationship" with users—faculty, researchers, and students—to deeply understand, architect, and translate complex computational workflows into scalable technical solutions. Rather than simply maintaining existing infrastructure, the Research Computing Engineer provides strategic leadership, engineers robust processes, and drives long-term planning to ensure the HPC ecosystem proactively evolves alongside the university's research mission. The ideal candidate is highly curious, creative, tenacious, and entirely self-directed. They bring an advanced technical toolkit combined with the leadership capacity to identify, define, and resolve complex systemic and human workflows independently and collaboratively.
Essential Duties and Responsibilities
Strategic Leadership, Planning, and Research Facilitation
- Lead the strategic roadmap and long-term capacity planning for SCU's HPC infrastructure, partnering with the Dean and academic stakeholders to forecast future computational demands.
- Own the full-cycle consultation process with researchers and faculty, translating cutting-edge academic requirements into scalable, robust technical solutions.
- Architect and implement proactive infrastructure enhancements to optimize application pipelines for emerging domains, including AI, Machine Learning, Data Science, and GPU-accelerated processing.
- Evaluate, recommend, and drive the adoption of emerging technologies and external integrations with national academic computing resources to expand institutional research capabilities.
- Provide high-level technical leadership and programming support to resolve complex, multi-disciplinary computational challenges across university departments.
Process Innovation, Governance, and Training
- Design, implement, and institutionalize standard operating procedures (SOPs) and automated workflows for user onboarding, resource allocation, and system governance.
- Develop lifecycle management processes for scientific software deployment, cluster usage auditing, and data management.
- Establish system performance metrics and reporting frameworks to showcase HPC utilization and research impacts to executive leadership.
- Design and spearhead comprehensive training programs, advanced workshops, and modern digital documentation to cultivate a self-sustaining research culture.
- Lead institutional initiatives to train users in modern code-management, AI-assisted coding, CI/CD, and version control best practices (e.g., Git/GitHub).
Full-Cycle HPC Infrastructure Architecture & Operations
- Own the deployment lifecycle, configuration, and optimization of specialized scientific software, compilers, containerized environments, and shared libraries.
- Lead the architecture, fine-tuning, and policy creation for workload managers and cluster schedulers (e.g., Slurm) to ensure optimal, equitable resource distribution.
- Lead the scaling and operational strategy for parallel storage and distributed file systems (e.g., BeeGFS, Lustre), ensuring total data integrity, high-throughput performance, and business continuity.
- Lead network design and execution within the HPC environment, overseeing high-speed fabrics (e.g., InfiniBand) and complex VLAN configurations.
Security Frameworks and Systems Stewardship
- Architect and enforce comprehensive security frameworks, including server hardening, access controls, and vulnerability mitigation protocols to safeguard sensitive research data.
- Proactively monitor, analyze, and optimize system telemetry to perform deep root-cause analysis on complex hardware and software bottlenecks.
- Stay current with emerging trends in HPC, AI, and cloud technologies to inform long-term infrastructure planning.
Qualifications
Knowledge, Skills, and Abilities
- Full-Cycle Ownership & Strategy: Demonstrated ability to independently design, implement, and govern enterprise-grade computational environments and workflows.
- Technical Mastery: Advanced, hands-on mastery of Linux systems administration, automated provisioning, and comprehensive package management systems.
- Scripting and Automation: Demonstrated experience writing and debugging complex scripts in Bash, Python, or Ansible.
- HPC Ecosystem Expertise: Deep knowledge of workload managers (Slurm), container technologies (Docker, Apptainer), and version control. Proven success implementing distributed file systems (BeeGFS, Lustre) and environment module systems.
- Cybersecurity Leadership: Advanced understanding of cybersecurity principles, encryption standards, and risk-mitigation strategies unique to open research cluster environments.
- Communication: Exceptional interpersonal and verbal communication skills; ability to explain complex technical concepts to non-technical users.
- Problem Solving: Strong analytical skills with a proactive approach to identifying and resolving technical and human issues.
Experience and Education
- Education: Bachelor's degree in Computer Science, Engineering, or a highly quantitative field required. Advanced degree (MS or PhD) strongly preferred to bridge the gap during high-level research consultations.
- Experience: 8–10 years of progressively responsible experience in Information Technology operations and system design, ideally within an academic, government lab, or corporate R&D research setting.
- Preferred Experience: 5+ years of experience explicitly leading, architecting, and supporting multi-node HPC cluster environments.
Pay
$115,200 - $129,600 per year
Schedule
- Hybrid eligible: Regular on-site presence required, typically at least 3–4 days a week, depending on task requirements.
- Standard work hours are 9 am – 5 pm Pacific, with occasional evening or weekend work required for system maintenance or outage mitigation.
- Telecommute employees are required to perform their work within one of the following states: California, Nevada, Oregon, Washington, Arizona, or Illinois.
Physical Demands and Work Environment
- Routinely perform server installation, troubleshooting, and repairs at the data center, including lifting or moving objects up to 50 pounds.
- Considerable time spent at a desk using a computer terminal.
- Ability to meet in-person with researchers and colleagues on the Santa Clara University campus.
- Occasional exposure to data center conditions, including equipment noise (>80dB), high voltage electricity, and varying temperatures.