AI/HPC Cluster Design Engineer
The Role
We are seeking an experienced AI systems engineer to design scalable AI/HPC clusters with specific focus on compliance with customer and/or design requirements. This role involves reviewing and selecting compute, storage, networking, and power delivery components and solutions to optimize performance and reliability across global deployments. You will collaborate with cross-functional teams to deliver cutting-edge infrastructure for AI and high-performance computing workloads.
The Person
An experienced systems engineer with a strong background in HPC, AI systems, and cluster engineering. You bring deep technical knowledge of compute, power, and networking components, a strategic mindset for system-level design, and the ability to collaborate across diverse technical domains. You thrive in fast-paced environments and are passionate about building efficient, scalable, and reliable compute platforms.
Key Responsibilities
- System Architecture & Design
- Design scalable AI/HPC clusters including compute, storage, and networking
- Evaluate and select CPUs, GPUs, accelerators, interconnects, and memory configurations for optimal cluster performance
- Network
- Design network topologies to maximize overall cluster performance
- Understand the network performance needs of different types of workloads
- Understand advantages and performance trade-offs of network topologies for AI/HPC clusters
- Storage
- Design and optimize storage solutions to maximize AI/HPC cluster performance
- Understand advantages and performance trade-offs of cluster storage solutions, e.g. Lustre, Ceph, etc.
- Collaboration
- Work across multiple organizations with subject matter experts from hardware, software, network, data center, and operations teams to deliver scalable, efficient, and reliable compute infrastructure
Qualifications
- Experience in HPC, AI systems/clusters, or data center engineering
- Strong understanding of rack and cluster design
- Knowledge of GPU/CPU architectures, PCIe, UALink, InfiniBand, and Ethernet networking
- Familiarity with AI/ML frameworks and workload characteristics
- Excellent problem-solving, communication, and documentation skills
Preferred Qualifications
- Experience in HPC, AI infrastructure, or data center systems engineering
- Experience designing power delivery solutions for racks and data centers
- Contributions to open-source HPC or AI infrastructure projects
Academic Credentials
Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science or related field.
This role is not eligible for Visa Sponsorship.