Senior AI GPU Deployment Engineer
About the role
We are seeking an experienced Senior AI GPU Deployment Engineer to plan, deploy, and operationalize large-scale GPU AI infrastructure environments. This role delivers production-grade GPU clusters supporting AI training, inference, and high-performance computing workloads in our hyperscale data centers.
The ideal candidate brings deep technical expertise in GPU infrastructure, network fabrics, storage, automation, and Linux systems administration, combined with strong execution and troubleshooting skills.
Responsibilities
- Deploy, integrate, and validate multi-rack GPU-based compute platform deployments
- Deploy fabric configuration engines (Subnet Manager), observability platforms (UFM), and validate interconnect and fabric performance (nccl)
- Collaborate with network engineering team on topology implementation and optimization and storage engineering team on deployment and integration of high-performance storage environments supporting AI workloads (e.g. VAST Data)
- Configure settings and manage firmware updates for GPUs, NICs, BMC, BIOS, and other components across large-scale clusters
- Contribute to infrastructure-as-code automation development for cluster provisioning and lifecycle management
- Contribute to improving and documenting repeatable deployment methodologies and scalable operational standards
- Query and analyze deployment outcomes using SQL for diagnostics and operational reporting
Requirements
- Bachelor's degree in Computer Science, Engineering, IT, or related field (or equivalent experience)
- 5+ years of infrastructure engineering or datacenter deployment experience
- 3+ years deploying large-scale AI, HPC, or GPU infrastructure
- Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments
Skills
Strong expertise with:
- GPU architectures
- InfiniBand (NDR/XDR) and Ethernet GPU fabrics (Spectrum-X)
- NVLink, NVSwitch, and GPU-direct technologies
- Canonical MaaS and automated provisioning systems
- VAST Data or similar high-performance storage platforms
- Linux systems administration for HPC/AI workloads
- Infrastructure-as-Code and configuration management (Ansible)
- Python, Shell, and SQL for infrastructure automation and diagnostics
Strong understanding of:
- RDMA, RoCE, and lossless Ethernet fabrics
- Cluster automation, observability, and lifecycle management
Benefits
- Competitive pay plus meaningful, lasting impact on the work you do
- Career growth in one of the world's fastest-growing industries
- Industry leadership by helping power the infrastructure behind AI and high-performance computing
- Entrepreneurial culture where your ideas matter and you can help shape our future
About Us
At 5C, we believe great people build great companies. We're more than a workplace—we're a team of builders, innovators, and problem-solvers united by a shared purpose: creating infrastructure that powers the technologies transforming our world. Your voice matters here, and we encourage fresh ideas at every level.