Senior AI Compute Engineer - NVIS
NVIDIA · Santa Clara, CA · 1 mo ago
EngineeringFull-time
About the role
NVIDIA is seeking a Senior AI Compute Engineer to join its Infrastructure Specialists team. The role involves working with academic, commercial, and government groups worldwide who utilize NVIDIA products for deep learning, data analytics, and AI Compute systems.
Responsibilities
- Deploy, manage, and validate AI Compute/HPC infrastructure in Linux-based environments for new and existing customers.
- Work with customers, partners, and internal teams to analyze, define, and implement large-scale AI Compute projects.
- Interact with customers during planning calls and provide hands-on implementation support.
- Handover-related documentation and perform knowledge transfers to support customers as they deploy sophisticated systems.
- Provide feedback to internal teams, including reporting bugs, documenting workarounds, and suggesting improvements.
Requirements
- 8+ years of in-depth support and deployment services, solving problems for hardware and software products.
- Knowledge and experience with Linux system administration, process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, network-routing/advanced networking (tuning and monitoring).
- Experience with cluster management and provisioning technologies for bare-metal servers (bonus credit for BCM (Base Command Manager)).
- Scripting proficiency (Bash, Python, Ansible, etc.).
- Excellent interpersonal skills and the ability to resolve customer issues promptly.
- Strong organizational skills and ability to prioritize tasks without direct supervision.
- Experience with schedulers such as SLURM, LSF, UGE, etc.
- Ability to travel to customer sites within the United States up to 20% of the time.
- Experience with benchmarking tools such as HPL, NCCL tests, MLPerf as well as Kubernetes experience.
Preferred Skills
- InfiniBand experience.
- Experience with GPU (Graphics Processing Unit) focused hardware/software.
- Experience with MPI (Message Passing Interface).
- Familiarity with storage technologies such as Lustre or GPFS.
- Familiarity with OEM GPU platforms.
Benefits
The base salary range for this position is $148,000 - $235,750 for Level 4 and $184,000 - $287,500 for Level 5. Additional compensation includes equity and benefits.
Pay
Base salary is determined based on location, experience, and the pay of employees in similar positions.
Schedule
Travel to customer sites within the United States may be required up to 20% of the time.