Jobs · Information Technology · Massachusetts

Senior HPC Support Engineer, InfiniBand - NVLink

NVIDIA · Westford, MA · Today
Information TechnologyFull-time

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

About the role

We are seeking a highly motivated Senior HPC Support Engineer focusing on InfiniBand and NVLink technology, passionate about data center and networking technologies, to provide comprehensive solutions for sophisticated installations, maintenance, or operations for a broad scope of groundbreaking networking products. As a primary point of contact for our customers, you will assist them with technical questions, debugging, and resolving their issues. As a member of our NVIDIA Experience (NVEX) Global Technical Support team, you are a conscientious, proficient communicator who is fundamentally interested in taking ownership in resolving issues while ensuring a high level of customer satisfaction is maintained and delivered. A significant part of the role is also to collaborate with Engineering, Marketing, and Support teams regularly on technical issues.

Responsibilities

  • Ability to resolve sophisticated customer concerns and technical issues through meticulous research, reproduction, and solving problems for customers installing our products and supporting systems using Linux Operating Systems (Multi-distro), with a focus on NVIDIA InfiniBand, NVLink, and GPU Technology and our End-to-End Solutions.
  • Responding to customer product support inquiries via telephone, email, or conference calls.
  • Resolving customer issues during installation, operation, maintenance, or product application or interoperability with other vendors.
  • Participate in multi-functional team meetings and provide feedback to engineering and marketing regarding product requirements, customer experience, support tools, etc.
  • As a technical resource, develop, re-define, and document standard methodologies to provide to internal teams (Support/R&D) for support process and improvements.

Requirements

  • 5+ years in providing in-depth Customer Support and debugging experience focused on large-scale networking and AI Infrastructure environments and products such as InfiniBand, Ethernet, and GPU, including infrastructure performance.
  • Strong interpersonal skills and ability to prioritize/multi-task easily with limited supervision.
  • Shown use of established Agentic AI technologies (Claude, Codex, Cursor) in day-to-day job responsibilities.
  • Excellent verbal and written English skills.
  • An academic degree from an accredited university or college in Networking, Computer Science/Engineering, or Electrical/IT (or equivalent experience).
  • Intellectual curiosity, positive attitude, flexibility, analytical ability, self-motivation, and team orientation, including professional-level communication skills, interpersonal skills, with a passion for solving problems.

Skills

  • Networking Technology, protocols, and routing including IP, L2, and L3 on a CCNP/CompTIA Networking+ and Cloud+ level.
  • Able to debug networking protocols using tools such as TCPDUMP and Wireshark or similar packet generation and analysis tools.
  • Configuration and operational expertise with traditional network switch/router and Open platforms.
  • Linux OS System Administration, Networking, and Performance on a LFCS/RHCSA level.
  • Containerized solutions experience on a level of DCA and/or CKA, Virtualization (KVM/ESXi), and Cloud Infrastructure (AWS/OCI) Technologies.

Ways to stand out

  • Knowledge and working experience with InfiniBand, RDMA/RoCEv2, and GPU Technology.
  • Clustering or HPC Data-Center technologies including Upper Layer Protocols (i.e., NCCL, MPI, Slurm/SchedMD).
  • Shell scripting (Bash/Python).
  • Linux, Networking, and NVIDIA AI Infrastructure and Operations Certifications such as CCNP, CCIE, JNCIE-DC/ENT, RHCE, LFCS, NCP-AII/AIO/AIN.

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 108,000 USD - 172,500 USD for Level 3, and 120,000 USD - 207,000 USD for Level 4. You will also be eligible for equity and benefits.

Benefits

NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family at www.nvidiabenefits.com.

Similar jobs