Senior Solution Engineer, Compute Systems
Full time position within the NVIDIA Experience (NVEX) Solutions Engineering team, supporting NVIDIA’s GPU-accelerated platforms including DGX, HGX, and MGX.
About the role
You will apply the latest AI technologies to triage customer issues, identify solutions, and keep customers delighted. The role involves direct customer support for hardware platform issues and AI/ML workloads in large datacenters, solving problems, and contributing to product and software tooling improvements. Strong Linux expertise, programming skills, and experience with multi-GPU platforms are required. AI is foundational to all solution engineering workflows, including customer support and software development.
Responsibilities
- Provide direct support to NVIDIA Enterprise customers to resolve or advance customer issues.
- Work with engineering teams on customer issues, providing logs, reproduction steps, and other triage information.
- Apply AI to create or update products and support tools.
- Take ownership and drive customer issues from inception to resolution.
- Document customer interactions to enhance the knowledge base.
- Apply agentic AI skills to solving customer issues and software development.
- Occasional work on weekends and holidays to support customers.
Requirements
- Minimum of a BS in Computer Engineering, Electrical Engineering, or equivalent experience.
- At least 10 years of engineering experience with multi-GPU platforms.
- Strong system software expertise (firmware, BIOS, kernel, driver, operating system).
- Solid understanding of Linux and the ability to troubleshoot, optimize, and customize Linux environments for AI/ML workloads.
- Containerized solutions experience with Docker, Kubernetes, and/or Slurm.
- Professional-level communication skills, including adjusting communication to the technical level of the audience and staying calm in high-pressure situations.
- Excellent follow-up and organizational skills, with a passion for problem-solving.
- Proficient in C/C++ programming for platform OS, firmware, BIOS, kernel, and drivers.
- Proficient in Python programming with the ability to build custom tools.
Skills
- Background with parallel programming or GPU acceleration (e.g., CUDA).
- Experience developing in GPU-accelerated, cloud, or virtualized environments.
- Experience analyzing software performance of distributed workloads.
- Clustering or HPC datacenter technologies, including Upper Layer Protocols (e.g., NCCL, MPI).
Pay
Base salary range: $168,000 – $270,250 USD (Level 4) or $200,000 – $322,000 USD (Level 5), depending on location, experience, and position level. Eligible for equity and benefits.