Jobs · Quality Assurance · California

Senior Software QA Test Development Engineer - Diagnostics

Thomas To · Santa Clara, CA · 1 wk ago
Quality AssuranceFull-time

NVIDIA is the world leader in GPU Computing, passionate about markets including gaming, automotive, vision, HPC, datacenters, and networking, in addition to our traditional OEM business. NVIDIA is also well positioned as the ‘AI Computing Company’, with GPUs powering Deep Learning software frameworks, analytics, data centers, and autonomous vehicles. We have some of the most experienced and dedicated people in the world working for us.

About the role

We are looking for an outstanding individual who thrives in a diverse work environment, has outstanding interpersonal skills, and possesses a strong sense of engagement and continuous process improvement. This candidate must have enterprise server integration, strong Linux experience, reliability testing with various telemetries, scale-out cluster testing, test plan development, a track record in developing AI tools and NLP, and DevOps/CI/CD experience to join our platform SWQA team.

Responsibilities

  • Responsible for the development and execution of NVIDIA HGX/DGX/MGX platform test plans on servers, OS, FW, and CUDA SW stack from design documents.
  • Install and test various system OS, server firmware, and SW stacks.
  • Drive support for root cause analysis on reliability and validation test failures to identify root cause(s) and achieve mitigation.
  • Build, develop, and debug server and OS-level automation front-end and back-end framework and tests.
  • Review partner and supplier test results and prescribe additional reliability testing on components, servers, and packaging as needed.
  • Work in an agile software development team with very high production quality standards.
  • Manage bug lifecycle and collaborate with inter-groups to drive for solutions.

Requirements

  • Bachelor’s Degree (or equivalent experience) in a STEM (Science, Technology, Engineering, Math, or Physics) field.
  • 5+ years of proven experience; or a master’s degree.
  • Proven years of OS and server-level automation, CI/CD process, and DevOps experience using Python, SHELL, Ansible, Jenkins, C/C++, Java, and JavaScript.
  • Strong server and Linux (Ubuntu, RedHat, CentOS, SuSE, Fedora, etc.) troubleshooting and debugging experience in bare-metal and KVM/VMWare/Hyper-V environments.
  • Good knowledge and hands-on experience in model testing, AI tools/frameworks (TensorFlow, PyTorch, Cursor, etc.), NLP, and LLM benchmarking.
  • Experience in using AI development tools for test plan creation, test case development, and test case automation.
  • Strong experience in FW, BMC/OpenBMC, network protocols, internal/external enterprise storage devices, PCIe buses and devices, IO sub-devices, CPU and memory, ACPI, UEFI spec, and Redfish (huge plus).
  • Proven years of experience in GitHub/GitLab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker (huge plus).

Skills

  • AI-related tools, LLM, and NLP.
  • Experience working with NVIDIA GPU hardware (strong plus).
  • Solid understanding of virtualization in Linux (KVM, Docker orchestrated with Kubernetes).
  • Background in parallel programming, ideally CUDA/OpenCL (plus).

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is:

  • 140,000 USD - 224,250 USD for Level 3.
  • 168,000 USD - 270,250 USD for Level 4.

You will also be eligible for equity and benefits.

Similar jobs