Jobs · Engineering · California

Senior Software QA Test Development Engineer - Diagnostics

NVIDIA · Santa Clara, CA · 1 mo ago
EngineeringFull-time

About the role

NVIDIA is the world leader in GPU Computing, passionate about markets including gaming, automotive, vision, HPC, datacenters, networking, and our traditional OEM business. As the ‘AI Computing Company,’ NVIDIA GPUs power Deep Learning software frameworks, analytics, data centers, and autonomous vehicles. We are seeking an outstanding individual who thrives in a diverse work environment, possesses strong interpersonal skills, and has a passion for continuous process improvement to join our platform SWQA team.

Responsibilities

  • Develop and execute NVIDIA HGX/DGX/MGX platform test plans on servers, OS, firmware, and CUDA software stack from design documents.
  • Install and test various system operating systems, server firmware, and software stacks.
  • Drive root cause analysis for reliability and validation test failures to identify and mitigate issues.
  • Build, develop, and debug server and OS-level automation for front-end and back-end frameworks and tests.
  • Review partner and supplier test results and prescribe additional reliability testing on components, servers, and packaging as needed.
  • Work in an agile software development team with high production quality standards.
  • Manage bug lifecycle and collaborate with cross-functional teams to drive solutions.

Requirements

  • Bachelor’s Degree (or equivalent experience) in a STEM field (Science, Technology, Engineering, Math, or Physics).
  • 5+ years of proven experience; or a master’s degree.
  • Proven experience in OS and server-level automation, CI/CD processes, and DevOps using Python, Shell, Ansible, Jenkins, C/C++, Java, or JavaScript.
  • Strong server and Linux (Ubuntu, RedHat, CentOS, SuSE, Fedora, etc.) troubleshooting and debugging experience in bare-metal and KVM/VMWare/Hyper-V environments.
  • Good knowledge and hands-on experience with model testing, AI tools/frameworks (TensorFlow, PyTorch, Cursor, etc.), NLP, and LLM benchmarking.
  • Experience using AI development tools for test plan creation, test case development, and automation.
  • Strong experience with firmware, BMC/OpenBMC, network protocols, enterprise storage devices, PCIe buses, IO sub-devices, CPU, memory, ACPI, UEFI spec, and Redfish (huge plus).
  • Proven experience with GitHub/GitLab/Gerrit, PXE, SLURM, Stack/Kubernetes/Docker (huge plus).

Skills

  • AI-related tools, LLM, and NLP.
  • Experience working with NVIDIA GPU hardware (strong plus).
  • Solid understanding of virtualization in Linux (KVM, Docker orchestrated with Kubernetes).
  • Background in parallel programming, ideally CUDA/OpenCL (plus).

Pay

Base salary range: $140,000 – $224,250 USD for Level 3, and $168,000 – $270,250 USD for Level 4. Eligible for equity and benefits.

Similar jobs