Jobs · Engineering · California

Senior HPC Storage Engineer

NVIDIA · Santa Clara, CA · 1 mo ago
EngineeringFull-time

What You'll Be Doing

Research and analyze existing internal distributed storage services.

Research, design, and implement scalable, next-gen distributed storage services for HPC workloads, optimizing both performance and cost-effectiveness to meet NVIDIA’s growing infrastructure needs.

Develop tooling to automate management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources.

Detail the general procedures and practices, perform technology evaluations, related to distributed file systems.

Collaborate across teams to better understand developers' workflows and capture their infrastructure requirements.

Influence and guide methodologies for building, testing, and deploying applications to ensure efficient performance and resource utilization.

Supporting our researchers to run their flows on our clusters including performance analysis and optimizations of deep learning workflows.

Root cause analysis and suggest corrective action for problems large and small scales.

What We Need To See

  • Bachelor’s degree in Computer Science, Electrical Engineering or related field or equivalent experience.
  • 8+ years of experience designing and/or operating large scale storage infrastructure.
  • Experience analyzing and tuning storage performance for a variety of workloads.
  • Proficient in Centos/RHEL and/or Ubuntu Linux distros including Python programming and bash scripting.
  • In depth understanding of container technologies like Docker, Enroot.

How to Stand Out

  • Distributed Storage Expertise: Extensive experience with parallel and distributed filesystems (Ceph, Weka.io, Vast, Lustre, GPFS) and Linux storage kernel development.
  • GPU & AI Infrastructure: Proficient with NVIDIA GPUs, CUDA programming, and NCCL, including performance benchmarking via MLPerf.
  • Hardware & Storage Engineering: Deep familiarity with storage hardware (HDDs, SSDs, NVMe), enclosures, and specialized appliances like Network Appliance.
  • Advanced Networking: Strong background in Software Defined Networking (SDN) and high-performance networking for AI/HPC clusters.
  • Deep Learning Frameworks: Practical experience applying industry-standard frameworks, specifically PyTorch and TensorFlow.

Pay & Benefits

NVIDIA offers highly competitive salaries and a comprehensive benefits package. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.

Contact Information

Applications for this job will be accepted at least until June 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar jobs

Senior Storage Engineer

Shield Consulting Solutions, Inc.Laurel, MD· 1 mo ago
Information Technology$190k–$200k/yrapply on theapplicantmanager.com

Senior Storage Engineer

Pacific LifeNewport Beach, CA· 1 mo ago
Information Technology$138k–$168k/yrapply on pacificlife.wd1.myworkdayjobs.com