Jobs · Engineering · California

Senior Storage Software Engineer - DGX Cloud

NVIDIA · Santa Clara, CA · Today
EngineeringFull-time

What You’ll Be Doing

  • Contribute to open-source file systems.
    • Contribute code to open-source parallel and distributed file systems, and distributed object storage.
      • Upstream fixes and features, and engage directly with the upstream communities and maintainers.
  • Serve as a hands-on storage software lead.
    • Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it.
      • Make the final technical calls on storage deliveries against measurable targets.
  • Triage and troubleshoot at scale.
    • Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery.
  • Validate architecture and capabilities.
    • Validate storage architecture, capabilities, performance, and durability.
      • Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets.
  • Recommend configuration, tuning, and guidelines.
    • Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them.
  • Partner broadly.
    • Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture.
  • Work AI-first.
    • Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations.

What We Need To See

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience.
    • Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale.
  • Contributions to open-source projects involving a distributed or parallel file system.
    • You are fully engaged in engineering tasks.
      • You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others.
  • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.
    • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.
  • Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python;
    • Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python;
      • Comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).
  • Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI).
    • Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI).
  • Strong written and verbal communication;
    • Strong written and verbal communication;
      • capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers.
  • Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build.
    • Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build.
  • 100% hands-on engineering.
    • 100% hands-on engineering.
      • You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating.

Ways To Stand Out From The Crowd

  • Maintainers or sustained contributions to widely used public projects.
    • Maintainers or sustained contributions to widely used public projects.
  • Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks.
    • Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks.
  • Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience.
    • Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience.
  • Kubernetes and CSI driver development for storage.
    • Kubernetes and CSI driver development for storage.
  • Hands-on experience with SPDK, libfabric, or FUSE performance optimization.
    • Hands-on experience with SPDK, libfabric, or FUSE performance optimization.

Similar jobs