Senior Storage Software Engineer - DGX Cloud
NVIDIA · Santa Clara, CA · Today
EngineeringFull-time
What You’ll Be Doing
- Contribute to open-source file systems.
- Contribute code to open-source parallel and distributed file systems, and distributed object storage.
- Upstream fixes and features, and engage directly with the upstream communities and maintainers.
- Contribute code to open-source parallel and distributed file systems, and distributed object storage.
- Serve as a hands-on storage software lead.
- Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it.
- Make the final technical calls on storage deliveries against measurable targets.
- Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it.
- Triage and troubleshoot at scale.
- Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery.
- Validate architecture and capabilities.
- Validate storage architecture, capabilities, performance, and durability.
- Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets.
- Validate storage architecture, capabilities, performance, and durability.
- Recommend configuration, tuning, and guidelines.
- Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them.
- Partner broadly.
- Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture.
- Work AI-first.
- Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations.
What We Need To See
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience.
- Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale.
- Contributions to open-source projects involving a distributed or parallel file system.
- You are fully engaged in engineering tasks.
- You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others.
- You are fully engaged in engineering tasks.
- Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.
- Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance.
- Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python;
- Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python;
- Comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath).
- Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python;
- Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI).
- Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI).
- Strong written and verbal communication;
- Strong written and verbal communication;
- capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers.
- Strong written and verbal communication;
- Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build.
- Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build.
- 100% hands-on engineering.
- 100% hands-on engineering.
- You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating.
- 100% hands-on engineering.
Ways To Stand Out From The Crowd
- Maintainers or sustained contributions to widely used public projects.
- Maintainers or sustained contributions to widely used public projects.
- Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks.
- Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks.
- Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience.
- Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience.
- Kubernetes and CSI driver development for storage.
- Kubernetes and CSI driver development for storage.
- Hands-on experience with SPDK, libfabric, or FUSE performance optimization.
- Hands-on experience with SPDK, libfabric, or FUSE performance optimization.