Senior Manager, Storage Engineering
NVIDIA AI · Santa Clara, CA · Yesterday
EngineeringFull-time
About the role
The Storage Deployment Manager leads petabyte-scale storage deployments across NVIDIA's on-premises data centers and major CSPs (AWS, Azure, GCP), owning the end-to-end lifecycle from design and procurement through physical installation, configuration, and production hand-off.
Responsibilities
- Lead petabyte-scale storage deployments across NVIDIA's on-premises data centers and major CSPs (AWS, Azure, GCP), owning the end-to-end lifecycle from design and procurement through physical installation, configuration, and production hand-off.
- Engineer and maintain automation pipelines that integrate deployed storage systems with interdependent tooling — including CMDB (asset tracking and dependency mapping), configuration management platforms, and observability stacks, ensuring a single source of truth across the entire storage fleet.
- Define and drive continuous improvement initiatives focused on data center efficiency: optimizing DC power consumption (PUE impact, drive density, power shelf utilization) and minimizing rack space footprint through high-density hardware selection and intelligent workload placement strategies.
- Develop and maintain self-service tools and dashboards that enable internal customers (EDA, Manufacturing, Software Engineering teams) to track real-time storage capacity availability, consumption trends, and projected growth, reducing friction and improving planning accuracy.
- Coordinate and partner with partner infrastructure teams — networking, compute, cloud, and Data center to present a unified capacity view, resolve cross-domain bottlenecks, and contribute to NVIDIA's holistic infrastructure capacity planning process.
- Manage vendor relationships and lead hardware refresh and EOL planning cycles across the storage portfolio, balancing performance requirements, cost efficiency, and supply chain constraints.
- Recruit, mentor, and grow a team of storage deployment engineers; establish engineering standards, runbooks, and on-call practices to ensure a high operational bar.
- Partner with architecture and security teams to evaluate new storage technologies, drive POCs, and translate findings into production-ready deployment standards.
- Own and engineer the full hardware and data lifecycle — from procurement, rack integration, and initial provisioning through capacity expansion, decommissioning, and secure data destruction — ensuring compliance, auditability, and zero unplanned data loss at every stage.
- Define and govern data lifecycle management policies in close collaboration with key stakeholders across Chip Design and Software Engineering teams, aligning retention, tiering, archival, and deletion standards to business workflows, improving overall storage efficiency, and reducing cost by eliminating stale or redundant data across the fleet.
Requirements
- BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
- 12+ years of overall experience in large-scale storage architecture, operations, production engineering, or infrastructure.
- 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
- Deep, protocol-level knowledge of enterprise storage systems spanning block (/NVMe-oF), file (NFS, SMB/CIFS), object (S3-compatible), and high-performance parallel file systems (GPFS/IBM Spectrum Scale, Lustre).
- Hands-on experience deploying and operating solutions from multiple major storage vendors, including NetApp (ONTAP, StorageGRID), Pure Storage (FlashArray, FlashBlade), Cloudian HyperStore, and DDN (EXAScaler, A³I).
- Solid understanding of storage hardware internals — drive types and endurance profiles (NVMe, SAS, SATA, QLC/TLC/MLC NAND), controller architectures, shelf and enclosure design, cabling standards, and failure domain planning.
- Working knowledge of bare-metal server hardware (rack units, HBA/NIC selection, BMC/iDRAC/iLO, firmware management) and data center network hardware relevant to storage connectivity (ToR switches, fiber and copper interconnects, SFP/QSFP optics).
- In-depth expertise in observability tooling: building and maintaining Prometheus exporters, Grafana dashboards, and alerting rulesets for storage fleet health, performance SLOs, and capacity burn-rate tracking.
- Strong configuration management skills using Ansible (playbook authoring, role design, inventory management, Ansible Tower/AAP) for automated provisioning and day-2 operations of storage systems at scale.
- Scripting and automation proficiency in Python and/or Bash; experience integrating with REST APIs to drive CMDB updates, provisioning workflows, and reporting pipelines.
- Promising track record of managing large-scale, multi-vendor storage environments (100 PB+) in a fast-paced, high-availability production setting.
Skills
- Deep knowledge of Kubernetes storage integrations — CSI driver operations, persistent volume lifecycle management, StorageClass design for stateful workloads, and experience running storage-intensive workloads on K8s at scale.
- HPC systems expertise: experience deploying and tuning storage for HPC clusters, including parallel file system performance optimization (stripe tuning, client-side caching, OST/MDT balancing) and co-designing storage architectures with HPC compute and fabric teams.
- Familiarity with modern storage technologies (NVMe, RDMA, DPUs) and their impact on system performance, including kernel-level concepts around I/O subsystems and volumes.
- Familiarity with HPC job schedulers such as IBM LSF or Slurm — understanding how job scheduling behavior (burst patterns, checkpoint I/O, scratch-space usage) drives storage design decisions, and experience implementing storage-aware scheduling policies.
- Experience building or scaling storage for AI/ML or HPC workloads, including hybrid or multi cloud setups (for example AWS S3, Azure Blob, or Google Cloud Storage), as well as on-prem infrastructure.
Benefits
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!
Pay
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 248,000 USD - 396,750 USD.
Schedule
You will also be eligible for equity and benefits.