Senior Manager, Storage Engineering
NVIDIA · Santa Clara, CA · Yesterday
EngineeringFull-time
What You’ll Be Doing
- Lead petabyte-scale storage deployments across NVIDIA's on-premises data centers and major CSPs (AWS, Azure, GCP), owning the end-to-end lifecycle from design and procurement through physical installation, configuration, and production hand-off.
- Engineer and maintain automation pipelines that integrate deployed storage systems with interdependent tooling — including CMDB (asset tracking and dependency mapping), configuration management platforms, and observability stacks, ensuring a single source of truth across the entire storage fleet.
- Define and drive continuous improvement initiatives focused on data center efficiency: optimizing DC power consumption (PUE impact, drive density, power shelf utilization) and minimizing rack space footprint through high-density hardware selection and intelligent workload placement strategies.
- Develop and maintain self-service tools and dashboards that enable internal customers (EDA, Manufacturing, Software Engineering teams) to track real-time storage capacity availability, consumption trends, and projected growth, reducing friction and improving planning accuracy.
- Coordinate and partner with partner infrastructure teams — networking, compute, cloud, and Data center to present a unified capacity view, resolve cross-domain bottlenecks, and contribute to NVIDIA's holistic infrastructure capacity planning process.
- Manage vendor relationships and lead hardware refresh and EOL planning cycles across the storage portfolio, balancing performance requirements, cost efficiency, and supply chain constraints.
- Recruit, mentor, and grow a team of storage deployment engineers; establish engineering standards, runbooks, and on-call practices to ensure a high operational bar.
- Partner with architecture and security teams to evaluate new storage technologies, drive POCs, and translate findings into production-ready deployment standards.
- Own and engineer the full hardware and data lifecycle — from procurement, rack integration, and initial provisioning through capacity expansion, decommissioning, and secure data destruction — ensuring compliance, auditability, and zero unplanned data loss at every stage.
- Define and govern data lifecycle management policies in close collaboration with key stakeholders across Chip Design and Software Engineering teams, aligning retention, tiering, archival, and deletion standards to business workflows, improving overall storage efficiency, and reducing cost by eliminating stale or redundant data across the fleet.
What We Need To See
- BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
- 12+ years of overall experience in large-scale storage architecture, operations, production engineering, or infrastructure.
- 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
- Deep, protocol-level knowledge of enterprise storage systems spanning block (/NVMe-oF), file (NFS, SMB/CIFS), object (S3-compatible), and high-performance parallel file systems (GPFS/IBM Spectrum Scale, Lustre).
- Hands-on experience deploying and operating solutions from multiple major storage vendors, including NetApp (ONTAP, StorageGRID), Pure Storage (FlashArray, FlashBlade), Cloudian HyperStore, and DDN (EXAScaler, A³I).
- Solid understanding of storage hardware internals — drive types and endurance profiles (NVMe, SAS, SATA, QLC/TLC/MLC NAND), controller architectures, shelf and enclosure design, cabling standards, and failure domain planning.
- Working knowledge of bare-metal server hardware (rack units, HBA/NIC selection, BMC/iDRAC/iLO, firmware management) and data center network hardware relevant to storage connectivity (ToR switches, fiber and copper interconnects, SFP/QSFP optics).
- In-depth expertise in observability tooling: building and maintaining Prometheus exporters, Grafana dashboards, and alerting rulesets for storage fleet health, performance SLOs, and capacity burn-rate tracking.
- Strong configuration management skills using Ansible (playbook authoring, role design, inventory management, Ansible Tower/AAP) for automated provisioning and day-2 operations of storage systems at scale.
- Scripting and automation proficiency in Python and/or Bash; experience integrating with REST APIs to drive CMDB updates, provisioning workflows, and reporting pipelines.
- Promising track record of managing large-scale, multi-vendor storage environments (100 PB+) in a fast-paced, high-availability production setting.