Jobs · Engineering

Principal Software Engineer – Large-Scale LLM Memory and Storage Systems

NVIDIA · Massachusetts, United States · 1 wk ago
RemoteRemoteEngineeringFull-time

The NVIDIA Dynamo Principal Systems Engineer position is dedicated to defining the vision and roadmap for memory management of large-scale Large Language Model (LLM) and storage systems. The role involves designing and evolving a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage. Key responsibilities include architecting and implementing deep integrations with leading LLM serving engines, focusing on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters. Additionally, the role requires co-designing interfaces and protocols that enable disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage.

About the role

  • Design and evolve a unified memory layer spanning GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage.
  • Architect and implement deep integrations with leading LLM serving engines, focusing on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters.
  • Co-design interfaces and protocols enabling disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage.
  • Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools.
  • Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives).

Responsibilities

  • Design and evolve a unified memory layer that spans GPU memory, pinned host memory, RDMA-accessible memory, SSD tiers, and remote file/object/cloud storage.
  • Architect and implement deep integrations with leading LLM serving engines, focusing on KV-cache offload, reuse, and remote sharing across heterogeneous and disaggregated clusters.
  • Co-design interfaces and protocols enabling disaggregated prefill, peer-to-peer KV-cache sharing, and multi-tier KV-cache storage.
  • Partner closely with GPU architecture, networking, and platform teams to exploit GPUDirect, RDMA, NVLink, and similar technologies for low-latency KV-cache access and sharing across heterogeneous accelerators and memory pools.
  • Mentor senior and junior engineers, set technical direction for memory and storage subsystems, and represent the team in internal reviews and external forums (open source, conferences, and customer-facing technical deep dives).

Requirements

  • Masters or PhD or equivalent experience with 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python.
  • Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency.
  • Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters.
  • Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput.
  • Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams.

Qualifications

  • Masters or PhD or equivalent experience with 15+ years of experience building large-scale distributed systems, high-performance storage, or ML systems infrastructure in C/C++ and Python.
  • Deep understanding of memory hierarchies (GPU HBM, host DRAM, SSD, and remote/object storage) and experience designing systems that span multiple tiers for performance and cost efficiency.
  • Hands-on experience with networked I/O and RDMA/NVMe-oF/NVLink-style technologies, and familiarity with concepts like disaggregated and aggregated deployments for AI clusters.
  • Strong skills in profiling and optimizing systems across CPU, GPU, memory, and network, using metrics to drive architectural decisions and validate improvements in TTFT and throughput.
  • Excellent communication skills and prior experience leading cross-functional efforts with research, product, and customer teams.

Skills

  • Experience with open-source LLM serving or systems projects focused on KV-cache optimization, compression, streaming, or reuse.
  • Publications or patents in areas such as LLM systems, memory-disaggregated architectures, RDMA/NVLink-based data planes, or KV-cache/CDN-like systems for ML.

Benefits

  • Competitive salaries and a comprehensive benefits package.
  • Opportunities to work with highly talented and innovative colleagues.
  • Opportunities to contribute to cutting-edge technology and make a significant impact.

Pay

Base salary range: $272,000 - $431,250 USD.

Schedule

NVIDIA offers flexible scheduling options to accommodate individual needs and preferences.

Contact

To apply for this position, please visit our careers page at [insert link]. Applications are accepted until January 13, 2026.

Similar jobs