Jobs · Engineering · California

Senior Deep Learning Communication Architect

NVIDIA · Santa Clara, CA · 2 wk ago
EngineeringFull-time

About the role

The software architecture group at NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to systems with hundreds of thousands of nodes.

Responsibilities

  • Identify and eliminate bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
  • Design and implement communication algorithms and protocols tailored for deep learning workloads, minimizing communication overhead and latency.
  • Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC, NVSHMEM).
  • Explore and evaluate new communication technologies and techniques to enhance the performance and scalability of deep learning systems.
  • Develop and implement solutions through building proofs-of-concept, conducting experiments, and performing quantitative modeling to validate and deploy new communication strategies.

Requirements

  • A Ph.D., Masters, or BS in Computer Science (CS), Electrical Engineering (EE), Computer Science and Electrical Engineering (CSEE), or a closely related field or equivalent experience.
  • 6+ years of experience in building DNNs, scaling of DNNs, parallelism of DNN frameworks, or deep learning training and inference workloads.
  • Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware.
  • Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and FSDP.
  • Understanding of emerging serving architectures like Disaggregated Serving and inference servers like Dynamo and Triton.
  • Proficiency in developing code for one or more deep neural network (DNN) training and Inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, SGLang.
  • Strong programming skills in C++ and Python.
  • Familiarity with GPU computing, including CUDA and OpenCL, and familiarity with InfiniBand and RoCE networks.

Qualifications

  • Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware.
  • Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and FSDP.
  • Understanding of emerging serving architectures like Disaggregated Serving and inference servers like Dynamo and Triton.
  • Proficiency in developing code for one or more deep neural network (DNN) training and Inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, SGLang.
  • Strong programming skills in C++ and Python.
  • Familiarity with GPU computing, including CUDA and OpenCL, and familiarity with InfiniBand and RoCE networks.

Skills

  • Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware.
  • Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and FSDP.
  • Understanding of emerging serving architectures like Disaggregated Serving and inference servers like Dynamo and Triton.
  • Proficiency in developing code for one or more deep neural network (DNN) training and Inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, SGLang.
  • Strong programming skills in C++ and Python.
  • Familiarity with GPU computing, including CUDA and OpenCL, and familiarity with InfiniBand and RoCE networks.

Benefits

Competitive salaries and a generous benefits package, we are widely considered to be one of the technology world’s most desirable employers.

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

Schedule

You will also be eligible for equity and benefits.

How to Apply

Applications for this job will be accepted at least until May 24, 2026.

Similar jobs