Senior Deep Learning Communication Architect
NVIDIA · Redmond, WA · 2 days ago
EngineeringFull-time
About the Role
The software architecture group at NVIDIA is seeking a Deep Learning Communication Architect to scale DNN models and training/inference frameworks to systems with hundreds of thousands of nodes. This role involves optimizing communication performance, designing efficient communication protocols, and collaborating with hardware and software teams to leverage high-speed interconnects and communication libraries.
Responsibilities
- Optimizing communication performance by identifying and eliminating bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
- Designing and implementing communication algorithms and protocols tailored for deep learning workloads to minimize overhead and latency.
- Collaborating with hardware and software teams to craft systems that effectively utilize high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC, NVSHMEM).
- Researching and evaluating new communication technologies to enhance the performance and scalability of deep learning systems.
- Developing proofs-of-concept, conducting experiments, and performing quantitative modeling to validate and deploy new communication strategies.
Requirements
- Ph.D., Masters, or BS in Computer Science (CS), Electrical Engineering (EE), Computer Science and Electrical Engineering (CSEE), or a closely related field, or equivalent experience.
- 6+ years of experience in building DNNs, scaling DNNs, parallelism of DNN frameworks, or deep learning training and inference workloads.
- Experience evaluating and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware.
- Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and FSDP.
- Understanding of emerging serving architectures like Disaggregated Serving and inference servers like Dynamo and Triton.
- Proficiency in developing code for DNN training and inference frameworks such as PyTorch, TensorRT-LLM, vLLM, SGLang.
- Strong programming skills in C++ and Python.
- Familiarity with GPU computing, including CUDA and OpenCL, and InfiniBand and RoCE networks.
Qualifications
- Prior contributions to DNN training and inference frameworks.
- Deep understanding and contributions to scaling LLMs on large-scale systems.
Pay
The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.