Principal Deep Learning Communication Architect
NVIDIA · Austin, TX · 1 wk ago
EngineeringFull-time
What You’ll Be Doing
- Architecture Leadership: Define the long-term technical roadmap for communication libraries across NVIDIA’s next-generation platforms.
- You will ensure the seamless scaling of models to clusters comprising hundreds of thousands of nodes.
- AI Communication Library Design: Lead the development of next-generation communication primitives and collective algorithms.
- This includes optimizing for heterogeneous interconnects such as NVLink, Spectrum-X (Ethernet), and Quantum-X (InfiniBand).
- Application-Communication Library Co-Design: Partner with application developers to architect and implement specialized communication primitives.
- You will ensure that AI and HPC libraries—including NCCL, NIXL, NVSHMEM, UCC, and UCX—evolve to meet the requirements of trillion-parameter and Agentic AI.
- Hardware/Software Co-Design: Collaborate with silicon architects and software engineers to influence hardware specifications for next-generation networking, ensuring they meet the evolving demands of trillion-parameter LLMs and Agentic AI.
- Quantitative Modeling: Develop high-fidelity analytical models and simulators to predict system behavior under emerging workloads.
What We Need To See
- Ph.D. or M.S. in Computer Science, Electrical Engineering, or a related field (or equivalent experience), with 12+ years of industry experience in high-performance computing (HPC) or distributed deep learning.
- Deep understanding of 3D parallelism (Data, Tensor, Pipeline) and advanced strategies including Context Parallelism, Expert Parallelism, and Zero Redundancy Optimizer (ZeRO) variants.
- Deep technical proficiency with NCCL, UCX, UCC, NVSHMEM, or MPI. Experience with RDMA, RoCE, and low-level InfiniBand verbs is required.
- Advanced knowledge of high-throughput inference engines and schedulers, specifically TensorRT-LLM, vLLM, SGLang, and NVIDIA Dynamo.
- Expert knowledge of the NVIDIA GPU memory hierarchy (HBM3e/HBM4, L2 cache) and CUDA programming models.
How To Stand Out
- Hands-on experience developing within Megatron-Core, DeepSpeed, or JAX/XLA, with an understanding of how these frameworks interact with low-level communication runtimes is a plus.
- Significant upstream contributions to major open-source projects (e.g., PyTorch Distributed, KServe, or Ray).
- A proven track record of deploying and optimizing models on NVIDIA platforms or similar rack-scale systems.
- A strong portfolio of patents or papers in top-tier systems/architecture venues (e.g., ISCA, ASPLOS, NeurIPS, SC).
Pay & Benefits
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD. You will also be eligible for equity and benefits.
Application Information
Applications for this job will be accepted at least until April 18, 2026.