Jobs · Design · California

Senior Deep Learning Communication Architect

Thomas To · Santa Clara, CA · Today
DesignFull-time

What You'll Be Doing

  • Optimizing communication performance: Identify and eliminate bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
  • Designing efficient communication protocols: Develop and implement communication algorithms and protocols tailored for deep learning workloads, minimizing communication overhead and latency.
  • Hardware and software co-craft: Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC, NVSHMEM).
  • Exploring innovative communication technologies: Research and evaluate new communication technologies and techniques to enhance the performance and scalability of deep learning systems.
  • Developing and implementing solutions: Build proofs-of-concept, conduct experiments, and perform quantitative modeling to validate and deploy new communication strategies.

What We Need To See

  • A Ph.D., Masters, or BS in Computer Science (CS), Electrical Engineering (EE), Computer Science and Electrical Engineering (CSEE), or a closely related field or equivalent experience.
  • 6+ years of experience in Building DNNs, Scaling of DNNs, Parallelism of DNN frameworks, or deep learning training and inference workloads.
  • Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware.
  • Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and FSDP.
  • Understanding of the emerging serving architectures like Disaggregated Serving and inference servers like Dynamo and Triton.
  • Proficiency in developing code for one or more deep neural network (DNN) training and Inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, SGLang.
  • Strong programming skills in C++ and Python.
  • Familiarity with GPU computing, including CUDA and OpenCL, and familiarity with InfiniBand and RoCE networks.

Ways To Stand Out From The Crowd

  • Prior contributions to one or more DNN training and Inference frameworks as part of your previous work experience.
  • Deep understanding and contributions to the scaling of LLMs on large-scale systems.

Pay

Base salary range: 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5. You will also be eligible for equity and benefits.

Similar jobs