Jobs · Engineering · North Carolina

Senior Software Architect - Deep Learning and HPC Communications

NVIDIA · Durham, NC · 3 wk ago
EngineeringFull-time

About the role

The GPU Communications Libraries and Networking team at NVIDIA is seeking a Senior Software Architect to help co-design next-generation data center platforms and scalable communications software. This role focuses on improving communication performance, designing and implementing new communication technologies, and exploring innovative solutions in hardware and software for our next generation platforms.

Responsibilities

  • Investigate opportunities to improve communication performance by identifying bottlenecks in today's systems.
  • Design and implement new communication technologies to accelerate AI and HPC workloads.
  • Explore innovative solutions in HW and SW for our next generation platforms as part of co-design efforts involving GPU, Networking, and SW architects.
  • Build proofs-of-concept, conduct experiments, and perform quantitative modeling to evaluate and drive new innovations.
  • Use simulation to explore performance of large GPU clusters (think scales of 100s of 1000s of GPUs).

Requirements

  • M.S./Ph.D. degree in CS/CE or equivalent experience.
  • 5+ years of relevant experience.
  • Excellent C/C++ programming and debugging skills.
  • Experience with parallel programming models (MPI, SHMEM) and at least one communication runtime (MPI, NCCL, NVSHMEM, OpenSHMEM, UCX, UCC).
  • Deep understanding of operating systems, computer and system architecture.
  • Solid in fundamentals of network architecture, topology, algorithms, and communication scaling relevant to AI and HPC workloads.
  • Strong experience with Linux.
  • Able and flexible to work and communicate effectively in a multi-national, multi-time-zone corporate environment.

Qualifications

  • Expertise in related technology and passion for what you do.
  • Experience with CUDA programming and NVIDIA GPUs.
  • Knowledge of high-performance networks like InfiniBand, RoCE, NVLink, etc.
  • Experience with Deep Learning Frameworks such as PyTorch, TensorFlow, etc.
  • Knowledge of deep learning parallelisms and mapping to the communication subsystem.
  • Experience with HPC applications.
  • Strong collaborative and interpersonal skills and a proven track record of effectively guiding and influencing within a dynamic and multi-functional environment.

Benefits

Base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is $184,000 - $287,500 for Level 4, and $224,000 - $356,500 for Level 5. You will also be eligible for equity and benefits.

Pay

Base salary range: $184,000 - $287,500 for Level 4, and $224,000 - $356,500 for Level 5.

Schedule

Not specified.

Skills

Expertise in related technology and passion for what you do. Experience with CUDA programming and NVIDIA GPUs. Knowledge of high-performance networks like InfiniBand, RoCE, NVLink, etc. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, etc. Knowledge of deep learning parallelisms and mapping to the communication subsystem. Experience with HPC applications. Strong collaborative and interpersonal skills and a proven track record of effectively guiding and influencing within a dynamic and multi-functional environment.

Benefits

Base salary range: $184,000 - $287,500 for Level 4, and $224,000 - $356,500 for Level 5. Eligible for equity and benefits.

Similar jobs