Jobs · Engineering · California

Senior Software Engineer, AI Networking

NVIDIA · Santa Clara, CA · 1 wk ago
EngineeringFull-time

About the role

NVIDIA seeks a senior software engineer to join the AI Networking co-design and benchmark R&D team. In this role, you will build and productize machine learning tools that use ML-based combinatorial optimization and design space exploration (DSE) techniques to optimize AI workloads across large GPU and CPU clusters, ensuring efficient and productive utilization of system resources at data center scale.

The role involves working on distributed Deep Learning, particularly within LLM training and inference stacks, with a strong focus on collective communication and networking. You will interact with diverse hardware and platforms, such as Host Channel Adapters (HCAs), Switches, CPUs, GPUs, and complete Systems, and engage across multiple software layers, including LLM applications, machine learning frameworks, and communication and computing libraries.

This position offers a distinct opportunity to support the core infrastructure powering the next generation of large-scale AI systems.

Responsibilities

  • Design and implement resource allocation and combinatorial optimization techniques (e.g., reinforcement learning, LLM agents for DSE, Bayesian optimization, and other multi-objective optimization techniques) to optimize LLM models at data center scale.
  • Research, develop, and deploy AI/ML techniques to optimize large-scale Deep Learning (LLM) training and inference on NVIDIA supercomputers and distributed systems, with a focus on high-performance networking and NVIDIA communication libraries.
  • Build and productionize ML-based tools for performance prediction and optimization, with a strong emphasis on networking aspects.
  • Develop and deploy a scalable, reliable data curation pipeline capable of handling complex data types, such as time series and PyTorch model graphs, to support the training of high-performance Machine Learning models.
  • Collaborate across hardware and software teams to deliver valuable performance analysis insights.
  • Lead performance test planning, establish performance targets for new technologies and solutions, and drive efforts to achieve those performance goals.

Requirements

  • PhD or Master's degree in Computer Science, Software Engineering, or equivalent experience.
  • 4+ years of experience applying machine learning techniques to computer architecture and system optimization problems. Desired experience involves ML at the intersection of at least two of the following areas: HPC, networking, and AI applications.
  • Hands-on experience developing and deploying various learning algorithms (e.g., reinforcement learning, offline RL, supervised learning) to tackle optimization challenges within computer architecture, system design, or networking domains.
  • Proficiency in building and using ML models with leading frameworks such as PyTorch, TensorFlow, or JAX.
  • Proven ability to apply GNNs/transformers-based optimization to PyTorch model graph and Kineto execution traces.
  • Expertise combining knowledge of NVIDIA GPUs, the CUDA library, and deep learning frameworks (TensorFlow/PyTorch) with networking concepts, including collective communication libraries (like NCCL) and protocols (such as RoCE and RDMA).
  • Strong programming capabilities in Python, Bash, and C++.
  • A collaborative teammate with effective communication and interpersonal abilities.

Skills

  • In-depth knowledge and experience with machine learning/reinforcement learning and frameworks.
  • Comprehensive understanding of computer architecture, system architecture, and networking.
  • Extensive experience in applying machine learning techniques such as GNNs or related graph-based models.
  • Knowledge of PyTorch, CUDA, and NCCL libraries.
  • Proven software engineering/development skills.

Pay

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is:

  • 152,000 USD - 241,500 USD for Level 3.
  • 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Benefits

NVIDIA offers competitive salaries and a comprehensive benefits package. Our teams are composed of some of the most forward-thinking and driven engineers in the industry.

Similar jobs