Jobs · Engineering · California

Lead Software Engineer, ML Network Stack - Annapurna Labs

Amazon Web Services (AWS) · Cupertino, CA · Yesterday
EngineeringFull-time

About the role

We are seeking an experienced engineer and technical leader to join our team that owns the network stack for EC2 distributed AI/ML systems. The team develops support for a variety of frameworks and communication libraries including NCCL, NVSHMEM, NIXL, NCCL GIN, and CUDA kernels. Solid knowledge of Linux, networking, and performant coding is important. Experience with embedded systems is valued, and experience with high-speed networking or HPC/RDMA interconnects is highly valued. If you like solving hard problems, want to work with HPC and ML customers, iterate fast and deliver meaningful solutions at scale, then come join us! This truly is a role at the forefront of AI/ML—you'll be working on features for the largest clusters, with the largest customers, for the largest AI models.

About the team

The organization you would be joining is Annapurna Labs, an integral part of AWS that develops hardware and software components that are critical building blocks for EC2 infrastructure. Every instance in EC2 is running some type of hardware designed by Annapurna Labs. We specialize in designing software, systems, and chips that optimize the AWS customer experience. Annapurna Labs' inventions are routinely mentioned in keynotes and all of re:Invent. Do you recognize Nitro, Trainium, EFA, ENA-X, Inferentia, Graviton? Join us and help create the future.

Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.

Key job responsibilities

  • Be the Leader which works across the NVIDIA ML communication stack for enabling NVIDIA GPUs to work with the AWS EC2 machines
  • Lead senior, mid-level, and junior SDEs and direct work to ensure the team delivers functions and features required for the latest and largest ML workloads

Basic qualifications

  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • Experience as a mentor, tech lead or leading an engineering team, or experience in development in the last 3 years
  • 5+ years of experience with programming language: C or C++

Preferred qualifications

  • Bachelor's degree in computer science or equivalent
  • Experience working with ML Communication Libraries or RDMA Networking
  • Knowledge of ML Networking, ML applications, and frameworks

Pay & benefits

Base salary range: $193,300.00 - $261,500.00 USD annually (USA, CA, Cupertino). Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

This response is AI-generated, for reference only.

Similar jobs