Jobs · Consulting · Washington

Software Engineer- AI/ML, Amazon Neuron Training

Amazon Web Services (AWS) · Seattle, WA · 1 wk ago
ConsultingFull-time

We are the Annapurna Labs team at Amazon Web Services (AWS), building AWS Neuron—the software development kit that accelerates deep learning and GenAI workloads on AWS Trainium, Amazon’s custom machine learning accelerator. Neuron includes an ML compiler, runtime, collectives library, and application framework that integrate seamlessly with PyTorch and JAX, enabling customers to train frontier-scale models on Trainium without rewriting their stack.

About the role

The Distributed Training team is at the forefront of training a wide range of models on AWS’s custom ML accelerators, supporting novel architectures while maximizing their training performance. Working across the stack—from PyTorch and JAX down to the hardware and software boundary—our engineers build the infrastructure that large-scale training depends on, develop new parallelism and numerics techniques, and tune high-performance kernels for the operations that dominate a training step. This ensures every compute unit is doing useful work on our customers’ most demanding workloads.

As part of this role, you will help lead the effort to optimize distributed training performance on Trainium, focusing on training throughput across the Neuron software stack. You will collaborate with the Neuron compiler and runtime teams to enable and tune large-scale training workloads on the latest Trainium instances, own the parallelism strategies these models depend on (including data, tensor, pipeline, expert, and context parallelism), and apply reduced-precision formats where they measurably improve performance.

You will profile end-to-end workloads to determine whether they are bound by compute, memory, collectives, or host overhead, then drive fixes to the appropriate layer by working with compiler, runtime, and collectives engineers. Your work will translate performance gaps into requirements that influence frameworks, and you will contribute upstream to the open-source frameworks our customers rely on. This role offers a rare opportunity to work at the intersection of machine learning, high-performance computing, and distributed systems, shaping the future of AI acceleration technology.

Responsibilities

  • Lead the optimization of distributed training performance on AWS Trainium, focusing on end-to-end training throughput across the Neuron software stack.
  • Collaborate with Neuron compiler, runtime, and collectives teams to enable and tune large-scale training workloads on Trainium instances.
  • Own and implement parallelism strategies (data, tensor, pipeline, expert, and context parallelism) for large-scale models.
  • Apply reduced-precision formats where they provide measurable performance gains.
  • Profile workloads to identify bottlenecks (compute, memory, collectives, or host overhead) and drive fixes to the appropriate layer of the stack.
  • Translate performance gaps into actionable requirements for frameworks and contribute upstream to open-source frameworks (e.g., PyTorch, JAX).
  • Work across the stack—from high-level frameworks to hardware boundaries—to maximize training efficiency at scale.
  • Influence future architecture designs by characterizing current performance gaps and defining requirements for next-generation Trainium hardware.

Requirements

  • 3+ years of non-internship professional software development experience.
  • 2+ years of non-internship experience in designing or architecting (design patterns, reliability, and scaling) new and existing systems.
  • Experience programming with at least one software programming language.

Preferred Qualifications

  • 3+ years of full software development life cycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations.
  • Bachelor’s degree in computer science or equivalent.

Benefits

  • Comprehensive health insurance (medical, dental, vision, prescription, Basic Life & AD&D, with optional supplemental life plans).
  • Employee Assistance Program (EAP) and mental health support, including a medical advice line.
  • Flexible Spending Accounts (FSA) for healthcare and dependent care.
  • Adoption and surrogacy reimbursement coverage.
  • 401(k) matching.
  • Paid time off and parental leave.
  • Sign-on payments and restricted stock units (RSUs).

Pay

  • USA, CA, Cupertino: $165,200 - $223,600 USD annually.
  • USA, WA, Seattle: $143,700 - $194,400 USD annually.

About the team

Our team combines deep hardware knowledge with ML expertise to push the limits of training efficiency at scale. We work across multiple technology layers—from frameworks and kernels to the compiler, runtime, and collectives teams—practicing hardware and software co-design. Performance gaps rarely sit in one layer, so we trace problems across the stack, determine where fixes belong, and collaborate with the owning team to implement them. We not only optimize current performance but also contribute to future architecture designs, as the gaps we characterize today become requirements for the next generation of Trainium.

We work closely with customers to enable their models and ensure they train efficiently, offering a rare opportunity to shape the direction of AI acceleration technology.

Work/Life Balance

Our team values work-life balance, focusing on the flow that brings energy to both professional and personal life. We believe this balance is critical to long-term happiness and fulfillment and offer flexibility in working hours to help you find your own equilibrium.

Similar jobs