Software Engineer- AI/ML, AWS Neuron Distributed Training - Performance Optimization
About the role
Annapurna Labs designs silicon and software that accelerates innovation. Customers choose us to create cloud solutions that solve challenges previously unimaginable. Our custom chips, accelerators, and software stacks enable tackling unprecedented technical challenges and delivering results that help customers change the world.
AWS Neuron is the complete software stack for the AWS Trainium and Inferentia cloud-scale machine learning accelerators and the Trn3/Trn2/Trn1 and Inf2/Inf1 servers that use them. This role is for a software engineer in the Distributed Training team for AWS Neuron, responsible for development, enablement, and performance tuning of a wide variety of ML model families, including large-scale multi-modal large language models (e.g., Llama, Qwen, gpt-oss, DeepSeek) and multi-modal generation models (e.g., Stable Diffusion, Flux, WAN).
The Distributed Training team collaborates with chip architects, compiler engineers, and runtime engineers to create, build, and tune distributed training solutions with AWS Trainium, maximizing training throughput, minimizing time-to-convergence, and pushing the boundaries of training efficiency. You will identify and resolve performance bottlenecks across the stack, from collective communications and memory utilization to compiler optimizations and kernel performance.
Responsibilities
- Lead efforts to optimize distributed training performance on Trainium, focusing on maximizing training throughput, model flops utilization, and efficiency across the Neuron software stack.
- Work across PyTorch, JAX, and the Neuron compiler and runtime to enable and tune large-scale training workloads on the latest Trainium instances.
- Collaborate with cross-functional teams to resolve performance bottlenecks in collective communications, memory utilization, compiler optimizations, and kernel performance.
About the team
Our team is dedicated to supporting new members with a broad mix of experience levels and tenures. We foster an environment that celebrates knowledge-sharing and mentorship, with senior members providing one-on-one mentoring and thorough, constructive code reviews. We prioritize your career growth, assigning projects that develop your engineering expertise and empower you to tackle more complex tasks in the future.
AWS values diverse experiences and encourages candidates from all backgrounds to apply, even if they don’t meet every qualification listed. We embrace alternative career paths and non-traditional experiences.
Qualifications
Basic Qualifications
- 3+ years of non-internship professional software development experience.
- 2+ years of non-internship design or architecture experience (design patterns, reliability, and scaling) of new and existing systems.
- Experience programming with at least one software programming language.
Preferred Qualifications
- 3+ years of full software development life cycle experience, including coding standards, code reviews, source control management, build processes, testing, and operations.
- Bachelor's degree in computer science or equivalent.
Benefits
Amazon offers comprehensive benefits, including:
- Health insurance (medical, dental, vision, prescription, Basic Life & AD&D, and optional supplemental life plans).
- Mental health support (EAP, Medical Advice Line).
- Flexible Spending Accounts.
- Adoption and surrogacy reimbursement coverage.
- 401(k) matching.
- Paid time off and parental leave.
Learn more about our benefits at amazon.jobs/en/benefits.
Pay
- USA, CA, Cupertino: $165,200.00 - $223,600.00 USD annually.
- USA, WA, Seattle: $143,700.00 - $194,400.00 USD annually.
Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on experience, qualifications, and location.