Machine Learning - Compiler Engineer , AWS Neuron, Annapurna Labs
About the role
Join the AWS Neuron Compiler team and be part of the AI revolution. AWS Neuron is the SDK that optimizes the performance of complex ML models executed on AWS Inferentia and Trainium, our custom chips designed to accelerate deep-learning workloads. In this role, you will build the next-generation Neuron compiler, transforming ML models written in frameworks like PyTorch, TensorFlow, and JAX for deployment on AWS Inferentia and Trainium-based servers in the Amazon cloud.
You will tackle hard compiler optimization problems to achieve optimum performance for a variety of ML model families, including large language models (e.g., Llama, Deepseek), stable diffusion, vision transformers, and multimodal models. A deep understanding of these models will guide your decisions in generating optimal compiler implementations. You will collaborate with internal and external stakeholders, contribute to pre-silicon design, and help bring new products/features to market, making the Neuron compiler highly performant and easy to use.
Responsibilities
- Design, implement, test, deploy, and maintain innovative software solutions to enhance the Neuron compiler’s performance, stability, and user interface.
- Collaborate with chip architects, runtime/OS engineers, scientists, and ML Apps teams to seamlessly deploy state-of-the-art ML models on AWS accelerators with optimal cost/performance benefits.
- Work with open-source software (e.g., StableHLO, OpenXLA, MLIR) to pioneer optimizations for advanced ML workloads on AWS hardware and software.
- Build innovative features to deliver the best possible experience for developers worldwide.
- Create compiler optimization and verification passes, expose features of AWS accelerators to developers, implement tools to analyze numerical errors, and resolve compiler defects.
- Participate in design discussions, code reviews, and communicate with internal (Neuron SDK and Amazon-wide teams) and external stakeholders (open-source communities).
Qualifications
- 3+ years of non-internship professional software development experience.
- 2+ years of non-internship experience in design or architecture (design patterns, reliability, and scaling) of new and existing systems.
- Experience programming with at least one software programming language.
- Experience in object-oriented languages like C++/Java is required.
- Experience with compilers or building ML models using ML frameworks on accelerators (e.g., GPUs) is preferred but not required.
- Experience with technologies like OpenXLA, StableHLO, or MLIR is a plus.
Preferred Qualifications
- Master’s degree or PhD in Computer Science or a related technical field.
- 3+ years of experience writing production-grade code in object-oriented languages such as C++/Java.
- Experience in compiler design for CPU/GPU/vector engines/ML-accelerators.
- Experience with open-source compiler toolsets like LLVM/MLIR.
- Experience with PyTorch, OpenXLA, StableHLO, JAX, TVM, deep learning models, and algorithms.
- Experience with modern build systems like Bazel/CMake.
About the team
Our team is dedicated to supporting new members with a broad mix of experience levels and tenures. We foster an environment that celebrates knowledge-sharing and mentorship, with senior members providing one-on-one mentoring and thorough, constructive code reviews. We prioritize your career growth, assigning projects that develop your engineering expertise and empower you to tackle more complex tasks in the future.
Benefits
- Comprehensive health insurance (medical, dental, vision, prescription, Basic Life & AD&D, with options for supplemental life plans).
- Employee Assistance Program (EAP), Mental Health Support, and Medical Advice Line.
- Flexible Spending Accounts (FSA) for healthcare and dependent care.
- Adoption and Surrogacy Reimbursement coverage.
- 401(k) matching.
- Paid time off and parental leave.
- Sign-on payments and restricted stock units (RSUs).
Pay
Base salary range: $165,200 - $223,600 USD annually (USA, CA, Cupertino). Final compensation will be determined based on experience, qualifications, and location.