Sr. Machine Learning - Compiler Engineer III, AWS Neuron, Annapurna Labs
About the role
This role is for a senior software engineer in the Compiler team for AWS Neuron. As part of this role, you will be responsible for building next generation Neuron compiler which transforms ML models written in ML frameworks (e.g, PyTorch, TensorFlow, and JAX) to be deployed AWS Inferentia and Trainium based servers in the Amazon cloud. You will be responsible for solving hard compiler optimization problems to achieve optimum performance for variety of ML model families including massive scale large language models like Llama, Deepseek, and beyond as well as stable diffusion, vision transformers and multi-model models. You will be required to understand how these models work inside-out to make informed decisions on how to best coax the compiler to generate optimal implementation instruction. You will leverage your technical communications skill to partner with other teams and will be involved in pre-silicon design, bringing new products/features to market, and many other exciting projects.
Key job responsibilities
- Design, implement, test, deploy and maintain innovative software solutions to transform Neuron compiler's performance, stability and user-interface
- Work side by side with chip architects, runtime/OS engineers, scientists and ML Apps teams to seamlessly deploy cutting edge ML models from our customers on AWS accelerators with optimal cost/performance benefits
- Become front-face of Neuron Compiler to work with open-source communities (e.g., StableHLO, OpenXLA, MLIR) and influence industry wide partners to pioneer optimizing cutting-edge ML workloads on AWS software and hardware
- Work on building innovative features that will deliver best possible experiences for our customers – developers across the globe
A day in the life
As you design and code solutions to help our team drive efficiencies in compiler architecture, you'll create compiler optimization and verification passes, build features surface features and peculiarities of AWS accelerators to developers, implement tools to analyze numerical errors, and resolve the root cause of compiler defects. You'll also participate in design discussions, code review, and communicate with internal (other Neuron SDK and Amazon wide teams) and external stakeholders (open-source communities and respond to Neuron compiler related questions in open forums, e.g. GitHub). Lastly, work in a startup-like development environment, where you're always working on the most important stuff.
Basic qualifications
- Bachelor's degree in computer science or equivalent
- 5+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- 3+ years of experience in developing compiler features and optimizations
- Proficiency with 1 or more of the following programming languages: C++ (preferred), Python
Preferred qualifications
- Master or PhD degree in computer science or equivalent
- Proficiency with compiler design, resource management, instruction scheduling, memory allocation, data transfer optimization, compute graph optimization, code generation, and Instruction Set Architecture
- Experience with LLVM and/or MLIR
- Experience in LLM, Vision or other deep-learning models
About the team
Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we're building an environment that celebrates knowledge-sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects that help our team members develop your engineering expertise so you feel empowered to take on more complex tasks in the future.
Pay
- USA, CA, Cupertino: 193,300.00 - 261,500.00 USD annually
- USA, WA, Seattle: 168,100.00 - 227,400.00 USD annually