Sr. SDM, AI Inference, Neuron SDK
About the role
AWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon Elastic Compute Cloud (EC2), to new product innovations that continue to set AWS’s services and features apart in the industry. As a Sr. SDM for the Inference Team, you will lead a strong team of managers and engineers to optimize models for best performance and robustness when deployed at scale on Trainium, Amazon's custom cloud-scale machine learning accelerators that power the latest AI models.
You will be responsible for the full development life cycle of model onboarding, low-level performance optimization, feature support, and efficient serving, including reliability and scalability.
Responsibilities
- Lead a team of managers and engineers to optimize inference models for performance and robustness on Trainium and Inferentia devices.
- Oversee the full development life cycle, including model onboarding, low-level performance optimization, feature support, and efficient serving.
- Collaborate with executive leadership, senior management, and technical leaders to define and deliver product directions to customers.
- Build massive-scale distributed training and inference solutions, developing the full stack of software, servers, and chips in partnership with teams across the Annapurna organization.
- Ensure reliability and scalability of AI model deployments at cloud scale.
Qualifications
Preferred qualifications include an established background in optimizing and serving AI models under demanding, fast-changing priorities, and a strong technical ability to understand and manage a vertically integrated system stack consisting of hardware, frameworks, serving stack, model, and workflows.