Senior Hardware Development Engineer AWS AI & ML, Accelerator Servers
About the role
AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. Our team designs, builds and operates Amazon's fleet of Accelerated Servers using Internal Amazon design silicon or specialized purpose accelerators (EC2.TRN, INF, G, F + more instance types). We solve systemic hardware issues and build hardware and software systems to detect and mitigate future recurrences so that our customers can experience the highest quality of service possible.
Key job responsibilities
- Design and Architecture: Own server architecture, board design, component selection, thermal and power design, and ODM technical reviews. Make trade-off decisions balancing performance, cost, and manufacturability. Lead design reviews with manufacturing partners and ensure designs scale to production volumes.
- Fleet Operations: Monitor production quality, analyze field data to inform future designs, and drive continuous improvement. Collaborate with operations teams to ensure your designs meet reliability targets in production environments.
- Architect, design, and own a new segment of accelerated servers for the AWS fleet, including defining board-level architecture, component selection, and managing manufacturing partnerships through development and production.
- Make critical design decisions on thermal, power, signal integrity, and mechanical integration while leading cross-functional teams from concept through data center deployment.
- Define requirements and conduct technical reviews to ensure designs meet AWS standards.
- Design for reliability and manufacturability, incorporating lessons learned from fleet operations into your architecture decisions.
- Include built-in diagnostics and telemetry in your designs to enable efficient validation and operations.
A day in the life
You will spend your time making design decisions—defining technical requirements, conducting design reviews with manufacturing partners, selecting components, and architecting thermal and power solutions. You will interface with customers to translate requirements into technical specifications and work with manufacturing partners to ensure your designs scale to production. You will collaborate with interdisciplinary teams including component engineers, firmware developers, test engineers, and integration specialists to deliver complete server solutions.
About the team
The Hardware Engineering AI / ML development team is a group of engineers and technical program managers directly responsible for launching hardware in the fleet. Located out of Seattle, Cupertino and Austin, we work on programs with global development teams (both internal and external to Amazon). Our servers are located in datacenters globally. The members of our team have a diverse set of technical backgrounds but all share a common trait of Bias for Action and strong Ownership. We enjoy applying a startup model of delivering fully functional solutions for our customers.
Basic qualifications
- Experience in developing functional specifications, design verification plans and functional test procedures
- Experience working with interdisciplinary teams to execute product design from concept to production
- Experience in server technologies such as thermal, mechanical, power, and signal integrity
- Bachelor's degree in electrical engineering or equivalent
Preferred qualifications
- Master's degree in electrical engineering, computer engineering, or equivalent
- Experience with the project management of technical projects
- 7+ years of server, storage, networking, or large-scale distributed systems experience
- 10 years hardware development with a focus on system / server development in compute and/or storage server architecture and design for large scale applications
Pay
- USA, CA, Cupertino: $183,000.00 - $247,600.00 USD annually
- USA, TX, Austin: $159,200.00 - $215,300.00 USD annually
- USA, WA, Seattle: $159,200.00 - $215,300.00 USD annually
Benefits
- Health insurance (medical, dental, vision, prescription)
- Basic Life & AD&D insurance and option for Supplemental life plans
- Employee Assistance Program (EAP)
- Mental Health Support
- Medical Advice Line
- Flexible Spending Accounts
- Adoption and Surrogacy Reimbursement coverage
- 401(k) matching
- Paid time off
- Parental leave
- Sign-on payments and restricted stock units (RSUs)