Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), Accelerator Servers
Amazon Web Services (AWS) · Seattle, WA · 1 wk ago
EngineeringFull-time
About the role
Build the backbone of Generative AI cloud at AWS and the future of the cloud for AI training and inference. Deliver industry-leading continuous price performance improvements in the cloud for AI model training for multi-billion variable LLMs. Design, deliver, and operate AWS cloud offerings that enable high performance and scalability in AI/ML and HPC workloads.
Key job responsibilities
- Own and lead the design, development, and root cause of a new segment of accelerated servers
- Work closely with customers to understand their technical needs and business goals
- Leverage server design experience and knowledge of various teams to architect solutions deployed at scale
- Work with an interdisciplinary team of component, firmware, test, qualification, and integration engineers
- Lead design and manufacturing partners to bring servers to the data center
- Oversee the fleet of servers after launch, monitoring quality and customer requirements fulfillment
- Interface with internal and external customers to understand project requirements and facilitate system development
- Learn operational challenges to the existing fleet to improve current customer experience and develop improved systems for future designs
- Work directly with vendors and ODM/JDM design teams to develop and manufacture products at scale
About the team
The team comprises Hardware Design Engineers, System Design Engineers, Software Development Engineers, and Technical Program Managers, all with the common goal of delivering the best Accelerated Server fleet possible to customers.
Basic qualifications
- Experience in developing functional specifications, design verification plans, and functional test procedures
- Bachelor's degree or above in electrical engineering, computer engineering, or equivalent
- Experience in English-language communication skills, both written and verbal
- Experience with design and innovation and research and development
- Knowledge of operating systems, hardware, storage, network, security, database administration, and cloud infrastructure
- Experience in server technologies such as thermal, mechanical, power, and signal integrity
- 5+ years of professional work (non-internship) experience
Preferred qualifications
- 5+ years of hardware design and validation of components, subsystems, and systems experience
- Experience in server technologies: board design, high-speed bus design and signal integrity, failure analysis, server components (CPU, GPU, SSDs, memory), BIOS, BMC, and networking
- Experience developing and executing test procedures for mechanical or electrical systems/components
- Experience working with ODMs/manufacturer through the product development and manufacturing lifecycle
- Experience building predictive failure detection or proactive remediation systems at fleet scale
- Experience with storage/compute/GPU/accelerator platforms including integration, diagnostics, or performance validation
- Familiarity with PCIe topology, NVLink, NVMe, and accelerator interconnects
- Experience with large-scale datacenter or cloud environment
Pay
- USA, CA, Cupertino: $157,300.00 - $212,800.00 USD annually
- USA, TX, Austin: $136,000.00 - $184,000.00 USD annually
- USA, WA, Seattle: $136,000.00 - $184,000.00 USD annually
- Package includes sign-on payments and restricted stock units (RSUs); final compensation determined based on factors including experience, qualifications, and location
Benefits
- Health insurance (medical, dental, vision, prescription)
- Basic Life & AD&D insurance and option for Supplemental life plans
- EAP, Mental Health Support, Medical Advice Line
- Flexible Spending Accounts
- Adoption and Surrogacy Reimbursement coverage
- 401(k) matching
- Paid time off
- Parental leave