Senior Hardware Development Engineer, Cloud AI/ML Server Team
Amazon Web Services (AWS) · Denver, CO · Today
EngineeringFull-time
About the role
The Hardware Engineering AI/ML Server platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Denver, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end — from server conception through fleet-scale operations.
Responsibilities
- Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale
- Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs
- Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)
- Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance
- Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems
- Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions
- Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes
- Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms
- Partner with test and automation teams to improve manufacturing yield and reduce test dwell times
- Cross-team collaboration with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs
- Drive ODM/JDM design partners through development milestones and production ramp
- Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready
- May require occasional (
Qualifications
- Bachelor's degree in electrical engineering, computer engineering, or equivalent
- 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms
- Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, or accelerator subsystems
- Experience leading hardware development through full product lifecycle (concept through production ramp)