Jobs · Engineering · Washington

Cloud Hardware Development Manager, AWS Gen AI & ML Servers

Amazon Web Services (AWS) · Seattle, WA · Today
EngineeringFull-time

About the role

The Hardware Engineering AI/ML Ultraserver platform team is a group of engineers and technical program managers directly responsible for launching and maintaining GPU-accelerated servers in the AWS fleet. Located in Seattle, Cupertino, and Austin, we work with internal engineering teams, ODMs, and design partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams.

Responsibilities

  • Lead hardware engineers who define the platforms running the world's largest AI workloads.
  • Set technical direction for your team's hardware portfolio working alongside product and leadership teams, making trade-offs between schedule, cost, reliability, and performance.
  • Guide system-level design decisions across thermal, mechanical, power delivery, signal integrity, and accelerator subsystems — including trade-offs on cooling architecture (liquid vs. air), power budget allocation, and PCIe/interconnect topology.
  • Drive design reviews, qualification gates, and go/no-go decisions with deep technical judgment.
  • Own fleet reliability and availability for your team's hardware; drive continuous improvement to reduce failure rates.
  • Hire, develop, and retain a team of hardware and system software engineers; build a high-performing team culture focused on engineering excellence and customer obsession.
  • Align with EC2 architecture teams on instance requirements, workload characterization, and platform roadmaps.
  • Partner with firmware, software, test automation, and datacenter operations teams to ensure hardware is debuggable, serviceable, and automation-ready.
  • Communicate technical strategy, program status, and risk posture to senior leadership.

Qualifications

  • Bachelor's degree or above in Electrical or Mechanical Engineering, or Bachelor's degree in computer science, engineering, mathematics or equivalent
  • 7+ years of hardware development experience for server, compute, networking, or storage platforms
  • 1+ years of experience managing hardware engineering teams
  • Experience driving hardware development programs (servers, racks, networking, or storage) through full product lifecycle including design, validation, manufacturing, and deployment
  • Experience working with ODMs or manufacturing partners through product development and production
  • Experience in one or more server technologies: thermal/mechanical design, power delivery, high-speed signal integrity, reliability, or accelerator subsystems

Preferred Qualifications

  • Master's degree or above in Electrical Engineering, Mechanical Engineering, or a related field
  • 5+ years of experience managing multi-discipline hardware development teams delivering products at datacenter scale
  • Experience leading ODM and silicon supplier partnerships across multiple geographies, establishing technical standards and driving quality frameworks
  • Track record of owning fleet quality metrics (annualized failure rates, system availability) and driving design improvements based on operational data
  • Knowledge of datacenter infrastructure constraints including networking, power, space, and cooling
  • Experience with GPU/accelerator server platforms, NPI processes (EVT, DVT, PVT), or hardware qualification programs
  • Demonstrated ability to communicate technical strategy and influence senior executives through written narratives
  • Track record of hiring, developing, and retaining high-performing hardware engineering talent

Similar jobs