Jobs · Quality Assurance · California

Cloud Hardware Development Engineer, Edge & High Performance Accelerator Servers for AI/ML

Amazon Web Services (AWS) · Cupertino, CA · 2 wk ago
Quality AssuranceFull-time

As a Cloud Hardware Development Engineer, you will be an end-to-end owner of Edge and/or Accelerator (AI/ML/GPU) Server Platforms, from New Product Introduction (NPI) through fleet health in production. You own the full lifecycle: design, development, qualification, launch, and ongoing operational excellence of servers running at scale in the AWS fleet.

Responsibilities

  • NPI — New Product Introduction
    • Own the end-to-end NPI lifecycle for Edge and/or Accelerator (AI/ML/GPU) Server Platforms — from architecture definition through design, qualification, manufacturing ramp, and launch
    • Lead technical solutions for complex server and rack system architectural challenges
    • Work with ODM/manufacturing partners to develop, validate, and manufacture server products at scale
    • Develop functional specifications, design verification plans, and test procedures
    • Drive qualification and readiness milestones, ensuring new platforms meet performance, reliability, and cost targets before fleet deployment
    • Identify and resolve technical risks early in the development cycle
  • Fleet Health, Diagnostics & Automation
    • Own fleet health for the server platforms you launch — reliability doesn’t end at ship
    • Design and implement predictive failure detection systems using telemetry, sensor data, error trending, and log correlation to identify hardware issues before they cause customer impact
    • Drive toward zero-touch operations — help build detection, diagnoses, and remediation of faults without human intervention
    • Debug complex system failures in time-sensitive settings — personally diving deep when the problem demands it
    • Perform root cause analysis correlating across firmware, kernel, driver, thermal, power, and physical layers
  • Systems Design & Technical Depth
    • Apply expertise across hardware, software, system design, x86 architecture, processes, and operations (compute, storage, network, GPU)
    • Design and implement solutions to address system-level issues at large scale
    • Decompose complex server system problems (testability, reliability, diagnostics) into deliverable tasks and features
    • Collaborate with hardware, software, manufacturing, supply chain, and product management teams
  • Cross-Team Collaboration
    • Work closely with internal customers to ensure new server hardware meets data path and control path requirements
    • Identify early any potential problems onboarding new servers into customer ecosystems
    • Collaborate across Hardware Engineering, component, firmware, test, qualification, and integration teams
    • Partner with datacenter operations to close the loop between field failures and design improvements

Qualifications

Basic Qualifications

  • Experience in developing functional specifications, design verification plans and functional test procedures
  • Bachelor's degree or above in electrical engineering, computer engineering, or equivalent
  • Experience in English-language communication skills, both written and verbal
  • Experience with design & innovation and research & development
  • Knowledge of operating systems, hardware, storage, network, security, database administration and cloud infrastructure
  • Experience in server technologies such as thermal, mechanical, power, and signal integrity
  • 5+ years of professional work (non-internship) experience

Preferred Qualifications

  • 5+ years of hardware design and validation of components, subsystems and systems experience
  • Experience working with Advanced Compute technologies including, but not limited to: Accelerated Compute, High Performance Compute, Visual/Spatial Compute, and/or IoT
  • Experience in server technologies: board design, high-speed bus design and signal integrity, failure analysis, server components (CPU, GPU, SSDs, memory), BIOS, BMC, and networking
  • Experience developing and executing test procedures for mechanical and/or electrical systems/components
  • Experience working with ODMs/manufacturer through the product development and manufacturing lifecycle
  • Experience building predictive failure detection or proactive remediation systems at fleet scale
  • Familiarity with PCIe topology, NVMe, and accelerator interconnects
  • Experience with large-scale datacenter or cloud environments

Pay

  • USA, CA, Cupertino: $157,300 - $212,800 USD annually
  • USA, TX, Austin: $136,000 - $184,000 USD annually
  • USA, WA, Seattle: $136,000 - $184,000 USD annually

Benefits

  • Sign-on payments and restricted stock units (RSUs)
  • Health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans)
  • Employee Assistance Program (EAP), Mental Health Support, Medical Advice Line
  • Flexible Spending Accounts
  • Adoption and Surrogacy Reimbursement coverage
  • 401(k) matching
  • Paid time off and parental leave

Similar jobs