Jobs · Manufacturing · Washington

Cloud Hardware Dev Engineer (AWS Generative AI & ML Servers), Accelerator Servers

Amazon Web Services (AWS) · Seattle, WA · Today
ManufacturingFull-time

About the role

Do you want to build the backbone of Generative AI cloud at AWS? Do you want to build the future of the cloud for AI training and inference? Want to do industry leading work delivering continuous price performance improvements in the cloud for AI model training for multi billion variable LLMs? Come Join us in designing, delivering and operating AWS cloud offerings that enable high performance and scalability in AI/ML and HPC workloads.

About the team

Utility Computing (UC) provides product innovations — from foundational services such as Amazon's Simple Storage Service (S3) and Amazon Elastic Compute Cloud (EC2), to consistently released new product innovations that continue to set AWS's services and features apart in the industry. As a member of the UC organization, you'll support the development and management of Compute, Database, Storage, Internet of Things (IoT), Platform, and Productivity Apps services in AWS, including support for customers who require specialized security solutions for their cloud services. You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You'll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

The Hardware Engineering Services team is comprised of both Hardware Design Engineers, System Design Engineers, Software Development Engineers and Technical Program Managers, all with the common goal of delivering the best Accelerated Server fleet possible to our customers.

Key job responsibilities

  • Own and lead the design, development and root cause of a new segment of accelerated servers.
  • Work closely with our customers to understand their technical needs and business goals, leveraging your experience with server design and the knowledge of various teams to architect the solutions that we will deploy at scale.
  • Work with an interdisciplinary team of component, firmware, test, qualification, and integration engineers, and lead our design and manufacturing partners to bring these servers to the data center.
  • After launch, oversee the fleet of servers you develop, monitoring their quality and how they are meeting the customer requirements.
  • Interface with our internal and external customers to understand project requirements and facilitate system development on top of your server design.
  • Learn operational challenges to our existing fleet with the goal of improving the current customer experience as well as developing improved systems for future designs.
  • Work directly with vendors and ODM/JDM design teams to develop and manufacture your product at scale.

Basic Qualifications

  • Experience in developing functional specifications, design verification plans and functional test procedures
  • Bachelor's degree or above in electrical engineering, computer engineering, or equivalent
  • Experience in English-language communication skills, both written and verbal
  • Experience with design & innovation and research & development
  • Knowledge of operating systems, hardware, storage, network, security, database administration and cloud infrastructure
  • Experience in server technologies such as thermal, mechanical, power, and signal integrity
  • 5+ years of professional work (non-internship) experience

Preferred Qualifications

  • 5+ years of hardware design and validation of components, subsystems and systems experience
  • Experience in server technologies: board design, high-speed bus design and signal integrity, failure analysis, server components (CPU, GPU, SSDs, memory), BIOS, BMC, and networking
  • Experience developing and executing test procedures for mechanical or electrical systems/components
  • Experience working with ODMs/manufacturer through the product development and manufacturing lifecycle
  • Experience building predictive failure detection or proactive remediation systems at fleet scale
  • Experience with storage/compute/GPU/accelerator platforms including integration, diagnostics, or performance validation
  • Familiarity with PCIe topology, NVLink, NVMe, and accelerator interconnects
  • Experience with large-scale datacenter or cloud environment

Pay

  • USA, CA, Cupertino: $157,300.00 - $212,800.00 USD annually
  • USA, TX, Austin: $136,000.00 - $184,000.00 USD annually
  • USA, WA, Seattle: $136,000.00 - $184,000.00 USD annually

Benefits

  • Health insurance (medical, dental, vision, prescription)
  • Basic Life & AD&D insurance and option for Supplemental life plans
  • EAP, Mental Health Support, Medical Advice Line
  • Flexible Spending Accounts
  • Adoption and Surrogacy Reimbursement coverage
  • 401(k) matching
  • Paid time off
  • Parental leave
  • Sign-on payments and restricted stock units (RSUs)
This response is AI-generated, for reference only.

Similar jobs