Jobs · Engineering · Colorado

Senior Accelerator Engineer, Cloud AI/ML server team

Amazon Web Services (AWS) · Denver, CO · Today
EngineeringFull-time

About the role

This is a role where you will deepen your GPU expertise at a pace and scale that doesn't exist outside of hyperscale cloud, working directly with the industry's leading GPU vendors, analyzing behaviors across hundreds of thousands of accelerators, and defining standards that govern production deployment.

Responsibilities

  • Own the technical relationship with vendors across all platforms in the portfolio: roadmap alignment, escalations, partnerships.

  • Own qualification of new GPU SKUs and baseboard assemblies during NPI bring-up -- define test plans, acceptance criteria, and production readiness gates.

  • Define and maintain GPU firmware qualification criteria across the org -- pass/fail gates, staged rollout policy, regression detection methodology.

  • Drive RMA strategy: build failure evidence packages, negotiate acceptance criteria with vendors, manage submission quotas and pipeline velocity.

  • Lead root-cause analysis on fleet-wide GPU failure modes (component errors, PCIE interface errors, thermal events, link degradation, manufacturing escapes) using telemetry, event log data, and vendor diagnostics.

  • Set GPU health standards: define the metrics, thresholds, and alerting that platform teams execute against.

  • Define GPU fleet health dashboards: identify relevant telemetry, failure rate trends, replacement pipeline status, firmware version distribution, qualification status.

  • Represent the organization in technical discussions with leading vendor’s engineering -- translate fleet-scale patterns into prioritized vendor action items.

  • Partner with server teams to ensure consistent GPU operational practices; provide expertise without owning their execution.

  • Present GPU fleet health, replacement pipeline status, and qualification progress to senior leadership (VP-level) regularly.

  • Mentor engineers on GPU failure analysis methodology.

Requirements

  • Bachelor's degree in electrical engineering, computer engineering, or equivalent.
  • Experience in developing functional specifications, design verification plans and functional test procedures.
  • 7+ years of hardware design and development experience for server, compute, or large-scale infrastructure platforms.

Qualifications

  • Direct experience with data center GPUs and associated tooling.
  • Experience developing or influencing server roadmap.
  • Track record of influencing vendor engineering priorities through failure evidence.
  • Familiarity with GPU thermal management, power delivery, and PCIe and interconnect architectures.
  • Experience with firmware lifecycle management at scale (qualification, staged rollout, regression detection, rollback).
  • Comfortable presenting to VP-level audiences.
  • Strong data analysis skills at fleet scale (statistical failure modeling, trend detection, threshold setting).

Similar jobs