Jobs · Marketing · California

Head of Platform Product Reliability

Etched · San Jose, CA · 4 wk ago
On-siteMarketing$2k/moFull-time

About the role

We are seeking a highly technical and execution-focused Head of Platform Product Reliability to lead reliability engineering across Etched's server, rack, and datacenter platform products. This role owns system-level product reliability from architecture through fleet deployment.

Key Responsibilities

  • Define and own the end-to-end reliability strategy for AI servers, accelerator platforms, rack systems, and datacenter infrastructure, from design requirements through field deployment
  • Establish reliability requirements, qualification standards, and validation methodologies that scale across product generations
  • Build and institutionalize reliability engineering processes spanning the full product lifecycle: EVT / DVT / PVT qualification gates and exit criteria, accelerated life testing (ALT) and accelerated stress testing (AST), environmental testing: temperature, humidity, altitude, contamination, HALT / HASS programs for design margin and production screening, vibration, shock, and transportation stress testing, power cycling, thermal cycling, and long-duration soak testing
  • Lead root-cause investigations for reliability failures surfaced during development, manufacturing, and field deployment, driving corrective actions across hardware, firmware, thermal, and mechanical domains
  • Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modeling, component derating methodologies, and reliability growth tracking
  • Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change
  • Work closely with ODMs, JDMs, contract manufacturers, and component suppliers to validate and enforce long-term platform reliability commitments
  • Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops, and monitoring frameworks that give Etched visibility into deployed system health at scale
  • Drive reliability signoff criteria and lead product release readiness reviews across engineering and program teams
  • Build and lead a high-performing product reliability engineering organization — hiring, developing, and retaining technical talent as the company scales

Qualifications

  • BS, MS, or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering, or a related technical field
  • 10+ years of reliability engineering experience in hardware-centric organizations, with meaningful time spent on complex systems rather than component-level work
  • Experience leading reliability programs for one or more of: AI accelerator or GPU-class compute systems, hyperscale or cloud server infrastructure, networking platforms, storage systems, or rack-scale infrastructure
  • Deep understanding of system-level failure mechanisms — including thermal, power delivery, mechanical, and connector/interconnect failure modes — and how design decisions affect long-term field reliability
  • Hands-on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies, and reliability statistics and modeling
  • A track record of driving cross-functional root-cause investigations in fast-moving hardware organizations where schedule pressure is real and accountability is high
  • Strong technical judgment — capable of making defensible tradeoffs between reliability targets, cost, schedule, and performance without losing sight of customer expectations
  • Excellent communication skills and the credibility to influence design decisions with engineering leads, program managers, and executive stakeholders

Nice-to-have Qualifications

  • Experience with liquid-cooled systems, high-density power delivery, or thermal management for high-power AI infrastructure
  • Direct experience supporting hyperscale or cloud datacenter deployments at scale, including customer-facing reliability commitments and SLA management
  • Demonstrated experience building a reliability organization from early-stage — establishing processes, tooling, and team norms in environments without established infrastructure
  • Familiarity with fleet telemetry systems, large-scale field reliability analytics, and data-driven approaches to proactive reliability management
  • Experience working closely with ODM or JDM partners in Taiwan or broader Asia, including NPI support and on-site qualification engagement
  • Background in high-speed digital systems, GPU compute platforms, or accelerator-based architectures — with an understanding of how these affect system-level reliability behavior

Similar jobs

Platform Reliability Engineer

Konica Minolta Business Solutions U.S.A., Inc.Ramsey, NJ· 6 days ago
Engineeringapply on careers.konicaminoltaus.com