Platform / System Debug Validation Engineer
At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. We’re looking for talent who want to leave the planet better than they found it—people who don’t shy away from humanity’s challenges but are determined to help solve them. AMD powers the next generation of supercomputing, high-performance computing, cloud, and AI. Every role contributes to technology that moves the world forward.
About the role
The Datacenter Platform Engineering Group (DPEG) is seeking an experienced system-level debug engineer to join a team responsible for Data Center deployments and ensuring the availability and uptime of a large number of systems. The ideal candidate will be familiar with system-level debug at all levels, possess RAS (Reliability, Availability, Serviceability) knowledge, and quickly root cause issues. Strong technical skills and the ability to direct junior engineers in solving problems and handling service-level tickets are essential.
Responsibilities
- Debug and triage complex system-level issues using industry tools for root cause analysis.
- Understand GPU and system-level hardware and software flows.
- Probe board components, check electrical and power currents, and validate system functionality.
- Apply RAS knowledge to diagnose issues at the GPU level.
- Provide leadership in driving root cause analysis and guide junior engineers.
- Document flows and methods for system bring-up, boot-up, initialization, and debug.
- Collaborate with internal teams to resolve complex issues and find optimal solutions.
- Serve as the go-to person for resolving Data Center issues, including electrical complexities, direct liquid cooling, and complex cluster networks.
- Communicate effectively with stakeholders via phone, chat, and email to drive issue resolution.
Requirements
- Hands-on experience debugging complex hardware and firmware issues.
- Understanding of GPU flow through different system layers and validation of components connecting to the GPU SOC (e.g., PCIe, VRMs, retimers, HBM, internal networking).
- Experience handling Data Center issues, including electrical complexities and direct liquid cooling.
- Proficiency with industry debug tools, oscilloscopes, and board-level power examination.
- Strong communication skills to work with various functional code stack owners.
- Hands-on experience with hardware in a Data Center environment.
Preferred Qualifications
- Experience in systems integration and debug.
- Proven ability to drive resolution of critical problems in a lab or Data Center environment.
- Relationships with external customers/partners to resolve Data Center issues, manufacturing failures, and deployment requirements.
- Significant experience in SoC and/or system debug of complex issues.
- Ability to develop and document debug capabilities for a given SOC and system.
Qualifications
Bachelor’s or Master’s degree in Electrical or Computer Engineering.
Location
Austin, TX. This role is not eligible for visa sponsorship.
Benefits
AMD offers a comprehensive benefits package. For details, visit AMD benefits at a glance.