Principal System Test Validation Engineer
At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. We’re looking for talent who want to leave the planet better than they found it—people who don’t shy away from humanity’s challenges but are determined to help solve them. AMD powers the next generation of supercomputing, high-performance computing, cloud, and AI. Every role at AMD contributes to technology that moves the world forward.
About the role
The technical leader will be responsible for System and Silicon validation of AMD EPYC Server & AMD Instinct products. The successful candidate will work as part of the post-silicon validation group, facilitating all aspects of validation and debug for system-level failures while collaborating with engineering teams across AMD. The role involves challenging system debug work, developing and executing validation strategies for optimal debug throughput on current products, and driving improvements to future debug methodology. The candidate should thrive in a global environment while maintaining a synergistic culture.
Responsibilities
- Lead system-level, rack-level, and cluster-scale validation strategy for AI and machine learning server platforms.
- Define and drive comprehensive validation plans covering customer deployment scenarios across scale-up and scale-out environments.
- Develop validation methodologies and test strategies spanning CPU, GPU, memory, BIOS, BMC, networking, storage, platform firmware, operating systems, and infrastructure components.
- Collaborate with architecture, hardware, firmware, software, and platform engineering teams to define validation requirements, identify coverage gaps, and improve overall product quality.
- Lead investigation and root cause analysis of complex hardware, firmware, software, networking, and system integration issues.
- Drive technical decision-making within the validation organization and influence cross-functional teams to resolve critical product risks and quality concerns.
- Define and enhance scalable validation frameworks, automation infrastructure, telemetry-driven workflows, and data-driven validation methodologies.
- Analyze large-scale validation data to identify systemic failures, reliability risks, performance bottlenecks, and product readiness concerns.
- Establish validation readiness criteria, quality metrics, and coverage strategies that improve execution efficiency and provide early risk identification.
- Mentor and provide technical leadership to engineers, helping improve validation methodologies, automation capabilities, debugging effectiveness, and system-level engineering practices.
- Drive execution across multiple validation programs while managing priorities, dependencies, schedules, and technical risks.
- Communicate validation status, readiness assessments, quality metrics, technical recommendations, and key risks to senior technical leadership and management.
- Support customer-focused validation efforts by ensuring test environments accurately represent real-world deployment conditions and hyperscale operating environments.
- Drive continuous improvements in validation processes, tools, and automation to improve organizational efficiency, test coverage, and product quality.
Requirements
- Extensive experience in validation roles involving debugging OS, FW, Silicon, and HW issues.
- Understanding of industry-standard busses and their software stack, such as PCIe, CXL.
- Strong knowledge of X86 architecture, SoC design, memory, RAS, and power management.
- Extensive knowledge of system architecture, technical debug, and validation strategy.
- Good understanding and experience in platform/system-level debug, Operating System, Device Drivers, and System BIOS interactions.
- Excellent communication and coordination skills.
- Detail-oriented, highly organized, able to prioritize, and juggle multiple work streams to tight deadlines.
- Strong analytical/problem-solving skills and pronounced attention to detail.
- Experience in technical program management.
- Thorough understanding of datacenter industry technologies and their software stack.
- Must be a self-starter and able to independently drive tasks to completion.
Qualifications
Bachelor’s or Master’s degree in electrical or computer engineering preferred.
Location
Austin, TX or Secaucus, NJ
Benefits
AMD offers a comprehensive benefits package. For details, see AMD benefits at a glance.