AI Systems Validation Engineer
AMD · Secaucus, NJ · 1 wk ago
On-siteInformation TechnologyFull-time
Key Responsibilities
- Lead system-level, rack-level, and cluster-scale validation strategy for AI and machine learning server platforms.
- Define and drive comprehensive validation plans covering customer deployment scenarios across scale-up and scale-out environments.
- Develop validation methodologies and test strategies spanning CPU, GPU, memory, BIOS, BMC, networking, storage, platform firmware, operating systems, and infrastructure components.
- Collaborate with architecture, hardware, firmware, software, and platform engineering teams to define validation requirements, identify coverage gaps, and improve overall product quality.
- Lead investigation and root cause analysis of complex hardware, firmware, software, networking, and system integration issues.
- Drive technical decision-making within the validation organization and influence cross-functional teams to resolve critical product risks and quality concerns.
- Define and enhance scalable validation frameworks, automation infrastructure, telemetry-driven workflows, and data-driven validation methodologies.
- Analyze large-scale validation data to identify systemic failures, reliability risks, performance bottlenecks, and product readiness concerns.
- Establish validation readiness criteria, quality metrics, and coverage strategies that improve execution efficiency and provide early risk identification.
- Mentor and provide technical leadership to engineers, helping improve validation methodologies, automation capabilities, debugging effectiveness, and system-level engineering practices.
- Drive execution across multiple validation programs while managing priorities, dependencies, schedules, and technical risks.
- Communicate validation status, readiness assessments, quality metrics, technical recommendations, and key risks to senior technical leadership and management.
- Support customer-focused validation efforts by ensuring test environments accurately represent real-world deployment conditions and hyperscale operating environments.
- Drive continuous improvements in validation processes, tools, and automation to improve organizational efficiency, test coverage, and product quality.
Preferred Experience
- 8+ years of experience in system-level validation, platform validation, post-silicon validation, large-scale infrastructure validation, or related engineering disciplines.
- Proven experience validating large-scale server, AI, cloud, HPC, or data center platforms.
- Demonstrated ability to lead investigation and resolution of complex hardware, firmware, software, networking, and infrastructure issues.
- Experience validating scale-up and scale-out architectures, including rack-scale and cluster-scale deployments.
- Deep understanding of computer architecture, server platform design, operating systems, networking, storage, and hardware/software interactions.
- Experience developing automated validation frameworks, orchestration systems, and large-scale test infrastructure.
- Strong proficiency in Python and Linux environments, with working knowledge of Bash and/or PowerShell.
- Experience leveraging telemetry, analytics, and data-driven methodologies to improve validation coverage, efficiency, and root cause analysis.
- Familiarity with AI, machine learning, GPU computing, or accelerated computing environments deployed in data center infrastructures.
- Demonstrated ability to lead cross-functional technical initiatives and influence engineering decisions across organizational boundaries.
- Experience mentoring engineers and driving technical excellence within validation and product engineering organizations.
- Experience supporting product readiness, customer deployment validation, and release qualification activities.
- Strong communication skills with the ability to convey technical risks, recommendations, and status to executive and technical leadership.