Data Center Engineer
About the Role
In this critical and highly technical role, you will execute AMD's Datacenter graphics hardware/software subsystem projects for AMD OEM partners and enterprise commercial end-customers. This position provides a unique opportunity to leverage your expertise in graphics, compute, datacenter technologies, virtualization, AI/Machine Learning, and program management to collaborate with customers utilizing AMD Instinct™ Accelerators. You must be a team player with a strong commitment to meeting deadlines and the ability to thrive in a fast-paced, multi-tasking environment. This role may require up to 75% travel.
Responsibilities
- Perform node and cluster-level software installation and validation for GPU/compute AI and Machine Learning projects.
- Resolve technical issues for customers utilizing AMD Instinct™ products.
- Provide technical guidance and support to customers for server graphics and compute projects related to AI and Machine Learning workloads.
- Build datacenter GPU dockers and containers for customer testing and deployment.
- Qualify and assess new software functionality to ensure compatibility with customer requirements.
- Assist development teams in identifying and resolving hardware/software technical issues throughout the product lifecycle, from initial hardware bring-up to end-of-life.
- Follow procedures to communicate, report, and escalate incidents to AMD Management.
- Collaborate with program managers to maintain project schedules, track action items, ensure deliverables are met, and provide project status updates to customers and AMD management.
- Develop a strong understanding of the client’s business to ensure impactful and effective task completion.
Requirements
- Bachelor’s or Master’s degree in Computer Engineering, Electrical Engineering, or Computer Science.
- Technical certifications in relevant software systems are highly desirable.
- Exceptional datacenter deployment and troubleshooting skills in AI GPU hardware, software, and networking.
- Highly analytical, detail-oriented, self-motivated, and maintain a positive, results-driven attitude.
Preferred Experience
- Datacenter customer support roles.
- Large-scale cluster deployment within hyperscale datacenters.
- Server architecture and functionality, including remote management, network topologies, and graphics software/hardware subsystems.
- Linux installation, setup, usage, and debugging.
- Virtual environments (e.g., VMWare, Citrix, KVM, Microsoft) and virtual machine setup/management.
- Datacenter GPU software stacks such as AMD ROCm™ or Nvidia CUDA.
- Validating multimode AI clusters using AMD tools (e.g., AGFHC, RCCL RDMA) or equivalents.
- AI/Machine Learning workloads, frameworks, and models.
- Strong debugging, problem-solving, and analytical skills.
- Excellent verbal and written communication skills for conveying technical information.
- Self-starter with attention to detail, organizational skills, and the ability to multitask in a fast-paced environment.
This role is not eligible for visa sponsorship.
Benefits
AMD benefits at a glance. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.