Jobs · Engineering · Texas

Sr. Data Center Systems Engineer

AMD · Austin, TX · 3 days ago
On-siteEngineeringFull-time

About the Role

Join AMD's Data Center Platform Engineering Group (DPEG) and play a critical role in enabling the reliability, deployment, and availability of next-generation data center platforms. In this highly collaborative, hands-on engineering role, you will work across hardware, firmware, software, and systems teams to diagnose complex technical challenges and ensure platform stability at scale. You will have the opportunity to drive root-cause investigations, validate cutting-edge server technologies, and partner with internal engineering teams, OEMs, ODMs, and vendors to resolve critical issues. This role offers exposure to advanced CPU, GPU, and platform technologies while providing the opportunity to influence product quality, system performance, and operational excellence.

The Person

The ideal candidate is a highly motivated and self-driven engineer with a passion for solving complex technical problems. They thrive in a fast-paced environment, possess strong debugging and analytical skills, and can effectively collaborate across multiple teams and disciplines. Success in this role requires excellent communication skills, strong ownership, attention to detail, and the ability to drive issues to resolution while balancing multiple priorities.

Key Responsibilities

  • Drive day-to-day platform support activities, including system availability, workload failures, thermal, power, networking, clustering, and capacity-related issues.
  • Develop and execute debug plans, methodologies, and strategies to identify root causes and resolve complex system-level issues.
  • Coordinate and track issue resolution efforts across engineering teams while ensuring timely communication and escalation.
  • Collaborate with hardware, firmware, software, reliability, and validation teams to define requirements and assess technical risks.
  • Develop and execute laboratory validation plans to ensure platform functionality, reliability, and performance.
  • Partner with OEMs, ODMs, vendors, and internal stakeholders to investigate issues and drive corrective actions.
  • Apply system-level reliability knowledge to improve platform quality and operational efficiency.
  • Drive continuous improvement initiatives related to debug processes, validation methodologies, automation, and overall engineering effectiveness.
  • Contribute to the development of scalable support and validation processes that improve long-term platform reliability and availability.

Preferred Experience

  • Experience debugging complex hardware, firmware, or system-level issues and driving root-cause investigations.
  • Knowledge of server, CPU, GPU, SoC, microcontroller, or data center platform architectures.
  • Experience with hardware validation, verification, and reliability methodologies.
  • Familiarity with Linux environments and system debug techniques.
  • Knowledge of BMC, BIOS, FPGA, and CPLD technologies.
  • Experience using issue tracking and project management tools such as JIRA.
  • Scripting and test automation experience.
  • Understanding of high-speed industry-standard interfaces and protocols, including PCIe.
  • Experience collaborating with OEMs, ODMs, suppliers, and external partners.
  • Familiarity with laboratory equipment and hardware bring-up environments.
  • Strong analytical, troubleshooting, and problem-solving abilities.
  • Excellent verbal and written communication skills.

Academic Credentials

  • Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, or a related technical field preferred.
  • Master's degree in a relevant engineering discipline desired.
  • Relevant industry certifications or specialized training in systems engineering, validation, or reliability engineering are beneficial.

Location: Rockdale, Texas, US

Similar jobs

Sr. Data Center Engineer

OpenTextUnited States· 2 wk ago
RemoteInformation Technology$88k–$132k/yrapply on careers.opentext.com