Jobs · Engineering · California

AI Systems Engineer - HPC

AMD · San Jose, CA · 1 wk ago
On-siteEngineeringFull-time

About the role

We are seeking an AI Systems Engineer to join our AMD IT compute platforms engineering team. The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers.

Responsibilities

  • Develop, implement, and maintain GPU-based clusters, ensuring optimal performance.
  • Administer ML/AI platforms – Distributed ML services, LLMs and AI inferencing, by managing deployments, resource allocation, monitoring, and security.
  • Automate system provisioning and Cluster management end to end.
  • Collaborate with cross-functional teams to address AI infrastructure requirements, support AI-related projects, and provide technical expertise.
  • Monitor and evaluate the performance of AI systems and clusters, ensuring that they adhere to industry best practices and meet company standards.
  • Use AI/ML to continuously improve internal processes and tools that are used in end-to-end delivery of your services in this team.

Requirements

  • Passion for learning and a curiosity to improve scalable HPC systems.
  • Significant experience working across a globally distributed organization.
  • Strong organizational, problem-solving, and troubleshooting skills, with the ability to manage multiple projects simultaneously.
  • Excellent verbal and written communication skills, with the ability to collaborate effectively with team members and stakeholders at all levels of the organization.

Qualifications

  • Bachelor's or master's degree in computer science or computer engineering preferred.
  • Experience in developing Python-based AI apps and UI.
  • HPC infrastructure engineering for AI/HPC domain.
  • SLURM and Kubernetes management.
  • Managing GPU clusters optimizing GPU-based services/tools/software.
  • Experience in creating web services with HPC backend (like AI).
  • Proficiency in RoCEv2, K8s, KVM, Ubuntu, Python, Shell, GPU drivers, and Cluster interconnect with 400G networking.
  • Demonstrated experience with AI workload schedulers and allocation optimization.
  • Automation/monitoring tools: Ansible / Saltstack, Terraform, Prometheus, Grafana.

Location

San Jose, CA

Benefits

AMD benefits at a glance are offered. For more details, refer to the official AMD benefits information.

Similar jobs