Jobs · Engineering · California

Data Center Operations Engineer

Cadence · San Jose, CA · 3 days ago
Engineering$57.21–$106.25/hrFull-time

Key Responsibilities

  • Provide hands-on operational support for all data center projects, deployments, and repair activities.
  • Participate in an on-call rotation and provide on-site or remote support during maintenance windows and incidents.
  • Troubleshoot and resolve operational issues related to Linux servers, GPU platforms, networking, and storage infrastructure.
  • Support customer and internal deployments, ensuring timely and successful bring-up of GPU servers and clusters.
  • Perform InfiniBand fabric bring-up, switch configuration, subnet management, and troubleshooting.
  • Conduct daily health checks of Linux systems and infrastructure components, proactively identifying and mitigating risks.
  • Install, configure, test, and maintain server hardware (rack and stack, labeling, HDDs, memory, CPUs, RAID batteries, NICs, etc.).
  • Install, configure, and troubleshoot networking equipment including routers, switches, and terminal servers for out-of-band management.
  • Review and validate equipment deployments against approved design documentation and standards.
  • Collaborate with global teams across time zones to support operational initiatives and continuous improvement efforts.
  • Contribute to process improvement initiatives and ensure adherence to documented policies, processes, and procedures.

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
  • Strong hands-on experience in Linux environments, including system administration, troubleshooting, and performance validation.
  • Proficiency with Linux command-line tools and shell scripting (Bash or equivalent).
  • Hands-on experience setting up and validating GPU servers in clustered environments.
  • Experience with end-to-end GPU testing in InfiniBand-based clusters.
  • Working knowledge of InfiniBand networking, including switch configuration and subnet management.
  • Solid understanding of networking fundamentals, including the OSI model and TCP/IP protocol suite (IP, ARP, ICMP, TCP, UDP, SMTP, FTP, TFTP).
  • Experience installing, configuring, and troubleshooting routers, switches, and terminal servers.
  • Familiarity with fiber and copper cabling, including IP and SAN deployments.
  • Experience managing incident tickets, maintaining acceptable ticket loads, and meeting SLAs.
  • Strong organizational skills with meticulous attention to detail in data center environments.
  • Able to follow and enforce documented escalation procedures and operational policies.
  • Strong verbal and written communication skills, with the ability to collaborate effectively with cross-functional and global teams.

Preferred Qualifications

  • Experience supporting HPC, AI, or large-scale GPU environments.
  • Exposure to data center monitoring.
  • Experience documenting operational processes and maintaining technical runbooks.
  • Familiarity with large-scale data center buildouts or refresh programs.

Similar jobs