Jobs · Engineering · California

Datacenter Field Engineer

Sciforium · San Jose, CA · Yesterday
On-siteEngineering$150k–$180k/yrFull-time

Key Responsibilities

  • On-Call Response: Serve as the primary point of contact for physical system outages, hardware failures, and network interruptions to minimize downtime.
  • Cluster Monitoring: Proactively monitor hardware health, including GPU thermals, power draw, and physical system loads, catching anomalies before they impact active workloads.
  • Vendor Liaison: Work closely with data center facility staff and third-party hardware vendors to coordinate RMA processes, physical repairs, part replacements, and routine maintenance.
  • Hardware Deployment: Rack, cable, and lead the physical bring-up of new GPU nodes, ensuring power and network connectivity are fully integrated into the existing cluster.

Linux & Network Administration

  • OS Management: Install, patch, and maintain Linux operating systems (Ubuntu/CentOS/RHEL) across the cluster bare-metal servers.
  • Security & Access: Configure and maintain edge and internal networking, including firewalls, VPNs, and strict SSH access controls to secure our infrastructure.
  • Identity & Storage Management: Administer LDAP/Active Directory for centralized user authentication and ensure network storage systems (NFS/GPFS/Lustre) are reliably mounted and properly permissioned.

Qualifications

  • Must-Haves: 3+ years of experience in Linux Systems Administration (deep knowledge of boot processes, systemd, disk management, etc.). Strong background in server hardware troubleshooting, specifically within high-density environments (power, cooling, PCIe topologies).
  • Experience managing networking security (VPNs, iptables/firewalld, VLANs) and directory services (LDAP/FreeIPA/Active Directory).
  • Proficiency in Bash scripting for essential system automation.

Nice-to-Haves

  • Experience using configuration management tools like Ansible, SaltStack, or Terraform for OS provisioning.
  • Familiarity with data center operations, cooling requirements for high-TDP accelerators (like NVIDIA H100 or AMD MI300).

Benefits

  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity

Career Equity

  • Equal opportunity

Similar jobs