Jobs · Information Technology · California

Systems Operations Support Engineer — Linux

Vast.ai · Los Angeles, CA · 4 wk ago
On-siteInformation Technology$90k–$160k/yrFull-time

About the role

This is a systems operations support role focused on deep-diving into escalated infrastructure issues that go beyond frontline triage. You’ll be the engineering resource our L1 support team relies on when tickets become complex, investigating and resolving issues across the full infrastructure stack—including hardware, BIOS and firmware, networking, Ubuntu, Docker, NVIDIA CUDA and GPUs, and KVM-based virtual machines.

  • Own complex escalations end-to-end: gathering evidence, reproducing issues, identifying the root cause, proposing solutions, and working with the appropriate teams to bring each issue to resolution.
  • The best engineers in this role don’t just resolve individual tickets—they identify recurring patterns, improve operational tooling, and build runbooks that prevent future incidents.
  • Collaborate directly with the engineering and host support teams on systemic infrastructure issues.

Key Responsibilities

  • Handle escalated support tickets involving GPU workload failures, container issues, networking problems, account infrastructure, and host-side configuration.
  • Provide managed support for supplier onboarding and ongoing machine management, acting as a technical resource through installation, configuration, and post-setup troubleshooting.
  • Assist clients and infrastructure suppliers working with TensorFlow, PyTorch, and other GPU-accelerated workloads.
  • Diagnose and resolve issues across Docker, NVIDIA CUDA/GPU drivers, and KVM virtualization environments.
  • Troubleshoot network-layer issues, including VLAN, DNS, DHCP, VPN, NAT, firewall rules, and connectivity failures on host machines.
  • Investigate performance issues involving GPU utilization, container resource constraints, thermal throttling, driver conflicts, and disk I/O bottlenecks.
  • Advisory suppliers on installation best practices, including hardware setup, driver configuration, BIOS/firmware settings, and network configuration for optimal performance.
  • Write and maintain internal runbooks, escalation guides, and knowledge base articles to reduce repeat escalations.
  • Create diagnostic and automation tooling in Python and Bash to reduce manual triage overhead.
  • Collaborate with the engineering and support teams to flag and document systemic or recurring platform issues.

Qualifications

  • Experienced with Linux, especially Ubuntu, and comfortable troubleshooting from the command line.
  • Enjoy debugging difficult problems and fixing broken systems.
  • Methodical and focused on finding root causes, not just temporary fixes.
  • Able to manage complex tickets independently.
  • A clear written communicator with an interest in AI infrastructure and GPU computing.

Must-Haves

  • Strong Linux systems operations experience with Ubuntu, RHEL/CentOS, or Debian, including networking, storage, services, and permissions.
  • Proficiency with Docker, including container debugging, Docker Compose, image management, cgroup limits, and Docker storage and filesystem troubleshooting.
  • Experience with virtualization platforms such as Proxmox VE, VMware, or similar hypervisors, including VM provisioning and troubleshooting.
  • Strong networking fundamentals, including VLANs, DNS, DHCP, NAT, VPNs, firewall rules, and L2/L3 troubleshooting.
  • Hands-on experience with NVIDIA GPU drivers, CUDA, and GPU workload troubleshooting.
  • Python and Bash scripting skills for automation and diagnostic tooling.
  • Strong written English communication that is clear, professional, and technically precise.
  • Experience providing technical support in a customer-facing or internal help desk environment.
  • Ability to prioritize across a concurrent queue of escalated tickets, triaging by severity and customer impact, balancing reactive resolution against proactive documentation and tooling work, and making clear judgment calls on when to escalate versus own resolution end-to-end.

Annual Salary Range

$90,000 – $160,000 + equity + benefits

Similar jobs

Linux Systems Engineer

Aura HealthFort Meade, MD· 1 wk ago
Information Technologyapply on portal.recruitrookie.com

Linux Systems Engineer

NTT DATA North AmericaSalt Lake City, UT· 2 wk ago
Information Technology$30/hrapply on www1.jobdiva.com

Linux Systems Engineer

CACI International IncSterling, VA· 3 wk ago
OTHR$104k–$218k/yrapply on searchcareers.caci.com

Linux Systems Engineer

CollaberaCharlotte, NC· 1 mo ago
Information Technology$60–$70/hrapply on collabera.com

Linux Systems Engineer

Abile Group, LLCSt Louis, MO· 1 mo ago
Information Technologyapply on careers-abilegroup.icims.com