Jobs · Engineering · California

System Software Engineer, Node & Cluster Management

MatX · Mountain View, CA · 1 wk ago
HybridEngineering$120k–$250k/yrFull-time

About the role

MatX's mission is to make the world’s best AI models run as efficiently as allowed by physics, bringing the world years ahead in AI quality and availability. We are seeking a System Software Engineer to join our team as we create best-in-class silicon for high-performance and sustainable GenAI. The MatX host system software team owns everything that makes our AI silicon and systems usable: from Linux kernel drivers up through node and cluster management. The team also co-owns the BMC/OpenBMC firmware stack, ensuring host software and out-of-band management are designed together.

We're looking for self-driven engineers who can take a hardware spec and a register map and start building — prototype drivers, low-level utilities, daemons, and tooling — with minimal hand-holding. Each engineer has a primary focus area, but ownership of overlapping components is shared, and you should expect to venture across the stack.

Responsibilities

  • Design and build the node-level management plane for MatX's AI systems: expose node health, inventory, telemetry, and control operations through HTTP/REST endpoints (e.g., Redfish-style or custom APIs).
  • Design and implement cluster management solutions and failover algorithms to minimize downtime.
  • Build the management CLI utilities that operators and internal engineers use daily — interacting with on-node management and telemetry daemons to query state, run diagnostics, update firmware, and recover devices.
  • Partner with BMC firmware engineers to present unified management and observability across in-band and out-of-band paths — so operators see one coherent node, whether data comes from host daemons or the BMC (e.g., unified Redfish-style views, firmware update orchestration, and recovery flows).
  • Extend node-level capabilities to cluster level: fleet-wide health aggregation, device inventory, alerting hooks, and integration points for customers' fleet-management systems.
  • Get hands-on with the low-level stack: prototype, debug, or unblock yourself by dropping below the API layer into telemetry daemons, driver interfaces, or raw device access utilities.
  • Build tooling and automation for managing lab systems during bring-up: provisioning, test orchestration, regression monitoring.
  • Define software contracts between on-node daemons, the BMC stack, and the management layer — shared-ownership boundaries you'll co-design.
  • Debug production-grade issues spanning management APIs, daemons, kernel drivers, BMC firmware, and hardware.
  • Help shape what "manageable at scale" means for a new hardware platform, from single node to full rack to cluster.

Requirements

  • BS or higher in Computer Science, Electrical Engineering, or equivalent practical experience, with 8+ years in systems software — deep low-level systems experience is required.
  • Strong hands-on Linux systems development experience, including low-level userspace software; comfortable reading and debugging kernel driver and daemon code.
  • Strong programming skills in C plus a systems language suited to services and tooling (Go, Rust, C++, and/or Python).
  • Experience designing and building HTTP/REST APIs and CLI tools for hardware or infrastructure management.
  • Solid understanding of device drivers, telemetry paths, PCIe device behavior, BMC-managed subsystems, and the instinct to investigate when something misbehaves.
  • Experienced debugging across API, daemon, kernel, firmware, and hardware boundaries.
  • Comfortable working with firmware engineers to align host-side and BMC-side management capabilities behind common interfaces.
  • Self-driven and pragmatic: able to stand up a working management endpoint against brand-new hardware with minimal specification.

Bonus Skills

  • Experience with Redfish, OpenBMC, gNMI, IPMI, or other datacenter hardware management standards.
  • Cluster/fleet management experience for GPU or accelerator infrastructure.
  • Experience with hardware bring-up, lab automation, or manufacturing/qualification test infrastructure.
  • Familiarity with firmware update orchestration, secure boot, or attestation flows.

Pay

The US base salary for this full-time position is determined based on role, experience, location, job-related skills, and relevant education and training. Career length is a guideline for compensation:

  • Early Career: $120,000 - $250,000 + equity
  • Mid Career: $175,000 - $362,500 + equity
  • Senior Career: $250,000 - $475,000 + equity

Benefits

  • Time off: 4 weeks PTO (accrued) + 12 company holidays + up to 3 weeks remote work.
  • Health: Company-subsidized medical (Kaiser or Anthem) for employees & dependents, Guardian Dental and Vision insurances, life insurance (employee only), plus HSA and FSA offerings via Lively.
  • Financial Wellbeing: Roth IRA/401K with up to 5% company contribution (even if you don't contribute), 100% company-paid life insurance (up to $300K), and long-term disability insurance.
  • Professional Development: $1,500 annual budget.
  • Team Meals: Onsite team lunch & dinner Monday - Friday, with ordering options via WeBox, Specialty’s, or reimbursement.
  • Commute: Company Uber account or reimbursement for train rides (100% covered).
  • MatX Extras: $50/month for a perk of your choice.
  • Cell & Internet Reimbursement: $35/month for cellular and $40/month for wifi.
  • Mental Wellbeing: 100% paid mental health benefit via SpringHealth and Guardian EAP.
  • Support for Parents: Up to 12 weeks paid parental leave, 10 weeks pregnancy disability leave, flexible return-to-work hours, and Benepass reproductive health & parental benefit.
  • AI Resources: Up to $20K/month plus a dedicated internal AI Tooling Team to support productivity.

Similar jobs