Jobs · Engineering · California

Senior Systems Engineer, OS Automation

CoreWeave · Sunnyvale, CA · 2 days ago
Engineering$153k–$242k/yrFull-time

About The Role

As a Senior Software Engineer on the Automation team, you will design, build, and operate the services, APIs, and libraries that sit behind our OS image, payload, and boot-configuration systems — the software platform other HAVOCK engineers and partner teams rely on to release, test, and ship node software quickly and safely. You'll work on a constraint-solver–based service that resolves compatibility between images, kernels, drivers, payloads, and hardware into a single validated configuration; an end-to-end test framework that validates OS images on real hardware; a library suite for declaratively configuring node storage; and natural-language tooling that lets stakeholders query and interact with our systems. This is a software- and platform-engineering role first, with a clear forward trajectory toward AI-assisted automation — log triage, regression detection, natural-language interfaces to infrastructure — but the core of the job is designing and shipping reliable services and APIs.

What You'll Work On

  • Own and evolve a boot-configuration service that models complex compatibility and dependency relationships between OS images, kernels, drivers, payloads, and instance types as a constraint-solved graph, exposed through clean, well-specified interfaces.
  • Extend our Kubernetes-native, end-to-end test framework that validates OS images and configuration on real hardware, plus the broader testing and validation story for the team.
  • Build and maintain a library suite for declaratively configuring node storage — filesystems, mount options, block-device selection — with configurable strictness.
  • Ship changes to our versioned, boot-time payload system (networking, storage, Kubernetes join) that's published as artifacts and consumed during node bring-up.
  • Grow our natural-language / chat interface that lets stakeholders query and interact with the team's systems.
  • Design and evolve versioned service contracts (gRPC / Connect-RPC, Protobuf) with strong correctness guarantees and robust validation.
  • Build tooling that meaningfully shortens the build-and-release loop.
  • Lower the barrier to entry for everyone who touches this software, and lay the groundwork — clean interfaces, structured build/test metadata — for future AI-assisted automation across build triage, regression detection, and natural-language infrastructure tooling.
  • Operate the services you build: participate in an on-call rotation for the team's services and own their reliability.

Who You Are

  • 3+ years of professional software engineering experience building and operating backend services, platforms, or developer/infrastructure tooling.
  • Strong proficiency in Go and/or Python, with the ability to work fluently across both.
  • Experience designing and maintaining APIs and service contracts (REST, gRPC, or similar), with an eye for clean, well-specified, versioned interfaces.
  • A demonstrated instinct for data modeling — representing relationships, constraints, and dependencies in code (graphs, constraint solving, relational models, or similar).
  • Solid testing discipline: you write services that are testable, and you build the automation that proves they work.
  • Comfort operating in a Kubernetes-based environment and reasoning about how software is built, packaged, deployed, and released.
  • A working understanding of how Linux systems boot and are configured (the OS image / cloud-init / provisioning lifecycle), even if you haven't owned it end to end.
  • A collaborative, software-development-lifecycle mindset (sprints, planning, code review, design docs) and the judgment to refactor toward simplicity.

Preferred Experience

  • Experience modeling complex problems in novel ways.
  • gRPC / Connect-RPC and Protobuf experience, including evolving service contracts safely over time.
  • Familiarity with bare-metal or node provisioning — PXE-style network boot, cloud-init, OS image building, firmware/driver enablement.
  • Fluency with NVIDIA GPU platforms.
  • Rust experience and/or workflow orchestration tools like Argo Workflows.
  • Linux packaging and repository management, configuration management (e.g., Ansible), and shell-based build pipelines.
  • Interest in applying LLMs, RAG, and predictive modeling to large-scale infrastructure automation.

Technical Stack

  • Languages: Go (primary), Python, bash/sh; Rust (test framework); Protobuf
  • APIs & RPC: gRPC, Connect-RPC, HTTP/2, mTLS
  • Orchestration & Infra: Kubernetes, Custom Resources, Helm, GitOps, workflow orchestration, Docker/containerd
  • Node bring-up: cloud-init, OS image builds, GPU drivers, network-boot tooling
  • Storage & Packaging: S3-compatible object storage, Linux package management, configuration management
  • CI/CD & Observability: CI/CD pipelines, Prometheus-style metrics, dashboards

Pay

The base salary range for this role is $153,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption

Similar jobs