Jobs · Engineering · Ohio

Site Reliability Engineer

Leidos · Columbus, OH · 2 wk ago
On-siteEngineering$87k–$157k/yrFull-time

About the Role

We're a small team in a fun company looking for someone to come in and help us build systems that stay reliable when things get complicated. We need a Site Reliability Engineer who has experience building, deploying, automating, and operating complex compute platforms. You'll work across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a particular emphasis on creating systems that are reproducible, maintainable, resilient, and easy to operate.

A major part of this role will involve Nix and NixOS. We want someone who appreciates declarative systems, reproducible environments, infrastructure-as-code, and reducing configuration drift. You may be building NixOS-based servers, improving deployment pipelines, developing reusable Nix modules, debugging distributed systems, or making sure a platform can be rebuilt predictably from scratch.

Critical components of the environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems. You don't need to be an expert at everything. Whatever part you bite off, though, will be yours to own. This is a great opportunity to apply deep systems expertise while learning adjacent areas.

Our team needs your help making sure our customers are repeatedly provided with operational insights from a multi-domain data environment—and that the infrastructure producing those insights is dependable, observable, reproducible, and recoverable.

Responsibilities

  • Own the reliability and operation of critical compute and data platforms
  • Design, deploy, and maintain Linux infrastructure, including NixOS-based systems
  • Build reproducible system configurations and deployment workflows using Nix and infrastructure-as-code
  • Develop reusable NixOS modules, packages, flakes, and system configurations
  • Improve platform resilience, fault tolerance, recoverability, and maintainability
  • Build monitoring, logging, alerting, and observability capabilities that help us understand system behavior before customers notice problems
  • Diagnose difficult failures across Linux, networking, storage, containers, and distributed systems
  • Automate repetitive operational tasks and eliminate configuration drift
  • Design and build security and maintainability components for deployed systems
  • Help define operational standards, deployment practices, upgrade strategies, and disaster-recovery procedures
  • Perform capacity planning and identify performance bottlenecks across compute, storage, and networking
  • Own solutions spanning CNO, analytics, security, infrastructure, and field deployments
  • Learn new skills and teach the rest of us what you know

Minimum Qualifications

  • Enthusiasm for learning new stuff
  • Bachelor's degree in Computer Science, Computer Engineering, a related field, or amazing equivalent skills
  • Strong experience administering and troubleshooting Linux systems
  • Experience with Linux networking, storage, and file systems
  • Experience automating system configuration and deployment
  • Experience with Nix and/or NixOS in production, lab, or significant personal environments
  • Python, Go, Bash, or similar scripting/programming experience of 2+ years
  • Ability to debug complex systems methodically across multiple layers of the stack

Nice-to-Have Qualifications

  • Deep experience with NixOS, including custom modules, overlays, flakes, packaging, and reproducible deployments
  • Experience operating fleets of Linux systems
  • Infrastructure-as-code experience with Nix, Terraform, Ansible, or similar tooling
  • Experience designing highly available or fault-tolerant systems
  • Monitoring and observability experience with tools such as Prometheus, Grafana, Loki, OpenTelemetry, Elasticsearch, or similar systems

Pay

$87,100.00 - $157,450.00

Similar jobs

Site Reliability Engineer

Cracker BarrelTennessee, United States· 1 mo ago
Engineeringapply on cbrlgroup.wd503.myworkdayjobs.com