Jobs · Engineering · Colorado

Site Reliability Engineer

Leidos · Boulder, CO · 2 wk ago
On-siteEngineering$87k–$157k/yrFull-time

Kudu Dynamics is a 100% employee-owned company, forged out of a decade of experience in computer network operations and staffed with talent who have built, overseen, and enhanced capabilities throughout the entire USG arsenal. Our team of hackers, engineers, makers, and shakers brings deep experience across research, development, deployment, and operations. Kudu Dynamics is uniquely qualified to anticipate tomorrow’s threats and build the next generation of capabilities. Our team has flexible work hours and work-from-home options. When we work from the offices, we enjoy home-roasted coffee and award-winning workspaces.

About the role

We’re looking for a Site Reliability Engineer who has experience building, deploying, automating, and operating complex compute platforms. You’ll work across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a particular emphasis on creating systems that are reproducible, maintainable, resilient, and easy to operate. A major part of this role will involve Nix and NixOS. We want someone who appreciates declarative systems, reproducible environments, infrastructure-as-code, and reducing configuration drift.

You may be building NixOS-based servers, improving deployment pipelines, developing reusable Nix modules, debugging distributed systems, or making sure a platform can be rebuilt predictably from scratch. Critical components of the environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems.

Responsibilities

  • Own the reliability and operation of critical compute and data platforms.
  • Design, deploy, and maintain Linux infrastructure, including NixOS-based systems.
  • Build reproducible system configurations and deployment workflows using Nix and infrastructure-as-code.
  • Develop reusable NixOS modules, packages, flakes, and system configurations.
  • Improve platform resilience, fault tolerance, recoverability, and maintainability.
  • Build monitoring, logging, alerting, and observability capabilities that help us understand system behavior before customers notice problems.
  • Diagnose difficult failures across Linux, networking, storage, containers, and distributed systems.
  • Automate repetitive operational tasks and eliminate configuration drift.
  • Design and build security and maintainability components for deployed systems.
  • Help define operational standards, deployment practices, upgrade strategies, and disaster-recovery procedures.
  • Perform capacity planning and identify performance bottlenecks across compute, storage, and networking.
  • Own solutions spanning CNO, analytics, security, infrastructure, and field deployments.
  • Learn new skills and teach the rest of us what you know.

Requirements

  • Enthusiasm for learning new stuff.
  • Bachelor’s degree in Computer Science, Computer Engineering, a related field, or amazing equivalent skills.
  • Strong experience administering and troubleshooting Linux systems.
  • Experience with Linux networking, storage, and file systems.
  • Experience automating system configuration and deployment.
  • Experience with Nix and/or NixOS in production, lab, or significant personal environments.
  • Python, Go, Bash, or similar scripting/programming experience of 2+ years.
  • Ability to debug complex systems methodically across multiple layers of the stack.

Nice-to-Have Qualifications

  • Deep experience with NixOS, including custom modules, overlays, flakes, packaging, and reproducible deployments.
  • Experience operating fleets of Linux systems.
  • Infrastructure-as-code experience with Nix, Terraform, Ansible, or similar tooling.
  • Experience designing highly available or fault-tolerant systems.
  • Monitoring and observability experience with tools such as Prometheus, Grafana, Loki, OpenTelemetry, Elasticsearch, or similar systems.

Pay

Pay Range $87,100.00 - $157,450.00

Benefits

Pay and benefits include competitive compensation, Health and Wellness programs, Income Protection, Paid Leave, and Retirement. More details are available at www.leidos.com/careers/pay-benefits.

Similar jobs

Site Reliability Engineer

Cracker BarrelTennessee, United States· 1 mo ago
Engineeringapply on cbrlgroup.wd503.myworkdayjobs.com