Site Reliability Engineer
About the Role
We're a small team in a fun company looking for someone to come in and help us build systems that stay reliable when things get complicated. We need a Site Reliability Engineer who has experience building, deploying, automating, and operating complex compute platforms. You'll work across infrastructure, Linux systems, networking, distributed storage, observability, security, and deployment automation, with a particular emphasis on creating systems that are reproducible, maintainable, resilient, and easy to operate.
A major part of this role will involve Nix and NixOS. We want someone who appreciates declarative systems, reproducible environments, infrastructure-as-code, and reducing configuration drift. You may be building NixOS-based servers, improving deployment pipelines, developing reusable Nix modules, debugging distributed systems, or making sure a platform can be rebuilt predictably from scratch.
Critical components of the environment may include high-performance compute, distributed Linux file systems, network design, security-in-depth, ML/AI infrastructure, high-bandwidth data processing, cloud deployment, and fielded systems. You don't need to be an expert at everything. Whatever part you bite off, though, will be yours to own. This is a great opportunity to apply deep systems expertise while learning adjacent areas.
Our team needs your help making sure our customers are repeatedly provided with operational insights from a multi-domain data environment—and that the infrastructure producing those insights is dependable, observable, reproducible, and recoverable.
Responsibilities
- Own the reliability and operation of critical compute and data platforms
- Design, deploy, and maintain Linux infrastructure, including NixOS-based systems
- Build reproducible system configurations and deployment workflows using Nix and infrastructure-as-code
- Develop reusable NixOS modules, packages, flakes, and system configurations
- Improve platform resilience, fault tolerance, recoverability, and maintainability
- Build monitoring, logging, alerting, and observability capabilities that help us understand system behavior before customers notice problems
- Diagnose difficult failures across Linux, networking, storage, containers, and distributed systems
- Automate repetitive operational tasks and eliminate configuration drift
- Design and build security and maintainability components for deployed systems
- Help define operational standards, deployment practices, upgrade strategies, and disaster-recovery procedures
- Perform capacity planning and identify performance bottlenecks across compute, storage, and networking
- Own solutions spanning CNO, analytics, security, infrastructure, and field deployments
- Learn new skills and teach the rest of us what you know
Minimum Qualifications
- Enthusiasm for learning new stuff
- Bachelor's degree in Computer Science, Computer Engineering, a related field, or amazing equivalent skills
- Strong experience administering and troubleshooting Linux systems
- Experience with Linux networking, storage, and file systems
- Experience automating system configuration and deployment
- Experience with Nix and/or NixOS in production, lab, or significant personal environments
- Python, Go, Bash, or similar scripting/programming experience of 2+ years
- Ability to debug complex systems methodically across multiple layers of the stack
Nice-to-Have Qualifications
- Deep experience with NixOS, including custom modules, overlays, flakes, packaging, and reproducible deployments
- Experience operating fleets of Linux systems
- Infrastructure-as-code experience with Nix, Terraform, Ansible, or similar tooling
- Experience designing highly available or fault-tolerant systems
- Monitoring and observability experience with tools such as Prometheus, Grafana, Loki, OpenTelemetry, Elasticsearch, or similar systems
Pay
$87,100.00 - $157,450.00