Jobs · Engineering · Utah

Site Reliability Engineer

Enzo Health · Lehi, UT · Yesterday
On-siteEngineeringFull-time

About the role

Enzo Health is hiring a Senior Site Reliability Engineer to join the Security and Site Reliability team. The role focuses on a stable, scalable AWS and Kubernetes platform, reliable Postgres operations, and safer delivery.

What You'll Do

  • Strengthen the production platform
  • Operate and improve our AWS and Kubernetes environments
  • Own infrastructure changes through Terraform, including modules, state, review standards, and drift control
  • Improve environment management, cluster practices, and deployment reliability
  • Make CI/CD and releases safer through clear checks, reliable rollbacks, environment consistency, and release observability
  • Improve capacity planning, resource controls, resilience, and cost visibility
  • Build automation that removes repetitive operational work and reduces avoidable failures
  • Improve Postgres reliability
  • Own the operational health of Postgres in production
  • Verify backups and regularly test restore procedures against agreed recovery targets
  • Improve monitoring for connections, storage, slow queries, locks, and other important failure signals
  • Guide safe schema migrations, database access, and production change procedures
  • Find performance and capacity risks before they affect customers
  • Maintain clear database runbooks for common failures and emergency work
  • Mature production operations
  • Improve dashboards, monitors, logs, traces, and alert routing
  • Validate existing service-level indicators and objectives, then close important coverage gaps
  • Participate in and improve the company-wide on-call rotation
  • Write practical runbooks and make escalation paths clear for engineers across the company
  • Lead or support incident response, recovery, and blameless post-incident reviews
  • Use incident and reliability data to set priorities and measure improvement
  • Work across Security and engineering
  • Work with Security to keep infrastructure changes consistent with SOC 2 and health care security requirements
  • Apply least-privilege access, secure defaults, secrets management, encryption, and auditable change practices
  • Support vulnerability remediation, disaster-recovery exercises, and secure production access
  • Help product teams include reliability and operational risk in technical decisions

Qualifications

  • At least 5 years of experience in site reliability, platform, infrastructure, or production engineering
  • Strong hands-on experience with AWS and production Kubernetes
  • Strong experience with Terraform and infrastructure as code
  • Experience operating Postgres in production, including backup and restore, performance, and safe migrations
  • Experience building and operating CI/CD and deployment systems
  • Experience with modern observability tools across metrics, logs, and traces
  • Strong programming or scripting skills for automation and operational tools
  • A strong record of diagnosing production failures and leading incidents through recovery
  • Able to write clear automation, runbooks, and technical standards
  • Able to work independently and partner well with application engineers

Nice to have

  • Experience with Kubernetes security controls, policy as code, and software supply-chain security
  • Experience in health care or another regulated environment
  • Experience supporting SOC 2 controls or audit evidence
  • Experience improving cloud cost and capacity efficiency

Similar jobs

Site Reliability Engineer

Cracker BarrelTennessee, United States· 1 mo ago
Engineeringapply on cbrlgroup.wd503.myworkdayjobs.com