Site Reliability Engineer
Enzo Health · Lehi, UT · Yesterday
On-siteEngineeringFull-time
About the role
Enzo Health is hiring a Senior Site Reliability Engineer to join the Security and Site Reliability team. The role focuses on a stable, scalable AWS and Kubernetes platform, reliable Postgres operations, and safer delivery.
What You'll Do
- Strengthen the production platform
- Operate and improve our AWS and Kubernetes environments
- Own infrastructure changes through Terraform, including modules, state, review standards, and drift control
- Improve environment management, cluster practices, and deployment reliability
- Make CI/CD and releases safer through clear checks, reliable rollbacks, environment consistency, and release observability
- Improve capacity planning, resource controls, resilience, and cost visibility
- Build automation that removes repetitive operational work and reduces avoidable failures
- Improve Postgres reliability
- Own the operational health of Postgres in production
- Verify backups and regularly test restore procedures against agreed recovery targets
- Improve monitoring for connections, storage, slow queries, locks, and other important failure signals
- Guide safe schema migrations, database access, and production change procedures
- Find performance and capacity risks before they affect customers
- Maintain clear database runbooks for common failures and emergency work
- Mature production operations
- Improve dashboards, monitors, logs, traces, and alert routing
- Validate existing service-level indicators and objectives, then close important coverage gaps
- Participate in and improve the company-wide on-call rotation
- Write practical runbooks and make escalation paths clear for engineers across the company
- Lead or support incident response, recovery, and blameless post-incident reviews
- Use incident and reliability data to set priorities and measure improvement
- Work across Security and engineering
- Work with Security to keep infrastructure changes consistent with SOC 2 and health care security requirements
- Apply least-privilege access, secure defaults, secrets management, encryption, and auditable change practices
- Support vulnerability remediation, disaster-recovery exercises, and secure production access
- Help product teams include reliability and operational risk in technical decisions
Qualifications
- At least 5 years of experience in site reliability, platform, infrastructure, or production engineering
- Strong hands-on experience with AWS and production Kubernetes
- Strong experience with Terraform and infrastructure as code
- Experience operating Postgres in production, including backup and restore, performance, and safe migrations
- Experience building and operating CI/CD and deployment systems
- Experience with modern observability tools across metrics, logs, and traces
- Strong programming or scripting skills for automation and operational tools
- A strong record of diagnosing production failures and leading incidents through recovery
- Able to write clear automation, runbooks, and technical standards
- Able to work independently and partner well with application engineers
Nice to have
- Experience with Kubernetes security controls, policy as code, and software supply-chain security
- Experience in health care or another regulated environment
- Experience supporting SOC 2 controls or audit evidence
- Experience improving cloud cost and capacity efficiency