Jobs · Engineering

Lead Site Reliability Engineer

Gifthealth · United States · 5 days ago
RemoteRemoteEngineering$123k–$154k/yrFull-time

Description

About Us At Gifthealth, we're revolutionizing the way people experience healthcare by simplifying the process of managing prescriptions and health services. Our mission is to provide a seamless, personalized, and efficient healthcare experience for all our customers. We're a dynamic, innovative, and customer-centric company dedicated to making a positive impact on people's lives.

Position Summary

Reporting to the Director of Engineering, the Lead Site Reliability Engineer (SRE) is a senior technical contributor responsible for building reliable, scalable software systems and the DevOps practices that support them. This role blends software engineering, operational excellence, and automation to improve the performance and resilience of Gifthealth’s applications. We are seeking a Lead SRE to play a key part in enabling fast, safe delivery of customer-facing features. This role partners closely with product and application engineers to embed reliability, observability, and operational ownership directly into the development lifecycle, ensuring alignment with organizational goals, operational excellence, and compliance standards.

Key Responsibilities

  • Designs, builds, and maintains reliable, scalable software systems supporting Ruby on Rails applications
  • Embs reliability, performance, and operational best practices into application code and development workflows
  • Owes DevOps practices including CI/CD reliability, deployment strategies, and release safety
  • Led incident response, debugging, and root cause analysis across application and platform layers
  • Implements and evolves observability (logging, metrics, tracing) within application and service code
  • Partners with engineering teams on architecture, capacity planning, and technical standards

Qualifications

  • Education: Bachelor’s degree in computer science, engineering, or related field OR equivalent professional experience in software engineering, SRE, or DevOps roles
  • Licensure/Certification: Cloud platform certifications (AWS, GCP, Azure) (Preferred) SRE or DevOps-focused certifications (Preferred)
  • Experience: 5+ years of experience in software engineering, SRE, or DevOps roles
  • Hands-on experience building and operating Ruby on Rails applications in production
  • Experience in owning production incidents and application-level reliability
  • Experience in high-growth or scaling engineering organizations (Preferred)
  • Experience working in regulated or customer-impact–sensitive environments (Preferred)
  • Knowledge, Skills, & Abilities: Knowledge of Ruby on Rails application architecture and production operations; software reliability engineering principles (SLOs, SLIs, error budgets); and modern DevOps and CI/CD practices (Required) Knowledge of security and compliance considerations in production systems (Preferred) Strong software engineering skills (Ruby and/or comparable backend languages) (Required) Debugging and performance optimization of production applications skills (Required) CI/CD pipelines, deployment automation, and release tooling skills (Required) Monitoring and observability tooling (Datadog, New Relic, Prometheus, etc.) skills (Required) Infrastructure as Code (Terraform or similar) skills (Preferred) Containerization and orchestration (Docker) skills (Preferred) Ability to write production-quality code that improves system reliability (Required) Ability to collaborate with product and engineering teams to influence design decisions (Required) Ability to troubleshoot complex, cross-system failures (Required) Ability to mentor engineers on operational ownership and reliability practices (Preferred) Ability to balance speed of delivery with long-term system health (Preferred)

Work Environment

Location: Remote
Schedule: 8:00 A.M. to 5:00 P.M. Monday through Friday with night and weekend hours on occasion as determined by the needs of the business.
Regular meetings with internal Backend and Full-Stack Engineers, Engineering Managers, and Product and Security teams. This role may also have meetings with external cloud and tooling vendor representatives.

Key Essential Functions

  • Must be able to remain in a stationary position for extended periods while writing or reviewing documentation
  • Must be able to work on a computer for the entire shift
  • Must be able to attend virtual meetings with cross-functional teams.

Salary Description

$123,000- $154,000

Similar jobs