Jobs · Engineering · New York

Lead Site Reliability Engineer

Kontakt.io · New York, NY · 2 wk ago
HybridEngineering$400/hrFull-time

About the role

We’re looking for a Lead Site Reliability Engineer to own the reliability, performance, and automation of our cloud-based, real-time platform. This role will focus on keeping our platform running smoothly 24/7, minimizing downtime, improving observability, incident response, and self-healing automation. You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.

Responsibilities

  • Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.
  • Design and implement self-healing, fault-tolerant systems to prevent failures before they happen.
  • Define SLIs, SLOs, and SLAs, ensuring proactive performance monitoring and incident resolution.
  • Architect and manage scalable cloud infrastructure (AWS) for massive real-time data processing.
  • Optimize containerized environments (Kubernetes, Docker) to support multi-region deployments.
  • Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.
  • Build and refine a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.
  • Lead incident response and on-call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR).
  • Conduct blameless postmortems and continuously improve system resilience.
  • Reduce manual intervention through automated deployment, scaling, and failover mechanisms.
  • Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.
  • Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.
  • Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
  • Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.

Requirements

  • 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
  • Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.
  • Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.
  • Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.
  • Hands-on experience with incident management, postmortems, and building resilient systems.
  • Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).
  • A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team.
  • Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2).

Bonus Points

  • Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.
  • Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines.
  • Prior experience leading on-call rotations and major incident management processes.

Benefits

  • Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors.
  • Full health, dental, and vision coverage.
  • 401k, paid time off, and paid parental leave.
  • All the tools you need to do your best work.

Schedule

Hybrid schedule of 3 days/week minimum from our New York City office.

Similar jobs