Lead Site Reliability Engineer
Kontakt.io · New York, NY · 2 wk ago
HybridEngineering$400/hrFull-time
About the role
We’re looking for a Lead Site Reliability Engineer to own the reliability, performance, and automation of our cloud-based, real-time platform. This role will focus on keeping our platform running smoothly 24/7, minimizing downtime, improving observability, incident response, and self-healing automation. You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.
Responsibilities
- Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.
- Design and implement self-healing, fault-tolerant systems to prevent failures before they happen.
- Define SLIs, SLOs, and SLAs, ensuring proactive performance monitoring and incident resolution.
- Architect and manage scalable cloud infrastructure (AWS) for massive real-time data processing.
- Optimize containerized environments (Kubernetes, Docker) to support multi-region deployments.
- Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.
- Build and refine a world-class monitoring, alerting, and logging system using Prometheus, Grafana, OpenTelemetry, and Datadog.
- Lead incident response and on-call operations, reducing mean time to detection (MTTD) and mean time to resolution (MTTR).
- Conduct blameless postmortems and continuously improve system resilience.
- Reduce manual intervention through automated deployment, scaling, and failover mechanisms.
- Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.
- Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.
- Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
- Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.
Requirements
- 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
- Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.
- Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.
- Strong background in monitoring, logging, and observability with Prometheus, OpenTelemetry, or similar tools.
- Hands-on experience with incident management, postmortems, and building resilient systems.
- Deep knowledge of CI/CD automation, GitOps, and infrastructure as code (Terraform, etc.).
- A mature leadership approach, with the ability to drive technical strategy while growing and mentoring a high-performance SRE team.
- Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC 2).
Bonus Points
- Experience with healthcare IT, including EHR data, FHIR, and HL7 interoperability.
- Expertise in real-time distributed systems, event-driven architectures, or large-scale data pipelines.
- Prior experience leading on-call rotations and major incident management processes.
Benefits
- Equity in a high-growth company scaling toward $400M+ ARR and backed by leading investors.
- Full health, dental, and vision coverage.
- 401k, paid time off, and paid parental leave.
- All the tools you need to do your best work.
Schedule
Hybrid schedule of 3 days/week minimum from our New York City office.