Sr. Site Reliability Engineer
Position Summary
Lead Site Reliability Engineering (SRE) initiatives by defining and improving SLIs, SLOs, error budgets, and reliability standards for clinical-grade applications and platform services.
Own production incident management, including high-severity incident response, root cause analysis, blameless post-mortems, and corrective action planning.
Design and maintain observability solutions across logs, metrics, alerting, and distributed tracing using modern monitoring platforms.
Optimize Kubernetes-based containerized environments, including service mesh, ingress, autoscaling, performance tuning, and capacity planning.
Develop automation, Infrastructure as Code (IaC), runbooks-as-code, and self-healing systems to improve operational efficiency and reduce manual support effort.
Partner with engineering teams throughout the SDLC to embed reliability best practices into architecture, code reviews, CI/CD pipelines, release management, and deployment processes.
Drive disaster recovery, chaos engineering, load testing, and resiliency initiatives to improve system availability, scalability, and reduce MTTR/MTTD.
Leverage expertise in AWS, Azure, or GCP, Linux, networking, distributed systems, Kubernetes, and observability tools, while supporting compliance requirements such as HIPAA and HITRUST in a healthcare environment.
Ability to participate in an on-call rotation
Desired Professional Skills and Experience
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering supporting high-availability production systems
- Bachelor's degree in Information Technology, Computer Science, Engineering, or a related field preferred, Master’s preferred
- Deep expertise in Kubernetes, Linux, and cloud platforms such as AWS (EKS, IAM, VPC), Azure, or GCP
- Strong hands-on experience with Python, Go, or Bash, Infrastructure as Code (Terraform, CloudFormation), and implementing SLIs, SLOs, and reliability engineering best practices in production environments
- Expertise with observability and monitoring platforms, including Datadog, Prometheus, Grafana, CloudWatch, and distributed tracing, as well as production incident management, troubleshooting, and performance optimization at scale
- Experience supporting healthcare IT and Enterprise Imaging environments, including PACS, EMRs, VNAs, Voice Recognition systems, and healthcare interoperability standards such as DICOM and HL7
- Hands-on experience with Kubernetes service mesh and ingress technologies (Istio, Linkerd), CI/CD tools (Jenkins, ArgoCD, GitHub Actions), container security platforms (CrowdStrike, Aqua, Sysdig), and modern SDLC practices
- Strong understanding of HIPAA, HITRUST, and regulated cloud environments, with experience in networking (TCP/IP, HTTP, gRPC, DNS), databases (PostgreSQL, MySQL, Oracle), and building resilient cloud-native solutions
- PREFERRED CERTIFICATIONS: AWS Certified Solutions Architect, AWS Certified DevOps Engineer, or Certified Kubernetes Administrator (CKA) / Certified Kubernetes Application Developer (CKAD)