Jobs · Engineering

Sr. Site Reliability Engineer

Mosaic Clinical Technologies · United States · 1 wk ago
RemoteRemoteEngineering$125k–$145k/yrFull-time

Position Summary

Lead Site Reliability Engineering (SRE) initiatives by defining and improving SLIs, SLOs, error budgets, and reliability standards for clinical-grade applications and platform services.

Own production incident management, including high-severity incident response, root cause analysis, blameless post-mortems, and corrective action planning.

Design and maintain observability solutions across logs, metrics, alerting, and distributed tracing using modern monitoring platforms.

Optimize Kubernetes-based containerized environments, including service mesh, ingress, autoscaling, performance tuning, and capacity planning.

Develop automation, Infrastructure as Code (IaC), runbooks-as-code, and self-healing systems to improve operational efficiency and reduce manual support effort.

Partner with engineering teams throughout the SDLC to embed reliability best practices into architecture, code reviews, CI/CD pipelines, release management, and deployment processes.

Drive disaster recovery, chaos engineering, load testing, and resiliency initiatives to improve system availability, scalability, and reduce MTTR/MTTD.

Leverage expertise in AWS, Azure, or GCP, Linux, networking, distributed systems, Kubernetes, and observability tools, while supporting compliance requirements such as HIPAA and HITRUST in a healthcare environment.

Ability to participate in an on-call rotation

Desired Professional Skills and Experience

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Production Engineering supporting high-availability production systems
  • Bachelor's degree in Information Technology, Computer Science, Engineering, or a related field preferred, Master’s preferred
  • Deep expertise in Kubernetes, Linux, and cloud platforms such as AWS (EKS, IAM, VPC), Azure, or GCP
  • Strong hands-on experience with Python, Go, or Bash, Infrastructure as Code (Terraform, CloudFormation), and implementing SLIs, SLOs, and reliability engineering best practices in production environments
  • Expertise with observability and monitoring platforms, including Datadog, Prometheus, Grafana, CloudWatch, and distributed tracing, as well as production incident management, troubleshooting, and performance optimization at scale
  • Experience supporting healthcare IT and Enterprise Imaging environments, including PACS, EMRs, VNAs, Voice Recognition systems, and healthcare interoperability standards such as DICOM and HL7
  • Hands-on experience with Kubernetes service mesh and ingress technologies (Istio, Linkerd), CI/CD tools (Jenkins, ArgoCD, GitHub Actions), container security platforms (CrowdStrike, Aqua, Sysdig), and modern SDLC practices
  • Strong understanding of HIPAA, HITRUST, and regulated cloud environments, with experience in networking (TCP/IP, HTTP, gRPC, DNS), databases (PostgreSQL, MySQL, Oracle), and building resilient cloud-native solutions
  • PREFERRED CERTIFICATIONS: AWS Certified Solutions Architect, AWS Certified DevOps Engineer, or Certified Kubernetes Administrator (CKA) / Certified Kubernetes Application Developer (CKAD)

Similar jobs