Site Reliability Engineer
Evlo AI · Minneapolis, MN · 4 days ago
RemoteRemoteEngineeringFull-time
About the role
The role owns the reliability, scalability, and performance of distributed cloud infrastructure supporting core production workloads at scale. The team works closely with platform and software engineering squads to automate infrastructure provisioning, eliminate toil, and maintain high availability standards.
Responsibilities
- Architect, deploy, and manage production infrastructure using Terraform and Kubernetes across multi-region cloud environments
- Design and implement comprehensive observability pipelines using Prometheus, Grafana, and distributed tracing tools to monitor system health and latency
- Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for all critical production services
- Automate incident response, root cause analysis, and remediation workflows to minimize downtime and prevent recurring outages
- Conduct thorough post-mortems and implement systematic preventive measures across the infrastructure stack
- Participate in an on-call rotation to handle production incidents and ensure rapid recovery from system anomalies
Requirements
- 3–6 years of experience in site reliability engineering, DevOps, or systems engineering in large-scale production environments
- Deep expertise in Kubernetes administration, containerization, and cloud-native networking principles
- Strong proficiency in infrastructure-as-code tooling, specifically Terraform and Ansible
- Hands-on experience with modern observability stacks (Prometheus, Grafana, Datadog) and log aggregation platforms (ELK, Loki)
- Proficiency in at least one scripting or programming language such as Python, Go, or Bash for automation tasks
- Bonus: Experience with service mesh technologies like Istio, chaos engineering practices, or holding AWS/GCP professional certifications