Jobs · Engineering

Site Reliability Engineer

Evlo AI · Atlanta, GA · Yesterday
RemoteRemoteEngineeringFull-time

About The Role

The role owns the reliability, scalability, and security of production infrastructure across multi-cloud environments, ensuring high availability for millions of global users. The team works closely with software engineering squads to automate deployments, minimize latency, and build robust observability pipelines that catch incidents before they impact clients.

Key Responsibilities

  • Design, build, and maintain production infrastructure using Terraform and Ansible within AWS and GCP environments
  • Manage Kubernetes clusters at scale, optimizing resource allocation, auto-scaling policies, and cluster security configurations
  • Develop and improve CI/CD pipelines using GitHub Actions and ArgoCD to ensure fast, repeatable, and safe code deployments
  • Implement comprehensive observability stacks using Prometheus, Grafana, Datadog, and PagerDuty for metrics, logs, and distributed tracing
  • Participate in an on-call rotation to troubleshoot and resolve production incidents, conducting thorough post-mortems and driving preventive remediation
  • Enforce security best practices, identity and access management (IAM) policies, and compliance standards across all infrastructure layers

Requirements

  • 3–6 years of experience in Site Reliability Engineering, DevOps, or systems engineering roles in high-growth production environments
  • Deep expertise in Linux systems administration, TCP/IP networking, DNS, and containerization technologies (Docker, Kubernetes)
  • Strong proficiency in Infrastructure as Code (Terraform, CloudFormation) and scripting languages such as Python or Bash
  • Proven track record of managing CI/CD pipelines and modern observability tooling in cloud-native ecosystems
  • Bachelor's degree in Computer Science, related technical field, or equivalent practical experience

Qualifications

  • Experience with service mesh technologies like Istio, Chaos Engineering practices, or holding AWS/GCP professional certifications

Similar jobs

Site Reliability Engineer

Akamai TechnologiesCambridge, MA· 1 wk ago
RemoteEngineering$76k–$136k/yrapply on fa-extu-saasfaprod1.fa.ocs.oraclecloud.com