Jobs · Engineering

Site Reliability Engineer

Evlo AI · Miami, FL · 5 days ago
RemoteRemoteEngineeringFull-time

About the role

The role focuses on building, maintaining, and scaling the core platform infrastructure that powers high-throughput production systems. It requires a deep understanding of distributed systems, automated infrastructure provisioning, and continuous delivery models to ensure high availability and sub-millisecond latency. The engineer will collaborate closely with product development teams to design self-service platform tools, architect cloud-native deployments, and establish rigorous observability practices. This position is critical to maintaining a 99.99% uptime SLA across all user-facing services.

Responsibilities

  • Manage and scale multi-region Kubernetes clusters on AWS using infrastructure-as-code tools such as Terraform
  • Design and implement robust CI/CD pipelines utilizing GitHub Actions, Jenkins, or ArgoCD to automate deployments and minimize time-to-production
  • Configure and maintain comprehensive observability stacks using Prometheus, Grafana, Jaeger, and Datadog to proactively detect and diagnose system anomalies
  • Lead incident response post-mortems and conduct root-cause analysis for high-severity production issues, translating findings into automated preventive measures
  • Optimize cloud spend and database performance across PostgreSQL, Redis, and Elasticsearch clusters
  • Develop internal automation tools and CLI applications in Go or Python to simplify infrastructure provisioning for development teams

Requirements

  • 3–6 years of experience in SRE, DevOps, or systems engineering roles managing highly available SaaS platforms
  • Strong proficiency in shell scripting and at least one high-level language: Go, Python, or Ruby
  • Deep technical expertise with container orchestration via Kubernetes and managing infrastructure as code with Terraform
  • Hands-on experience with cloud networks, including VPC configurations, CDN routing, and load balancing on AWS or GCP
  • Proven track record of managing production databases and configuring caching layers at scale

Qualifications

  • Bonus: Certified Kubernetes Administrator (CKA), AWS Solutions Architect certification, or experience implementing service mesh topologies using Istio

Similar jobs

Site Reliability Engineer

Akamai TechnologiesCambridge, MA· 5 days ago
RemoteEngineering$76k–$136k/yrapply on fa-extu-saasfaprod1.fa.ocs.oraclecloud.com