Jobs · Engineering

Site Reliability Engineer

Scale.jobs · New York, NY · 1 wk ago
RemoteRemoteEngineeringFull-time

About The Role

The role focuses on building and maintaining the foundational infrastructure that powers high-throughput, low-latency production services. The engineer in this role will partner closely with product engineering teams to design resilient architectures, eliminate single points of failure, and automate operational overhead across multi-region cloud deployments. The ideal candidate is passionate about system observability, chaos engineering, and infrastructure-as-code. This position requires a strong software engineering mindset applied directly to complex infrastructure challenges to guarantee system reliability and performance at scale.

Key Responsibilities

  • Design, provision, and maintain scalable cloud infrastructure on AWS or GCP using Terraform to ensure high availability and disaster recovery readiness
  • Deploy and manage containerized services using Kubernetes, implementing service meshes, ingress controllers, and autoscaling policies
  • Build and maintain robust CI/CD pipelines using GitHub Actions or GitLab CI to automate testing, security scanning, and blue-green deployments
  • Develop and refine observability frameworks using Prometheus, Grafana, and OpenTelemetry to establish proactive alerting, SLOs, and SLIs
  • Participate in a blameless on-call rotation, conducting root cause analyses (RCAs) and implementing post-mortem action items to prevent recurrence of production issues
  • Write custom tooling and automation scripts in Python, Go, or Bash to eliminate manual operations and streamline release management

What We Are Looking For

  • 3-6 years of experience in an SRE, DevOps, or systems engineering role managing production environments serving high-volume traffic
  • Strong proficiency with Infrastructure as Code (IaC), specifically Terraform, and container orchestration platforms like Kubernetes
  • Solid software engineering skills in at least one language such as Go, Python, or Ruby, with a focus on writing clean, maintainable automation code
  • Deep understanding of Linux systems internals, networking protocols (TCP/IP, DNS, HTTP/S), and cloud security best practices
  • Experience configuring production monitoring, logging, and tracing tools (e.g., Datadog, Prometheus, ELK stack)

Bonus

  • Experience with service mesh technologies (Istio, Linkerd)
  • Advanced database administration (PostgreSQL, Redis)
  • Kubernetes operator development

Similar jobs

Site Reliability Engineer

QualityAIIllinois, United States· Yesterday
Quality Assurance$110k–$120k/yrapply on careers.quality-ai.com