Jobs · Engineering · California

Senior/Staff Site Reliability Engineer

Glocomms · San Francisco County, CA · 6 days ago
HybridEngineeringFull-time

What You'll Do

  • Own and operate highly available AWS infrastructure
  • Design, scale, and troubleshoot Kubernetes (EKS) environments
  • Improve reliability through monitoring, automation, and observability
  • Build and maintain CI/CD pipelines and GitOps workflows
  • Manage infrastructure as code with Terraform
  • Lead incident response, root cause analysis, and long-term remediation
  • Partner with engineering teams to improve platform scalability, performance, and reliability

What We're Looking For

  • Strong AWS infrastructure experience (networking, security, IAM, scaling)
  • Deep Kubernetes expertise, particularly EKS
  • Terraform experience in production environments
  • Coding experience with Go and/or Python
  • Experience with GitHub Actions, ArgoCD, and GitOps practices
  • Strong observability background using Datadog or similar tools
  • Experience defining and managing SLIs/SLOs
  • Prominent track record supporting large-scale production systems

Ideal Background

  • 5+ years in SRE, Platform Engineering, DevOps, or Infrastructure Engineering
  • Experience operating cloud-native environments at scale
  • Passion for automation, reliability, and operational excellence

Similar jobs