Senior/Staff Site Reliability Engineer
Glocomms · San Francisco County, CA · 6 days ago
HybridEngineeringFull-time
What You'll Do
- Own and operate highly available AWS infrastructure
- Design, scale, and troubleshoot Kubernetes (EKS) environments
- Improve reliability through monitoring, automation, and observability
- Build and maintain CI/CD pipelines and GitOps workflows
- Manage infrastructure as code with Terraform
- Lead incident response, root cause analysis, and long-term remediation
- Partner with engineering teams to improve platform scalability, performance, and reliability
What We're Looking For
- Strong AWS infrastructure experience (networking, security, IAM, scaling)
- Deep Kubernetes expertise, particularly EKS
- Terraform experience in production environments
- Coding experience with Go and/or Python
- Experience with GitHub Actions, ArgoCD, and GitOps practices
- Strong observability background using Datadog or similar tools
- Experience defining and managing SLIs/SLOs
- Prominent track record supporting large-scale production systems
Ideal Background
- 5+ years in SRE, Platform Engineering, DevOps, or Infrastructure Engineering
- Experience operating cloud-native environments at scale
- Passion for automation, reliability, and operational excellence