Senior Site Reliability Engineer (Cloud Platform)
Salve.Lab · Denver, CO · 2 wk ago
RemoteRemoteEngineeringFull-time
Key Responsibilities
- Maintain the reliability, availability, and performance of production and pre-production environments.
- Monitor platform health and improve alerting, automation, and operational processes.
- Respond to production incidents, participate in root cause analysis, and implement long-term improvements.
- Design, build, and enhance observability solutions using metrics, logs, traces, and dashboards.
- Partner with software engineers to improve application reliability throughout the development lifecycle.
- Develop and maintain operational documentation, troubleshooting guides, and runbooks.
- Automate repetitive operational tasks to improve efficiency and reduce manual intervention.
- Participate in on-call rotations while continuously improving incident response processes.
- Promote reliability engineering principles, operational excellence, and continuous improvement across engineering teams.
Requirements
- Bachelor's or Master's degree in Engineering, Computer Science, or a related field.
- Strong experience operating Kubernetes or other container orchestration platforms.
- Experience supporting large-scale production services.
- Hands-on experience with AWS.
- Experience with Prometheus, Grafana, and ELK.
- Strong scripting skills (Bash, Python, or Go).
- Experience administering Linux-based production environments.
- Experience with Infrastructure as Code or configuration management tools such as Terraform or Ansible.
- Solid understanding of networking fundamentals (TCP/IP, DNS, load balancing, routing).
- Excellent troubleshooting, communication, and collaboration skills.
- A proactive mindset with a passion for automation and reliability.
Nice to Have
- Experience with SIP or VoIP technologies.
- Familiarity with MySQL or PostgreSQL.
- Experience with Redis or other NoSQL databases.