Senior Site Reliability Engineer (Cloud Platform)
Jobgether · United States · 1 mo ago
RemoteRemoteEngineeringFull-time
About the role
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer (Cloud Platform) based in the United States. This is an exciting opportunity for an experienced Site Reliability Engineer to play a key role in building and evolving highly available cloud platforms that support mission-critical services at scale.
Responsibilities
- Maintain the reliability, performance, and availability of production and pre-production cloud environments.
- Design, implement, and optimize observability solutions using metrics, logging, tracing, and monitoring platforms.
- Respond to production incidents, participate in root cause analysis, and implement preventive improvements to enhance system resilience.
- Collaborate with software engineering teams to improve application reliability and integrate SRE best practices into development workflows.
- Automate operational processes and repetitive tasks to improve efficiency and reduce manual intervention.
- Develop and maintain operational documentation, runbooks, troubleshooting guides, and incident response procedures.
- Contribute to innovative global products and help shape resilient cloud architectures used by customers worldwide.
Requirements
- Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
- Strong hands-on experience managing Kubernetes and containerized production environments.
- Proven expertise supporting large-scale cloud infrastructures and mission-critical services.
- Solid experience with AWS cloud services and cloud-native architectures.
- Proficiency with observability and monitoring tools such as Prometheus, Grafana, and ELK Stack.
- Strong scripting and automation skills using Python, Bash, Go, or similar languages.
- Experience with Linux system administration and infrastructure automation tools such as Terraform or Ansible.
- Strong understanding of networking concepts, including TCP/IP, DNS, routing, and load balancing.
- Excellent troubleshooting, communication, and collaboration abilities with a proactive, automation-first mindset.
Qualifications
- Nice-to-have experience with SIP/VoIP technologies, relational databases (MySQL/PostgreSQL), and NoSQL solutions such as Redis.
Skills
- Hands-on experience with Kubernetes and containerized production environments.
- Proficiency with AWS cloud services and cloud-native architectures.
- Strong scripting and automation skills using Python, Bash, Go, or similar languages.
- Experience with Linux system administration and infrastructure automation tools such as Terraform or Ansible.
- Strong understanding of networking concepts, including TCP/IP, DNS, routing, and load balancing.
- Excellent troubleshooting, communication, and collaboration abilities with a proactive, automation-first mindset.
Benefits
- Long-term full-time B2B collaboration opportunity.
- Fully remote and flexible work environment.
- Professional development support, including technical training and continuous learning opportunities.
- Exposure to innovative cloud technologies and globally impactful projects.
- Collaborative engineering culture focused on mentorship, knowledge sharing, and continuous improvement.
- Modern Apple equipment provided.
- Inclusive and supportive workplace that values diversity and encourages authentic contributions.