Sr. Cloud Operations Reliability Engineer (SRE)
Ladders · United States · 1 mo ago
RemoteRemoteEngineering$120k–$150k/yrFull-time
Responsibilities
- Own service reliability and operational health, establishing SLOs and monitoring strategies
- Lead incident response and post-incident processes, troubleshooting complex production issues
- Design and implement automation and operational tooling to enhance production readiness
- Build observability solutions, ensuring rapid problem detection and response
- Conduct performance analysis and capacity planning for cloud services
- Collaborate with development teams to ensure deployment reliability and operational best practices
- Support disaster recovery planning and ensure operational readiness documentation is up-to-date
Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field
- 10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or related fields
- Extensive hands-on experience with production cloud environments on platforms like GCP or AWS
- Proficiency with Infrastructure as Code tools (Terraform, CloudFormation) and version control practices
- Strong knowledge of incident management and post-incident review processes
- Experience with Kubernetes operations and container orchestration
- Familiarity with application performance monitoring and distributed tracing
Benefits
- Mentorship opportunities to guide junior engineers
- Involvement in shaping operational standards and practices
- Opportunity to lead and own critical reliability initiatives
- Access to ongoing professional development and technology training
- Engagement with cross-functional teams for knowledge sharing
- Contribution to business continuity and disaster recovery planning