Jobs · Engineering

Sr. Cloud Operations Reliability Engineer (SRE)

Ladders · United States · 1 mo ago
RemoteRemoteEngineering$120k–$150k/yrFull-time

Responsibilities

  • Own service reliability and operational health, establishing SLOs and monitoring strategies
  • Lead incident response and post-incident processes, troubleshooting complex production issues
  • Design and implement automation and operational tooling to enhance production readiness
  • Build observability solutions, ensuring rapid problem detection and response
  • Conduct performance analysis and capacity planning for cloud services
  • Collaborate with development teams to ensure deployment reliability and operational best practices
  • Support disaster recovery planning and ensure operational readiness documentation is up-to-date

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field
  • 10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or related fields
  • Extensive hands-on experience with production cloud environments on platforms like GCP or AWS
  • Proficiency with Infrastructure as Code tools (Terraform, CloudFormation) and version control practices
  • Strong knowledge of incident management and post-incident review processes
  • Experience with Kubernetes operations and container orchestration
  • Familiarity with application performance monitoring and distributed tracing

Benefits

  • Mentorship opportunities to guide junior engineers
  • Involvement in shaping operational standards and practices
  • Opportunity to lead and own critical reliability initiatives
  • Access to ongoing professional development and technology training
  • Engagement with cross-functional teams for knowledge sharing
  • Contribution to business continuity and disaster recovery planning

Similar jobs

Cloud Engineer (SRE)

Samaritan's PurseUnited States· 3 wk ago
RemoteEngineeringapply on careers.samaritanspurse.org