Senior Site Reliability Engineer
About the role
Western National Insurance Group is a private mutual insurance company with over 120 years of experience serving customers' property-and-casualty insurance needs in the Midwestern, Northwestern, and Southwestern United States. Known as “The Relationship Company®,” we define success as a measure of the relationships we’ve built over time. In everything that we do, we know that delivering a friendly and helpful interaction makes for a better experience for everyone involved. That’s the power of “nice”. At Western National, nice is something we work to bring to every person and organization with whom we partner and serve. Western National is seeking a Site Reliability Engineer III to join our team! The individual in this role will have the opportunity to design, implement, manage, and monitor on-premise and cloud-based application platforms. This individual monitors, resolves, and proactively prevents system outages and performance issues while partnering with product teams to improve application deployment, infrastructure stability, and reliability.
Responsibilities
- Designs, deploys, and maintains scalable and secure platform infrastructure across cloud and on-premise environments.
- Participates in work refinement and estimation processes with product teams.
- Utilizes modern software delivery technologies to assist product teams with application deployments.
- Leverages automation tools to ensure processes are repeatable and efficient.
- Solicits and performs peer reviews of application and infrastructure designs and implementations in accordance with team, product, and organizational standards.
- Recommends and designs solutions that improve application deployment and stability and trains product teams on how to use the solutions.
- Creates and maintains documentation, including support processes, team processes, system documentation, and key decisions and discussions.
- Ensures timely delivery of projects and tasks and communicates delays and roadblocks to the team and leadership as needed.
- Supports noncore responsibilities as needed to ensure timely work delivery, including analysis, testing, and meeting facilitation.
- Monitors and resolves system outages and performance issues of mixed impact with an emphasis on proactive resolution.
- Collaboratively implements software application infrastructure modifications of mixed impact using software development best practices.
- Prepares and executes fault tolerance and load testing to support robust application infrastructure design.
- Researches and pilots opportunities to improve efficiency, effectiveness, and reliability across systems, job functions, team-level processes, and business processes with an emphasis on quality and user experience.
- Participates in and leads the troubleshooting, resolution, and communication of urgent production issues.
- Recommends improvements to company guidelines to align with industry-standard system security best practices in partnership with the Security Office.
- Provides technical expertise and leads projects in collaboration with other product teams.
- Assists teams in adapting to new infrastructure and deployment technologies in accordance with industry standards.
- Applies knowledge, experience, and creative resources to identify solutions to complex problems and effectively communicates the advantages and disadvantages of proposed solutions.
- Consistently acts according to our customer experience standards, including responding quickly, maintaining a positive attitude, building rapport, demonstrating empathy, managing the customer’s expectations, using the proper communication channel for the situation, and taking ownership to ensure the customer’s issue is resolved.
- Performs special projects and other duties as assigned.
Requirements
- Strong mathematical, analytical, and problem-solving skills.
- Ability to work independently and carry assignments through completion.
- Strong attention to detail and quality.
- Demonstrated advanced organizational and prioritization skills with the ability to manage multiple priorities of varying impact and risk within specified timeframes.
- Proven ability to collaborate effectively with product team members and other teams.
- Demonstrated intermediate knowledge and application of relevant project methodologies at the product-team level.
- Demonstrated ability to communicate clearly and effectively, both verbally and in writing, with technical and nontechnical audiences across all levels of the organization.
- Strong expertise with general AWS cloud services.
- Extensive experience with infrastructure-as-code concepts and tools, including Terraform, Helm charts, and YAML.
- Strong expertise managing workloads in Kubernetes and/or EKS.
- Strong understanding of system networks, networking, and related security best practices.
- Strong proficiency with cloud-based CI / CD pipelines and commonly associated services.
- Strong Linux operating system experience.
- Experience working with virtualization technologies.
- Experience deploying