Site Reliability Engineer
Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them. Chart your own path and grow your career while helping more customers achieve financial freedom.
Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.
Responsibilities
- Establish key indicators (SLIs) measuring the performance of services and build proactive monitor and alerts
- Support projects of varying complexity and impact across multiple disciplines and teams
- Conduct in-depth analysis of problems to identify relevant findings and root causes
- Play a key role in ensuring the high availability, resilience, and scalability of containerized applications in production
- Document critical systems and create runbooks for incidents
- Lead capacity planning and right-sizing exercises
- Maintain and optimize infrastructure as code (IaC)
- Troubleshoot and resolve complex system and deployment issues
- Manage observability within Kubernetes, specifically EKS
- Collaborate with development teams to support releases and create highly scalable, resilient, and maintainable services
- Work in a GitOps driven environment
Requirements
- Experience maintaining high availability and resiliency within AWS infrastructure components, including EKS, EC2, RDS, S3, VPC, and others
- Proficiency with Infrastructure as Code frameworks such as Terraform and CloudFormation
- Demonstrated experience with containerization and orchestration technologies such as Docker and Kubernetes
- Experience with technologies, systems, networks, and potential gaps that can impact an organization’s ability to effectively detect and respond to production incidents
- Strong problem-solving abilities and a desire to learn
Qualifications
- Bachelor’s degree in Computer Science, Information Systems or equivalent experience
- AWS, Kubernetes, or relevant certifications
- Experience with observability suites and APM tooling such as DataDog, AppDynamics, New Relic, etc.
- Strong programming skills in one or more languages such as shell, Go, Python, etc.
- Experience supporting Java Spring Boot applications
- Production experience in Kubernetes, especially EKS
Benefits
- Medical, dental, vision and life insurance
- Retirement savings – 401(k) plan with generous company matching contributions (up to 6%), financial advisory services, potential company discretionary contribution, and a broad investment lineup
- Tuition reimbursement up to $5,250/year
- Business-casual environment that includes the option to wear jeans
- Generous paid time off upon hire – including a paid time off program plus ten paid company holidays and three floating holidays each calendar year
- Paid volunteer time — 16 hours per calendar year
- Leave of absence programs – including paid parental leave, paid short- and long-term disability, and Family and Medical Leave (FMLA)
- Business Resource Groups (BRGs) – BRGs facilitate inclusion and collaboration across our business internally and throughout the communities where we live, work and play. BRGs are open to all.
Pay
Base Salary Range $87,400.00 - $123,400.00
The salary range above shows the typical minimum to maximum base salary range for this position in the location listed. Non-sales positions have the opportunity to participate in a bonus program. Sales positions are eligible for sales incentives, and in some instances a bonus plan, whereby total compensation may far exceed base salary depending on individual performance. Actual compensation offered may vary from posted hiring range based upon geographic location, work experience, education, licensure requirements and/or skill level and will be finalized at the time of offer.
Schedule
For remote and hybrid positions you will be required to provide reliable high-speed internet with a wired connection as well as a place in your home to work with limited disruption. You must have reliable connectivity from an internet service provider that is fiber, cable or DSL internet. Other necessary computer equipment will be provided. You may be required to work in the office if you do not have an adequate home work environment and the required internet connection.