Senior Cloud Site Reliability Engineer
About the role
The State of Wisconsin Investment Board (SWIB) manages more than $178 billion in assets, including those of the fully-funded Wisconsin Retirement System (WRS). SWIB operates at a level more often seen in top-tier global asset managers than in typical public pension funds and is a home for top talent. Approximately 61 percent of SWIB’s investment professionals are Chartered Financial Analyst (CFA) charterholders.
The City of Madison, the state capitol and home of Wisconsin’s flagship university, makes regular appearances on lists of best places to live, eat, and play. SWIB offers a modern workspace, hybrid work options, and competitive compensation and benefits. Serving over 703,000 WRS beneficiaries, SWIB is driven by a clear mission: securing the financial future of those who serve Wisconsin.
We are seeking a highly skilled and experienced Senior Site Reliability Engineer to oversee the build and transformation of SWIB’s cloud native technology stack. This role will serve in a critical capacity to ensure all aspects of the Software Development Lifecycle (SDLC) are built from the ground up with modern tools and techniques, including software, data, and infrastructure. The ideal candidate will have a strong background in financial services, exceptional leadership skills, and the ability to manage platforms that require continuity on a 24x6 basis effectively. This role will serve as a thought leader within the technology organization, driving change and transformation across all teams.
Responsibilities
- Function as subject matter expert in delivering and operating reliable, performant, robust, and secure applications, with an emphasis on multi-region and multi-cloud patterns.
- Work with development teams to design, document, create, and maintain highly available systems.
- Design, implement, and manage centralized monitoring solutions that provide expedient actionable feedback to the development teams.
- Partner with application development teams to create flows, processes, automation, and tooling.
- Facilitate the evaluation, adoption, and integration of AI into the software delivery platform.
- Create an environment of continuous experimentation and learning.
- Contribute to the evolution of cloud-focused architecture to increase its flexibility and ease of use.
- Follow technology trends/tools and recommend improvements to technology when appropriate.
- Mentor new or less senior members of the team.
- Share experience, knowledge, and ideas to improve processes and productivity.
- Provide tier 2 and 3 escalations for related issues and questions.
- Establish and monitor key performance indicators (KPIs) and service level agreements (SLAs) to ensure the support team meets or exceeds performance expectations.
- Conduct regular performance reviews and provide ongoing training and development opportunities for the support team.
- Drive continuous improvement initiatives to enhance support processes, reduce incidents, and improve overall application reliability and user satisfaction.
- Manage vendor relationships and ensure third-party support services align with organizational needs and standards.
- Maintain comprehensive documentation of support processes, incidents, and resolutions.
- Stay current with industry trends and emerging technologies to ensure the SRE function remains cutting-edge and effective.
Requirements
- Enablement mindset and attitude – you win when teams win.
- Excellent verbal and written communication skills.
- Bachelor’s Degree in Computer Science, or a related field, or equivalent work experience.
- 8+ years of professional Site Reliability Engineering experience (or equivalent demonstrated impact).
- Strong background in designing, implementing, and delivering complex technical architectures.
- Hands-on experience with the Cloud in a production environment (AWS preferred).
- Solid experience implementing Infrastructure-as-Code (Terraform or OpenTofu preferred).
- Hands-on experience building and running CI/CD infrastructure (GitLab preferred).
- Hands-on development experience in a modern programming language (Python preferred).
- Experience implementing and operating container orchestration platforms such as Kubernetes, EKS, Elastic Container Service (ECS).
- Deep understanding of information security concepts.
- Experience implementing and operating monitoring tools such as Sentry, Prometheus, and Datadog.
- Experience with or strong interest in AI technologies.
- Familiarity with version control systems such as Git.
- Working experience with agile methodologies.
- Experience with cloud-performant microservices and event-driven architectures is a plus.
- Ability to work under pressure and manage multiple priorities in a fast-paced environment.
Benefits
- Competitive total cash compensation, based on AON (formerly McLagan) industry benchmarks.
- Comprehensive benefits package.
- Educational and training opportunities.
- Tuition reimbursement.
- Challenging work in a professional environment.
- Hybrid work environment.
This position requires U.S. work authorization. Pursuant to our Hybrid Remote Work Policy, all staff have the flexibility to work remotely but are required to have a weekly presence in our offices, the frequency of which is dependent on their distance from the office. Staff are not required to reside locally; however, we offer relocation reimbursement to the Dane County area per our policy.