Manager, Site Reliability Engineer
About the role
At Forge, we know our team is our greatest asset. As technology innovators in the private market, our vision is to deliver a richer future for everyone. We live that vision through our values of being bold, accountable, and humble. We experience the value that our vision brings to the world every day, helping the teams behind the greatest innovations of our generation, from space travel to artificial intelligence, and more. With liquidity solutions, exclusive data and insights, a custody offering, and a vibrant marketplace, Forge's goal is to build the best-in-class technology infrastructure to power a global private market that is transparent, accessible, and seamless for companies, their employees, and investors. Through Forge, employees can sell their private shares, employers can reward shareholders with pre-IPO liquidity and individual and institutional investors can participate in private unicorn growth. Forge's differentiated global marketplace addresses rising demand among individual and institutional investors for exposure to private company stocks and is building a growing network effect. Our ability to offer these powerful financial solutions has generated incredible interest from investors, demand from customers, and a need to grow our team to meet the needs of more companies, teams, and innovators in this way. The Role: As an engineering organization, we pride ourselves on engineering as a creative activity. Engineering managers enable engineers to do their best work by maintaining a culture and environment where engineers can achieve autonomy, mastery, and purpose. The Manager, Site Reliability Engineering will lead Forge's SRE team responsible for keeping Forge systems highly available for customers, while partnering closely with Platform, Engineering, Security, Compliance, and Product teams to improve reliability, observability, incident response, and operational maturity. This is an opportunity for a hands-on technical leader who can coach engineers, improve production operations, and help Forge build and run secure, scalable, and highly reliable products.
Responsibilities
- Manage Forge's Site Reliability Engineering team responsible for keeping Forge systems highly available for customers.
- Drive strong incident management practices in partnership with engineering teams, including response, mitigation, follow-up, and post-incident learning.
- Build, improve, and manage observability infrastructure in partnership with Platform Engineering, including monitoring, alerting, dashboards, and operational metrics.
- Improve monitoring coverage and alert quality to reduce noise, shorten time to detect, and support faster response and mitigation.
- Champion reliability best practices across engineering, including service ownership, operational readiness, disaster recovery, and production support standards.
- Contribute to technical design, architecture, automation, infrastructure, and overall team delivery.
- Collaborate with engineering teams to troubleshoot production issues, identify recurring problems, and improve system reliability.
- Hire, coach, mentor, and manage performance for SRE team members while supporting career development and team health.
- Partner with Security, Compliance, and Risk partners to ensure reliability and infrastructure practices meet the needs of a regulated business.
Qualifications
Required:
- 5+ years of experience leading a Site Reliability Engineering, DevOps, Cloud Operations, or similar reliability-focused function.
- 10+ years of total software engineering, infrastructure, platform, cloud, or production operations experience.
- Bachelor's degree in Computer Science, Engineering, or a closely related field, or equivalent practical experience.
- Experience building, operating, and maintaining large-scale cloud infrastructure and distributed systems.
- Hands-on experience with observability, monitoring, alerting, incident response,