Senior Manager, Site Reliability Engineer - Remote
Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives.
About the Role
We are seeking an experienced Senior Manager to lead enterprise Site Reliability Engineering (SRE), DevOps, IT Service Management (ITSM), and Operational Excellence initiatives across Optum bank. This leader will be responsible for improving service reliability, operational resiliency, deployment automation, observability, incident management, and production readiness for critical banking platforms. The ideal candidate combines strong technical expertise with operational leadership experience, driving engineering excellence, automation, reliability, and continuous improvement while ensuring technology services meet business, customer, regulatory, and operational expectations. This role will also help identify and implement emerging automation and AI-enabled operational capabilities that improve service health, reduce operational toil, and accelerate engineering productivity.
Responsibilities
Leadership & Engineering Management
- Lead and develop multidisciplinary teams responsible for Site Reliability Engineering, DevOps, Platform Engineering, ITSM, and Operational Excellence
- Establish and execute enterprise reliability, availability, resiliency, and operational maturity strategies
- Drive engineering excellence through automation, observability, operational readiness, and continuous improvement practices
- Partner with Technology, Operations, Security, Infrastructure, Risk, and Business leaders to improve service reliability and customer experience
- Build and mentor high-performing teams while fostering accountability, innovation, operational ownership, and learning
- Manage staffing, capacity planning, talent development, succession planning, and organizational growth
- Establish operational metrics, governance standards, and service review processes to improve service performance and risk management
SRE, DevOps & Platform Engineering
- Lead enterprise SRE practices including SLI/SLO adoption, error-budget management, reliability engineering, and operational maturity assessments
- Drive DevOps transformation initiatives, emphasizing automation, deployment standardization, CI/CD pipelines, Infrastructure-as-Code, and GitOps practices
- Establish production readiness standards and operational acceptance criteria for new technology deployments
- Improve platform resiliency through capacity planning, disaster recovery, fault tolerance, and resilience testing
- Drive reduction of operational toil through automation and self-healing capabilities
- Partner with application and infrastructure teams to improve system scalability, availability, and performance
- Lead initiatives to improve deployment frequency, reduce change failure rates, and accelerate service recovery times
ITSM & Operational Excellence
- Establish and mature Incident, Problem, Change, Release, and Service Request Management processes
- Lead major incident management programs and executive communications during critical service disruptions
- Drive root-cause analysis and problem-management practices to eliminate recurring incidents
- Improve operational scorecards, service health reviews, and reliability reporting for executive stakeholders
- Ensure compliance with regulatory, audit, risk, and operational governance requirements
- Partner with Technology and Business leaders to improve service quality and customer outcomes through data-driven operational improvements
- Champion a culture of operational excellence and continuous service improvement
Intelligent Automation & AI-Enabled Operations
- Identify opportunities to leverage AI and automation to improve operational effectiveness and engineering productivity
- Lead implementation and evaluation of solutions involving: AIOps, Intelligent alert correlation, Automated incident triage, Root cause analysis assistance, Knowledge management copilots, Agentic operational workflows
- Partner with enterprise AI teams to evaluate emerging technologies that improve reliability and operational efficiency
- Drive responsible adoption of AI-enabled engineering and operational practices
- Support proof-of-concept initiatives that demonstrate measurable reductions in operational effort and incident resolution times
Cross-Functional Leadership
- Collaborate with Engineering, Infrastructure, Security, Architecture, Risk, Compliance, and Operations teams to prioritize reliability and operational improvements
- Serve as a trusted advisor on reliability engineering, operational excellence, and automation strategies
- Drive alignment between technology and business stakeholders to improve service quality and operational outcomes
- Influence technology investment decisions that improve platform stability, resiliency, and operational efficiency
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Technology, or related field
- 10+ years of experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, Infrastructure Engineering, or Technology Operations
- 5+ years of experience leading engineering or operational teams
- Proven experience supporting large-scale, business-critical production environments
Skills
Required
- SRE principles and practices
- DevOps and CI/CD
- ITSM processes
- Cloud platforms (Azure, AWS)
- On-prem environments
- Infrastructure-as-Code
- Container platforms (Kubernetes, OpenShift)
- Observability and monitoring platforms
- Incident Management
- Problem Management
- Change Management
- Disaster Recovery
- Business Continuity
- Service Reliability Programs
- Operational transformations and continuous-improvement initiatives
Preferred
- Experience in banking, financial services, healthcare, or other highly regulated industries
- Experience implementing enterprise observability solutions such as Datadog, Splunk, Grafana, Prometheus, or OpenTelemetry
- Experience with cloud-native architectures and platform engineering practices
- Experience deploying AIOps, ChatOps, or intelligent automation solutions
- Familiarity with:
- Agentic AI workflows
- LLM-powered operational tooling
- Knowledge management platforms
- AI-enabled incident management solutions
- Experience establishing SLO frameworks and reliability governance programs
Leadership Competencies
- Strategic thinker with solid operational and technology acumen
- Proven ability to build, lead, and inspire high-performing engineering and operational teams
- Solid understanding of service reliability, operational risk, resiliency, and governance
- Exceptional stakeholder management and executive communication skills
- Data-driven decision maker with a focus on measurable outcomes
- Solid execution and delivery leadership in complex enterprise environments
- Collaborative leader focused on continuous improvement, operational excellence, and customer outcomes
Benefits
- Comprehensive benefits package
- Incentive and recognition programs
- Equity stock purchase
- 401k contribution (all benefits are subject to eligibility requirements)
Pay
The salary for this role will range from $112,700 to $193,200 annually based on full-time employment.