Manager, Incident Response
NBCUniversal · New York, NY · 1 mo ago
RemoteRemoteEngineeringFull-time
Responsibilities
- Manage the delivery of the Major Incident Management function, hands on as required.
- Manage the Major Incident Team Leads and offshore team to ensure internal SLOs are met.
- Ensure restoration of normal service operations as quickly as possible to minimize the impact on business operations.
- Define, manage and ensure the delivery of clear communication strategies for major incidents, keeping all key stakeholders in formed to agreed and appropriate levels throughout the incident lifecycle.
- Provide updates to end users and leadership by way of outage notifications to keep them informed of progress to resolution and/or workarounds that have been implemented.
- Analyze workflows to identify operational bottlenecks and recommend targeted AI or machine learning solutions that automate repetitive tasks and shorten cycle times.
- Identify and implement continual service improvement opportunities across the Major Incident, Change, Problem management and wider ITSM and AIOps related functions.
- Serve as the escalation point for Major Incident functions and incidents.
- Develop internal dashboards and reports on Major incidents and lead their review with Executive Management.
- Partner with Service Management Ops colleagues to deliver on Enterprise Technology strategies.
- Partner with the providers of core functions to ensure compliance to business related SLAs and contractual obligations.
- Promote the Major Incident process and act as a champion of this function within the wider Crisis Management community.
- Collaborate with Technology Resolver teams and Business Stakeholders on service improvements within the ITSM area.
Qualifications
- An experienced Major Incident Management leader/process owner and strategist.
- 10+ years in Major Incident Management or similar ITSM roles.
- 8+ years managing a team.
- Strong practical ITIL/ITSM skill set with operational experience.
- Demonstrated knowledge of incident management practices, activities, techniques and tools within a large, complex organization.
- Be technically fluent in modern AI tools (such as generative AI platforms, predictive analytics, and RPA) and have a foundational understanding of how APIs integrate with existing enterprise software.
- Proven ability to identify operational inefficiencies and implement AI, automation or data-driven solutions that improve performance.
- Demonstrated experience of managing 3rd parties & vendors within a Service Delivery of Infrastructure remit.
- Candidate must have excellent communications and facilitation skills to lead outage calls as well as to draft periodic updates to business stakeholders on progress to service restoration.