Executive Director, Solutions Engineering
Job Summary
The Executive Director, Solutions Engineering leads a critical function within Disney Entertainment and ESPN Product & Technology, responsible for strengthening the operational health, stability, and resilience of the services that support consumer, platform, and enterprise experiences.
Responsibilities And Duties Of The Role
Lead the Operational Readiness function and own the support standards that prepare services for production.
Direct proactive launch planning and change management so new features and platform changes are reviewed, risk assessed, and supportable before go live.
Own launch readiness reviews, change governance and the associated review forums, load tests, runbooks, rollback plans, and a forward view of upcoming changes shared with stakeholders.
Lead the global 24x7 Support Center, including continuous monitoring of service health and leadership of major incidents.
Set standards for incident command, escalation, and triage, and serve as the senior point of coordination during high severity events.
Own problem management, ensuring recurring issues are driven to root cause and corrective actions are tracked to completion.
Translate complex technical status into clear and timely updates for technical teams, business partners, and senior or executive leaders.
Serve as the unifying voice of operational communication across the business and technology organizations.
Advance the adoption of AI enabled operational tools that move the organization from a reactive model to a proactive one.
Leverage anomaly detection, event correlation, and intelligent automation across operational telemetry to surface issues earlier and reduce noise.
Drive earlier detection of service impacting events, faster and better coordinated response, reduced manual effort for operations teams, and improved service stability over time.
Lead other executives, managers, and technical professionals, both directly and through their organizations.
Develop and execute the operational strategy for P+T, set performance and service standards, and hold accountability for the structure, budget, and results of the function.
Build a high performing leadership team, establish clear operating standards and policies, and drive talent development, succession, and year over year operational maturity.
Required Education, Experience/Skills/Training
Minimum 12 years of progressive experience in technology operations, service delivery, or site reliability, including direct accountability for the operational posture of a large and complex technology environment.
10 or more years of leadership experience, with a proven track record of leading a function through multiple levels of managers and directors, in addition to senior technical professionals, both directly and indirectly.
Demonstrated experience leading mission critical, global technology operations and serving as the executive escalation point during major incidents.
Proven ability to set and enforce ITIL aligned incident, problem, and change management practices at enterprise scale.
Experience owning launch and operational readiness governance for major product launches and platform changes.
Ability to build and deliver operational communications and reporting for audiences that range from technical teams to senior and executive leaders.
Strong executive presence, with the ability to influence and negotiate with great latitude on outcomes and to present and defend complex or delicate matters with sensitivity to the audience.
Demonstrated ability to remain calm and decisive under pressure and to bring clarity, structure, and urgency to high severity events.
A proactive orientation that favors prevention over reactive firefighting, with the judgment to balance speed of restoration against operational risk and sound governance.
Experience developing functional strategy and operating plans, with accountability for the structure, budget, and results of a function.
Preferred Qualifications
An advanced degree, such as a Master of Business Administration or a technical master's degree.
Experience in a large, consumer facing technology environment, such as media, streaming, or entertainment.
Experience leading a global, always on Support Center, network operations center, or command center.
Hands on experience adopting AI enabled detection, event correlation, and intelligent automation to improve operational outcomes.
Experience applying site reliability engineering practices, including automation, observability, resilience, and the reduction of manual effort.
Experience managing operating budgets, vendor relationships, and sourcing strategy for an operations function.
ITIL 4 certification at the Foundation level or higher.
A site reliability, DevOps, or change management certification.
Relevant cloud platform certifications.
Experience with Supporting the operational health of large scale, consumer facing streaming services, including live and on demand video delivery.
Incident management and major incident command, including incident bridges, escalation, and post incident review.
Problem management and root cause analysis that drives corrective action through to completion.
Change management and release readiness, including change review forums and a forward schedule of upcoming changes.
Continuous monitoring, alerting, and observability across metrics, logs, traces, and events.
AIOps capabilities, such as anomaly detection, event correlation, and intelligent automation.
ITSM platforms, such as ServiceNow.
Operational reporting and executive communication during service impacting events.
Operating a global, always on support or command center model.
Service level objectives and the measurement of service health, availability, and reliability.
Oversight of vendors and managed service partners within an operations function.