Director of Cloud, Enterprise Operations
Mastercard · St Louis, MO · 1 mo ago
HybridFull-time
Role
The role is primary based in Dublin and leads a globally distributed organization spanning EMEA, the Americas and APAC, operating within a matrixed reporting environment where influence, alignment and clarity matter as much as formal authority.
About the role
This is a hands-on leadership role at the intersection of three priorities: operating a resilient, secure and cost-efficient cloud platform; measurably improving the day-to-day experience of the software builders who depend on it; and bringing AI into operations.
Responsibilities
- Lead, grow and retain a multi-region team of SREs, cloud operations engineers, automation engineers and engineering managers; own hiring plans, capability mapping, succession and performance management for the Dublin hub and its extended global teams.
- Establish follow-the-sun operating rhythms: handovers, on-call rotations, incident command, and escalation paths that work across EMEA, NAM and APAC and that do not concentrate burden in any single geography.
- Set a clear engineering culture: blameless post-incident review, documented decisions, psychological safety, and an expectation that operational work is engineered away rather than absorbed.
- Coach managers and senior individual contributors; build technical career paths that make deep engineering expertise as promotable as people leadership.
- Represent Cloud Operations to senior stakeholders, audit, regulators and customers, translating technical risk into business terms.
- Own the developer-facing quality of cloud services: onboarding time, environment provisioning latency, self-service coverage, golden-path adoption, and the friction engineers encounter between commit and production.
- Partner with Platform Engineering to productize operational capability, treating internal platforms as products with roadmaps, users, feedback loops and adoption goals rather than as ticket queues.
- Reduce cognitive load on product teams by shifting security, compliance and resilience controls left into templates, modules and guardrails that are safe by default.
- Drive automation of provisioning, configuration, patching, scaling, failover, cost optimization and compliance evidence collection across AWS and Azure.
- Own operational capacity, resilience and disaster recovery posture for cloud workloads, including regular chaos and failover exercises.
- Own the cloud Run cost model with FinOps practice, support chargeback accuracy, unit economics, and continuous efficiency work.
- Enable AI and GenAI use cases on the cloud platform: managed inference, GPU and accelerator capacity, model hosting, vector and retrieval services, MLOps pipelines, and the governance and cost controls that make them safe to consume at enterprise scale.
- Design and deliver agentic SRE workflows with AI agents that triage alerts, correlate telemetry, propose or execute remediation, draft incident timelines, and maintain runbooks.
- Using explicit guardrails, human-in-the-loop approval gates, audit trails and rollback paths. Define the evaluation and safety model for operational AI: how agent actions are scoped, permissioned, tested, measured for accuracy, and retired when they underperform.
- Build the data foundation that makes this possible: high-quality observability, structured incident and change history, and machine-readable runbooks.
Requirements
- Substantial experience leading cloud operations, SRE or platform engineering organizations at enterprise scale, including experience leading managers and multi-region teams.
- Demonstrated success operating within a matrix reporting environment, delivering outcomes through partners and dotted-line teams across time zones and cultures.
- Deep, current hands-on technical grounding in both AWS and Azure, including networking, identity, containers and orchestration (Kubernetes), infrastructure as code (Terraform), CI/CD, and observability tooling.
- Expert or Professional-level certification in AWS (for example AWS Certified Solutions Architect – Professional, or AWS Certified DevOps Engineer – Professional) and Expert-level certification in Azure (for example Microsoft Certified: Azure Solutions Architect Expert, or Azure DevOps Engineer Expert).
- All About You
- All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
- Abide by Mastercard’s security policies and practices;
- Ensure the confidentiality and integrity of the information being accessed;
- Report any suspected information security violation or breach, and;
- Complete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.