Jobs · Florida

Associate Director Cloud Operations

Full-time

Impact You Will Have

The Cloud Operations team is responsible for the availability, security, and governed change of DTCC's production AWS and Azure environments. The team operates the infrastructure that supports financial market services: clearing, settlement, equities, fixed income, and derivatives markets. Cloud Operations owns observability, incident response, change authority, problem management, and operational readiness for cloud-hosted workloads. The team builds governed execution paths so application teams can operate safely at speed and drives every production incident to root cause and verified closure.

Primary Responsibilities

  • Own incident response on your shift in a 24x7 hybrid cloud environment that supports critical financial market operations.
  • Drive complex, cross-domain incidents to resolution across AWS, Azure, on-prem, and hybrid infrastructure, making mitigation decisions under pressure with incomplete information.
  • Manage engineers on your shift rotation, developing them into independent decision-makers who can run incidents.
  • Lead your team's build work on operational tooling and automation between incidents, with a focus on eliminating recurring manual work.
  • Prioritize work based on incident patterns, translating recurring failures into automation that prevents recurrence.
  • Maintain operational visibility through daily change posture awareness and pattern detection.
  • Partner across Platform Engineering, Application Teams, and Delivery Engineering to ensure operational excellence drives improvements across the organization.

Qualifications

  • Minimum of 8 years of related experience
  • Bachelor's degree preferred or equivalent experience
  • 3+ years managing people in an operational environment
  • Deep experience in cloud operations, SRE, or infrastructure engineering
  • Infrastructure knowledge spanning AWS, Azure, VMware, hybrid networking, storage, Kubernetes, Linux, and Windows
  • Proficiency in Python, Terraform, infrastructure-as-code, and CI/CD pipelines
  • Working knowledge of AI-assisted operations tooling such as Amazon Kiro, AWS AgentCore, or agentic automation frameworks
  • Observability platform experience such as Grafana, Prometheus, Datadog, CloudWatch, or Dynatrace
  • Incident management expertise across the full ITSM lifecycle, including root cause analysis that drives architectural change
  • People leadership in an operational environment, including shift rotation design, on-call models, and performance development
  • Experience operating in a regulated or financial services environment preferred

Talents Needed for Success

  • Experience in cloud operations, SRE, or infrastructure engineering
  • Infrastructure knowledge spanning AWS, Azure, VMware, hybrid networking, storage, Kubernetes, Linux, and Windows
  • Proficiency in Python, Terraform, infrastructure-as-code, and CI/CD pipelines
  • Working knowledge of AI-assisted operations tooling such as Amazon Kiro, AWS AgentCore, or agentic automation frameworks
  • Observability platform experience such as Grafana, Prometheus, Datadog, CloudWatch, or Dynatrace
  • Incident management expertise across the full ITSM lifecycle, including root cause analysis that drives architectural change
  • People leadership in an operational environment, including shift rotation design, on-call models, and performance development

Pay and Benefits

Competitive compensation, including base pay and annual incentive
Comprehensive health and life insurance and well-being benefits
Pension
Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.

Schedule

DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).

Similar jobs