Senior Application Support Engineer / Site Reliability Engineer
The Depository Trust & Clearing Corporation (DTCC) · Dallas, TX · Today
EngineeringFull-time
The Impact You Will Have in This Role As a Senior Application Support Engineer, you will help power DTCC's global financial markets infrastructure by ensuring the reliability, availability, and performance of Institutional Trade Processing (ITP) platforms that support cross-border equity and debt trade processing and settlement. Leveraging Site Reliability Engineering (SRE) principles, you will support a portfolio of 40+ mission-critical applications across a modern ecosystem of AWS, OpenShift Container Platform (OCP), Kafka, IBM MQ, and distributed systems. You will play a key role in driving operational excellence, improving resiliency, reducing operational risk, and advancing automation across critical trade processing platforms. Working closely with globally distributed teams across Application Development, Infrastructure, Cloud, Network, and Operations, you will help deliver stable, scalable, and highly available services that support DTCC's mission-critical business functions. Your Primary Responsibilities Production Reliability & Incident Management Ensure the availability, stability, and performance of mission-critical applications. Lead incident response, troubleshooting, and root cause analysis (RCA) activities. Drive preventative solutions that improve resiliency and reduce recurring issues. Operational Excellence & Resiliency Support application recovery, failover, disaster recovery (DR), and business continuity activities. Maintain operational readiness across a portfolio of enterprise applications. Support critical processing schedules and service-level commitments. Change, Release & Automation Support application deployments, releases, and vendor upgrades. Drive automation initiatives that improve efficiency and reduce operational risk. Enhance monitoring, alerting, observability, and operational tooling. Collaboration & Governance Partner with Development, Infrastructure, Cloud, Database, Network, Security, and Product teams to ensure platform reliability. Support audit, risk, compliance, and operational control requirements. Build strong partnerships across globally distributed teams to drive continuous improvement. Qualifications Bachelor’s degree preferred or equivalent practical experience 6–8 years of experience supporting enterprise applications in complex production environments Shift information – Second Shift (EST) Monday – Friday 3:00 PM – 11:00 PM EST Talents Needed For Success Core Technologies Java/J2EE Oracle, SQL, DB2 IBM MQ, Kafka Splunk, Grafana, AutoSys ServiceNow and ITIL Linux/Unix AWS Required Experience Strong understanding of Site Reliability Engineering (SRE) principles, including reliability, observability, automation, and incident prevention. Experience supporting enterprise-scale, distributed applications in production environments. Experience with cloud technologies and modern infrastructure platforms. Understanding of messaging systems, middleware, and integration technologies (IBM MQ, Kafka, IIB, or similar). Exposure to container platforms such as OpenShift, Kubernetes, or Docker. Strong troubleshooting, analytical, and problem-solving skills. Preferred Skills Kafka and event-driven architecture experience. Mainframe exposure (JCL, batch processing, job monitoring). Python or other scripting languages. Financial Services, Capital Markets, or Securities Processing experience. AI-driven automation or operational efficiency initiatives. Enterprise integration and distributed processing platforms.