Principal Site Reliability Engineer (Cloud, Observability & Automation)
Are you ready to make an impact at DTCC? Work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development. At DTCC, we are at the forefront of innovation in the financial markets, committed to helping employees grow and succeed. We foster a thriving internal community and create a workplace that reflects the world we serve.
Pay and Benefits
- Competitive compensation, including base pay and annual incentive
- Comprehensive health and life insurance and well-being benefits, based on location
- Pension / Retirement benefits
- Paid Time Off and Personal/Family Care, and other leaves of absence to support physical, financial, and emotional well-being
DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays, and a third day unique to each team or employee).
About the Role
The Enterprise Application Support (EAS) team supports critical applications across the ITP and ECS business lines, ensuring the reliability, scalability, and performance of enterprise platforms. As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across mission-critical systems. This is a hands-on technical leadership role focused on improving system performance, reducing operational risk, accelerating recovery, and advancing SRE best practices through modern cloud, observability, automation, and AI-powered technologies.
Responsibilities
- Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications
- Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms
- Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs
- Lead major incident response, root cause analysis, and continuous service improvement initiatives
- Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies
- Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle
- Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives
- Identify operational risks and deliver strategic reliability improvements across the technology ecosystem
- Collaborate with technical and business stakeholders to improve service reliability and operational outcomes
Requirements
- Bachelor's degree in Computer Science, Engineering, or equivalent experience
- 8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines
Skills
- Strong hands-on experience with AWS and cloud-native architectures
- Proficiency in Python, Java, Go, or similar programming languages
- Strong Linux/Unix systems administration and troubleshooting experience
- Expertise in observability and monitoring platforms including Splunk, Grafana, Dynatrace, and ITSI
- Experience leading major incident management and root cause investigations in complex production environments
- Strong understanding of distributed systems, resiliency engineering, performance tuning, automation, and operational excellence
- Excellent communication and stakeholder management skills with the ability to influence technical and business partners
Preferred Qualifications
- Experience with AI-assisted engineering tools such as Amazon Q, Kiro, or similar technologies
- Experience designing and measuring SLOs, SLIs, and operational KPIs
- Experience supporting large-scale enterprise applications in financial services or other highly regulated environments
The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations.