Senior Manager, Site Reliability Engineering – Paylo Platform
Role Overview
PDI Technologies is looking for a Senior Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI's payments, loyalty, and fuel-pricing product suite. This role owns the reliability, infrastructure, and operational strategy for a portfolio of high-traffic, customer- and partner-facing platforms that power payment transactions, fuel pricing, loyalty and rewards, and offer/coupon redemption for convenience retail and fuel customers around the world. This is a hands-on, leadership-first role. You will manage a team of three SRE Managers/Leads who together lead approximately 20 engineers, while staying technically engaged yourself — reviewing architecture, unblocking hard infrastructure problems, and setting the technical bar across the organization. You will bring strong, current, hands-on expertise across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and you will be a strong, visible people leader who can coach managers and represent SRE to senior engineering and business stakeholders.
Key Responsibilities
- Directly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an approximately 20-person SRE organization supporting the Paylo product suite.
- Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.
- Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.
- Foster a blameless, learning-oriented culture around incidents, on-call, and operational excellence.
- Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.
- Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure-as-code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.
- Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.
- Own adoption and standards for GitOps-based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi-cluster promotion.
- Own the Infrastructure-as-Code strategy across teams (Terraform, OpenTofu), including module standards, state management, drift detection, and remediation.
- Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue-green/canary deployments and automated rollback.
- Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.
- Own the observability strategy across all supported products, with deep, hands-on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.
- Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.
- Own the incident management program end to end, including on-call structure, escalation paths, severity definitions, postmortems, and follow-through on remediation actions.
- Drive root-cause analysis and long-term reliability investments that reduce Sev1/Sev2 frequency and recurrence.
- Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment- and transaction-related systems.
- Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.
- Track and report cost, capacity, and operational KPIs to senior leadership.
Required Qualifications
- 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people-leadership role.
- Proven experience managing managers — you have directly led team leads/managers, not just individual contributors, and are comfortable operating at the scale of approximately 20 total reports.
- Strong, hands-on expertise across AWS and Azure — you can architect, troubleshoot, and operate multi-cloud infrastructure yourself, not just direct others to do so.
- Strong, hands-on expertise with Kubernetes and Helm — cluster operations, troubleshooting at scale, and chart design/maintenance.
- Strong, hands-on expertise with Argo CD/Argo Workflows for GitOps-based continuous delivery.
- Strong, hands-on expertise with Infrastructure as Code (Terraform, OpenTofu), including module design and state management.
- Strong, hands-on expertise with Jenkins for CI/CD pipeline design, administration, and automation.
- Strong, hands-on expertise with Datadog (or equivalent enterprise observability platform), including designing monitoring/alerting strategy, dashboards, and APM/tracing at scale.
- Demonstrated track record of driving incident management, on-call, and postmortem programs for high-traffic, customer-facing systems.
- Excellent communication and stakeholder-management skills; able to represent SRE to engineering leadership and business partners with equal credibility.
- A strong, visible leadership style — someone who sets clear direction, holds teams accountable, and builds trust across the organization.
Preferred Qualifications
- Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations.
- Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect, Microsoft Certified: Azure Solutions Architect, or HashiCorp Terraform Associate.