Jobs · Quality Assurance · Texas

Senior Manager, Site Reliability Engineering – Paylo Platform

PDI Technologies · Temple, TX · 1 mo ago
HybridQuality AssuranceFull-time

Role Overview

PDI Technologies is looking for a Senior Manager, Site Reliability Engineering to lead the SRE organization supporting Paylo, PDI's payments, loyalty, and fuel-pricing product suite. This role owns the reliability, infrastructure, and operational strategy for a portfolio of high-traffic, customer- and partner-facing platforms that power payment transactions, fuel pricing, loyalty and rewards, and offer/coupon redemption for convenience retail and fuel customers around the world. This is a hands-on, leadership-first role. You will manage a team of three SRE Managers/Leads who together lead approximately 20 engineers, while staying technically engaged yourself — reviewing architecture, unblocking hard infrastructure problems, and setting the technical bar across the organization. You will bring strong, current, hands-on expertise across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and you will be a strong, visible people leader who can coach managers and represent SRE to senior engineering and business stakeholders.

Key Responsibilities

  • Directly manage and develop 3 SRE Managers/Leads and own the overall health, growth, and performance of an approximately 20-person SRE organization supporting the Paylo product suite.
  • Set the vision, priorities, and operating cadence for the SRE function; translate business and product priorities into a reliability roadmap your managers can execute against.
  • Build a strong bench by hiring, coaching, and developing managers and senior engineers while creating clear career paths and succession plans.
  • Foster a blameless, learning-oriented culture around incidents, on-call, and operational excellence.
  • Partner closely with engineering directors, product managers, and business stakeholders across the Paylo organization to align reliability investments with business risk and customer impact.
  • Stay technically engaged day to day by participating in architecture and design reviews, troubleshooting complex production issues, and directly contributing to infrastructure-as-code, Kubernetes manifests/Helm charts, and CI/CD pipelines when needed.
  • Set and enforce engineering standards for multi-cloud infrastructure across AWS and Azure and for container orchestration on Kubernetes at scale.
  • Own adoption and standards for GitOps-based continuous delivery using Argo CD/Argo Workflows, including deployment strategy, rollout policy, and multi-cluster promotion.
  • Own the Infrastructure-as-Code strategy across teams (Terraform, OpenTofu), including module standards, state management, drift detection, and remediation.
  • Own CI/CD pipeline architecture and standards built on Jenkins, driving build/deploy automation, pipeline reliability, and progressive delivery practices such as blue-green/canary deployments and automated rollback.
  • Evaluate and guide adoption of new infrastructure tooling and patterns as the platform evolves across AWS and Azure.
  • Own the observability strategy across all supported products, with deep, hands-on expertise in Datadog (APM, infrastructure monitoring, log management, dashboards, and alerting) as the standard platform for metrics, tracing, and alerting.
  • Define and drive adoption of SLIs/SLOs, error budgets, and reliability KPIs across the organization, holding managers and teams accountable to them.
  • Own the incident management program end to end, including on-call structure, escalation paths, severity definitions, postmortems, and follow-through on remediation actions.
  • Drive root-cause analysis and long-term reliability investments that reduce Sev1/Sev2 frequency and recurrence.
  • Ensure appropriate resilience, disaster recovery, and capacity planning practices are in place given the sensitivity of payment- and transaction-related systems.
  • Partner with Security and Compliance to maintain awareness of PCI DSS and related compliance requirements and ensure the SRE organization supports audit and compliance readiness.
  • Track and report cost, capacity, and operational KPIs to senior leadership.

Required Qualifications

  • 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure/Platform Engineering, including 4+ years in a people-leadership role.
  • Proven experience managing managers — you have directly led team leads/managers, not just individual contributors, and are comfortable operating at the scale of approximately 20 total reports.
  • Strong, hands-on expertise across AWS and Azure — you can architect, troubleshoot, and operate multi-cloud infrastructure yourself, not just direct others to do so.
  • Strong, hands-on expertise with Kubernetes and Helm — cluster operations, troubleshooting at scale, and chart design/maintenance.
  • Strong, hands-on expertise with Argo CD/Argo Workflows for GitOps-based continuous delivery.
  • Strong, hands-on expertise with Infrastructure as Code (Terraform, OpenTofu), including module design and state management.
  • Strong, hands-on expertise with Jenkins for CI/CD pipeline design, administration, and automation.
  • Strong, hands-on expertise with Datadog (or equivalent enterprise observability platform), including designing monitoring/alerting strategy, dashboards, and APM/tracing at scale.
  • Demonstrated track record of driving incident management, on-call, and postmortem programs for high-traffic, customer-facing systems.
  • Excellent communication and stakeholder-management skills; able to represent SRE to engineering leadership and business partners with equal credibility.
  • A strong, visible leadership style — someone who sets clear direction, holds teams accountable, and builds trust across the organization.

Preferred Qualifications

  • Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations.
  • Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect, Microsoft Certified: Azure Solutions Architect, or HashiCorp Terraform Associate.

Similar jobs