Director - Infra Engineering - Platform Security and Lifecycle Management
American Express · New York, NY · 6 days ago
HybridEngineeringFull-time
Responsibilities
- Define and execute the enterprise strategy and roadmap for Enterprise PaaS platform using Redhat Openshift and Data Middleware platform upgrades and lifecycle management.
- Lead major version upgrades, cluster modernization, and infrastructure refresh initiatives across development, test, and production environments.
- Establish standards, reference architectures, and best practices for platform lifecycle and security management.
- Leverage Gen AI/Agentic AI to drive a fully automated pipeline for platform provisioning, upgrades, patching, and configuration management.
- Own end-to-end upgrade planning, governance, risk assessment, and execution.
- Establish upgrade readiness processes, validation frameworks, rollback strategies, and post-upgrade monitoring.
- Carefully coordinate with application teams to ensure platform compatibility and minimize business disruption during upgrades.
- Manage lifecycle risks associated with OpenShift, Kubernetes, operating systems, middleware, and supporting infrastructure.
- Track and report platform currency and technology lifecycle compliance metrics.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or related field (Master’s preferred).
- 8+ years of experience in Platform Engineering & Operations, Site Reliability Engineering (SRE), Platform lifecycle management with a proven track record of leading teams in managing large-scale cloud infrastructure with a focus on automation, reliability and resilience.
- Deep hands-on experience with any Kubernetes platform (multi-cloud preferred).
- Strong experience with: Infrastructure as Code (Terraform, CloudFormation, ARM), Infrastructure automation tools like Ansible, Container platforms (OpenShift/Kubernetes), Monitoring tools (Prometheus, OTEL, LOKI), CI/CD pipelines (Jenkins, GitHub Actions), Open source based messaging, caching, and database technologies like Kafka, Redis, Elastic, Strong understanding of cloud networking, security, and architecture.
- Experience managing large-scale, mission-critical production environments.
- Relevant certifications preferred.
- Experience with DevOps practices and methodologies, including CI/CD pipelines, configuration management, and infrastructure as code.
- Experience with observability tools such as Prometheus, Splunk, ELK, Dynatrace.
- Strong analytical and problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.
- Excellent communication and leadership skills, with the ability to effectively collaborate with cross-functional teams and influence decision-making at all levels of the organization.