Director - Infra Engineering - Platform Security and Lifecycle Management
American Express · Palo Alto, CA · 5 days ago
HybridEngineeringFull-time
Responsibilities
- Define and execute the enterprise strategy and roadmap for Enterprise PaaS platform using Redhat Openshift and Data Middleware platform upgrades and lifecycle management.
- Leverage Gen AI/Agentic AI to drive a fully automated pipeline for platform provisioning, upgrades, patching, and configuration management.
- Own end-to-end upgrade planning, governance, risk assessment, and execution.
- Establish upgrade readiness processes, validation frameworks, rollback strategies, and post-upgrade monitoring.
- Coordinate with application teams to ensure platform compatibility and minimize business disruption during upgrades.
- Manage lifecycle risks associated with OpenShift, Kubernetes, operating systems, middleware, and supporting infrastructure.
- Track and report platform currency and technology lifecycle compliance metrics.
Security & Compliance
- Establish and maintain a strong security posture across clusters and supporting infrastructure.
- Lead vulnerability management, container image security, platform hardening, patch management, and remediation efforts.
- Ensure compliance with enterprise security policies, regulatory requirements, and industry standards.
- Partner with Information Security and Risk teams to manage security assessments, audits, and remediation activities.
- Manage security exceptions and risk acceptance processes while driving long-term remediation strategies.
Capacity & Performance Management
- Own capacity planning and forecasting for Platform, including compute, memory, storage, and network resources.
- Develop predictive capacity models to support business growth and application onboarding.
- Establish monitoring and reporting processes for cluster utilization, performance, and scalability.
- Optimize infrastructure consumption and platform efficiency while maintaining service-level objectives.
- Lead resource optimization initiatives to improve workload density and reduce infrastructure costs.
Leadership & Team Management
- Build and lead high-performing cloud native DevOps engineers, Kubernetes administrators, security specialists, and capacity planners.
- Foster a culture of operational excellence, continuous learning, innovation, and accountability.
- Mentor leaders and technical experts within the organization.
- Drive workforce planning, succession planning, and talent development initiatives.
Stakeholder & Vendor Management
- Serve as the senior technology leader for Private Cloud platform services.
- Partner with application development, architecture, security, infrastructure, and business leaders to align platform capabilities with organizational priorities.
- Manage relationships with key vendors, managed service providers, and strategic partners.
- Present platform health, upgrade status, security posture, risks, and capacity forecasts to executive leadership.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or related field (Master’s preferred).
- 8+ years of experience in Platform Engineering & Operations, Site Reliability Engineering (SRE), Platform lifecycle management with a proven track record of leading teams in managing large-scale cloud infrastructure with a focus on automation, reliability and resilience.
- Deep hands-on experience with any Kubernetes platform(multi-cloud preferred).
- Strong experience with: Infrastructure as Code (Terraform, CloudFormation, ARM), Infrastructure automation tools like Ansible, Container platforms (OpenShift/Kubernetes), Monitoring tools (Prometheus, OTEL, LOKI), CI/CD pipelines (Jenkins, GitHub Actions), Open source based messaging, caching, and database technologies like Kafka, Redis, Elastic.
- Strong understanding of cloud networking, security, and architecture.
- Experience managing large-scale, mission-critical production environments.
- Relevant certifications preferred.
- Experience with DevOps practices and methodologies, including CI/CD pipelines, configuration management, and infrastructure as code.
- Experience with observability tools such as Prometheus, Splunk, ELK, Dynatrace.
- Strong analytical and problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.
- Excellent communication and leadership skills, with the ability to effectively collaborate with cross-functional teams and influence decision-making at all levels of the organization.