Jobs · Engineering · Arizona

Site Reliability Engineer Lead (SRE) – Internal Kubernetes Container Platform (IKCP)

Bank of America · Chandler, AZ · 1 wk ago
EngineeringFull-time

About the Role

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We drive Responsible Growth and deliver for our clients, teammates, communities, and shareholders every day. This role serves as a technical lead within the platform organization, ensuring the reliability, scalability, performance, security, and operational excellence of the enterprise Internal Kubernetes Container Platform (IKCP). The IKCP Site Reliability Engineer Lead partners closely with Engineering, Architecture, Product Management, Security, Infrastructure, and Central Operations teams to deliver a highly available platform-as-a-product experience for application teams.

Responsibilities

  • Own platform reliability objectives, including service availability, resiliency, recoverability, and operational health.
  • Lead critical incident response, root cause analysis, and problem management activities.
  • Serve as a senior escalation point for L3 platform support and on-call operations.
  • Develop and maintain operational runbooks, recovery procedures, and standard operating practices.
  • Drive production readiness reviews for new platform capabilities and services.
  • Ensure platforms meet enterprise resiliency and availability objectives.
  • Conduct resilience exercises and continuous improvement activities following recovery testing.
  • Execute platform upgrades, patching strategies, cluster modernization, and release orchestration.
  • Improve platform scalability, performance, and resource utilization across production and non-production environments.
  • Support platform modernization initiatives including OpenShift virtualization, VKS, and cloud-native technologies.
  • Collaborate with Product, Architecture, Engineering, and Operations teams to improve developer experience and platform adoption.
  • Design and implement enterprise observability solutions leveraging monitoring, logging, tracing, and alerting platforms.
  • Automate operational processes using Infrastructure-as-Code, GitOps, CI/CD, and scripting frameworks.
  • Reduce operational toil through self-healing, intelligent automation, and proactive remediation capabilities.
  • Drive operational efficiency through automation of cluster provisioning, upgrades, compliance, and day-2 operations.
  • Perform platform capacity planning and trend analysis.
  • Forecast infrastructure growth requirements and optimize platform resource consumption.
  • Conduct performance tuning for clusters, workloads, networking, and storage services.
  • Support enterprise-scale growth while maintaining platform stability and customer experience.
  • Partner with security teams to implement platform security controls and governance requirements.
  • Support vulnerability remediation, image compliance, platform hardening, and policy enforcement.
  • Implement and maintain RBAC, Network Policies, and container security controls.
  • Drive compliance with enterprise standards, vulnerability management processes, and audit requirements.
  • Collaborate with Development and Infrastructure teams to implement monitoring capabilities.
  • Develop and maintain reliability scripts, tools, and libraries for common instrumentation, automation, and operational needs.
  • Partner to implement code changes to utilize common reliability libraries and tools.
  • Participate in architecture community of practice meetings and communications.
  • Identify vulnerabilities and opportunities for reliability improvement, such as investigating low-level error rates and monitoring noise.
  • Define solutions to reduce manual support effort and improve system reliability.
  • Engage as a subject matter expert in major incident triage efforts and failure scenario modeling.
  • Diagnose root causes for major incident/problem management investigations.

Requirements

  • 8+ years of infrastructure, cloud, platform engineering, or SRE experience.
  • 5+ years managing Kubernetes and/or OpenShift production environments.
  • Experience operating large-scale mission-critical distributed systems.
  • Experience supporting enterprise production environments with 24x7 operational responsibilities.

Skills

  • Kubernetes, OpenShift, Rancher, VKS container orchestration platforms.
  • Linux administration and troubleshooting.
  • Terraform, Ansible, GitOps, ArgoCD, Helm.
  • CI/CD platforms such as Jenkins, GitHub, GitLab, Bitbucket, or equivalent.
  • Monitoring and observability tools such as Dynatrace, Prometheus, Grafana, Splunk, ELK, OpenTelemetry.
  • Infrastructure as Code and automation frameworks.
  • Networking fundamentals, load balancing, ingress, DNS, and service mesh concepts.
  • Storage platforms, backup technologies, and disaster recovery solutions.
  • Scripting in Python, Go, Bash, or similar languages.
  • Automation, collaboration, influence, production support, result orientation.
  • Analytical thinking, application development, architecture, solution design.
  • Stakeholder management, adaptability, DevOps practices, project management.
  • Risk management, solution delivery process.

Qualifications

  • BS/MS degree in Computer Science, Engineering, Information Systems, or related technical discipline, or equivalent experience.
  • OpenShift Administration or Kubernetes certifications (desired).
  • Experience running large-scale enterprise container platforms (desired).
  • Experience with virtualization technologies including VMware, VCF, and OpenShift Virtualization (desired).
  • Experience implementing cloud-native security controls and platform governance (desired).
  • Knowledge of platform engineering, developer experience, and platform-as-a-product operating models (desired).
  • Experience with vulnerability management and container security scanning solutions (desired).
  • Drives operational excellence and continuous improvement.
  • Demonstrates strong ownership and accountability.
  • Influences cross-functional teams without direct authority.
  • Communicates effectively with senior technical and business leaders.
  • Champions automation-first and reliability-first engineering culture.

Schedule

1st shift (United States of America), 40 hours per week.

Similar jobs