Jobs · Engineering · Texas

Site Reliability Engineer Lead

Bank of America · Plano, TX · 6 days ago
EngineeringFull-time

About the role

Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the reliability, scalability, and performance of critical Infrastructure Automation platforms. This role will lead the design and implementation of SRE practices across a federated technology ecosystem, ensuring operational excellence through automation, observability, and resilient architecture. The ideal candidate will bring deep expertise in distributed systems, cloud-native infrastructure, SaaS application support, and DevOps/SRE principles, along with strong leadership and collaboration skills to influence cross-functional engineering and Production management teams and drive continuous improvement in service reliability.

Responsibilities

  • SRE Strategy & Governance: Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols. Establish governance models for reliability engineering across distributed teams. Champion a culture of observability, proactive monitoring, and continuous feedback loops.
  • Reactive & Proactive Problem Management: Lead root cause analysis (RCA) and post-incident reviews to identify systemic issues and prevent recurrence. Implement proactive problem detection using telemetry, anomaly detection, and trend analysis. Collaborate with engineering and operations teams to eliminate toil and reduce incident frequency and impact.
  • Capacity & Performance Management: Develop and maintain capacity models to ensure systems scale efficiently with business demand. Monitor performance trends and lead optimization efforts across infrastructure and applications. Partner with finance and engineering teams to align capacity planning with cost and growth objectives.
  • Platform Reliability & Automation: Drive automation of operational tasks including deployments, scaling, and recovery. Integrate reliability tooling with CI/CD pipelines, ITSM platforms (e.g., ServiceNow), and observability systems.
  • Incident Management & Operational Excellence: Oversee major incident response, escalation, and communication processes. Develop and maintain runbooks, playbooks, and escalation protocols. Drive continuous improvement through blameless retrospectives and operational reviews.
  • Technical Leadership: Serve as a senior technical advisor and thought leader in SRE and platform engineering. Mentor and guide SRE teams and partner with engineering leaders across the enterprise. Provide input on staffing, tooling strategy, and budget planning for reliability initiatives.
  • Managerial Responsibilities: Demonstrate leadership in fostering an inclusive environment, managing processes and data, advocating for enterprise decisions, managing risk, coaching and developing talent, stewarding financial resources, building enterprise talent bench strength, and driving business outcomes.

Requirements

  • 10+ years of experience in systems engineering, DevOps, or SRE roles in large-scale environments.
  • Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing.
  • Proven experience with observability stacks (e.g., Dynatrace, Grafana, Splunk, OpenTelemetry).
  • Expertise in infrastructure-as-code and automation tools (e.g., Terraform, Ansible, Python).
  • Strong knowledge of cloud platforms and container orchestration (Kubernetes).
  • Demonstrated success in leading incident response and driving systemic improvements.
  • Experience with capacity planning, performance tuning, and cost optimization.
  • Excellent communication and stakeholder management skills, including executive engagement.

Desired Qualifications

  • Experience with ITIL/ITSM processes and integration with platforms like ServiceNow.
  • Familiarity with security and compliance in regulated industries (e.g., financial services).
  • Background in performance engineering and infrastructure analytics.
  • Experience developing dashboards and metrics for operational health and reliability.

Skills

  • Influence
  • Risk Management
  • Solution Design
  • Stakeholder Management
  • Technical Strategy Development
  • Analytical Thinking
  • Application Development
  • Collaboration
  • Result Orientation
  • Solution Delivery Process
  • Agile Practices
  • Architecture
  • Automation
  • Data Management
  • DevOps Practices

Schedule

1st shift (United States of America), 40 hours per week.

Similar jobs