Jobs · Engineering · North Carolina

Lead Platform Reliability Engineer

Wells Fargo · Charlotte, NC · 1 mo ago
Engineering$119k–$224k/yrFull-time

Wells Fargo is seeking a highly experienced infrastructure engineer to join the CTO Platform organization. This role is designed for engineers who possess deep technical expertise in one core platform discipline (Network, Middleware, Database, or Storage) and have demonstrated experience collaborating across at least one additional infrastructure domain (Enterprise Tools, Cloud, Observability Tools). The engineer will identify systemic issues spanning multiple domains and apply modern Site Reliability Engineering (SRE) practices to improve the availability, resiliency, observability, scalability, and operational excellence of critical enterprise platforms.

About the Role

As part of the Platform Reliability Engineering (PRE) team, you will leverage your domain expertise to drive automation, deliver engineering solutions, and strengthen platform stability at scale. You will serve as the reliability engineering expert for your primary domain while partnering across adjacent technology disciplines.

Responsibilities

  • Lead the investigation and resolution of complex production incidents, identifying root causes and implementing long-term corrective actions.
  • Apply SRE principles including service level indicators (SLIs), service level objectives (SLOs), error budgets, and reliability engineering practices to improve platform health.
  • Conduct capacity analysis, forecasting, and utilization reviews to identify future scaling risks and prevent service degradation before customer impact.
  • Perform deep performance analysis across infrastructure layers, identifying bottlenecks, contention points, latency drivers, and resource inefficiencies.
  • Identify and remediate configuration drift, operational debt, and platform hygiene issues that impact long-term reliability.
  • Drive proactive reliability improvements through observability, automation, performance optimization, and resiliency engineering.
  • Design and implement automation solutions that eliminate operational toil, reduce manual intervention, and improve recovery capabilities.
  • Define and enhance enterprise observability standards through metrics, logging, tracing, alerting, and service health monitoring.
  • Partner closely with engineering, infrastructure, application, cloud, and operations teams to improve platform performance and availability.
  • Lead blameless post-incident reviews and convert recurring operational issues into measurable engineering improvements.
  • Identify reliability risks and communicate technical recommendations to engineering leaders and senior stakeholders.
  • Mentor engineers and technical teams on reliability engineering, operational excellence, automation, and platform best practices.

Requirements

  • 5+ years of Systems Engineering, Infrastructure Engineering, Platform Engineering, Technology Architecture, or equivalent experience demonstrated through work experience, military experience, training, or education.
  • 5+ years supporting and engineering enterprise-scale production environments.
  • 5+ years of hands-on expertise in one of the following technology domains:
    • Network Engineering (routing, switching, load balancing, DNS, network observability, performance analysis).
    • Middleware Engineering (WebSphere, Tomcat, JBoss, Kafka, MQ, application platforms, integration technologies).
    • Database Engineering (Oracle, SQL Server, PostgreSQL, MongoDB, database performance, replication, HA/DR).
    • Storage Engineering (SAN/NAS technologies, storage virtualization, backup/recovery, performance and capacity management).

Qualifications

  • Strong experience applying SRE principles, including SLI/SLO development, error budgets, incident analysis, and reliability measurement.
  • Experience supporting highly available, mission-critical production environments.
  • Proven success troubleshooting complex issues spanning multiple technology domains in large-scale distributed environments.
  • Experience with capacity planning, resiliency engineering, fault tolerance, disaster recovery, and performance optimization.
  • Hands-on experience with observability and monitoring platforms such as Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, or similar technologies.
  • Experience building dashboards, alerts, service health indicators, and operational reporting.
  • Strong automation and scripting experience using Python, Bash, PowerShell, or similar technologies.
  • Experience developing operational tooling, API integrations, self-healing capabilities, and automated remediation solutions.
  • Familiarity with Git-based development practices, CI/CD pipelines, infrastructure automation, and Infrastructure as Code tools such as Ansible or Terraform.
  • Ability to diagnose and resolve issues that span multiple infrastructure layers.
  • Experience influencing technical direction across infrastructure and engineering organizations.
  • Experience leading major incident reviews and driving sustainable operational improvements.
  • Demonstrated success mentoring engineers and promoting reliability engineering best practices.
  • Strong communication skills with the ability to translate technical concepts into business-focused outcomes.

Schedule

This position offers a hybrid schedule.

Pay

The base pay range for this position is $119,000.00 - $224,000.00. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work location. Employees may also be eligible for incentive opportunities.

Benefits

Wells Fargo provides eligible employees with a comprehensive set of benefits, including:

  • Health benefits.
  • 401(k) Plan.
  • Paid time off.
  • Disability benefits.
  • Life insurance, critical illness insurance, and accident insurance.
  • Parental leave.
  • Critical caregiving leave.
  • Discounts and savings.
  • Commuter benefits.
  • Tuition reimbursement.
  • Scholarships for dependent children.
  • Adoption reimbursement.

Similar jobs