Platform Reliability Engineer - Principal Engineer
About the role
This role is designed for highly experienced infrastructure engineers who possess deep technical expertise in one of the following technology domains: Enterprise Tools / Applications (API Management, Scheduling, Observability, Telemetry, CICD, Testing) and have demonstrated experience collaborating across at least one additional infrastructure domains (Network, Database or Storage).
Responsibilities
- Lead the investigation and resolution of complex production incidents, identifying root causes and implementing long-term corrective actions.
- Apply SRE principles including service level indicators (SLIs), service level objectives (SLOs), error budgets, and reliability engineering practices to improve platform health.
- Perform capacity analysis, forecasting, and utilization reviews to identify future scaling risks and prevent service degradation before customer impact occurs.
- Identify and remediate configuration drift, operational debt, and platform hygiene issues that impact long-term reliability.
- Drive proactive reliability improvements through observability, automation, performance optimization, and resiliency engineering.
- Design and implement automation solutions that eliminate operational toil, reduce manual intervention, and improve recovery capabilities.
- Define and enhance enterprise observability standards through metrics, logging, tracing, alerting, and service health monitoring.
- Partner closely with engineering, infrastructure, application, cloud, and operations teams to improve platform performance and availability.
- Lead blameless post-incident reviews and convert recurring operational issues into measurable engineering improvements.
- Identify reliability risks and communicate technical recommendations to engineering leaders and senior stakeholders.
- Mentor engineers and technical teams on reliability engineering, operational excellence, automation, and platform best practices.
- Act as an advisor to leadership to develop or influence applications, network, information security, database, operating systems, or web technologies for highly complex business and technical needs across multiple groups.
- Lead the strategy and resolution of highly complex and unique challenges requiring in-depth evaluation across multiple areas or the enterprise, delivering solutions that are long-term, large-scale and require vision, creativity, innovation, advanced analytical and inductive thinking.
- Maintain knowledge of industry best practices and new technologies and recommends innovations that enhance operations or provide a competitive advantage to the organization.
- Strategically engage with all levels of professionals and managers across the enterprise and serve as an expert advisor to leadership.
Requirements
- 7+ years of Engineering experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 5+ years supporting and engineering enterprise-scale production environments
- 5+ years of experience with hands-on expertise in one of the following technology domains: Enterprise Tools / Applications (API Management, Scheduling, Observability, Telemetry, CICD, Testing) and secondary skills in: Network Engineering (routing, switching, load balancing, DNS, network observability, performance analysis), Database Engineering (Oracle, SQL Server, PostgreSQL, MongoDB, database performance, replication, HA/DR), Storage Engineering (SAN/NAS technologies, storage virtualization, backup/recovery, performance and capacity management)
Desired Qualifications
- Strong experience applying SRE principles, including SLI/SLO development, error budgets, incident analysis, and reliability measurement
- Proven success troubleshooting complex issues spanning multiple technology domains in large-scale distributed environments
- Experience supporting highly available, mission-critical production environments
- Experience with capacity planning, resiliency engineering, fault tolerance, disaster recovery, and performance optimization
- Hands-on experience with observability and monitoring platforms such as Grafana, Splunk, Prometheus, AppDynamics, Cribl, ThousandEyes, Dynatrace, or similar technologies
- Experience building dashboards, alerts, service health indicators, and operational reporting
- Strong automation and scripting experience using Python, Bash, PowerShell, or similar technologies
- Experience developing operational tooling, API integrations, self-healing capabilities, and automated remediation solutions
- Familiarity with Git-based development practices, CI/CD pipelines, infrastructure automation, and Infrastructure as Code tools such as Ansible or Terraform
- Ability to diagnose and resolve issues that span multiple infrastructure layers
- Ability to influence technical direction across infrastructure and engineering organizations
- Proven success mentoring engineers and promoting reliability engineering best practices
- Strong communication skills with the ability to translate technical concepts into business-focused outcomes
Pay Range
$159,000.00 - $305,000.00
Benefits
Health benefits
401(k) Plan
Paid time off
Disability benefits
Life insurance, critical illness insurance, and accident insurance
Parental leave
Critical caregiving leave
Discounts and savings
Commuter benefits
Tuition reimbursement
Scholarships for dependent children
Adoption reimbursement
Job Expectations
This position offers a hybrid schedule.
This position does not offer Visa sponsorship.
Recruitment and Hiring Requirements
Third-Party recordings are prohibited unless authorized by Wells Fargo.
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Recruitment and Hiring Requirements
Wells Fargo Recruitment and Hiring Requirements