Site Reliability Engineering Lead
Truist · Atlanta, GA · 2 wk ago
Full-time
Regular, full-time position. Fluency in English required. 1st shift (United States of America).
About the Role
The Site Reliability Engineering Lead role focuses on enhancing the reliability and operational excellence of enterprise platforms across hybrid cloud and on-premises environments. This senior technical leader drives improvements in automation, observability, and incident management while collaborating across multiple business and technology teams.
Responsibilities
- Implements software architecture and engineering approaches for complex initiatives, contributing to technical plans and achieving operational targets.
- Adopts and refines advanced software engineering standards, practices, and governance mechanisms to improve quality, reliability, and delivery.
- Collaborates with senior engineers, product partners, and architecture teammates to shape technology approaches and inform local roadmaps.
- Leads the end-to-end technical design and implementation of scalable, secure, and highly available software solutions.
- Independently troubleshoots and resolves complex technical issues, designing innovative architectures and performance improvements.
- Provides technical guidance, coaching, and training to other engineers, delegating and reviewing work to raise the technical bar.
- Evaluates emerging technologies and techniques, building prototypes and solution concepts for new features or capabilities.
- Contributes to long-term technical goals and plans through well-reasoned recommendations and design proposals.
- Leads large or complex initiatives, coordinating technical work across teams to ensure high-quality outcomes.
Key Responsibilities
- Incident & Problem Management Leadership: Lead major and high-severity incident responses, diagnose root causes, and drive multi-team resolutions. Establish incident playbooks, escalation paths, and communication frameworks.
- Reliability Engineering & Automation: Architect and deliver automation solutions to reduce toil, MTTR, and increase service resilience. Implement intelligent alerting and SLO/SLI adoption.
- Observability & Operational Excellence: Enhance telemetry coverage using platforms like Dynatrace and Splunk. Define standardized observability practices and ensure operational readiness.
- Cross-Functional Leadership & Influence: Partner with Delivery, Architecture, Security, and Risk teams to embed reliability into design and execution. Lead workshops and maturity assessments.
- Standardization & Documentation: Develop and enforce runbooks, response playbooks, and automated recovery patterns. Contribute to enterprise SRE frameworks.
- Mentorship & Technical Development: Coach and mentor SREs to build technical depth and operational discipline. Provide thought leadership in SRE methodologies.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or related field.
- Minimum of 7 years of professional experience in software development.
- Deep knowledge of multiple programming languages, software architecture, and design principles.
- Deep understanding of software development lifecycle, testing, deployment, and security practices.
- 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
- Hands-on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
- Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
- Strong understanding of observability platforms (Splunk, Dynatrace) and event-driven monitoring.
- Proven leadership in major incident management and cross-team technical coordination.
- Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
- Excellent communication skills, including executive-level situational awareness during critical incidents.
- Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
Preferred Qualifications
- Advanced degree in Computer Science or related technical discipline.
- Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
- Deep expertise in cloud-native architectures, microservices, and DevOps.
- Strong familiarity with Agile frameworks and CI/CD.
- Financial services or regulated industry experience.
- Experience enabling large-scale SRE transformations or modernization initiatives.
- Familiarity with chaos engineering, resilience assessments, and service failure modeling.
- Exposure to hybrid-cloud and multi-cloud operational frameworks.
- Experience contributing to or leading Center for Enablement functions or Communities of Practice.
Location
Candidate must be willing to work onsite Monday - Friday at either office in Charlotte, NC; Raleigh, NC; or Atlanta, GA.
Benefits
All regular teammates (working 20+ hours per week) are eligible for benefits, including:
- Medical, dental, vision, life insurance, and disability coverage.
- Tax-preferred savings accounts and a 401k plan.
- No less than 10 days of vacation (prorated based on hire date and full-time/part-time status) during the first year, plus 10 sick days and paid holidays.
- Depending on the position, eligibility for a defined benefit pension plan, restricted stock units, and/or a deferred compensation plan.