Site Reliability Engineering Lead
Regular, full-time position requiring English fluency. Work hours are 1st shift (United States). Candidate must be willing to work onsite Monday–Friday at either office in Charlotte, NC; Raleigh, NC; or Atlanta, GA.
About the role
This senior technical leadership position drives reliability and operational excellence for enterprise platforms spanning hybrid-cloud and on-premises environments. The Site Reliability Engineering Lead enhances automation, observability, and incident management while collaborating across business and technology teams to advance enterprise-wide reliability frameworks.
Responsibilities
- Implement software architecture and engineering approaches for complex initiatives, contributing to technical plans and operational targets.
- Adopt and refine advanced software engineering standards, practices, and governance mechanisms, influencing quality, reliability, and delivery across teams.
- Collaborate with senior engineers, product partners, and architecture teammates to shape technology approaches and solution patterns that inform roadmaps and priorities.
- Lead end-to-end technical design and implementation of scalable, secure, and highly available software solutions, producing reusable patterns for other technical professionals.
- Independently troubleshoot and resolve complex technical issues, designing innovative architectures and performance improvements that advance business objectives.
- Provide ongoing technical guidance, coaching, and training to engineers, delegating and reviewing work to raise the technical bar through design reviews and knowledge sharing.
- Evaluate emerging technologies and techniques, building prototypes and solution concepts that contribute to new features, products, or capabilities.
- Contribute to long-term technical goals and plans through well-reasoned recommendations, design proposals, and implementation experience.
- Lead large or complex initiatives, coordinating and delegating technical work across teams to ensure cohesive, high-quality outcomes with limited supervision.
- Lead major and high-severity incident response efforts, diagnosing root causes and driving multi-team technical resolution.
- Drive problem management to closure, ensuring systemic fixes replace recurring operational risks.
- Establish and maintain standardized incident playbooks, escalation paths, and communication frameworks.
- Architect and deliver automation solutions that eliminate toil, reduce mean time to recovery (MTTR), and increase service resilience.
- Implement intelligent alerting, anomaly detection, and event correlation using AI and AIOps tools.
- Guide and enforce adoption of service-level objectives (SLOs) and service-level indicators (SLIs) across product teams, ensuring metrics inform decision-making.
- Enhance telemetry coverage across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
- Define and standardize enterprise observability practices, dashboards, and key performance indicators (KPIs).
- Ensure operational readiness through resiliency testing, chaos engineering, and failure-mode validation.
- Partner with Delivery, Architecture, Security, and Risk teams to embed reliability and resilience into design and execution.
- Act as a change agent to elevate operational maturity and drive transformative improvements across the organization.
- Lead workshops, maturity assessments, and enablement sessions through the SRE Center of Enablement (C4E) and Communities of Practice.
- Develop, maintain, and enforce runbooks, response playbooks, and automated recovery patterns.
- Contribute to enterprise SRE frameworks, templates, and maturity models.
- Promote consistent adoption of best practices across domains and lines of business.
- Coach and mentor Associate, Professional, and Senior SREs to build technical depth and operational discipline.
- Provide thought leadership in SRE methodologies, cloud-native operational patterns, and automated reliability engineering.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or related field.
- Minimum of 7 years of professional experience in software development.
- Deep knowledge of multiple programming languages, software architecture, and design principles.
- Deep understanding of software development lifecycle, testing, deployment, and security practices.
- 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations.
- Deep hands-on experience with distributed systems, container orchestration (Kubernetes), and cloud-native operational tooling.
- Proficiency with automation and scripting languages (Python, Go, PowerShell, Ansible).
- Strong understanding of observability platforms (e.g., Splunk, Dynatrace) and event-driven monitoring.
- Proven leadership in major incident management and cross-team technical coordination.
- Strong grasp of networking, Linux/Unix internals, and modern infrastructure patterns.
- Excellent communication skills, including executive-level situational awareness during critical incidents.
- Demonstrated ability to influence technical roadmaps and drive adoption of reliability best practices.
Preferred Qualifications
- Advanced degree in Computer Science or related technical discipline.
- Professional certifications such as Certified Software Development Professional (CSDP) or equivalent.
- Deep expertise in cloud-native architectures, microservices, and DevOps practices.
- Strong familiarity with Agile frameworks, continuous integration/continuous deployment (CI/CD), and enterprise innovation management.
- Financial services or regulated industry experience.
- Experience enabling large-scale SRE transformations or modernization initiatives.
- Familiarity with chaos engineering, resilience assessments, and service failure modeling.
- Exposure to hybrid-cloud and multi-cloud operational frameworks.
- Experience contributing to or leading Center for Enablement functions or Communities of Practice.
Benefits
All regular teammates (not temporary or contingent workers) working 20 hours or more per week are eligible for benefits. Eligibility for specific benefits may vary by division.
- Medical, dental, and vision insurance.
- Life insurance, disability, and accidental death and dismemberment coverage.
- Tax-preferred savings accounts and a 401k plan.
- No less than 10 days of vacation (prorated based on date of hire and full-time/part-time status) during the first year of employment, along with 10 sick days and paid holidays.
- Depending on position and division, eligibility may extend to a defined benefit pension plan, restricted stock units, and/or a deferred compensation plan.