Senior Manager, Site Reliability Engineering
Ladders · San Francisco, CA · 2 wk ago
On-siteQuality Assurance$227k–$325k/yrFull-time
Responsibilities
- Lead and mentor a Site Reliability Engineering team, fostering innovation and technical excellence
- Define and drive the multi-year technical strategy for observability and automation platforms
- Own the availability, performance, and efficiency of critical user-facing services, evolving incident response practices
- Streamline existing processes and collaborate to enhance production release standards
- Manage SRE budget, tooling, and vendor relationships for observability and AI platforms
- Influence cross-functional engineering practices to ensure reliability, scalability, and observability of services from inception
- Lead the integration of AI into the SRE function, focusing on predictive maintenance and automation
Qualifications
- 8+ years of experience in a technical field with engineering leadership in SRE, DevOps, or Production Engineering
- Deep understanding of SRE principles, including SLIs, SLOs, error budgets, and capacity planning
- Exceptional communication and negotiation skills for articulating complex concepts to diverse stakeholders
- Strong technical background as a hands-on software engineer or site reliability engineer, particularly with AWS and Kubernetes
- Familiarity with modern SRE tools like Terraform, Kubernetes, Prometheus, Datadog, and incident tooling
Benefits
- Comprehensive medical, dental, and vision insurance
- 401(k) plan with employer contributions
- Flexible time off policy for salaried employees
- Generous parental leave program offering 12 weeks of paid leave
- Monthly wellness reimbursement for employees
Pay
$227,200 – $324,500 annually
Schedule
N/A
Location
San Francisco, CA - US based candidates only, no visa sponsorship available