Manager, Site Reliability Engineering (Auth0)
Okta · Northern Virginia, VA · Yesterday
Hybrid$182k/yrFull-time
About the role
The SRE Leadership Team at Okta is seeking a Manager, Site Reliability Engineer to lead the team's technical direction, drive complex initiatives, and champion reliability best practices.
Responsibilities
- Lead the SRE team's technical direction, translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams
- Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems
- Build infrastructure resilience, designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency
- Champion reliability best practices, establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts
- Mentor and develop SRE talent, elevating team capabilities through pair programming, design discussions, and code reviews while fostering a culture of continuous learning
- Represent reliability as a senior technical leader in architectural reviews and strategic planning, ensuring reliability is a core consideration in major engineering decisions
Requirements
- 3+ years of hands-on team leadership in SRE or software engineering roles within cloud-native environments, combined with 8+ years of total industry experience
- Deep expertise in cloud platforms (AWS, Azure) and infrastructure as code (Terraform), with proven experience managing cloud-native architectures including containers, Kubernetes, microservices, and databases
- Strong programming skills in Go or Python, with a track record of building and maintaining production-grade tools, automation, and infrastructure solutions
- Data-driven mindset grounded in SRE principles: blameless culture, systematic problem-solving, and the ability to apply software engineering approaches to operational challenges
- Exceptional communication skills—both verbal and written—enabling you to drive clarity during high-pressure incidents and articulate complex concepts to diverse stakeholders
- Promised ability to build and lead high-performing teams in globally distributed, remote-first environments with strong interpersonal and collaboration skills
- Strategic vision and technical depth, combining leadership acumen with hands-on technical excellence and a passion for mentoring senior engineers and shaping team direction
Qualifications
- Experience leading reliability initiatives that directly improved system uptime and reduced incident response times at scale
- Contributions to open-source infrastructure or observability tooling
- Experience designing and implementing comprehensive incident response programs and runbook automation
Skills
- Experience working in federal environments
- U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee)
Benefits
Okta offers competitive compensation, including an annual base salary range between $182,000 USD - $250,800 USD for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York, and Washington. Additional benefits include:
- Equity (where applicable)
- Bonus
- Health, dental, and vision insurance
- 401(k)
- Flexible spending account
- Paid leave (including PTO and parental leave)
To learn more about our Total Rewards program, please visit: https://rewards.okta.com/us.