Senior Manager, Site Reliability Engineering
Oracle · United States · 2 wk ago
RemoteRemoteQuality Assurance$122k–$264k/yrFull-time
About the role
This role supports team members in designing and architecting infrastructure and service, sharing guidance on practices for reliability and functionality. It involves directing the team to ensure accurate forecasting and adequate resources, monitoring data, and implementing standards for automation.
Responsibilities
- Supports team members in designing and architecting infrastructure and/or service, sharing guidance on practices and terms for reliability and functionality.
- Supervises team members and provides direction to ensure accurate forecasting of demands for infrastructure and response to capacity needs, ensuring systems have sufficient resources to handle current and future workloads and identifying resource gaps.
- Maintains a collaborative relationship with the software development team to develop infrastructures, ensuring features are reliable and scalable according to deployment requirements.
- Monitors data collection, triage, technical analysis, and redirection, ensuring team members maintain and optimize operations and infrastructure reliability.
- Aids in incident response activities to ensure service reliability.
- Makes sure team members monitor health and performance reports and take appropriate actions based on trends in data.
- Implements standards for identifying and recommending opportunities for automation and assesses potential benefits to enhance operational efficiency.
- Takes a proactive role in reviewing and offering feedback on design, automation tools, or scripts, acting as a leader during implementation.
- Shares strategies for conducting testing on automations to ensure they perform tasks correctly and produce expected results.
- Reviews and provides feedback on release notes and ensures team members communicate comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.
- Proactively anticipates and articulates the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.
- Serves as a senior management point and shares expectations for documentation.
- Sets expectations for experimenting with new technology, executing improvements, building site reliability knowledge, and providing clear data.
Qualifications
- Experience in designing and architecting infrastructure and service.
- Knowledge of practices and terms for reliability and functionality.
- Ability to supervise and provide direction for accurate forecasting and resource allocation.
- Collaborative relationship with the software development team to develop reliable, scalable infrastructures.
- Experience in monitoring data collection, triage, technical analysis, and redirection.
- Ability to aid in incident response activities to ensure service reliability.
- Experience in monitoring health and performance reports and taking appropriate actions based on trends in data.
- Proven ability to implement standards for identifying and recommending opportunities for automation.
- Experience in reviewing and offering feedback on design, automation tools, or scripts, acting as a leader during implementation.
- Experience in sharing strategies for conducting testing on automations to ensure they perform tasks correctly and produce expected results.
- Experience in reviewing and providing feedback on release notes and ensuring team members communicate comprehensive information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers and immediate and related teams.
- Proactive approach to anticipating and articulating the potential impact of infrastructure, feature, and tool changes, considering their impact across team operations.
- Experience serving as a senior management point and sharing expectations for documentation.
- Experience setting expectations for experimenting with new technology, executing improvements, building site reliability knowledge, and providing clear data.