Senior Manager, Site Reliability Engineering
hackajob · Reston, VA · 2 wk ago
On-siteQuality Assurance$122k–$264k/yrFull-time
Oracle is seeking a senior infrastructure leader to guide a team responsible for capacity planning, incident response, automation, and service reliability across Oracle’s cloud platforms.
Responsibilities
- Design and architect infrastructure and services, providing guidance on reliability and scalability best practices.
- Forecast infrastructure demands, identify resource gaps, and ensure systems have sufficient capacity for current and future workloads.
- Maintain collaboration with software development teams to build reliable, scalable infrastructure aligned with deployment requirements.
- Lead prototyping initiatives to test new applications, infrastructures, or onboarding processes.
- Monitor data collection, triage, technical analysis, and incident redirection to optimize operations and infrastructure reliability.
- Support team members in incident response, root cause analysis, and maintenance (e.g., software installs, version upgrades, security updates).
- Oversee comprehensive health and performance reporting, ensuring appropriate actions are taken based on data trends.
- Enforce procedures for provisioning infrastructure, applications, and services, and decommissioning unused resources.
- Identify and recommend automation opportunities to enhance operational efficiency, review automation tools/scripts, and lead implementation.
- Develop and test automation solutions to ensure they perform tasks correctly and produce expected results.
- Review and provide feedback on release notes, ensuring clear communication about service scale, capacity, security, and performance attributes.
- Anticipate and articulate the impact of infrastructure, feature, or tool changes across team operations.
- Serve as a senior escalation point for complex incidents, ensuring effective investigation, debugging, and resolution to meet service level objectives (SLOs).
- Guide team members in documenting incidents and performing root cause analyses to prevent recurrence.
- Enforce adherence to service level agreements (SLAs) with customers.
- Set expectations for evaluating cutting-edge tools and technologies to optimize infrastructure performance and reliability while adhering to security standards.
- Prioritize initiatives to improve performance bottlenecks, deployments, and resource efficiency.
- Develop and maintain knowledge of site reliability trends, sharing insights with team members, management, and other stakeholders.
- Contribute to business development decisions using team analyses and data.
- Manage multiple medium- to large-scale projects, ensuring timelines, deliverables, and budgets are met.
- Provide direction to teams on project work, setting priorities aligned with business needs.
- Drive cross-functional partnerships to align expectations and objectives across teams.
- Coach team members to develop strategic relationships with business leaders, stakeholders, and external partners.
- Promote inclusivity by actively seeking and listening to diverse perspectives.
- Guide teams in addressing complex operational and technical issues, analyzing data to identify solutions.
- Review and provide insights into unresolved or critical issues to help teams identify potential solutions.
- Model continuous learning to deepen expertise and integrate best practices into strategic planning.
- Leverage feedback to drive personal and team skill improvements.
- Identify skill gaps across teams and empower members to pursue learning opportunities.
- Drive collaboration on developing and implementing ideas to increase process efficiency and effectiveness.
- Encourage adoption of new approaches and methods, fostering a culture of continuous improvement.
- Provide performance feedback and coaching aligned with organizational guidelines and expectations.
- Discuss development goals with team members and align individual goals with broader organizational objectives.
- Develop and manage the talent acquisition pipeline, including leading candidate interviews and monitoring promotion eligibility.
Requirements
- 9 years of experience in software engineering, infrastructure management, or a related field
- OR a Bachelor’s Degree in Computer Science, Engineering, or a related field AND 5 years of relevant experience
- OR a Master’s Degree in Computer Science, Engineering, or a related field AND 3 years of relevant experience
- OR a Doctorate in Computer Science, Engineering, or a related field AND 1 year of relevant experience
- 5 years of experience in automation
- 5 years of experience in programming and/or scripting
Preferred Qualifications
- 11 years of experience in software engineering, infrastructure management, or a related field
- OR a Bachelor’s Degree in Computer Science, Engineering, or a related field AND 7 years of relevant experience
- OR a Master’s Degree in Computer Science, Engineering, or a related field AND 5 years of relevant experience
- OR a Doctorate in Computer Science, Engineering, or a related field AND 3 years of relevant experience
- 2 years of experience in a leadership role with direct reports
- 2 years of experience working with operating budgets and/or project financials
- 7 years of experience in automation
- 7 years of experience in programming and/or scripting
This position is located in Reston, VA or Austin, TX.
Pay
Hiring range: $121,500 to $264,100 per annum. Eligible for bonus, equity, and compensation deferral.
Benefits
- Medical, dental, and vision insurance, including expert medical opinion
- Short-term and long-term disability insurance
- Life insurance and AD&D
- Supplemental life insurance (Employee/Spouse/Child)
- Health care and dependent care Flexible Spending Accounts
- Pre-tax commuter and parking benefits
- 401(k) Savings and Investment Plan with company match
- Flexible Vacation for salaried employees (13 days annually for first three years, 18 days thereafter; accrual prorated for part-time employees working 20–34 hours per week)
- 11 paid holidays
- Paid sick leave: 72 hours upon hire, refreshes annually, with a maximum carryover cap of 112 hours
- Paid parental leave
- Adoption assistance
- Employee Stock Purchase Plan
- Financial planning and group legal services
- Voluntary benefits, including auto, homeowner, and pet insurance