Jobs · Quality Assurance · Virginia

Senior Manager, Site Reliability Engineering

hackajob · Reston, VA · 2 wk ago
On-siteQuality Assurance$122k–$264k/yrFull-time

Oracle is seeking a senior infrastructure leader to guide a team responsible for capacity planning, incident response, automation, and service reliability across Oracle’s cloud platforms.

Responsibilities

  • Design and architect infrastructure and services, providing guidance on reliability and scalability best practices.
  • Forecast infrastructure demands, identify resource gaps, and ensure systems have sufficient capacity for current and future workloads.
  • Maintain collaboration with software development teams to build reliable, scalable infrastructure aligned with deployment requirements.
  • Lead prototyping initiatives to test new applications, infrastructures, or onboarding processes.
  • Monitor data collection, triage, technical analysis, and incident redirection to optimize operations and infrastructure reliability.
  • Support team members in incident response, root cause analysis, and maintenance (e.g., software installs, version upgrades, security updates).
  • Oversee comprehensive health and performance reporting, ensuring appropriate actions are taken based on data trends.
  • Enforce procedures for provisioning infrastructure, applications, and services, and decommissioning unused resources.
  • Identify and recommend automation opportunities to enhance operational efficiency, review automation tools/scripts, and lead implementation.
  • Develop and test automation solutions to ensure they perform tasks correctly and produce expected results.
  • Review and provide feedback on release notes, ensuring clear communication about service scale, capacity, security, and performance attributes.
  • Anticipate and articulate the impact of infrastructure, feature, or tool changes across team operations.
  • Serve as a senior escalation point for complex incidents, ensuring effective investigation, debugging, and resolution to meet service level objectives (SLOs).
  • Guide team members in documenting incidents and performing root cause analyses to prevent recurrence.
  • Enforce adherence to service level agreements (SLAs) with customers.
  • Set expectations for evaluating cutting-edge tools and technologies to optimize infrastructure performance and reliability while adhering to security standards.
  • Prioritize initiatives to improve performance bottlenecks, deployments, and resource efficiency.
  • Develop and maintain knowledge of site reliability trends, sharing insights with team members, management, and other stakeholders.
  • Contribute to business development decisions using team analyses and data.
  • Manage multiple medium- to large-scale projects, ensuring timelines, deliverables, and budgets are met.
  • Provide direction to teams on project work, setting priorities aligned with business needs.
  • Drive cross-functional partnerships to align expectations and objectives across teams.
  • Coach team members to develop strategic relationships with business leaders, stakeholders, and external partners.
  • Promote inclusivity by actively seeking and listening to diverse perspectives.
  • Guide teams in addressing complex operational and technical issues, analyzing data to identify solutions.
  • Review and provide insights into unresolved or critical issues to help teams identify potential solutions.
  • Model continuous learning to deepen expertise and integrate best practices into strategic planning.
  • Leverage feedback to drive personal and team skill improvements.
  • Identify skill gaps across teams and empower members to pursue learning opportunities.
  • Drive collaboration on developing and implementing ideas to increase process efficiency and effectiveness.
  • Encourage adoption of new approaches and methods, fostering a culture of continuous improvement.
  • Provide performance feedback and coaching aligned with organizational guidelines and expectations.
  • Discuss development goals with team members and align individual goals with broader organizational objectives.
  • Develop and manage the talent acquisition pipeline, including leading candidate interviews and monitoring promotion eligibility.

Requirements

  • 9 years of experience in software engineering, infrastructure management, or a related field
  • OR a Bachelor’s Degree in Computer Science, Engineering, or a related field AND 5 years of relevant experience
  • OR a Master’s Degree in Computer Science, Engineering, or a related field AND 3 years of relevant experience
  • OR a Doctorate in Computer Science, Engineering, or a related field AND 1 year of relevant experience
  • 5 years of experience in automation
  • 5 years of experience in programming and/or scripting

Preferred Qualifications

  • 11 years of experience in software engineering, infrastructure management, or a related field
  • OR a Bachelor’s Degree in Computer Science, Engineering, or a related field AND 7 years of relevant experience
  • OR a Master’s Degree in Computer Science, Engineering, or a related field AND 5 years of relevant experience
  • OR a Doctorate in Computer Science, Engineering, or a related field AND 3 years of relevant experience
  • 2 years of experience in a leadership role with direct reports
  • 2 years of experience working with operating budgets and/or project financials
  • 7 years of experience in automation
  • 7 years of experience in programming and/or scripting

This position is located in Reston, VA or Austin, TX.

Pay

Hiring range: $121,500 to $264,100 per annum. Eligible for bonus, equity, and compensation deferral.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short-term and long-term disability insurance
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Flexible Vacation for salaried employees (13 days annually for first three years, 18 days thereafter; accrual prorated for part-time employees working 20–34 hours per week)
  • 11 paid holidays
  • Paid sick leave: 72 hours upon hire, refreshes annually, with a maximum carryover cap of 112 hours
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal services
  • Voluntary benefits, including auto, homeowner, and pet insurance

Similar jobs