Jobs · Engineering · Utah

Site Reliability Engineer III- Operations

Woodbury School of Business · Orem, UT · 1 wk ago
EngineeringFull-time

About the Role

At Utah Valley University, you'll have the opportunity to build and support the technology solutions that help students, faculty, and staff succeed every day. In this role, you will design, implement, and maintain reliable, scalable systems and automation that improve service delivery, strengthen system performance, and support the university's ongoing digital transformation efforts. Working with modern cloud, infrastructure, and site reliability practices, you will play a key role in ensuring critical services remain secure, available, and responsive for the campus community. This position offers a dynamic mix of engineering, operations, and innovation. You will collaborate with technical teams and university leaders to enhance system reliability, develop monitoring and automation solutions, support technology upgrades, and resolve complex challenges. Ideal for a technology professional who enjoys continuous learning and problem-solving, this role provides the chance to make a meaningful impact while working in a collaborative environment that values expertise, initiative, and service excellence.

Responsibilities

  • Performs day-to-day administration, maintenance, upgrades, and operation of existing and recently developed systems, including Virtualization infrastructure on-premises and in the cloud, Microsoft Windows System administration, and Linux operating systems, as well as application systems and technologies.
  • Ensures standard operating procedures, runbooks, disaster recovery, and service catalog definitions are mature and ready for production.
  • Responsible for lifecycle of product set maintenance, availability, reliability, and performance reporting to decision makers and developers.
  • Perform tasks, as needed, to augment work needed on systems by their engineers to ensure timely achievement of project plans and goals.
  • Advocate, contribute, recommend, and facilitate ever-improving standards and best practices through successful adoption of change within UVU’s Digital Transformation department.
  • Ensure that the underlying infrastructure is running smoothly, and that systems and tools are working as expected.
  • Analyze day-to-day functions and the processes of systems and network management software to ensure they are performing within predetermined specifications.
  • In support of core systems availability and reliability, integrate diverse monitoring solutions for emerging and existing IT infrastructure using automation and API tools in on-premises and cloud architectures.
  • Engineer centralized, enterprise-wide alerting and key performance indicators that give timely, actionable information to subject matter experts, stakeholders, and leadership.
  • Conduct post-incident reviews, documenting findings and acting on lessons learned. Following the incident resolution, revisit the issue and determine the cause.
  • Build or optimize the incident lifecycle to bolster the reliability of services.
  • Maintain documentation and runbooks to ensure that teams get information when they need it.
  • Develop operational tools and processes, builds reliable systems, ensures compliance with operational standards, and provides support to operational staff.
  • Write and develop code to automate processes, such as analyzing logs, testing production environments, and responding to issues, allowing developers and engineers to focus on bug fixes and building new features.
  • Provides leadership, communications, development, engineering, automation, and feedback necessary for enterprise planning and architecture.
  • Participate in after-hours and weekend on-call rotation and provide training to other on-call staff.
  • Provide remote hands for systems and application administrators that need physical and virtual support within on-premises and cloud facilities.
  • Perform other job-related duties as assigned.

Qualifications

Graduation from an accredited institution with a bachelor's degree in Information Technology or a related field plus three years of work experience in IT; OR a combination of education and experience in a related field totaling seven years. Examples include:

  • Bachelor's degree in Information Technology or a related field + 3 years of related work experience.
  • Two years of completed college coursework in Information Technology or a related field + 5 years of related work experience = 7 years.
  • Technical certification program or college coursework equivalent to 1 year + 6 years of related work experience = 7 years.
  • 7 years of progressively responsible related work experience.

Licenses / Certifications

  • Site Reliability Engineering (SRE) Professional Certificate
  • Microsoft Certified: Azure Fundamentals, Administrator, Developer, or Associate-level certifications
  • Amazon Web Services (AWS) Cloud Practitioner or Associate-level certifications
  • Information Technology Infrastructure Library (ITIL) certification
  • The Open Group Architecture Framework (TOGAF) certification
  • Docker Certified Associate (DCA)
  • Certified Kubernetes Administrator (CKA)

Knowledge

  • ITIL Change, Incident, and Problem Management.
  • TCP/IP, firewall management, and operating system configuration.
  • Industry trends, tools, and processes.
  • Agile and iterative development processes (e.g., Scrum and Kanban).
  • Automation and containerization technologies such as Docker, Kubernetes, Ansible, Terraform, and SaltStack.
  • ITSM platforms such as Jira Service Management, ServiceNow, or others.
  • Engineering practices: availability, reliability, scalability, and disaster recovery.
  • Various automation tools for enhancing system reliability and scalability.

Skills

  • Recognize key design, implementation, and process issues and proactively craft and automate solutions.
  • System engineering and design for NOC/SOC purposes.
  • Scripting languages such as Perl, PowerShell, Bash, and Python.
  • Common programming languages including JavaScript, HTML5, CSS, JQuery, JSON, and PHP.
  • Design, implementation, and maintenance of Active Directory and/or LDAP directories.
  • TCP/IP, application network protocols, firewall management, operating system configuration, anti-virus software, and relational databases.
  • Monitoring solutions such as Prometheus, PRTG, Site24x7, TestCafe, Selenium, Splunk, New Relic, Azure Monitor, and AWS CloudWatch.
  • Major cloud providers such as Azure, AWS, and Google Cloud.
  • Alert management/on-call tools such as PagerDuty, VictorOps, and Opsgenie.
  • Instant communication and team collaboration platforms like MS Teams, Slack, or Jitsi.
  • IT project planning and development.

Abilities

  • Expert ability to read, write, and interpret technical documentation, runbooks, procedure manuals, and knowledge-base articles pertaining to network systems and application management.
  • Complete Root Cause Analysis (RCA) investigations and write post-incident reports.
  • Improve team practices through code reviews, handoffs of work and incidents.
  • Participate in an on-call (PagerDuty) rotation to respond to incidents that impact availability, and provide support for service engineers with customer incidents.
  • Debug production issues and build monitoring that alerts on symptoms rather than on outages.
  • Turn issues into repeatable actions and automation.
  • Conduct and direct research into IT issues and products, as required.
  • Communicate technical ideas and concepts to a non-technical audience.

Similar jobs