Jobs · Information Technology · Pennsylvania

Systems Reliability Engineer

Members 1st Federal Credit Union · Enola, PA · 1 wk ago
Information Technology$60k/yrFull-time

Responsibilities

  • Triage and resolve Tier II support issues escalated by the Help Desk, including application access issues, system performance degradation, service outages, and infrastructure-level incidents.
  • Perform system-level operational tasks to restore service, including IIS resets, application pool recycling, service restarts, and execution of approved PowerShell scripts in accordance with security and change controls.
  • Maintain IT operational runbooks using both traditional documentation and AI-assisted querying and revision techniques.
  • Collaborate with Tier III, infrastructure, security, and application teams to escalate, coordinate, and resolve complex technical incidents.
  • Support change management efforts, including review, testing coordination, execution of approved changes, and post-change validation.
  • Analyze trends in incidents and requests to provide insights that support service reliability, operational maturity, and continuous improvement.
  • Create, execute, and maintain IT operational runbooks to ensure consistent, repeatable support and response procedures.
  • Document support procedures, troubleshooting steps, and resolutions in ServiceNow Knowledge Base articles.
  • Capture successful support outcomes and lessons learned in shared knowledge repositories to improve Tier I efficiency and enable self-service.
  • Identify and analyze business needs, gather requirements, and help define scope and objectives for system enhancements or operational improvements.
  • Research business requirements and document relationships between users, business processes, data, applications, and devices to support impact analysis.
  • Translate business requirements into clear, actionable application or operational requirements.
  • Make recommendations for technology-based solutions or process improvements using new or existing tools.
  • Assist with product refinement and backlog prioritization by contributing operational insights and support data.
  • Use AI-enabled monitoring and event intelligence tools to receive, interpret, and prioritize operational alerts across applications, infrastructure, security, and integrations.
  • Leverage AI-assisted dashboards and visualizations to assess system health, service dependencies, and operational risk.
  • Query AI tools to identify upstream and downstream impacts during incidents, changes, or performance issues.
  • Use AI-assisted analysis to validate, enhance, and continuously improve operational runbooks, identifying gaps and outdated procedures.
  • Collaborate with engineering and platform teams to validate AI-suggested remediation actions prior to execution in regulated production environments.
  • Provide feedback to improve AI operational tooling accuracy, relevance, and governance over time.
  • Participate in on-call services providing after-hours support.

Skills

  • Strong analytical and troubleshooting skills, with experience managing and resolving complex incidents and participating in incident and problem management activities using ITIL-aligned practices.
  • Solid understanding of software engineering concepts and delivery disciplines, including iterative development, CI/CD pipelines, unit testing, and how applications are built, monitored, supported, and maintained in production environments.
  • Hands-on technical experience supporting Windows Server environments, IIS administration, PowerShell scripting, and foundational system administration, including working knowledge of TCP/IP networking, DNS, load balancing, and SSL/TLS certificate management.
  • Working knowledge of ITSM processes and tools, particularly incident, request, change, and problem management within platforms such as ServiceNow.
  • Experience leveraging AI-assisted tools for IT Operations, including operational monitoring, alert analysis, impact assessment, incident triage, runbook execution, and documentation updates, with demonstrated judgment in validating AI-generated insights before execution in production environments.
  • Ability to query and prompt AI tools effectively to support operations, troubleshooting, and continuous improvement initiatives.
  • Strong time management, evaluation, and adaptability skills, with the ability to stay current on emerging IT tools, trends, and operational best practices.
  • Experience with AI-enabled ITSM, AIOps, or observability platforms (for example ServiceNow AI features, AIOps tools, or intelligent monitoring platforms) preferred.
  • Familiarity with service dependency mapping, topology visualization, or blast-radius analysis tools preferred.
  • Experience improving operational workflows through automation, AI augmentation, or decision-support tooling preferred.

Qualifications

  • 1-3 years of related experience.

Similar jobs

Systems Engineer – Reliability

General Dynamics Information TechnologyOmaha, NE· 1 mo ago
Information Technology$98k–$127k/yrapply on gdit.wd5.myworkdayjobs.com