Systems Reliability Engineer
Members 1st Federal Credit Union · Enola, PA · 1 wk ago
Information Technology$60k/yrFull-time
Responsibilities
- Triage and resolve Tier II support issues escalated by the Help Desk, including application access issues, system performance degradation, service outages, and infrastructure-level incidents.
- Perform system-level operational tasks to restore service, including IIS resets, application pool recycling, service restarts, and execution of approved PowerShell scripts in accordance with security and change controls.
- Maintain IT operational runbooks using both traditional documentation and AI-assisted querying and revision techniques.
- Collaborate with Tier III, infrastructure, security, and application teams to escalate, coordinate, and resolve complex technical incidents.
- Support change management efforts, including review, testing coordination, execution of approved changes, and post-change validation.
- Analyze trends in incidents and requests to provide insights that support service reliability, operational maturity, and continuous improvement.
- Create, execute, and maintain IT operational runbooks to ensure consistent, repeatable support and response procedures.
- Document support procedures, troubleshooting steps, and resolutions in ServiceNow Knowledge Base articles.
- Capture successful support outcomes and lessons learned in shared knowledge repositories to improve Tier I efficiency and enable self-service.
- Identify and analyze business needs, gather requirements, and help define scope and objectives for system enhancements or operational improvements.
- Research business requirements and document relationships between users, business processes, data, applications, and devices to support impact analysis.
- Translate business requirements into clear, actionable application or operational requirements.
- Make recommendations for technology-based solutions or process improvements using new or existing tools.
- Assist with product refinement and backlog prioritization by contributing operational insights and support data.
- Use AI-enabled monitoring and event intelligence tools to receive, interpret, and prioritize operational alerts across applications, infrastructure, security, and integrations.
- Leverage AI-assisted dashboards and visualizations to assess system health, service dependencies, and operational risk.
- Query AI tools to identify upstream and downstream impacts during incidents, changes, or performance issues.
- Use AI-assisted analysis to validate, enhance, and continuously improve operational runbooks, identifying gaps and outdated procedures.
- Collaborate with engineering and platform teams to validate AI-suggested remediation actions prior to execution in regulated production environments.
- Provide feedback to improve AI operational tooling accuracy, relevance, and governance over time.
- Participate in on-call services providing after-hours support.
Skills
- Strong analytical and troubleshooting skills, with experience managing and resolving complex incidents and participating in incident and problem management activities using ITIL-aligned practices.
- Solid understanding of software engineering concepts and delivery disciplines, including iterative development, CI/CD pipelines, unit testing, and how applications are built, monitored, supported, and maintained in production environments.
- Hands-on technical experience supporting Windows Server environments, IIS administration, PowerShell scripting, and foundational system administration, including working knowledge of TCP/IP networking, DNS, load balancing, and SSL/TLS certificate management.
- Working knowledge of ITSM processes and tools, particularly incident, request, change, and problem management within platforms such as ServiceNow.
- Experience leveraging AI-assisted tools for IT Operations, including operational monitoring, alert analysis, impact assessment, incident triage, runbook execution, and documentation updates, with demonstrated judgment in validating AI-generated insights before execution in production environments.
- Ability to query and prompt AI tools effectively to support operations, troubleshooting, and continuous improvement initiatives.
- Strong time management, evaluation, and adaptability skills, with the ability to stay current on emerging IT tools, trends, and operational best practices.
- Experience with AI-enabled ITSM, AIOps, or observability platforms (for example ServiceNow AI features, AIOps tools, or intelligent monitoring platforms) preferred.
- Familiarity with service dependency mapping, topology visualization, or blast-radius analysis tools preferred.
- Experience improving operational workflows through automation, AI augmentation, or decision-support tooling preferred.
Qualifications
- 1-3 years of related experience.