Incident & Event Management Specialist (2nd Shift)
About the Role
As an Eyes on Glass (EOG) IT Incident and Event Management Specialist, you will support Veteran Affairs (VA) Event Management by improving monitoring and triage processes that help keep enterprise systems running reliably. You will work within the Enterprise Command Center (ECC) Major Incident and Event Management (MIEM) team to detect and investigate issues across enterprise applications and technology stacks, while helping evaluate and modernize VA enterprise systems. This opportunity is well-suited for a hands-on technical SME who enjoys fast-paced incident triage and enterprise-scale troubleshooting.
Responsibilities
- Detect and investigate issues across enterprise-level applications and technology stacks
- Assist in major incident triage by correlating and delivering real-time performance data reported across monitoring platforms
- Leverage comprehension of workflow systems and application processes within multiple system environments to escalate critical alerts to stakeholders
- Work with system, application and network administrators to tailor alert thresholds, reducing false positives
- Escalate issues with monitoring tools to the proper support teams
- Provide feedback to improve Knowledge Base Articles
- Provide shift support during weekends, holidays, or off-hours, as required
Requirements
- Bachelor's degree in Computer Science or Engineering and 5+ years of experience in a professional work environment, or 13+ years of experience in a professional work environment in lieu of a degree
- 5+ years of experience with systems administration and operations, including maintaining responsibility for overseeing operational performance, and performing performance trend analysis and alert generation for failed or failing services in a professional IT work environment
- 5+ years of experience with application/platform support including monitoring tool use/configuration, maintenance support to include change control and experience participating and/or running major incident response
- 5+ years of experience deploying, maintaining, and troubleshooting complex applications at an enterprise scale while working with cross-functional teams
- 5+ years monitoring and troubleshooting experience with two or more of the following APM tools: DynaTrace, ScienceLogic, Elastic, AppDynamics, LogicMonitor or SolarWinds
- 5+ years enterprise IT experience in Unix/Linux, Windows and mainframe
- 1+ years of experience in service virtualization, AWS or Azure Cloud technologies, containers, and SaaS and PaaS implementation
- Experience with tracking and reporting team scheduling considerations as well as navigating client communications regarding team performance
- Experience with using Microsoft Office, including Word, Excel, Teams and PowerPoint
- Ability to obtain and maintain a Public Trust clearance
Desired Skills
- Proficiency in enterprise ITSM tools with emphasis on ServiceNow or other similar tools
- Experience with virtual team management
- Prior experience supporting IT operations or systems within a healthcare environment
Schedule
On-site in Austin, TX. This position is scheduled for Second Shift, working Sunday through Thursday from 2:30PM to 11:00PM CT.
Pay
Pay range: $80,000-$100,000 (annual). This pay range is based on qualifications and experience, and assumes all role requirements are met.
Why Apply
You will play a key role in enterprise incident and event management, working alongside major incident, monitoring, problem management, and DevOps teams to improve detection, triage, and operational performance. If you enjoy hands-on troubleshooting, real-time monitoring, and collaborating across teams to drive better outcomes, this role offers meaningful, high-impact work.