NOC Lead
Western Alliance Bank · Columbus, OH · 1 wk ago
Information TechnologyFull-time
About the role
As a Network Operations Center Lead - Staff Engineer II you'll provide expertise in a suite of technical tools and products across network operations to ensure production services are safe, secure, compliant, stable, and reliable. You'll identify operational support needs, break/fix gaps, service restoration improvements, and production readiness risks while leading day-to-day technical execution across incidents, requests, escalations, and recurring operational issues. You'll facilitate dialogue across NOC L1, NOC L2, fulfillment, observability, Critical Response Coordination, and engineering teams.
What You'll Do
- Provide day-to-day technical leadership for shift execution, queue health, escalation quality, operational readiness, and incident response while developing effective presentations and narratives for technical teams and leadership.
- Lead operational support planning for production environments, including break/fix response, service restoration, escalation readiness, and technical recovery execution.
- Design and implement engineering principles aimed at reliability practices and sound recovery procedures.
- Design transaction specific application performance monitoring metrics that can be captured and passed to business partners.
- Develop and maintain technical documentation, including system configurations and procedures while also ensuring compliance with IT policies, procedures, and industry standards.
- Develop desktop procedures when needed that others follow.
- Provide day-to-day technical leadership for NOC L1, NOC L2, and fulfillment activities across distributed shifts.
- Drive queue discipline, shift handoff quality, escalation quality, and follow-through for incidents, service requests, operational tasks, and aging work.
- Serve as a senior technical escalation point for network incidents, service degradation, monitoring gaps, and high-impact operational issues.
- Coordinate with Critical Response Coordinators during major incidents by providing technical triage, impact clarity, remediation status, and next-step ownership.
- Partner with engineering, observability, security, automation, and service management teams to identify recurring issues, reduce alert noise, improve monitoring quality, and strengthen operational readiness.
- Coach analysts and engineers on troubleshooting approach, documentation expectations, ticket hygiene, escalation standards, and operational accountability.
- Support network readiness activities including health checks, patching, vulnerability remediation, disaster recovery exercises, and operational validation.
- Prepare clear operational updates, leadership summaries, post-incident notes, and improvement recommendations that explain risk, progress, constraints, and decisions.
- Complete work in a timely and accurate manner while providing exceptional customer service.
What You'll Need
- 7+ years of related experience in IT Networking, Network Operations, Network Security, Infrastructure Support, IT App Support, IT Development, or similar field.
- Bachelor's degree in Computer Science, Information Technology, Information Systems, Engineering, or a related technical or business-technology field required.
- Recommended industry certifications include CCNP Enterprise or equivalent industry network certification, Palo Alto NSP or equivalent network security certification, and ITIL Foundation.
- Previous leadership experience preferred.
- Intermediate to advanced knowledge of general Financial Services or Banking is preferred.
- Advanced knowledge of applicable regulatory and legal compliance obligations, rules and regulations, industry standards and practices.
- Advanced proven experience leading operational support teams through incident response, break/fix troubleshooting, production escalations, service restoration, and recurring issue remediation.
- Additional experience coordinating operational work across ticket queues, monitoring alerts, change activity, vulnerability remediation, and platform support responsibilities.
- Advanced ability to understand the operational impact of production issues and align break/fix priorities with service stability, customer impact, risk reduction, and business continuity.
- Capable of leading and motivating cross-functional operational teams during escalations, resolving support conflicts, driving technical follow-through, and identifying risks that affect service reliability, recovery readiness, and operational execution.
- Proficient in governance patterns tied to incident handling, escalation quality, operational validation, change readiness, and technical support standards.
- Strong technical knowledge across enterprise routing, switching, firewall, DNS, DHCP, IPAM, load balancing, SD-WAN, cloud networking, monitoring, and network troubleshooting.
- Experience with ServiceNow, SolarWinds, LiveNX or LiveAction, ThousandEyes, NetBrain, AppDynamics, F5, Cloudflare DNS, Infoblox, Palo Alto, Cisco ASA / FTD / Firepower, Cisco FMC, Cisco Catalyst, Cisco Nexus, Cisco Meraki, Azure Firewall, and Azure Load Balancers preferred.
- Strong working knowledge of ITIL practices including incident, request, change, problem, knowledge, and major incident management.
- Advanced speaking and writing communication skills.
- May require up to 25% travel.
Benefits
- Competitive salaries
- Ownership stake in the company
- Medical and dental insurance
- Time off
- 401k matching program
- Tuition assistance program
- Employee volunteer program
- Wellness program
- Opportunity to bolster business knowledge and learn the ins and outs of how successful companies operate and manage their finances