Manager IT Infrastructure Systems
Valeris · Morrisville, NC · 1 wk ago
Information Technology$50k/yrFull-time
About the role
Network Operations Center (NOC) and Systems Administration is responsible for leading 24x7 IT operations to ensure the availability, performance, and security of enterprise infrastructure across on-premises, cloud, and SaaS environments. This role oversees centralized monitoring, incident response, and service management while providing technical leadership for system administration monitoring spanning Azure, identity management, M365, and core infrastructure platforms. The position drives operational excellence through ITIL-aligned processes, automation, and standardized runbooks, while delivering clear communication during incidents and maintaining accountability for service levels, system health, and continuous improvement.
Responsibilities
- NOC Operations & Service Monitoring:
- Lead 24x7 monitoring operations for network, infrastructure, cloud, and SaaS systems
- Ensure visibility across: Network (LAN/WAN, firewalls, load balancers); Servers, compute, storage, databases; Azure IaaS/PaaS and SaaS platforms (M365, etc.)
- Establish service-level monitoring aligned to business services and dependencies
- Oversee alerting strategy (thresholds, anomaly detection, service impact)
- Drive end-user experience monitoring and SaaS observability
- Incident, Problem & Change Management (ITIL):
- Own escalation model and execution of: Incident management; Problem management (root cause analysis); Change coordination (as applicable)
- Ensure all incidents include: Business impact; Root cause; Resolution actions
- Lead P1/P2 incident response, including bridge calls and leadership communications
- Drive MTTR reduction and SLA adherence
- Systems Administration Leadership:
- Oversee Enterprise System Administration Monitoring Across: Cloud Platform (Azure) – VMs, storage, networking, subscriptions, IAM; Identity & Access (Entra/AD) – SSO, MFA, RBAC, conditional access; M365 Administration – Exchange, SharePoint, Teams, Intune; Infrastructure – VMware, storage, backups, patching (Tanium/Veeam); Security & Certificates – PKI, certificate lifecycle, Key Vault; Monitoring & Tooling – Datadog, Tanium, automation tooling
- Runbooks, Automation & Knowledge Management:
- Establish and maintain standardized runbooks for all operational processes
- Integrate with ServiceNow knowledge base and workflows
- Drive automation and self-healing capabilities, including: Service restarts; System cleanup and recovery actions
- Ensure continuous update and governance of operational documentation
- Operational Governance & Reporting:
- Define and enforce operational KPIs and dashboards, including: System health and availability; Incident and outage metrics; SLA compliance and MTTR
- Deliver real-time operational dashboards and executive reporting
- Identify trends, recurring issues, and prevention strategies
- Communication & Stakeholder Management:
- Own incident communication framework (P1/P2/P3 severity levels)
- Ensure business-impact messaging (users impacted, workaround, ETA)
- Coordinate with: Application teams; Cloud/DevOps teams; Vendors and SaaS providers
- Provide post-incident reports and continuous improvement actions
- Process Improvement & ITIL Maturity:
- Standardize operational processes across NOC and Systems teams
- Improve: Monitoring coverage; Alert quality and noise reduction; Ticket lifecycle management
- Align to ITIL best practices and governance model
- Drive accountability across infrastructure operation
- Team Leadership & Development:
- Lead NOC and Systems Administration
- Define roles, responsibilities, and escalation paths
- Build a culture of: Accountability; Operational excellence; Continuous improvement
- Identify skill gaps and support training for modern cloud/NOC operations
Qualifications
- 8+ years in IT Infrastructure / Operations (with leadership experience)
- Strong expertise across: Azure (IaaS/PaaS, governance); Networking (firewalls, routing, load balancing); Enterprise systems (AD/Entra, M365, VMware)
- Experience with monitoring platforms (e.g., Datadog)
- Knowledge of automation and scripting (PowerShell, Python preferred)
- Strong ITIL experience (Incident, Problem, Change)
- Experience managing NOC or enterprise operations teams
- Proven ability to drive operational transformation and standardization
- Strong communication skills for executive and technical audiences
Benefits
- Medical, dental, and vision plans, including HSA- and FSA-eligible options, with Valeris contributing toward premium costs
- Additional health support, including telehealth and Employee Assistance Program (EAP) services
- Company match on Health Savings Account contributions
- Free Basic Life and AD&D coverage equal to your annual earnings, with a minimum of $50,000 and a maximum of $300,000
- Company-paid Short-Term Disability coverage, with the option to purchase Long-Term Disability
- 401(k) Retirement Savings Plan with 100% match on the first 5% you contribute, with immediate vesting
- Paid Time Off (PTO) and Sick Leave to support work-life balance
- Nine paid holidays plus two floating holidays
- Opportunities for advancement in a company that supports personal and professional growth
- A challenging, stimulating work environment that encourages new ideas
- Work for a company that values diversity and makes deliberate efforts to create an inclusive workplace
- A mission-driven, inclusive culture where your work makes a meaningful impact