Senior Infrastructure Operations Engineer
About the Role
We are seeking a Sr. Infrastructure Operations Engineer to join a dynamic team providing high-quality engineering, implementation, and support across Windows platforms, virtualization, storage, SQL Server, and cloud infrastructure—both on-premises and in Azure. This role serves as the senior technical escalation point for the Infrastructure Operations team, accountable for the design, resilience, and documented recoverability of the firm’s virtualization, storage, SQL, and cloud platforms.
As the firm’s designated Operations SQL Subject Matter Expert (SME), you will provide technical leadership for SQL Server operational support, availability, performance, maintenance, and recovery. This role requires strong communication skills, teamwork, and the ability to independently lead projects from planning to completion. A key objective is to strengthen the team’s overall capability by eliminating single points of failure through high-quality documentation, repeatable procedures, and cross-training.
Responsibilities
- Primary owner for Azure: administer and support the firm’s Azure estate, including Landing Zone, IaaS/PaaS, Entra ID / AAD Connect synchronization, SSO, Conditional Access, and MFA; lead workload migrations from on-premises to Azure.
- Primary owner and designated SME for SQL Server platform: end-to-end ownership of SQL Server installation, patching, version/edition upgrades, Always On Availability Groups, failover clustering, DR runbooks, backups, restores, and validated recovery testing. Serve as first-touch triage for all SQL-related tickets and route application-owned SQL work to the appropriate team.
- Primary owner for VMware and VMware SRM: administer vSphere/vCenter across multiple sites; own DR replication and recovery plans in Site Recovery Manager, including scheduled failover testing and validated recovery documentation.
- Primary owner for storage: administer Pure Storage and HPE 3PAR arrays—provisioning, capacity and performance management, firmware lifecycle, and replication.
- Primary owner for Dell server hardware: manage Dell OpenManage and Dell blade chassis—firmware lifecycle, hardware fault remediation, and vendor RMA coordination.
- Primary owner for monitoring: manage SolarWinds—node coverage accuracy, alert tuning, and reduction of alert noise to ensure alerts are actionable.
- Primary owner for enterprise application certificates: full lifecycle management, including a maintained renewal calendar to prevent expiry-driven outages.
- Secondary owner for Active Directory and Exchange On-Premises: provide depth behind existing primary owners.
- Serve as the senior technical escalation point for the Infrastructure Operations team: take ownership of major incidents, drive root cause analysis (RCA) to closure, and publish written RCAs.
- Create and maintain high-quality operational documentation and develop repeatable procedures to institutionalize critical expertise across the team.
- Participate in a weekly on-call rotation (8 AM Mondays, rotating every 4-5 weeks).
Requirements
- Broad hands-on knowledge of Windows Server (2012/2016/2019/2022), including in-place upgrades and failover-cluster methodologies (e.g., Windows Server 2012 end-of-life refresh for 14 SQL cluster nodes).
- Strong hands-on experience with Microsoft SQL Server (2014 through 2022), including planning and executing SQL version upgrades and failover-cluster node-evict/rebuild procedures.
- Hands-on experience with Azure (PaaS, IaaS) and workload migrations to Azure.
- Administration experience with virtualization environments (Azure and on-prem VMware) across multiple sites; HyperV experience a plus.
- Strong hands-on experience with Active Directory (Group Policies, ACL management) and PowerShell scripting.
- Strong hands-on experience with Office 365, Exchange on-prem, and migrations to Microsoft Teams, SharePoint, and OneDrive.
- Familiarity with tools such as 1Password, Barracuda, Atlassian Jira, LanSweeper, Dell OpenManage Enterprise, Intune, Pure Storage, SolarWinds, Varonis, and Veeam is a plus.
- Ability to conduct research into new technologies and products as required.
- High level of analytical and problem-solving abilities, with the capacity to prioritize tasks in a high-pressure environment.
- Strong project management skills, self-motivation, and the ability to work both independently and collaboratively.
- Excellent interpersonal, written, and verbal communication skills, with a focus on documentation (Visio diagram experience a plus).
- Strong customer service skills when interacting with management, peers, and office administrators.
Qualifications
- Bachelor’s degree in Computer Science or a related field, or equivalent practical experience.
- Minimum of 7 years of hands-on infrastructure engineering experience, including at least 3 years in a senior or escalation capacity.
- Demonstrated ability to act as a final technical escalation point—calm, methodical, and decisive under major-incident pressure.
- Ability to produce documentation that another engineer can execute without assistance (a primary measure of success in this role).
- Relevant certifications preferred: Microsoft Azure Administrator (AZ-104) or Azure Solutions Architect (AZ-305); Microsoft Certified: Azure Database Administrator Associate (DP-300) or equivalent SQL Server credential; VMware VCP-DCV; ITIL Foundation.
- College or University degree in a related field or 10+ years of equivalent work experience.
- Certifications in ITIL, Agile, and/or other technical acumen (Microsoft, Azure, O365) are a plus.
Benefits
- Outstanding benefits package, including a 401(k) match and generous PTO plan.
- Medical, dental, vision, disability, and life insurance.
- Access to corporate discount plans and other employee perks.
- Ample opportunities for professional development and career advancement.
Pay
Salary Range: $120,000 USD - $165,000 USD. A variety of factors are considered in making compensation decisions, including but not limited to experience, education, licensure and/or certifications, geographic location, market demands, and other business and organizational needs.
Schedule
- Full-time position with participation in a weekly on-call rotation (8 AM Mondays, rotating every 4-5 weeks).