Server Engineer Lead
About the role
The Server Engineer - Lead is a senior technical contributor responsible for architecting, operating, and optimizing the organization's enterprise compute and virtualization ecosystem. This role ensures the security, stability, scalability, and lifecycle management of a large fleet of Intel-based server hardware and virtual infrastructure built on VMware and other hypervisors that support critical manufacturing, ERP, engineering, and enterprise business applications across data center and hybrid cloud environments.
Responsibilities
- Lead the architecture, implementation, lifecycle management, and operational support of the enterprise compute environment, including Intel-based server platforms and virtualization platforms based on VMware and other hypervisors.
- Manage server hardware standards, platform design, firmware strategy, host lifecycle management, and infrastructure refresh planning across the enterprise.
- Oversee the health, performance, availability, and capacity of virtualization environments, including hypervisor hosts, management platforms, clusters, distributed resource scheduling, and high availability configurations.
- Design and maintain resilient compute and virtualization solutions that support mission-critical enterprise workloads with strong emphasis on uptime, recoverability, and operational consistency.
- Partner with storage, networking, backup, security, and application teams to ensure end-to-end performance and reliability of hosted workloads.
- Establish and maintain operational standards for provisioning, patching, upgrades, host remediation, and configuration consistency across the compute estate.
- Serve as a senior escalation point for complex server hardware and virtualization incidents, leading root-cause analysis and restoration efforts for critical platform issues.
- Drive proactive management of system health, resource utilization, hardware events, performance bottlenecks, and operational risks across the compute environment.
- Lead capacity planning and performance optimization efforts for compute, memory, clustering, and virtualization resources to support current and future business demand.
- Oversee platform resilience through design and support of high-availability, fault-tolerant, and disaster recovery aligned infrastructure services.
- Ensure enterprise operational readiness for maintenance events, lifecycle transitions, incident response, and business continuity requirements.
- Architect and maintain automated workflows using tools such as Ansible, Terraform, PowerCLI, and scripting languages such as PowerShell or Python for provisioning, configuration management, patching, and compliance activities.
- Build and enhance automation for hypervisor host deployment, cluster configuration, lifecycle management, and policy enforcement.
- Codify operational procedures and infrastructure standards into repeatable, auditable automation workflows to reduce manual effort and improve consistency.
- Support infrastructure change through automated validation, testing, and deployment processes that improve quality and reduce risk.
- Lead engineering and administration of virtualization platforms, including core hypervisor services, cluster design, host profiles, virtual networking coordination, and integration with enterprise storage and backup platforms.
- Maintain deep expertise in Vmware, as well as other hypervisors and technologies that make up the enterprise compute portfolio.
- Maintain familiarity with adjacent technologies supporting the virtual infrastructure ecosystem, including hyperconverged platforms, disaster recovery tooling, monitoring systems, VMware Aria Operations, container platforms, and hybrid cloud extensions where applicable.
- Maintain familiarity with public cloud compute services and how enterprise workloads, recovery strategies, and management practices may extend into Azure, AWS, or similar environments as part of a broader infrastructure portfolio.
- Collaborate with platform, cloud, and application teams to support evolving infrastructure patterns, including integration with private cloud and container-hosting platforms where virtual infrastructure is foundational.
- Provide technical leadership on modernization opportunities that improve efficiency, scalability, recoverability, and operational simplicity across the enterprise compute platform.
- Maintain and enhance platform visibility through enterprise monitoring, alerting, and performance analytics tools, including VMware Aria Operations, Grafana, and related ecosystem tooling, to support rapid issue detection and response.
- Establish dashboards, alert thresholds, and operational reporting for server hardware health, virtualization performance, resource consumption, capacity trends, and availability.
- Use telemetry, trend analysis, and platform insights to inform capacity decisions, lifecycle planning, and service improvement initiatives.
- Partner with enterprise monitoring and operations teams to improve actionable insight across the compute and virtualization landscape.
- Serve as a subject matter expert and technical leader for enterprise server hardware and virtualization technologies.
- Mentor junior engineers and help establish best practices for compute operations, virtualization engineering, automation, lifecycle management, and operational monitoring.
- Collaborate cross-functionally with infrastructure, security, architecture, and application stakeholders to align platform capabilities with business priorities.
- Contribute to strategic planning, roadmaps, standards development, and investment recommendations for enterprise compute and virtualization services.
Requirements
- Five (5) or more years of experience in the field or in a related area.
- Monitoring, troubleshooting, customer service, problem solving, cross team collaboration, risk analysis, analytical, operating systems, hardware, infrastructure design, scripting.
- Strong communication, time management, problem solving, teamwork, leadership, mentoring, project management, business acumen, requirements gathering, planning.
Qualifications
- Experience: 7+ years administering and architecting enterprise server infrastructure and virtualization environments at scale.
- Technical Expertise: Deep understanding of Intel-based enterprise server hardware, VMware vSphere, ESXi, vCenter, VMware Aria Operations, and comparable virtualization platforms, including clustering, virtualization performance tuning, and infrastructure lifecycle management.
- Platform Operations: Strong experience managing large virtualized environments supporting mission-critical enterprise applications in a highly available and regulated setting.
- Automation: Proficiency with Ansible, Terraform, PowerCLI, and scripting languages such as PowerShell or Python for infrastructure automation and operational efficiency.
- Hardware Lifecycle Management: Experience with firmware baselines, hardware compatibility, server provisioning, vendor interoperability, and compute platform refresh strategy.
- Ecosystem Knowledge: Strong understanding of integration points across compute, storage, networking, backup, disaster recovery, identity, monitoring, and container platforms.
- Container Familiarity: Familiarity with container management platforms and how virtual infrastructure supports solutions such as OpenShift, Kubernetes, or similar enterprise container ecosystems.
- Public Cloud Familiarity: Working familiarity with public cloud compute services and adjacent infrastructure patterns in Azure, AWS, or similar environments, with understanding of how they complement enterprise data center operations.
- Monitoring and Reliability: Experience with enterprise monitoring, alerting, and operational visibility platforms used to manage compute and virtualization health and performance, including VMware Aria Operations or similar platforms.
- Leadership: Demonstrated ability to lead complex infrastructure initiatives, mentor technical staff, and drive cross-functional operational improvements.
- Soft Skills: Strong analytical, communication, and problem-solving skills with a focus on platform stability, scalability, and continuous improvement.
Skills
- Strong communication and collaboration skills.
- Experience with enterprise server hardware and virtualization technologies.
- Proficiency with automation tools such as Ansible, Terraform, and PowerCLI.
- Experience with monitoring and operational visibility platforms such as VMware Aria Operations.
- Strong analytical and problem-solving skills.
Benefits
The salary range for this position is $110k-$130k depending on experience, with a 7 1/2 annual bonus target based on company performance.
Pay
The salary range for this position is $110k-$130k depending on experience, with a 7 1/2 annual bonus target based on company performance.
Schedule
Hours: Monday-Friday 8AM-4PM with on call. Must be able to work at the office 3 days a week and travel 10-15% of the time to local data centers.