Head of Infrastructure Operations (US)
Nscale · Charlotte, NC · 3 wk ago
RemoteRemoteInformation Technology$180k–$277k/yrFull-time
About the role
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. The Head of Infrastructure Operations will lead the end-to-end operational management of Nscale's data center portfolio across a defined region (EMEA, APAC, Americas).
Responsibilities
- Own the strategic vision and execution of data center infrastructure operations across the region, ensuring alignment with Nscale's business objectives and growth plans.
- Establish and maintain operational standards, processes, and procedures that drive efficiency, safety, and reliability across all sites.
- Lead the development and implementation of operational roadmaps that support capacity planning, infrastructure scaling, and service delivery milestones.
- Drive continuous improvement initiatives to optimize costs, reduce downtime, and enhance operational maturity.
- Build, mentor, and lead high-performing teams across multiple data center sites, specifically operations staff.
- Establish clear accountability structures, performance metrics, and development pathways for direct reports and broader teams.
- Foster a culture of ownership, safety, and excellence where team members are empowered to make decisions and drive impact.
- Conduct regular performance reviews, provide constructive feedback, and support career progression.
- Oversee Datacenter Leads in their execution of day-to-day Infrastructure Operational procedures, from routine inspections to the handling of ITSM tickets ensuring all SLAs are met.
- Support the Datacenter provider (Nscale or Colo) to ensure optimum performance of the facility, including physical infrastructure, power distribution, cooling systems, security, and environmental controls.
- Maintain accurate asset inventory for all AI Infrastructure and supporting hardware and tooling.
- Support the physical security programme, maintaining audit trails, incident documentation and physical security protocols across all sites.
- Coordinate with the wider Nscale teams to ensure infrastructure layouts, rack elevations, and reference architectures are implemented correctly and optimised for efficient operations.
- Establish SLOs/SLIs for data center availability, performance, and incident response.
- Lead incident response and root-cause analysis for operational failures; own remediation and prevention strategies.
- Ensure full compliance with health and safety regulations, environmental standards, and industry best practices.
- Support ongoing certifications and audits (ISO 27001, ISO 22237, SOC 2, Cyber Essentials Plus, ISO 22301).
- Maintain comprehensive documentation for compliance, audit readiness, and regulatory requirements.
- Manage relationships with critical vendors, contractors, and service providers.
- Oversee vendor performance, SLAs, and contract compliance; escalate issues and drive resolution.
- Conduct procurement activities for equipment, services, and maintenance contracts with cost and quality discipline.
- Coordinate with the Supply Chain team to ensure smooth hardware deployment and logistics flow.
- Partner closely with Infrastructure Engineering, Network Engineering, and Security teams to ensure operational readiness and alignment.
- Work with the Deployment Supply Chain team to support hardware intake, staging, and deployment timelines.
- Collaborate with Finance and Commercial teams on capacity planning, cost optimization, and customer commitments.
- Support project delivery teams in commissioning new sites and scaling existing facilities.
- Engage with senior leadership on operational metrics, risk management, and strategic initiatives.
- Establish KPIs and KRIs for operational health (uptime, energy efficiency, cost per rack, incident rates, etc.).
- Implement monitoring and alerting systems to track infrastructure performance and environmental conditions.
- Produce regular operational reports for senior leadership, including performance metrics, risks, and improvement initiatives.
- Use data-driven insights to identify optimization opportunities and inform decision-making.
Qualifications
- 10+ years of experience in data center operations, infrastructure management, or facilities management at scale.
- Proven track record leading regional or multi-site operations in a high-growth, fast-paced environment.
- Experience managing teams across multiple locations and coordinating complex operational initiatives.
- Demonstrated success in scaling operations, improving efficiency, and maintaining high reliability standards.
- Background in hyperscale, cloud, or HPC data center environments (preferred).
- Deep understanding of data center infrastructure, including power systems, cooling, networking, and security.
- Familiarity with ISO 22237 (data center design and operations) and ISO 27001 Annex A.11 (physical security).
- Knowledge of monitoring systems, environmental controls, and infrastructure automation.
- Understanding of GPU/HPC infrastructure and the unique operational requirements of AI cloud platforms.
- Familiarity with compliance frameworks (SOC 2, ISO 27001, Cyber Essentials Plus, ISO 22301).
Skills
- Exceptional leadership capability with the ability to inspire, develop, and hold teams accountable.
- Strong stakeholder management skills; comfortable influencing senior leaders and cross-functional partners.
- Excellent communication and presentation skills; able to translate complex operational concepts for diverse audiences.
- Problem-solving mindset with the ability to operate in ambiguous, fast-moving environments.
- Bias toward ownership, pragmatism, and delivering results with urgency.
- Disciplined, organized, and methodical approach to operational management and compliance.
- Proven ability to establish processes, standards, and controls that scale with business growth.
- Strong attention to detail and commitment to accuracy in documentation and reporting.
- Proactive approach to risk management, safety, and continuous improvement.