Jobs · Management

Operations and Support Lead

Hydra Host · Miami, FL · 1 mo ago
RemoteRemoteManagementFull-time

About the role

Hydra Host operates mission-critical AI infrastructure where customer success depends on operational excellence. This role combines Customer Support Leadership with Infrastructure Operations Management, serving as the operational hub between customers, engineering, deployment teams, hardware vendors, and AI Factory partners. You'll own customer-facing operational support while building the internal processes that keep our NeoCloud platform running efficiently. This is an operations leadership role responsible for the daily execution, reliability, and continuous improvement of Hydra Host's AI Factory platform.

Responsibilities

  • Lead Customer Support Operations
    • Design and build Hydra Host’s customer support organization from the ground up.
    • Interface with enterprise AI customers, Machine Learning engineering teams, GPU infrastructure customers, and AI Factory operators.
    • Build a high-performing support organization that delivers exceptional customer experiences while maintaining enterprise-grade service levels.
  • Own Operational Excellence
    • Drive the day-to-day operational health of Hydra Host's NeoCloud platform by coordinating activities across engineering, infrastructure, vendors, and customer-facing teams.
    • Ensure infrastructure operates reliably while continuously improving operational efficiency.
  • Lead Major Incident Management
    • Serve as the Incident Commander during production-impacting events.
    • Coordinate engineering, networking, infrastructure, vendors, and customers to resolve GPU cluster failures, network outages, hardware failures, firmware issues, storage performance degradation, infrastructure capacity constraints, and customer-impacting production incidents.
    • Own customer communications throughout incident response while driving rapid resolution and post-incident improvements.
  • Build Scalable Support Operations
    • Design and implement ticketing systems, escalation procedures, knowledge management, support automation, AI-powered support tools, on-call rotations, operational playbooks, and customer communication standards.
    • Create a support organization capable of scaling alongside Hydra Host's rapid growth.
  • Drive AI Factory Operations
    • Partner with deployment, engineering, and data center teams to coordinate infrastructure deployments, rack turn-up, GPU cluster readiness, network activation, capacity planning, maintenance scheduling, and production acceptance.
    • Conduct operational readiness reviews to ensure AI Factory infrastructure is deployed efficiently and operates reliably.
  • Vendor & Partner Management
    • Own operational relationships with hardware manufacturers, GPU vendors, data center operators, construction partners, logistics providers, network carriers, and service providers.
    • Track vendor performance, manage escalations, enforce SLAs, and ensure timely issue resolution.
  • SLA & Service Delivery
    • Establish, monitor, and improve operational KPIs including SLA compliance, MTTR, incident response times, customer satisfaction, infrastructure uptime, vendor performance, capacity utilization, and operational readiness.
    • Provide executive reporting on operational performance and identify opportunities for continuous improvement.
  • Cross-Functional Leadership
    • Collaborate daily with Infrastructure Engineering, Network Engineering, Platform Engineering, Customer Success, Deployment Program Managers, AI Factory Partners, Executive Leadership, and Enterprise Customers.
    • Act as the operational bridge between technical teams and customer-facing organizations.
  • Build the Organization
    • Recruit, mentor, and develop the Operations & Support organization as Hydra Host grows.
    • Establish the culture, processes, and operational standards that define world-class infrastructure operations.

Requirements

  • 5+ years leading customer support, technical operations, infrastructure operations, or service delivery organizations.
  • Experience supporting enterprise infrastructure, cloud platforms, AI infrastructure, NeoCloud providers, or large-scale data center environments.
  • Experience managing production incidents in mission-critical environments.
  • Strong understanding of servers, networking, storage, and enterprise infrastructure.
  • Experience building operational processes that scale rapidly.
  • Experience working with third-party vendors and infrastructure providers.
  • Excellent communication skills with both technical and executive stakeholders.
  • Proven ability to lead cross-functional initiatives across engineering, operations, and customer organizations.
  • Strong organizational and project management skills.
  • Bias toward ownership, accountability, and continuous improvement.

Preferred Qualifications

  • Experience operating NeoCloud or AI Factory infrastructure.
  • Experience supporting NVIDIA GPU environments (HGX, DGX, H100, H200, Blackwell).
  • Experience with bare-metal cloud infrastructure.
  • Familiarity with InfiniBand, RoCE, high-performance networking, and AI storage platforms.
  • Experience with ITIL Service Management, Incident Management, Change Management, and Problem Management.
  • Experience implementing AI-driven support automation and operational tooling.
  • Experience with CRM and service management platforms including Zendesk, Jira Service Management, Linear, HubSpot, Intercom, or similar platforms.
  • Previous experience managing technical support or operations teams.

Similar jobs

Operations Support Lead

Texas Department of TransportationAustin, TX· 4 wk ago
Information Technologyapply on fa009.taleo.net

Operations Lead

24 Hour FitnessLewisville, TX· 1 mo ago
Managementapply on rr.jobsyn.org