Manager, Infrastructure
About the role
RZR runs its own metal — four owned-and-operated data centers across Santa Clara, Ashburn, Amsterdam, and Hong Kong, housing approximately 1,300 servers and 1.15MW of capacity, placed next to the major ad exchanges, serving 5–6M+ bid requests per second at ~20ms. This infrastructure is our competitive moat, not a cost center. As Manager, Infrastructure, you will own day-to-day and quarter-to-quarter operation of that entire footprint: the team, hardware lifecycle, capacity planning, incident response, and vendor relationships. You will take over these functions directly from the Head of Cybersecurity & Infrastructure, freeing him to focus on security and multi-entity IT. This is a hands-on manager role. There is no "scale up" button here — latency, packets-per-second, and procurement lead times are the job. The right person combines deep bare-metal operational instincts with the leadership presence to run a distributed, experienced global team from day one.
Responsibilities
- Team Leadership & Operations
- Manage and develop the InfraOps team across US and APAC time zones, including regional DC owners (SV/VA and NL/HK) and network engineering
- Own the weekly DevOps check-in cadence, alert reviews, and 24/7 on-call coverage model
- Drive P1/P2 incident response end to end — accountability for MTTR reduction, runbook coverage, and alert hygiene
- Capacity Planning & Hardware Lifecycle
- Own capacity planning and hardware lifecycle across all four data centers: Dell and Supermicro procurement through VARs, GPU expansion for on-prem ML training and inference, colo power and space management, and remote-hands logistics with Equinix and Digital Realty
- Run the annual cloud-vs-colo evaluation alongside leadership, with full ownership of the recommendation
- Technical Platform Oversight
- Oversee the core infrastructure stack: Ubuntu/systemd fleet, FreeIPA, Ansible/Salt configuration management, MAAS provisioning, and Zabbix monitoring
- Manage the spine-leaf Mellanox/NVIDIA network via Netris, including 100G Google peering and transit blend (Lumen/Cogent/Zayo)
- Support the stateful data tier — Aerospike, Kafka, ClickHouse, Hadoop/HDFS — across capacity limits, evictions, migrations, and low-latency tuning
- Vendor & Budget Management
- Own colo and vendor relationships and budgets: Equinix and Digital Realty invoices, transit contracts, VAR procurement, and Netris licensing
- Partner with the security/compliance function on SOC 2 Type 2 evidence, infrastructure hardening, and access reviews
- Operate comfortably within a multi-entity environment (RZR/Skillz/Firy/Beamable shared IT) with comfort in M&A-flavored ambiguity
Requirements
- Must-Have
- 6–8 years in infrastructure or data center operations with 2+ years managing engineers — this role takes over a functioning global team on day one
- Bare-metal and colo depth: capacity planning, hardware procurement (Dell/Supermicro), IBX/remote-hands workflows, and the physical logistics of running owned cages
- Network fundamentals at scale: spine-leaf architecture, BGP/peering (100G-class), transit blends, and low-latency tuning (NIC/IRQ, packets-per-second thinking)
- Deep Linux operations: systemd, netplan, FreeIPA/Chrony, Ansible and/or Salt, Zabbix, running fleets of hundreds-plus servers
- Experience operating large stateful distributed systems — Aerospike, Cassandra, Scylla, Kafka, or ClickHouse — under sub-50ms latency budgets and hard capacity limits
- Demonstrated P1/P2 incident ownership: on-call program management, postmortems, and alert hygiene discipline
- Nice-to-Have
- Hybrid cloud experience alongside owned metal: AWS (IAM, Route 53, GuardDuty, S3); Kubernetes exposure a plus
- SOC 2 or compliance evidence experience; Okta and Vanta familiarity; security-minded infrastructure approach
- Experience in adtech, RTB, or other high-QPS, latency-sensitive environments
- Netris or other SDN controller experience; MAAS provisioning familiarity
Schedule
This is a remote role open to candidates based in the United States. The role requires periodic travel to our data center sites (Santa Clara, CA; Ashburn, VA; Amsterdam; Hong Kong) and SF HQ for operational reviews, team time, and site work.
Benefits
- Own a genuinely rare infrastructure environment — four global owned-and-operated data centers, 5–6M+ QPS real-time bidding at ~20ms. This is infrastructure that is the company's competitive moat, running 4–10x faster bid response than cloud DSPs. You will not find this kind of physical infrastructure problem at most companies.
- Full-stack ownership with no cloud-bill anxiety — real hardware decisions, GPU expansion for on-prem ML, and an annual cloud-vs-colo evaluation you will help drive. Zero marginal experimentation cost.
- Inherit a functioning, experienced global team — regional DC owners with clear ownership and an established ops cadence. Build on it, not rescue it.
- Direct line to leadership — reporting to the Head of Cybersecurity & Infrastructure with visibility to the SVP of Engineering. Your goals map straight to company OKRs.
- Company momentum — RZR rebranded in March 2026 and is expanding aggressively across mobile, CTV, and influencer. Your infrastructure is what makes all of it possible.