GPU DC East-West Network SRE Expert (SME)
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer provides comprehensive Bitcoin mining solutions and builds AI computational infrastructure to support the AI revolution. The company handles complex computing processes including equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities for high-demand artificial intelligence workloads. Headquartered in Singapore, Bitdeer has deployed data centers across the United States, Norway, Bhutan, and Ethiopia.
About The Role
You will maintain the high-performance fabric that unifies 10,000 GPUs into a single system, transforming InfiniBand and RoCE telemetry into actionable insights for congestion and link-failure prediction. Bitdeer is building an AI-operated GPU cloud where East-West bandwidth directly impacts training job success—preventing costly interruptions to multi-million-dollar training runs.
Responsibilities
- Operate and optimize InfiniBand fabrics (fat-tree, rail-optimized, and dragonfly topologies) for GPU clusters ranging from 100 to 10,000 GPUs.
- Manage RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
- Utilize UFM (Unified Fabric Manager) for InfiniBand fabric monitoring, diagnostics, and subnet management.
- Monitor and tune InfiniBand and RoCE performance, including adaptive routing, congestion control (DCQCN/ECN), and traffic isolation.
- Optimize NCCL communication through topology detection, ring/tree algorithm selection, and GDR configuration.
- Oversee firmware lifecycle for InfiniBand switches and HCAs.
- Diagnose faults such as link flaps, symbol errors, packet drops, routing anomalies, and credit stalls.
- Coordinate with Nvidia/Mellanox support for escalations, bug resolution, and RMAs.
- Feed InfiniBand and RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform’s collection pipeline.
- Collaborate with the platform team to define Link and Straggler predictors, identifying patterns in counters that indicate impending failures (e.g., "bad optic 30 minutes from failure").
- Convert incidents into labeled examples for the fault-prediction engine and routine mitigations into automated workflows for remediation.
Requirements
- 5+ years of experience in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics.
- Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale.
- Strong understanding of InfiniBand subnet management, partitioning, and QoS.
- Experience with RoCEv2 deployment, including PFC, ECN, and DCQCN configuration.
- Proficiency with UFM or equivalent InfiniBand fabric management tools.
- Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices.
- Experience diagnosing InfiniBand/RoCE network issues using tools like ibdiagnet, perfquery, and ibstat.
- Understanding of NCCL and how GPU communication maps to network topology.
- Telemetry-driven operations mindset—experience building dashboards/alerts on RDMA counters at scale or ability to articulate requirements for a fabric-health model.
- Runbook-as-code approach—ability to convert diagnostics into automation for future use.