SRE L1 Support/Cloud Platform Ops Engineer
About the Company
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. The company provides comprehensive Bitcoin mining solutions, including equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities for high-demand artificial intelligence applications. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
About the Role
You are the first human in the loop — the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations. NeoCloud is building an AI-operated GPU cloud. This role focuses on front-line monitoring and incident response for NeoCloud's US GPU data centers during the 8AM–8PM PST shift. You execute standard operating procedures (SOPs), escalate complex cases, and provide ground truth to train the AIOps platform.
Responsibilities
- Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
- Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
- Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
- Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
- Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
- Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
- Perform physical data center tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
- Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
- Maintain and update operational runbooks based on recurring issues.
- Assist with hardware deployment, firmware updates, and inventory management under SME guidance.
- Feed the AIOps substrate: tag and describe novel incidents to train the platform, ensuring runbooks evolve toward automation.
- Provide structured handoff notes to improve platform learning.
Requirements
- 2+ years in NOC, data center operations, or IT support role.
- Basic Linux system administration (command line, log analysis, service management).
- Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent).
- Experience with ticketing systems (ServiceNow, Jira Service Management).
- Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement.
- Strong communication skills for shift handoffs, incident documentation, and escalation.
- Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation).
- Curiosity about automation — you notice patterns and advocate for improvements.
- Comfort with structured data — understanding that ticket documentation trains the platform.
Schedule
8AM–8PM PST shift schedule (12-hour shifts with rotation).