Jobs · Information Technology · Texas

K8 Site Reliability SME

Bitdeer (NASDAQ: BTDR) · Austin, TX · 2 wk ago
Information TechnologyFull-time

About the Role

You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role, you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

Responsibilities

  • Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
  • Feed the AIOps substrate: make the control plane safe for automated action by ensuring your CRDs serve as the schema for the platform's predictors and remediators.
  • Turn every human intervention into an autonomous workflow for the next quarter.

What Success Looks Like

  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.

Requirements

  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S.
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S.
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees.
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for operator/CRD development.
  • AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or can design one.
  • Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.

Similar jobs

Reliability SME

ANDRITZGeorgia, United States· 3 wk ago
OTHRapply on andritz-ag.contactrh.com