K8 Site Reliability SME
Bitdeer (NASDAQ: BTDR) · Austin, TX · 2 wk ago
Information TechnologyFull-time
About the Role
You run the control plane where AIOps meets tenants — where topology-aware scheduling, self-healing, and agent-driven remediation actually execute. NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role, you design, deploy, and operate that control plane — and you make sure the AIOps substrate can reach in and remediate without a human on the pager.
Responsibilities
- Production Kubernetes clusters optimized for GPU workloads at scale (100–10,000 GPUs).
- Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
- Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
- Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
- AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
- Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
- Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
- Terraform providers and modules for infrastructure-as-code across GPU clusters.
- SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
- Incident management: runbook automation, escalation, post-incident reviews.
- Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
- GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
- Feed the AIOps substrate: make the control plane safe for automated action by ensuring your CRDs serve as the schema for the platform's predictors and remediators.
- Turn every human intervention into an autonomous workflow for the next quarter.
What Success Looks Like
- Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
- BMaaS live for external tenants with self-service onboarding.
- Cluster availability and job-completion SLOs published and met.
Requirements
- 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S.
- Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S.
- Experience with topology-aware scheduling and GPU-specific resource management.
- Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees.
- Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
- Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
- Strong SRE background: SLI/SLO frameworks, incident management, capacity planning.
- Experience with Prometheus, Grafana, and alerting at scale.
- Strong programming skills in Go or Python for operator/CRD development.
- AIOps aptitude — you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or can design one.
- Runbook-as-code mindset — every SRE playbook you write should be executable by the platform.