Principal Kubernetes Control Plane Engineer
About the role
We are seeking a Principal Kubernetes Control Plane Engineer to architect the foundational control plane for our AI-native NeoCloud platform. You will be responsible for building highly available, automated systems to manage the lifecycle of thousands of Kubernetes control planes, ensuring zero-downtime upgrades and seamless multi-cluster federation. This role is central to our mission of providing a "NeoCloud" experience, requiring a strategic mindset to balance complex distributed systems design with robust, production-grade engineering. You will lead the development of the "chassis" that supports our high-performance AI workloads, ensuring our infrastructure is scalable, resilient, and ready for enterprise-grade demands.
Responsibilities
- Design and implement scalable control plane architectures using tools like Cluster API to enable fleet-wide automation.
- Develop custom Kubernetes Operators and Custom Resource Definitions (CRDs) to streamline lifecycle management and orchestration.
- Optimize etcd performance and ensure rock-solid state management across distributed cloud regions.
- Implement secure, multi-tenant isolation mechanisms at the API server layer to support enterprise-grade security and tenancy.
- Architect solutions for zero-downtime upgrades and multi-cluster federation to provide a seamless user experience for our AI customers.
- Collaborate with AI Scheduling and Fabric engineering teams to ensure the control plane integrates tightly with GPU orchestration layers.
- Lead design reviews and mentor engineers to uphold high standards of distributed systems architecture across the team.
- Drive platform reliability by identifying and mitigating failure modes in the control plane chassis before they impact customer training or inference jobs.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or related field.
- 8+ years of software engineering experience, with deep, hands-on expertise in Go.
- Extensive experience contributing to or extending the Kubernetes ecosystem (e.g., Operators, API Server, etcd, Cluster API).
- Proven track record of operating, debugging, and scaling large-scale distributed systems in production environments.
- Deep understanding of multi-tenancy models, cluster federation, and API governance in cloud-native environments.
- Familiarity with infrastructure automation (e.g., Terraform, CI/CD pipelines) and managed cloud services.
- Strong technical leadership skills; ability to influence architectural decisions and align cross-functional teams.
- Excellent communication skills, with the ability to translate complex system requirements into manageable engineering milestones.
- Experience working in high-velocity, high-growth engineering environments is highly preferred.
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure, providing comprehensive Bitcoin mining solutions and building AI computational infrastructure. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.