Research Scientist-Model Efficiency
About the Company
Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud. Bitdeer provides comprehensive Bitcoin mining solutions, including designing industry-leading ASIC chips, manufacturing mining rigs, and managing complex computing processes across the value chain—equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities for artificial intelligence applications.
Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.
About Bitdeer AI Lab
Bitdeer AI Lab is a frontier AI lab under Bitdeer, focused on exploring the frontiers of artificial intelligence with a long-term vision. Our mission is to turn energy into intelligence that people can afford to use. We work on inference economics from the ground up, from power and datacenters to the software that turns them into tokens, ensuring that the cost of serving models drives what gets built.
Responsibilities
This role focuses on making models cheaper and faster to serve without sacrificing meaningful quality. Techniques in scope include:
- Quantization, sparsity and pruning
- Speculative decoding and MTP
- Serving-time attention and KV-cache methods
You will implement and adapt published methods on our models and hardware, develop custom optimizations where existing approaches fall short, and build rigorous evaluation frameworks to validate claims like “lossless at 2× throughput.”
Qualifications
- Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, or a related field
- Hands-on experience in LLM inference, model optimization, or ML systems
- Strong programming ability in Python and deep familiarity with PyTorch; experience with C++, CUDA, or Triton is a plus
- Implementation-level expertise in at least one area of model efficiency (e.g., quantization, sparsity, speculative decoding, or serving-time attention)
- Strong understanding of transformer internals and how accuracy loss manifests in model behavior
- Rigorous evaluation practice—task-level metrics, controlled comparisons, and honest baselines
- Experience taking efficiency methods into production serving or equivalent research depth
- Familiarity with inference engines such as vLLM, SGLang, or TensorRT-LLM and their efficiency features
- Publications at top-tier systems or ML venues, or substantial open-source contributions
- Deep enthusiasm for cutting-edge AI infrastructure and efficient inference, with a strong ownership mentality and solid engineering discipline
Benefits
- A culture that values authenticity and diversity of thought
- An inclusive and respectable environment with open workspaces and a startup spirit
- Opportunities to network with industrial pioneers and enthusiasts in a fast-growing company
- Direct impact on the future of the digital asset industry
- Involvement in new projects and process/system development
- Personal accountability, autonomy, fast growth, and learning opportunities
- Attractive welfare benefits and developmental opportunities, including training and mentoring