Staff MaaS Backend Engineer
About the Company
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure. Bitdeer provides comprehensive Bitcoin mining solutions and builds AI computational infrastructure to support the AI revolution. The company handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities for high-demand artificial intelligence applications. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
About the Role
We are seeking a Staff Backend Engineer to re-architect our Model-as-a-Service (MaaS) system into a commercial, globally distributed, multi-tenant token service. This service must sustain millions of monthly active users, high sustained token throughput per GPU, and invoice-grade accounting. The role involves co-owning the MaaS system design with the principal architect and owning the implementation end-to-end—code, migrations, and service reliability. The mandate is to incrementally evolve the platform with each step being independently shippable, reversible, and non-disruptive to existing tenants.
Responsibilities
- Architecture, Strategy, and Leadership: Co-own the end-to-end MaaS system design, author decision records, and defend technical trade-offs. Drive platform maturity by delivering operable, measurable capabilities. Set Go and API standards, mentor engineers, and align cross-functional teams.
- Inference Gateway and API Surface: Own wire compatibility for major formats (OpenAI, Anthropic), supporting streaming, tool calling, and structured output. Evolve the routing tier for load-aware, model-aware, and prefix-cache-aware endpoint selection with circuit breaking and fallback mechanisms. Manage versioning and deprecation as a published contract.
- Performance Optimization and Model Lifecycle: Optimize token throughput, KV cache tiering, and time-to-first-token (TTFT) latency at p95/p99 levels. Integrate deployment tooling, LoRA multiplexing, and cold-start-aware autoscaling with Kubernetes. Treat regressions in cost-per-million-tokens or latency as critical incidents.
- Global Topology and Reliability (SLOs): Scale the platform to a globally distributed architecture with regional inference pools, capacity-aware failovers, and an active-active control plane. Define, publish, and defend strict Service Level Objectives (SLOs) based on actual system performance. Ensure resilience through load testing, on-call runbooks, and predictable load-shedding.
- Security, Identity, and Multi-Tenant Isolation: Enforce fail-closed authorization, multi-tenant isolation, and zero-retention data paths across network, cache, and storage layers. Manage API keys and OAuth credentials, enforce rate limits and quotas, and design abuse controls for untrusted content.
- Token Metering and Billing Correctness: Build an idempotent, exactly-once metering system for uncached, cached, output, and reasoning tokens. Maintain a high-volume usage ledger with prepaid spend caps and perfect reconciliation with invoicing. Treat metering or billing defects as critical revenue incidents.
- Safe Migrations and Observability: Execute zero-regression upgrades using strangler-style replacements, shadow traffic, and stateful dual-writes. Deliver end-to-end request tracing and cost telemetry. Enforce log hygiene to prevent retention of prompts, completions, or PII outside consented policies.
Qualifications
- 8+ years of backend engineering experience, including 3+ years owning a high-traffic, multi-tenant API platform for paying customers.
- Expertise in scaling distributed systems through multi-region active-active deployments, caching, backpressure, and performance engineering.
- Proven track record of safely executing zero-downtime brownfield migrations for stateful subsystems (e.g., metering, ledgers) without regressions.
- Hands-on proficiency with Go-based services, production Kubernetes (including Envoy and GPU-aware scheduling), and datastores like PostgreSQL, Redis, and Kafka.
- Systems-level understanding of LLM serving, including streaming, KV/prefix caching, and trade-offs between TTFT and throughput.
- Experience building highly reliable, exactly-once metering and billing systems for billions of events under partial failure conditions.
- Strong multi-tenant security disciplines, including fail-closed authorization, verified identities, and cross-tenant isolation.
- Operational maturity for revenue-bearing platforms, including on-call responsibilities, blameless incident reviews, and structural improvements.