Jobs · Engineering · California

SRE Platform Software Engineer (Early Career / Temporary)

Bitdeer (NASDAQ: BTDR) · San Jose, CA · 1 wk ago
EngineeringFull-time

About the Company

Bitdeer is a world-leading technology company specializing in AI and Bitcoin mining infrastructure. The company provides comprehensive Bitcoin mining solutions, including equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities for high-demand artificial intelligence applications. Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

About the Role

Join the team building and operating NeoCloud’s SRE platform—the multi-region substrate that observes, protects, and operates a global GPU rental fleet across self-built and OEM-rented data centers. As an Early Career Platform Software Engineer, you will work alongside senior engineers to turn architect-approved designs into production-ready code. This is a "build + run" role where you will not only write code but also help operate critical services that other squads, cloud teams, and tenants depend on, participating in a mentored on-call rotation as you grow.

Responsibilities

  • Build & Maintain SRE Microservices: Collaborate with senior mentors to write, test, and deploy features across core platform components (e.g., collection agents, telemetry pipelines, alert engines, or cluster health services).
  • GitOps & Automation: Deliver infrastructure and application updates using modern GitOps practices, declarative configuration, and automated CI/CD pipelines.
  • Observability & Health: Help track, analyze, and optimize system metrics, logs, and traces to ensure high availability across our GPU infrastructure.
  • Operational Readiness: Learn and participate in the on-call rotation for services built by your squad, writing clear runbooks and incident post-mortems.
  • Testing & Quality: Write rigorous unit, integration, and end-to-end tests to ensure platform resilience before shipping to production.

Requirements

  • Bachelor’s degree in Computer Science, Computer Engineering, or a related technical field (or equivalent practical experience / internships).
  • 0–2 years of hands-on software development experience.
  • Proficiency in Go (preferred), Java, or Rust, along with solid scripting abilities in Python or Bash.
  • Strong foundational knowledge of data structures, algorithms, object-oriented design, and distributed systems concepts (e.g., APIs, concurrency, networking basics).
  • Hands-on exposure to Docker, Kubernetes, and Linux fundamentals through coursework, personal projects, open-source contributions, or internships.
  • A mindset focused on quality—experience writing unit and integration tests for your own code.
  • Strong technical writing skills for documenting design choices, runbooks, and clear Pull Request descriptions.

Nice-to-Haves

  • Prior internship or project work involving Kubernetes Operators, Helm, or GitOps tools (ArgoCD / Flux).
  • Exposure to time-series databases or observability tools (Prometheus, OpenTelemetry, Grafana, Loki).
  • Basic familiarity with hardware, GPU/AI infrastructure (NVIDIA DCGM, CUDA), or high-performance computing concepts.
  • Familiarity with infrastructure-as-code tools like Terraform or Ansible.

Similar jobs