Jobs · Customer Service · California

Observability Backend Engineer - Distributed Systems (Mandarin required)

Applied Intelligence Consulting (Singapore) · San Francisco, CA · 2 wk ago
On-siteCustomer ServiceFull-time

Location: Palo Alto, California. This is a 5 days onsite role.

About The Role

Our client is a rapidly scaling global consumer internet platform serving hundreds of millions of users across content, community, e-commerce, and advertising ecosystems. As the company expands internationally and deepens its investment in AI-driven infrastructure, observability is a critical backbone for platform reliability and performance. We are seeking a backend engineer with deep hands-on experience building observability systems at scale—not simply operating monitoring tools. You will help design and build next-generation observability infrastructure spanning metrics, logging, tracing, and profiling, supporting large-scale distributed systems as well as emerging AI infrastructure and AI-native workloads.

Responsibilities

  • Build and evolve large-scale observability systems across metrics, logging, tracing, and profiling.
  • Design and develop monitoring platforms, distributed tracing systems, logging services, real-time alerting systems, and computation engines for streaming analytics and time-series workloads.
  • Own architecture design and product-level implementation of core observability capabilities.
  • Build systems designed for high throughput, high concurrency, low latency, high availability, and reliability at significant scale.
  • Develop and improve service governance and observability capabilities across large-scale distributed and microservices environments.
  • Work with cloud-native observability technologies such as OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, and eBPF.
  • Drive AI infrastructure observability, AI application observability, and 'Observability + AI' capabilities.
  • Improve incident detection, troubleshooting, root-cause analysis, and overall platform stability through better telemetry and observability infrastructure.

Requirements

  • 2–7 years of relevant backend engineering, infrastructure, or observability experience.
  • Strong backend software engineering fundamentals and proficiency in Java or Go.
  • Deep hands-on experience building observability, telemetry, monitoring, or reliability platforms—ideally at large scale.
  • Strong understanding of distributed systems, concurrent programming, performance optimization, and high-concurrency system design.
  • Practical experience with one or more of OpenTelemetry, Prometheus, VictoriaMetrics, ELK, ClickHouse, SkyWalking, CAT, or eBPF.
  • Good understanding of Kubernetes and cloud-native infrastructure.
  • Strong foundational knowledge of Linux, networking, storage systems, and message queues.
  • Experience designing systems for high throughput, high availability, low latency, and large telemetry/data volumes.
  • Strong coding ability, ownership, and ability to solve complex infrastructure problems end-to-end.
  • Fluent Mandarin Chinese is mandatory for this position. Candidates must be able to communicate effectively in Mandarin with engineering and product teams in China, while also being comfortable working in an English-speaking global engineering environment.

Preferred Qualifications

  • Experience building observability infrastructure for very large-scale consumer internet or distributed systems environments.
  • Experience with AI infrastructure observability or AI application observability.
  • Familiarity with AI-related technologies or ecosystems such as PyTorch, Spring AI, or Langfuse.
  • Experience with cross-region or global infrastructure and international platform environments.
  • Open-source contributions in observability, infrastructure, or distributed systems.

Location & Immigration

This role is based in Palo Alto, California only. H-1B transfers can be supported for eligible US candidates. Green Card sponsorship/support can also be provided.

Why This Role

  • Build foundational observability infrastructure supporting a global consumer platform with hundreds of millions of users.
  • Work at the intersection of distributed systems, cloud-native observability, and AI infrastructure.
  • Solve engineering problems involving massive telemetry volumes, high concurrency, performance, reliability, and global-scale infrastructure.
  • Shape next-generation observability platforms rather than simply consuming existing monitoring tools.
  • Gain direct exposure to AI observability—an emerging and highly specialized infrastructure domain.

Similar jobs