Database/Distributed Storage Engineer – AI Infra (Mandarin Required)
Location: Palo Alto, California
Fluent Mandarin Chinese and professional English are mandatory. Candidates must be able to conduct technical discussions and collaborate effectively in Mandarin with engineering teams in China, as well as operate professionally in English.
About The Role
Our client is a rapidly scaling global consumer internet platform serving hundreds of millions of users worldwide. This is not a traditional DBA, DevOps, or database operations role. The position focuses on architecting and building the next generation of large-scale storage and database infrastructure, with a strong emphasis on applying LLMs and AI Agents to database operations and platform engineering. We are seeking engineers who have built or developed core database, distributed storage, KV/cache, or storage platform systems, and who are interested in evolving these systems toward AI-native, self-diagnosing, self-healing, and self-optimizing infrastructure. The role combines deep systems engineering with emerging AI infrastructure work, including distributed storage, databases, backend engineering, and LLM Agents/AIOps.
Responsibilities
- Architect and develop a large-scale storage and database cloud platform.
- Build and own core storage/database platform modules supporting global-scale workloads.
- Drive architectural evolution toward AI-native storage infrastructure.
- Build self-service platforms for database and storage management.
- Develop LLM-powered infrastructure Agents enabling natural-language interaction with database systems.
- Build autonomous infrastructure capable of self-diagnosis, self-healing, and self-optimization.
- Apply AI/LLMs to:
- Intelligent SQL optimization
- Anomaly detection
- Root Cause Analysis (RCA)
- Capacity forecasting
- Automated scaling
- Alert noise reduction
- Integrate emerging AI development and Agent paradigms into database/storage platform workflows.
- Improve the developer and operator experience of managing large-scale database infrastructure.
Note: These AI/AIOps responsibilities are central to the role, not just an additional layer.
Requirements
Database / Distributed Storage – Core Requirement
- 3 - 12+ years of hands-on engineering experience with distributed storage and/or database systems.
- Strong understanding of the architecture and implementation of at least two of the following:
- Distributed KV stores / caching systems
- MySQL / relational database systems
- Distributed databases
- Object storage
- Table / wide-column storage
- Graph databases
- Experience building, developing, or significantly extending database/storage infrastructure (not just administering or operating databases).
- Strong understanding of distributed systems fundamentals, scalability, reliability, and performance.
- Strong backend engineering skills in Go, Java, or Python.
- Distributed storage/database experience and knowledge across at least two storage categories is explicitly required.
AI / LLM Engineering
- Experience applying modern AI capabilities to infrastructure or developer platforms. Relevant experience includes:
- Building applications using LLM APIs
- Retrieval-Augmented Generation (RAG)
- Prompt Engineering
- Function Calling / Tool Use
- LLM Agents
- Multi-Agent systems
- LangChain / LlamaIndex / AutoGen or similar frameworks
- AIOps / intelligent operations
- AI-driven infrastructure automation
- Experience shipping an AI/LLM application or capability into production is particularly relevant.
Particularly Relevant Backgrounds
- Engineers who have worked on core database or distributed storage infrastructure at large-scale technology companies, including:
- Distributed Databases
- Database Engines / Database Platforms
- Distributed KV Systems
- Storage Engines
- Cloud Database Infrastructure
- Object / Table Storage
- Caching Infrastructure
- Database Control Plane / Management Platforms
- Database Reliability & Automation
- AI for Databases / AIOps
- Candidates who combine deep storage/database systems expertise with hands-on AI/LLM development are especially relevant.
Nice To Have
- AIOps / intelligent database operations experience
- LLM Agent development
- Infrastructure automation
- Open-source database/storage contributions
- Technical blogs or publications
- Experience turning AI capabilities into production infrastructure products
Why This Role
This is an opportunity to work on infrastructure supporting hundreds of millions of users globally, operating at the intersection of two technically challenging areas: Distributed Storage & Databases × AI/LLM Infrastructure. The longer-term goal is to rethink how database infrastructure can become self-diagnosing, self-healing, self-optimizing, and increasingly autonomous, rather than simply adding an AI interface to an existing platform. You will work alongside experienced AI and infrastructure engineers, helping to shape the next generation of intelligent database platforms.