Jobs · Information Technology · California

Staff, Backend Engineer - Catalog

DataHub · Palo Alto, CA · 3 wk ago
Information Technology$225k–$300k/yrFull-time

About the Role

We're looking for an exceptional Staff Backend Engineer to lead development of DataHub's Platform framework – the core that connects diverse data systems and powers our metadata collection capabilities.

  • Build scalable, fault-tolerant ingestion systems for enterprise-scale metadata
  • Create clean, intuitive APIs for our connector ecosystem
  • Develop event-driven architectures for real-time metadata processing
  • Implement schema mapping between diverse systems and DataHub's unified model
  • Design versioning systems for AI assets (training data, model weights, embeddings)

Responsibilities

  • Build and maintain production-grade distributed systems
  • Design and scale distributed systems handling live traffic at scale (100+ QPS)
  • Set up monitoring and alerting for services
  • Design indexing, storage, and data architectures to make large-scale data accessible to online services
  • Develop in a tight loop with LLMs and apply best practices for scalable LLM development

Requirements

  • 8+ years building production-grade distributed systems
  • Advanced Python and API design expertise
  • Experience with high-scale data processing or integration frameworks
  • Strong systems knowledge and distributed architecture experience
  • Proven track record solving complex technical challenges
  • Hands-on experience developing services that serve live traffic at scale

Qualifications

Languages:

  • One of Java/Scala/Kotlin/C#/Go - very strong nice-to-have / borderline must-have
  • Python/TypeScript/Node.js - nice-to-have

Technical Skills:

  • AWS
  • Kubernetes/Docker
  • CI/CD deployment pipelines
  • Microservice Architecture

Bonus Points

  • Experience with DataHub or similar metadata/ETL frameworks (Airflow, Airbyte, dbt)
  • Open-source contributions
  • Experience building and maintaining services that make calls to LLMs to serve live traffic
  • Experience fine-tuning LLM-powered applications exposed to end users
  • Early-stage startup experience

About DataHub

DataHub is an AI & Data Context Platform adopted by over 3,000 enterprises, including Apple, CVS Health, Netflix, and Visa. Innovated jointly with a thriving open-source community of 13,000+ members, DataHub's metadata graph provides in-depth context of AI and data assets with best-in-class scalability and extensibility. The company's enterprise SaaS offering, DataHub Cloud, delivers a fully managed solution with AI-powered discovery, observability, and governance capabilities.

Organizations rely on DataHub solutions to accelerate time-to-value from their data investments, ensure AI system reliability, and implement unified governance, enabling AI & data to work together and bring order to data chaos.

The Challenge

As AI and data products become business-critical, enterprises face a metadata crisis:

  • No unified way to track the complex data supply chain feeding AI systems
  • Engineering teams struggling with data discovery, lineage, and governance
  • Organizations needing machine-scale metadata management, not just human-browsable catalogs

This is where infrastructure meets impact. The metadata layer you'll build will directly power the next generation of AI systems at massive scale. Your code will determine how safely and effectively thousands of organizations deploy AI, affecting millions of users worldwide.

Why Join Us

DataHub is at a rare inflection point: we've achieved product-market fit, earned the trust of leading enterprises, and secured backing from top-tier investors like Bessemer Venture Partners and 8VC. The context platform market is expected to grow from $1B to $9B in the next five years—and we're leading the way. By joining our team, you'll:

  • Tackle high-impact challenges at the heart of enterprise AI infrastructure
  • Ship production systems that power real-world use cases at global scale
  • Collaborate with a high-caliber team of builders who've scaled some of the most influential data tools in the world
  • Build the next generation of AI-native data systems, including conversational agents, intelligent classification, automated governance, and more

Location and Compensation

Bay Area (hybrid, 3 days in Palo Alto office)

Salary Range: $225,000 to $300,000

Benefits and Perks

  • Competitive Compensation: We offer salaries that reflect your skills, experience, and the impact you make.
  • Equity for everyone: Every team member receives an ownership stake in the company.
  • Remote Work: All roles are remote unless otherwise specified. Hybrid roles require 3 days in the Palo Alto office.
  • Location flexibility: Receive a monthly coworking stipend to support your ideal setup.
  • Comprehensive health coverage: We cover 99% of medical, dental, and vision premiums for employees, and 65% for dependents.
  • Flexible savings accounts: FSAs to help cover planned or unexpected healthcare costs, including Dependent Care FSAs.
  • Support for every path to parenthood: Through Carrot Fertility, we provide inclusive fertility benefits and family-forming support for all U.S. employees.
  • Time off that works for you: Unlimited PTO and sick leave policy designed for flexibility, rest, and real life.

Similar jobs

Staff Backend Engineer

AtomsSan Francisco, CA· 2 mo ago
Engineering$224k–$284k/yrapply on job-boards.greenhouse.io