Jobs · Engineering · California

Senior Database Reliability Engineer

Scribe · San Francisco, CA · 1 wk ago
On-siteEngineering$145k–$230k/yrFull-time

About the role

We're hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of Scribe's data tier. Our engineering org is doubling — which means the guardrails, automation, and standards you put in place today will carry a much larger team through the next phase of growth. This is a senior IC role with real ownership: you'll set the bar for how engineers across the company interact with our databases, not just keep the lights on.

What You'll Do

  • Own database reliability across Aurora, OpenSearch, Redis, and our CDC pipeline — including schema design reviews, migration safety (locks, backfills, concurrent index builds, NOT VALID constraints), and incident response for the data tier
  • Make the Django ORM a strength at scale: catch N+1 patterns in review, extend QuerySet conventions and physical schema standards, and build the CI checks and AGENTS.md scaffolding that encode those standards so they scale beyond any single reviewer
  • Operate and evolve the CDC pipeline from Aurora through DMS to S3 Parquet to Snowflake – including replication slot hygiene, schema evolution safety, and automated checks that catch migrations likely to break downstream consumers before they ship
  • Drive multi-AZ resilience within our single-region architecture — Aurora writer/reader placement, failover behavior, RTO/RPO, ElastiCache and OpenSearch AZ topology, RabbitMQ survivability
  • Build self-service tooling and dashboards that give product and platform teams visibility into their own query footprint, reducing the review burden as the engineering org grows
  • Contribute to onboarding and knowledge-sharing as a large incoming class of engineers joins — write docs, run internal sessions on "what your ORM query is really doing," and feed that knowledge back into AI review tooling

What We're Looking For

  • Deep PostgreSQL expertise in practice: read EXPLAIN (ANALYZE, BUFFERS) fluently, understand MVCC, bloat, lock contention, and vacuum behavior, and tune Aurora Serverless V2 for latency and throughput
  • Work with an ORM (Django, SQLAlchemy, ActiveRecord, or similar) at production scale – predict the SQL a query generates, spot N+1 issues on sight, and know when joins beat batched IN queries and when they don't
  • Run CDC pipelines in production, ideally with AWS DMS — comfort with logical replication, slot hygiene, schema evolution, and Parquet-based data lakes feeding Snowflake, BigQuery, or Redshift
  • Hands-on experience with pganalyze (or Datadog DBM / pg_stat_statements pipelines), CloudWatch, and Honeycomb (or another high-cardinality tracing tool); comfortable with OpenTelemetry
  • Write real automation — Python, Go, or similar — and use Terraform or comparable IaC to manage infrastructure
  • Write AI coding and review tools in a team setting: write and maintain AGENTS.md files, configure review agents, iterate on prompts

Nice to Have

  • Event sourcing on Postgres, or experience with alternate CDC tooling (Debezium, Fivetran, Airbyte)
  • pgbouncer or RDS Proxy at scale with Django connection handling
  • Deep Honeycomb usage: SLOs, BubbleUp, Triggers, derived columns
  • Snowflake from the producer side: staging, Snowpipe, external tables on Parquet

Location

San Francisco (hybrid, 3 days per week in-office) or, Remote based permanently in PST (Pacific Standard Time).

Cash

Compensation varies by location. All full-time employees receive equity in Scribe. Final offers depend on experience and scope.

Bonus

At Scribe, we celebrate our differences and are committed to creating a workplace where all employees feel supported and empowered to do their best work. Scribe is proud to be an Equal Opportunity Employer.

Cash Range: $145K - $230K

Similar jobs