Senior Database Reliability Engineer
About the role
We're hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of Scribe's data tier. Our engineering org is doubling — which means the guardrails, automation, and standards you put in place today will carry a much larger team through the next phase of growth. This is a senior IC role with real ownership: you'll set the bar for how engineers across the company interact with our databases, not just keep the lights on.
What You'll Do
- Own database reliability across Aurora, OpenSearch, Redis, and our CDC pipeline — including schema design reviews, migration safety (locks, backfills, concurrent index builds, NOT VALID constraints), and incident response for the data tier
- Make the Django ORM a strength at scale: catch N+1 patterns in review, extend QuerySet conventions and physical schema standards, and build the CI checks and AGENTS.md scaffolding that encode those standards so they scale beyond any single reviewer
- Operate and evolve the CDC pipeline from Aurora through DMS to S3 Parquet to Snowflake – including replication slot hygiene, schema evolution safety, and automated checks that catch migrations likely to break downstream consumers before they ship
- Drive multi-AZ resilience within our single-region architecture — Aurora writer/reader placement, failover behavior, RTO/RPO, ElastiCache and OpenSearch AZ topology, RabbitMQ survivability
- Build self-service tooling and dashboards that give product and platform teams visibility into their own query footprint, reducing the review burden as the engineering org grows
- Contribute to onboarding and knowledge-sharing as a large incoming class of engineers joins — write docs, run internal sessions on "what your ORM query is really doing," and feed that knowledge back into AI review tooling
What We're Looking For
- Deep PostgreSQL expertise in practice: read EXPLAIN (ANALYZE, BUFFERS) fluently, understand MVCC, bloat, lock contention, and vacuum behavior, and tune Aurora Serverless V2 for latency and throughput
- Work with an ORM (Django, SQLAlchemy, ActiveRecord, or similar) at production scale – predict the SQL a query generates, spot N+1 issues on sight, and know when joins beat batched IN queries and when they don't
- Run CDC pipelines in production, ideally with AWS DMS — comfort with logical replication, slot hygiene, schema evolution, and Parquet-based data lakes feeding Snowflake, BigQuery, or Redshift
- Hands-on experience with pganalyze (or Datadog DBM / pg_stat_statements pipelines), CloudWatch, and Honeycomb (or another high-cardinality tracing tool); comfortable with OpenTelemetry
- Write real automation — Python, Go, or similar — and use Terraform or comparable IaC to manage infrastructure
- Write AI coding and review tools in a team setting: write and maintain AGENTS.md files, configure review agents, iterate on prompts
Nice to Have
- Event sourcing on Postgres, or experience with alternate CDC tooling (Debezium, Fivetran, Airbyte)
- pgbouncer or RDS Proxy at scale with Django connection handling
- Deep Honeycomb usage: SLOs, BubbleUp, Triggers, derived columns
- Snowflake from the producer side: staging, Snowpipe, external tables on Parquet
Location
San Francisco (hybrid, 3 days per week in-office) or, Remote based permanently in PST (Pacific Standard Time).
Cash
Compensation varies by location. All full-time employees receive equity in Scribe. Final offers depend on experience and scope.
Bonus
At Scribe, we celebrate our differences and are committed to creating a workplace where all employees feel supported and empowered to do their best work. Scribe is proud to be an Equal Opportunity Employer.
Cash Range: $145K - $230K