> Markdown version of [/jobs/ext/2703947-database-reliability-engineer](https://www.wearedevelopers.com/jobs/ext/2703947-database-reliability-engineer). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Database Reliability Engineer - **Company:** Scribe Inc. - **Location:** San Francisco, CA, United States (Remote available) - **Experience:** Expert - **Contract:** Permanent contract - **Skills:** Artificial Intelligence, Amazon Web Services, Amazon S3, BigQuery, Databases, Data Infrastructure, Django Web Framework, Python (Programming Language), PostgreSQL, Message Broker, RabbitMQ, Redis, SQL Databases, SQLAlchemy, Parquet, Datadog, Amazon ElastiCache, Snowflake, Data Lakes, Debezium, Apache Kafka, Event Sourcing, Cloudwatch, Amazon Simple Queue Service (SQS), Terraform, Serverless Computing - **Published:** September 4, 2026 - **Apply:** https://startup.jobs/senior-database-reliability-engineer-scribe-com-8027127 ## About the Role * Deep PostgreSQL expertise in practice: read EXPLAIN (ANALYZE, BUFFERS) fluently, understand MVCC, bloat, lock contention, and vacuum behavior, and tune Aurora Serverless V2 for latency and throughput * Work with an ORM (Django, SQLAlchemy, ActiveRecord, or similar) at production scale - predict the SQL a query generates, spot N+1 issues on sight, and know when joins beat batched IN queries and when they don't * Run CDC pipelines in production, ideally with AWS DMS - comfort with logical replication, slot hygiene, schema evolution, and Parquet-based data lakes feeding Snowflake, BigQuery, or Redshift * Hands-on experience with pganalyze (or Datadog DBM / pg_stat_statements pipelines), CloudWatch, and Honeycomb (or another high-cardinality tracing tool); comfortable with OpenTelemetry * Work with OpenSearch, Redis, and at least one production message broker (SQS, RabbitMQ, or Kafka) at scale * Write real automation - Python, Go, or similar - and use Terraform or comparable IaC to manage infrastructure * Use AI coding and review tools in a team setting: write and maintained AGENTS.md files, configure review agents, iterate on prompts Nice to Have * Event sourcing on Postgres, or experience with alternate CDC tooling (Debezium, Fivetran, Airbyte) * pgbouncer or RDS Proxy at scale with Django connection handling * Deep Honeycomb usage: SLOs, BubbleUp, Triggers, derived columns * Snowflake from the producer side: staging, Snowpipe, external tables on Parquet * Experience scaling data infrastructure through rapid engineering headcount growth * SOC 2 Type II, GDPR, or similar compliance work ## Description Scribe is where exceptional people come to do the best work of their careers. Our Workflow AI platform automatically captures and optimizes how work gets done - 94% of the Fortune 500 use it, and 45% are paying customers. We hit $100M ARR in May 2026 and have grown to over 5 million daily active users across 600,000 businesses. We're Series C and valued at $1.3 billion. We're builders who hold a high bar, move fast, and care deeply about each other and our customers., We're hiring a Senior Database Reliability Engineer to own the reliability, performance, and scalability of Scribe's data tier. Our engineering org is doubling - which means the guardrails, automation, and standards you put in place today will carry a much larger team through the next phase of growth. This is a senior IC role with real ownership: you'll set the bar for how engineers across the company interact with our databases, not just keep the lights on. Our stack is Django on PostgreSQL (Aurora Serverless V2), OpenSearch, Redis (ElastiCache), SQS, and RabbitMQ, with a CDC pipeline running Aurora to DMS to S3 Parquet to Snowflake. Engineers ship through the ORM, not raw SQL - which makes migration safety, index design, and query review genuinely high-stakes work. ️ What You'll Do * Own database reliability across Aurora, OpenSearch, Redis, and our CDC pipeline - including schema design reviews, migration safety (locks, backfills, concurrent index builds, NOT VALID constraints), and incident response for the data tier * Make the Django ORM a strength at scale: catch N+1 patterns in review, extend QuerySet conventions and physical schema standards, and build the CI checks and AGENTS.md scaffolding that encode those standards so they scale beyond any single reviewer * Operate and evolve the CDC pipeline from Aurora through DMS to S3 Parquet to Snowflake - including replication slot hygiene, schema evolution safety, and automated checks that catch migrations likely to break downstream consumers before they ship * Build and improve observability across pganalyze, CloudWatch, and Honeycomb, with Django-side instrumentation that ties slow ORM queries back to specific users, flags, and deploys * Drive multi-AZ resilience within our single-region architecture - Aurora writer/reader placement, failover behavior, RTO/RPO, ElastiCache and OpenSearch AZ topology, RabbitMQ survivability * Build self-service tooling and dashboards that give product and platform teams visibility into their own query footprint, reducing the review burden as the engineering org grows * Contribute to onboarding and knowledge-sharing as a large incoming class of engineers joins - write docs, run internal sessions on "what your ORM query is really doing," and feed that knowledge back into AI review tooling, San Francisco (hybrid, 3 days per week in-office) or, Remote based permanently in PST (Pacific Standard Time). ## Related Videos - [Parquet, Delta, Iceberg & Ducklake - An introduction for developers](https://www.wearedevelopers.com/videos/100075-parquet-delta-iceberg-ducklake-an-introduction-for-developers) - [Optimizing Discovery: PostgreSQL's Role in Transforming GetYourGuide's Search](https://www.wearedevelopers.com/videos/1647-optimizing-discovery-postgresql-s-role-in-transforming-getyourguide-s-search) - [Reducing LLM Calls with Vector Search Patterns - Raphael De Lio (Redis)](https://www.wearedevelopers.com/videos/1714-reducing-llm-calls-with-vector-search-patterns-raphael-de-lio-redis) - [Debugging in the Dark](https://www.wearedevelopers.com/videos/1658-debugging-in-the-dark) - [Accelerating Authentication Architecture: Taking Passwordless to the Next Level](https://www.wearedevelopers.com/videos/733-accelerating-authentication-architecture-taking-passwordless-to-the-next-level) - [Tracking vehicles at scale](https://www.wearedevelopers.com/videos/1999-tracking-vehicles-at-scale) ## Related Articles - [How We Built a Worry-Free System That Runs for 10+ Years – And What We’d Do Again](https://www.wearedevelopers.com/magazine/751-how-we-built-a-worry-free-system-that-runs-for-10-years-and-what-we-d-do-again) - [Making Data Warehouses Fast: A Developer’s Story](https://www.wearedevelopers.com/magazine/107-making-data-warehouses-fast-a-developer-s-story) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Dev Digest 168: Hacking Postgres, Blocking Meta and Fixing CSS](https://www.wearedevelopers.com/magazine/588-dev-digest-168-hacking-postgres-blocking-meta-and-fixing-css) - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 139 - Soft and hard queries](https://www.wearedevelopers.com/magazine/487-dev-digest-139-soft-and-hard-queries)