ben.barrera
back to blog
TechnicalHybrid

Building Bartie: a small Postgres CDC engine

Postgres WAL to Kafka to a second database. How Bartie handles acknowledgments, replay, and snapshot backfill.

I built Bartie to understand how streaming replication works, inspired by Artie and their post Introducing Artie Transfer. It reads changes from a Postgres write-ahead log, sends them through Redpanda, and applies them to a second Postgres database. I kept it to one source and one destination so I could follow each change through the pipeline and see what happens when a process crashes.

See it run

Bartie: architecture, live replication, and code walkthroughWatch on YouTube

The demo uses Artie's terra dataset: animals, observations, and watering holes.

The pipeline

Two Go processes with Redpanda between them:

  • The reader decodes Postgres logical replication messages and publishes change events.
  • The writer consumes those events in batches and merges them into the destination.
Bartie pipeline: source Postgres sends WAL to a reader, which publishes JSON events to Redpanda. A writer applies staging merges to destination Postgres. Acknowledgments follow durable writes, and cdcctl compares both databases.

Open the full-size architecture diagram

Every change travels in the same event envelope, including rows from the initial snapshot. The Kafka key is the table plus the primary key, so changes to one row stay in order.

Two rules for acknowledging progress

The reader tells Postgres it has consumed a WAL position only after Redpanda has confirmed the events. The writer commits a Kafka offset only after the destination transaction has committed.

A crash between those steps means some changes are delivered twice. That is safe, because applying the same row state again gives the same result.

Why the writer uses staging and merge

The writer flushes every 500 events or two seconds. For each table, it:

  1. Folds the batch into one final event per primary key.
  2. Loads those events into a temporary staging table.
  3. Runs a MERGE to insert, update, or delete destination rows.
  4. Commits that table's transaction.

Backfill needs a precise starting point

Bartie creates a replication slot with an exported snapshot, copies existing rows from that snapshot, then streams from the same WAL position. Writes that happen during the copy are not lost.

The design has two limits. A crash mid-backfill means starting over, and a long backfill holds WAL on the source until streaming catches up.

Checking the copy after a crash

  • crash-test.sh kills the writer and reader during a stream of mixed writes, restarts them, and verifies.
  • backfill-test.sh writes to the source while the initial copy runs, then verifies.

cdcctl verify compares row counts and checksums of every table in both databases. Counts alone are not enough, because a table can have the right number of rows and the wrong data.

Streaming versus re-copying the table

On a laptop, a change reaches the destination about 1.1 seconds after it commits on the source (recorded results). Most of that is the writer's two-second flush window.

The alternative is re-copying the table on a schedule, which can never be fresher than one full copy takes:

RowsFull-table copyStreaming CDC
10,0000.20 s1.1 s
100,0000.55 s1.1 s
1,000,0002.51 s1.1 s
10,000,00044.2 s1.1 s

At small sizes the copy wins. Past about a million rows, the copy grows with the table while CDC only touches the change. The benchmark ran on a single-broker laptop, so the numbers show the trend, not CDC latency at scale.

Where the project stops

Bartie requires primary keys and does not yet support schema evolution, TRUNCATE, or resumable backfill.

Explore Bartie on GitHub, or read the follow-up, Super Bartie.