ben.barrera
back to blog
TechnicalHybrid

Super Bartie: what I built around my CDC engine

A follow-up to Bartie. I added an API for checking on the pipeline, a second destination that stores vectors, an MCP server, and a deployment.

Super Bartie is a follow-up to Bartie, the small Postgres CDC engine I built to understand how Artie works. Bartie ran from a terminal. For this version I left the reader and writer alone and built the parts around them: an API for checking on the pipeline, a second destination that stores vectors, an MCP server, and a deployment I can link to.

See it run

Super Bartie demoWatch on YouTube

Try it

On the live demo you can watch rows get replicated, change a row yourself, pause the writer, and ask questions about the data. The second page shows the source and destination next to each other.

How it fits together

Super Bartie architecture. Top: source Postgres sends WAL to the reader, which publishes events to Redpanda. Two consumers read them: the writer merges rows into replica tables and the vecwriter upserts embeddings into a vector table, which is copied to a snapshot every five minutes. Bottom: the web page and Claude Code call the control API, which runs keyword and vector search and passes the top rows to a model.

Open the full-size diagram

What I added

1. An HTTP API

With Bartie, the only way to see what was happening was to read logs. I added an HTTP API so the pipeline can be checked and controlled from outside:

EndpointWhat it does
StatusWhether the pipeline is running or paused, which tables it copies, and which processes are up
UsageLatency per table, plus reader lag, backlog, and merge time, which show which part is slow
Error logWhat failed and in which process
VerifyCompares checksums of every table in both databases
Pause and resumeStops and restarts the writer

I copied the URL layout from Artie's public API, so the two are easy to compare.

2. A second destination for vectors

I added a second consumer that reads the same Redpanda topic as the writer. For each changed row it builds a sentence, embeds it, and stores it in pgvector. The idea came from Artie's post on real-time data for AI.

The demo keeps two copies: the live table, which is about two seconds behind the source, and a snapshot that refreshes every five minutes. If you change a row and ask the same question against both, only the live copy has the new answer.

3. An MCP server

Artie's MCP server generates its tools from an OpenAPI spec instead of defining each one by hand. I built mine the same way with a spec for my API and used their tool names where the endpoints matched.

I also added two tools they do not have. pipeline_verify compares checksums of every table in both databases, and destination_ask searches the vector table.

Running it

Everything runs with Docker Compose on one VM: the source Postgres, Redpanda, the destination Postgres with pgvector, and the Go binaries. Caddy handles TLS. A cron job resets the data every night, which also means the backfill runs from scratch once a day.

The code is on GitHub: Super Bartie, and the engine it is built on, Bartie.