In progress

Cross-Database / Cloud Migration (CDC, Near-Zero Downtime)

Database Management

The original pitch was a full cross-database/cloud migration with a near-zero-downtime cutover. What actually got built and proven so far is the CDC replication layer itself: real Kafka + Zookeeper + Kafka Connect running Debezium's MySQL connector into a JDBC sink writing self-hosted Postgres, kept in sync with MySQL's live binlog. Verified against project1_jobs only, with real integration bugs hit and fixed along the way. The migration and cutover this project is named for -- the other 9 databases, and an actual switch of a live app to Postgres or a managed cloud database -- has not happened.

DocsLast updated September 9, 2026

Scope, honestly

This project is titled for the end goal (cross-database/cloud migration, near-zero-downtime cutover), but what exists today is the CDC pipeline that a real migration would be built on top of -- not the migration itself. No app has been cut over. No cloud database is in use. 9 of the 10 target databases still need their Postgres schema hand-translated before this could even start on them.

What's actually running

Debezium (MySQL connector) via Kafka Connect, on real Kafka + Zookeeper (not Redpanda), streaming project1_jobs from self-hosted MySQL into self-hosted Postgres through a JDBC sink connector. End state as of the last verification: both connectors RUNNING, replication lag 0, row counts matched exactly against the live MySQL source.

Real bugs hit and fixed getting here

  • debezium-connector-mysql:2.6.2 didn't exist on Confluent Hub -- had to find the actual available version via the Hub API instead of guessing.
  • MySQL 8.4 renamed SHOW MASTER STATUS to SHOW BINARY LOG STATUS; Debezium only supports that from 3.0+, which needs Java 17. The base Kafka Connect image ships Java 11 -- installed Java 17 alongside it and pointed Connect's launcher at it via JAVA_HOME, rather than swapping base images and risking a wire-protocol mismatch with the already-running broker.
  • The JDBC sink needed schemas.enable=true on both the key and value converters -- delete.enabled + pk.mode=record_key require a typed Struct key, not schema-less JSON.
  • Postgres id columns had to be GENERATED BY DEFAULT AS IDENTITY, not GENERATED ALWAYS -- a migration has to preserve the source's real IDs.
  • Cross-topic FK ordering race: Kafka gives no ordering guarantee across topics, so jobs rows could arrive before their companies row existed yet, and Postgres FK constraints turned that into a hard failure during a burst of ~390 new rows. Recovered by diffing MySQL vs. Postgres IDs and backfilling ~340 missing parent rows.

What's deliberately deferred, not forgotten

The FK-ordering race above is a real structural gap (smaller per-row transactions, or errors.tolerance=all + a dead-letter queue, would fix it) -- left open on purpose since this pipeline exists for CDC practice, not as a live migration path, so the fix isn't worth doing until the actual migration work resumes. The other 9 databases need their MySQL schemas hand-translated to Postgres equivalents before CDC can even be pointed at them -- that translation work, not automation, is the actual point of doing it.

Pending databases

project1_jobs is the only one of the 10 real app databases actually running through CDC. Still pending, each needing its MySQL schema hand-translated to Postgres before CDC can even be pointed at it:

  • analytics
  • bookstack_db
  • budget_db
  • metabase_app
  • movielens
  • portfolio
  • project3_perf
  • sandbox_db
  • todo_db