Task A — make the container backup-capable + validated dump + mechanism fingerprint

SUPABASE_DB_URL is set in Coolify on the API service, and this WSL2 box can't
reach db.<ref>.supabase.co — so the backup runs INSIDE the API container, which
has the env + Supabase network. Made that real:
- Dockerfile: install postgresql-client (pg_dump/pg_restore) + rsync + bash in
  the runner image.
- backup-db.sh: added an integrity fingerprint on every run — pg_restore --list
  must parse the archive AND find ledger_entries, else the run FAILS + pages
  (stronger than the size check; catches a corrupt/structureless dump).
- BACKUP-RUNBOOK.md: rewritten for the container-exec reality — host cron does
  `docker exec <api> sh /app/scripts/backup-db.sh` (inherits env + network +
  pg_dump), or a Coolify Scheduled Task. Full restore-fingerprint steps included.

MECHANISM FINGERPRINT (run locally, docker + pg16): seeded a ledger_entries
table (137 rows) → ran backup-db.sh (dump + validate: 22 archive objects,
ledger_entries present) → pg_restore into a scratch DB → 137 rows restored,
exact match. The dump/validate/restore path is proven end-to-end; it's the same
pg_dump/pg_restore that run in the container against Supabase.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
Kev
2026-07-18 23:08:31 -04:00
parent ccb9668f0c
commit c2c43cdc92
3 changed files with 50 additions and 22 deletions
+36 -20
View File
@@ -4,29 +4,35 @@ Supabase free tier has **zero** backups (no scheduled, no PITR). `scripts/backup
is the safety net: a nightly full-database `pg_dump`, 14 days kept locally, a
weekly copy pushed off-box, ntfy alert on any failure.
## What Kev needs to set (one env var)
## Where it runs
**`SUPABASE_DB_URL`** — the Supabase **direct** connection string (session mode).
Supabase → Project → Settings → Database → **Connection string → URI**, the
`db.<ref>.supabase.co:5432` one (NOT the `:6543` transaction pooler — `pg_dump`
needs a real session). Paste it in Coolify as an env var; never commit it.
`SUPABASE_DB_URL` is set in Coolify **on the VYNDR API service** — so the backup
runs **inside that container**, which already has the env, the Supabase network,
and now `pg_dump`/`pg_restore`/`rsync` (added to the Dockerfile). The Hetzner box
can't reach `db.<ref>.supabase.co` directly and doesn't hold the connection
string; the container is the right place.
Optional: `BACKUP_REMOTE` (off-box rsync target for the weekly copy — see below).
**`SUPABASE_DB_URL`** must be the **direct** connection string (session mode):
Supabase → Settings → Database → **Connection string → URI**, the
`db.<ref>.supabase.co:5432` one (NOT the `:6543` pooler — `pg_dump` needs a real
session). Already set. Optional: `BACKUP_DIR` (mount a Coolify **persistent
volume** here so dumps survive redeploys — e.g. `/var/backups/vyndr`),
`BACKUP_REMOTE` (off-box rsync target, below).
## Install the cron (on the Hetzner box)
## Install the cron (host → docker exec into the API container)
Find the API container name (`docker ps | grep vyndr`), then a host cron:
```bash
# 1. Ensure the client is present
sudo apt-get install -y postgresql-client rsync
# 2. Nightly at 03:10 UTC. Pass the env the script needs (or source an env file).
# Nightly at 03:10 UTC. BACKUP_REMOTE can also be set in Coolify instead.
sudo crontab -e
# add:
10 3 * * * SUPABASE_DB_URL='postgresql://...' BACKUP_REMOTE='u123456@u123456.your-storagebox.de:vyndr-backups/' /path/to/vyndr/scripts/backup-db.sh >> /var/log/vyndr-backup.log 2>&1
# add (replace <api-container>):
10 3 * * * docker exec -e BACKUP_REMOTE='u123456@u123456.your-storagebox.de:vyndr-backups/' <api-container> sh /app/scripts/backup-db.sh >> /var/log/vyndr-backup.log 2>&1
```
(If the cron runs inside the Coolify container instead, the env vars are already
present — just schedule `scripts/backup-db.sh`.)
`docker exec` inherits the container's env (`SUPABASE_DB_URL`) + network + the
newly-installed `pg_dump`. (If you'd rather, Coolify's **Scheduled Tasks** can run
`sh /app/scripts/backup-db.sh` on the API service on the same cadence.)
## Off-box target (simplest reliable pick)
@@ -38,14 +44,24 @@ rather keep it off Hetzner entirely — swap the `rsync` line for `rclone copy`.
## Restore / FINGERPRINT (proves it's a real backup, not just a file)
A dump only counts once a restore of it succeeds. Load one into a scratch DB:
**Built-in (runs on every backup):** the script validates each dump with
`pg_restore --list` — a dump that isn't a valid archive, or that doesn't contain
`ledger_entries`, is treated as a FAILURE and paged. So every successful run has
already proven the archive parses and holds the ledger.
**Full restore proof (run once to fingerprint):** from the host, restore the
newest dump into a throwaway postgres and count the ledger:
```bash
# spin a throwaway local postgres, restore the newest dump, count a known table
# 1. Produce a dump on demand (writes into the container's BACKUP_DIR)
docker exec <api-container> sh /app/scripts/backup-db.sh
# 2. Copy the newest dump out of the container
newest=$(docker exec <api-container> sh -lc 'ls -t /var/backups/vyndr/vyndr-*.dump | head -1')
docker cp "<api-container>:${newest}" /tmp/vyndr-latest.dump
# 3. Restore into a scratch postgres + count a known table
docker run -d --name vyndr-restore-test -e POSTGRES_PASSWORD=x -p 55432:5432 postgres:15
sleep 5
newest=$(ls -t /var/backups/vyndr/vyndr-*.dump | head -1)
pg_restore --no-owner --no-privileges -d "postgresql://postgres:x@localhost:55432/postgres" "$newest"
sleep 6
pg_restore --no-owner --no-privileges -d "postgresql://postgres:x@localhost:55432/postgres" /tmp/vyndr-latest.dump
psql "postgresql://postgres:x@localhost:55432/postgres" -c "select count(*) from public.ledger_entries;"
docker rm -f vyndr-restore-test
```