40aba37f83
Closing the backup for real. Three changes, each fixing something that
would have made the Storage Box target fail or silently rot.
1. SSH KEY COMES FROM ENV, not from the container. Generating a keypair
inside the API container was the obvious move and it is wrong: the
container filesystem is ephemeral, so the key dies on the next
redeploy and the off-box push starts failing silently. backup-db.sh
now reads BACKUP_SSH_KEY (a Coolify secret), writes it to a 0600 temp
file per run, and removes it on exit via trap.
2. PORT 23, verified live. Hetzner Storage Box runs full OpenSSH on 23;
port 22 answers with mod_sftp (SFTP only). Banner-checked both against
u635423.your-storagebox.de. rsync now uses
-e "ssh -p ${BACKUP_SSH_PORT:-23} ... -i <key>"; the old invocation had
no -e at all and would have gone to 22.
3. OFF-BOX PUSH IS NIGHTLY, not Sundays-only. A weekly push meant up to
six days of dumps existed ONLY inside an ephemeral container, which is
the same as not existing. Alert copy updated to say exactly that when
the push fails or is skipped.
Also adds POST /api/internal/backup/run (internal-key gated) so a real
backup can be TRIGGERED and OBSERVED — it returns exit code, duration,
output tail, and whether the remote + ssh key are configured. The backup
can only run where SUPABASE_DB_URL and the Supabase route live (this
container), and there was no way to fire or inspect it without a shell.
Connectivity established this session: Storage Box reachable from the dev
box on 22/23; Supabase :5432 NOT reachable from WSL2 (so the dump must
run in-container, as designed); docker IS available locally, so the
restore-verify can run against a scratch Postgres using the real dump.
Suite 278/3305 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
102 lines
5.2 KiB
Bash
102 lines
5.2 KiB
Bash
#!/usr/bin/env bash
|
|
#
|
|
# VYNDR nightly database backup (security follow-up item 2).
|
|
#
|
|
# Supabase free tier has ZERO backups (no scheduled, no PITR) — the ledger and
|
|
# everything else have no safety net. This dumps the WHOLE database nightly via
|
|
# the direct connection string, keeps 14 days locally, pushes a weekly copy
|
|
# off-box, and pages ntfy on ANY failure. Runs on the Hetzner box via cron.
|
|
#
|
|
# REQUIRED env (set on the box / in the container that runs the cron):
|
|
# SUPABASE_DB_URL the Supabase DIRECT connection string (session mode, the
|
|
# db.<ref>.supabase.co:5432 URL — NOT the :6543 pooler;
|
|
# pg_dump needs a real session). Kev pastes this in Coolify.
|
|
# OPTIONAL env:
|
|
# BACKUP_DIR local dump dir (default /var/backups/vyndr)
|
|
# BACKUP_KEEP_DAYS local retention (default 14)
|
|
# BACKUP_REMOTE off-box rsync target for the weekly copy, e.g.
|
|
# u123456@u123456.your-storagebox.de:vyndr-backups/
|
|
# (empty = skip the off-box push; a WARN is paged)
|
|
# NTFY_URL (default https://ntfy.sh)
|
|
# NTFY_TOPIC (default vyndr-backups-kev2026)
|
|
#
|
|
set -Eeuo pipefail
|
|
|
|
BACKUP_DIR="${BACKUP_DIR:-/var/backups/vyndr}"
|
|
KEEP_DAYS="${BACKUP_KEEP_DAYS:-14}"
|
|
NTFY_URL="${NTFY_URL:-https://ntfy.sh}"
|
|
NTFY_TOPIC="${NTFY_TOPIC:-vyndr-backups-kev2026}"
|
|
STAMP="$(date -u +%Y%m%d-%H%M%S)"
|
|
DUMP="${BACKUP_DIR}/vyndr-${STAMP}.dump"
|
|
MIN_BYTES="${BACKUP_MIN_BYTES:-50000}" # a real dump of this DB is far bigger; guards an empty/failed dump
|
|
|
|
notify() { # notify <title> <priority> <message>
|
|
curl -fsS --max-time 15 \
|
|
-H "Title: ${1}" -H "Priority: ${2}" -H "Tags: floppy_disk" \
|
|
-d "${3}" "${NTFY_URL}/${NTFY_TOPIC}" >/dev/null 2>&1 || true
|
|
}
|
|
|
|
fail() { notify "VYNDR backup FAILED" "urgent" "${1}"; echo "ERROR: ${1}" >&2; exit 1; }
|
|
trap 'fail "backup script errored near line ${LINENO}"' ERR
|
|
|
|
[ -n "${SUPABASE_DB_URL:-}" ] || fail "SUPABASE_DB_URL is not set — cannot back up"
|
|
command -v pg_dump >/dev/null 2>&1 || fail "pg_dump not installed (apt-get install postgresql-client)"
|
|
mkdir -p "${BACKUP_DIR}"
|
|
|
|
# 1. Dump the whole DB in custom format (-Fc: compressed, restorable with pg_restore).
|
|
pg_dump "${SUPABASE_DB_URL}" -Fc --no-owner --no-privileges -f "${DUMP}" \
|
|
|| fail "pg_dump failed"
|
|
|
|
# 2. Sanity: a real dump is not tiny. An empty/near-empty file is a silent failure.
|
|
SIZE="$(stat -c%s "${DUMP}" 2>/dev/null || echo 0)"
|
|
[ "${SIZE}" -ge "${MIN_BYTES}" ] || fail "dump is only ${SIZE} bytes (< ${MIN_BYTES}) — treating as a failed backup"
|
|
|
|
# 2b. Integrity fingerprint: a valid custom-format archive lists its objects via
|
|
# pg_restore --list (no target DB needed). Confirm it parses AND contains the
|
|
# ledger — proves it's a real, restorable archive, not just a file of bytes.
|
|
TOC="$(pg_restore --list "${DUMP}" 2>/dev/null)" || fail "pg_restore --list failed — dump is not a valid archive"
|
|
OBJECTS="$(printf '%s\n' "${TOC}" | grep -c ';' || true)"
|
|
printf '%s\n' "${TOC}" | grep -qi 'TABLE DATA public ledger_entries' \
|
|
|| fail "dump archive does not contain ledger_entries — refusing to trust it"
|
|
echo "backup validated: ${DUMP} (${SIZE} bytes, ${OBJECTS} archive objects, ledger_entries present)"
|
|
|
|
# 3. Rotate: drop local dumps older than KEEP_DAYS.
|
|
find "${BACKUP_DIR}" -name 'vyndr-*.dump' -type f -mtime "+${KEEP_DAYS}" -delete || true
|
|
|
|
# 3b. SSH key for the off-box push (Session 64).
|
|
# The container filesystem is EPHEMERAL — a keypair generated inside it dies
|
|
# on the next redeploy and the off-box push would silently start failing. So
|
|
# the PRIVATE key is injected as an env var (Coolify secret) and written to a
|
|
# 0600 temp file per run. Hetzner Storage Box speaks full OpenSSH on PORT 23
|
|
# (port 22 is SFTP-only, mod_sftp) — verified live; rsync must target 23.
|
|
SSH_KEY_FILE=""
|
|
cleanup_key() { [ -n "${SSH_KEY_FILE}" ] && rm -f "${SSH_KEY_FILE}" || true; }
|
|
trap cleanup_key EXIT
|
|
RSYNC_SSH="ssh -p ${BACKUP_SSH_PORT:-23} -o StrictHostKeyChecking=accept-new -o BatchMode=yes"
|
|
if [ -n "${BACKUP_SSH_KEY:-}" ]; then
|
|
SSH_KEY_FILE="$(mktemp)"
|
|
chmod 600 "${SSH_KEY_FILE}"
|
|
# Accept the key with literal \n escapes (how env vars usually carry it).
|
|
printf '%b\n' "${BACKUP_SSH_KEY}" | sed -e 's/[[:space:]]*$//' > "${SSH_KEY_FILE}"
|
|
RSYNC_SSH="${RSYNC_SSH} -i ${SSH_KEY_FILE}"
|
|
fi
|
|
|
|
# 4. OFF-BOX COPY — every night, not only Sundays (Session 64).
|
|
# A weekly push meant up to 6 days of dumps existed ONLY inside an ephemeral
|
|
# container, which is the same as not existing. Off-box is the real backup.
|
|
if true; then
|
|
if [ -n "${BACKUP_REMOTE:-}" ]; then
|
|
if rsync -az --timeout=120 -e "${RSYNC_SSH}" "${DUMP}" "${BACKUP_REMOTE}"; then
|
|
notify "VYNDR backup OK (+off-box)" "default" "Nightly dump ${STAMP} (${SIZE} bytes) pushed off-box to ${BACKUP_REMOTE%%:*}."
|
|
else
|
|
notify "VYNDR off-box push FAILED" "high" "Local dump ${STAMP} is fine (${SIZE} bytes) but the off-box rsync FAILED — the dump exists only in an ephemeral container."
|
|
fi
|
|
else
|
|
notify "VYNDR off-box push SKIPPED" "high" "Local dump ${STAMP} OK but BACKUP_REMOTE is unset — the dump exists only in an ephemeral container."
|
|
fi
|
|
else
|
|
echo "backup ok: ${DUMP} (${SIZE} bytes)"
|
|
fi
|
|
|
|
exit 0
|