Files
vyndr/scripts/backup-db.sh
T
builtbykev d3ffa1b8c2 Retention: model_snapshots live + base64 SSH key support
RETENTION (Phase 2, priority zero). History starts compounding tonight.

migration 025 model_snapshots — APPLIED to prod. Append-only, one row per
graded prop PER SIDE PER CYCLE, with a unique index on
(snapshot_id, player_key, stat, line, side) so a retried cycle cannot
duplicate. RLS on, service-role writes only.

What it captures that the ledger never did:
- features jsonb — the model's INPUTS. Without these a backtest can only
  grade our own homework; with them any future model can be replayed
  against the exact conditions this one faced.
- REFUSALS (refused + refusal_reason). The ledger drops them, so a gate
  refusing props that would have WON is invisible — unmeasurable lost
  edge. Captured via a new onGraded hook in gradeSlateService that fires
  with BOTH sides before any filtering.
- grade_11, the pre-collapse grade. The 4-letter map throws away the
  entire live C-/C/C+/B- range.
- model_version + code_sha on every row. ledger_entries mixes pre/post-fix
  grades with no marker and cannot be separated retroactively.
- p_win / ev_pct / fair_odds / takeable / value — none of which any
  permanent store held.

Wiring: analyzeViaEngine1 attaches _features/_grade_11 (underscore =
internal); gradeSlateService fires onGraded then STRIPS them so they never
reach a cache or API payload; snapshotService builds rows and persists
best-effort. Retention reuses the LEDGER's dateET/gameIdFor helpers so
rows share the ledger's natural key exactly — otherwise the settle pass
could never join outcomes onto them. Rows are written BEFORE the empty-
slate early return: a slate that refused everything is exactly the case
worth recording.

CONTRACT HELD: retention is injectable and every path is caught. persist()
returns errors, never throws; a missing Supabase client is SKIPPED, not an
error. A retention failure can never break a snapshot.

BACKUP: backup-db.sh now accepts BACKUP_SSH_KEY as base64 (recommended —
survives env-var newline mangling, which is how injected SSH keys usually
break silently) OR raw PEM, detected by decoding and looking for the PEM
header. Verified both forms detect correctly against a real generated key.

Suite 279/3325 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 23:01:12 -04:00

124 lines
6.2 KiB
Bash

#!/usr/bin/env bash
#
# VYNDR nightly database backup (security follow-up item 2).
#
# Supabase free tier has ZERO backups (no scheduled, no PITR) — the ledger and
# everything else have no safety net. This dumps the WHOLE database nightly via
# the direct connection string, keeps 14 days locally, pushes a weekly copy
# off-box, and pages ntfy on ANY failure. Runs on the Hetzner box via cron.
#
# REQUIRED env (set on the box / in the container that runs the cron):
# SUPABASE_DB_URL the Supabase DIRECT connection string (session mode, the
# db.<ref>.supabase.co:5432 URL — NOT the :6543 pooler;
# pg_dump needs a real session). Kev pastes this in Coolify.
# OPTIONAL env:
# BACKUP_DIR local dump dir (default /var/backups/vyndr)
# BACKUP_KEEP_DAYS local retention (default 14)
# BACKUP_REMOTE off-box rsync target for the weekly copy, e.g.
# u123456@u123456.your-storagebox.de:vyndr-backups/
# (empty = skip the off-box push; a WARN is paged)
# NTFY_URL (default https://ntfy.sh)
# NTFY_TOPIC (default vyndr-backups-kev2026)
#
set -Eeuo pipefail
BACKUP_DIR="${BACKUP_DIR:-/var/backups/vyndr}"
KEEP_DAYS="${BACKUP_KEEP_DAYS:-14}"
NTFY_URL="${NTFY_URL:-https://ntfy.sh}"
NTFY_TOPIC="${NTFY_TOPIC:-vyndr-backups-kev2026}"
STAMP="$(date -u +%Y%m%d-%H%M%S)"
DUMP="${BACKUP_DIR}/vyndr-${STAMP}.dump"
MIN_BYTES="${BACKUP_MIN_BYTES:-50000}" # a real dump of this DB is far bigger; guards an empty/failed dump
notify() { # notify <title> <priority> <message>
curl -fsS --max-time 15 \
-H "Title: ${1}" -H "Priority: ${2}" -H "Tags: floppy_disk" \
-d "${3}" "${NTFY_URL}/${NTFY_TOPIC}" >/dev/null 2>&1 || true
}
fail() { notify "VYNDR backup FAILED" "urgent" "${1}"; echo "ERROR: ${1}" >&2; exit 1; }
trap 'fail "backup script errored near line ${LINENO}"' ERR
[ -n "${SUPABASE_DB_URL:-}" ] || fail "SUPABASE_DB_URL is not set — cannot back up"
command -v pg_dump >/dev/null 2>&1 || fail "pg_dump not installed (apt-get install postgresql-client)"
mkdir -p "${BACKUP_DIR}"
# 1. Dump the whole DB in custom format (-Fc: compressed, restorable with pg_restore).
pg_dump "${SUPABASE_DB_URL}" -Fc --no-owner --no-privileges -f "${DUMP}" \
|| fail "pg_dump failed"
# 2. Sanity: a real dump is not tiny. An empty/near-empty file is a silent failure.
SIZE="$(stat -c%s "${DUMP}" 2>/dev/null || echo 0)"
[ "${SIZE}" -ge "${MIN_BYTES}" ] || fail "dump is only ${SIZE} bytes (< ${MIN_BYTES}) — treating as a failed backup"
# 2b. Integrity fingerprint: a valid custom-format archive lists its objects via
# pg_restore --list (no target DB needed). Confirm it parses AND contains the
# ledger — proves it's a real, restorable archive, not just a file of bytes.
TOC="$(pg_restore --list "${DUMP}" 2>/dev/null)" || fail "pg_restore --list failed — dump is not a valid archive"
OBJECTS="$(printf '%s\n' "${TOC}" | grep -c ';' || true)"
printf '%s\n' "${TOC}" | grep -qi 'TABLE DATA public ledger_entries' \
|| fail "dump archive does not contain ledger_entries — refusing to trust it"
echo "backup validated: ${DUMP} (${SIZE} bytes, ${OBJECTS} archive objects, ledger_entries present)"
# 3. Rotate: drop local dumps older than KEEP_DAYS.
find "${BACKUP_DIR}" -name 'vyndr-*.dump' -type f -mtime "+${KEEP_DAYS}" -delete || true
# 3b. SSH key for the off-box push (Session 64).
# The container filesystem is EPHEMERAL — a keypair generated inside it dies
# on the next redeploy and the off-box push would silently start failing. So
# the PRIVATE key is injected as an env var (Coolify secret) and written to a
# 0600 temp file per run. Hetzner Storage Box speaks full OpenSSH on PORT 23
# (port 22 is SFTP-only, mod_sftp) — verified live; rsync must target 23.
SSH_KEY_FILE=""
cleanup_key() { [ -n "${SSH_KEY_FILE}" ] && rm -f "${SSH_KEY_FILE}" || true; }
trap cleanup_key EXIT
RSYNC_SSH="ssh -p ${BACKUP_SSH_PORT:-23} -o StrictHostKeyChecking=accept-new -o BatchMode=yes"
if [ -n "${BACKUP_SSH_KEY:-}" ]; then
SSH_KEY_FILE="$(mktemp)"
chmod 600 "${SSH_KEY_FILE}"
# Accept EITHER form (Session 64):
# 1. base64 (recommended — `base64 -w0`; survives any env-var mangling of
# newlines, which is the usual way an injected SSH key silently breaks)
# 2. raw PEM with literal \n escapes
# Detect base64 by trying to decode and checking for the PEM header.
DECODED="$(printf '%s' "${BACKUP_SSH_KEY}" | base64 -d 2>/dev/null || true)"
case "${DECODED}" in
*"PRIVATE KEY"*)
printf '%s\n' "${DECODED}" > "${SSH_KEY_FILE}"
echo "ssh key: base64-decoded"
;;
*)
printf '%b\n' "${BACKUP_SSH_KEY}" | sed -e 's/[[:space:]]*$//' > "${SSH_KEY_FILE}"
echo "ssh key: used raw (not base64)"
;;
esac
chmod 600 "${SSH_KEY_FILE}"
RSYNC_SSH="${RSYNC_SSH} -i ${SSH_KEY_FILE}"
fi
# 4. OFF-BOX COPY — DEFERRED (Session 64).
# The dump now lands on a PERSISTENT VOLUME (BACKUP_DIR=/app/backups), so it
# already survives redeploys — the container-ephemeral risk is closed. Storage
# Box SSH auth is not working yet, so the off-box push is explicitly DEFERRED:
# it must never fail the backup. A durable on-box dump is a real backup; a
# failing rsync on top of it is a follow-up, not an incident.
#
# Set BACKUP_OFFBOX=1 (with BACKUP_SSH_KEY) to re-enable. Until then we log
# and page at LOW priority, and we never call a deferred push a failure.
if [ "${BACKUP_OFFBOX:-0}" = "1" ] && [ -n "${BACKUP_REMOTE:-}" ] && [ -n "${BACKUP_SSH_KEY:-}" ]; then
if rsync -az --timeout=120 -e "${RSYNC_SSH}" "${DUMP}" "${BACKUP_REMOTE}"; then
echo "off-box push OK -> ${BACKUP_REMOTE%%:*}"
notify "VYNDR backup OK (+off-box)" "default" "Nightly dump ${STAMP} (${SIZE} bytes) pushed off-box."
else
# NOT a failure: the durable on-box dump succeeded.
echo "off-box push FAILED (deferred; on-box dump is durable)"
notify "VYNDR off-box push deferred" "low" "Dump ${STAMP} (${SIZE} bytes) is durable on the persistent volume; the off-box rsync failed and is deferred."
fi
else
echo "off-box push DEFERRED (BACKUP_OFFBOX!=1 or remote/key unset) — on-box dump is durable at ${DUMP}"
fi
echo "backup ok: ${DUMP} (${SIZE} bytes)"
exit 0