Phase 1: ship the backup cron as CODE + a manual regrade trigger

FOUNDATION-FIRST re-order, phase 1 (tooling + safety).

BACKUP (highest-severity open item) — INSTALLED, not re-proven.
src/backupScheduler.js runs scripts/backup-db.sh nightly from inside the
API container, armed at boot in server.js. The container already has
SUPABASE_DB_URL, pg_dump and the Supabase route, so deploy == installed:
no host crontab, no Coolify click. Arming is deliberately opt-OUT (armed
whenever SUPABASE_DB_URL exists; BACKUP_CRON=0 kills it) because the S62
design was opt-in and nobody ever opted in — the DB went unbacked every
night for weeks. A failed run pages high-priority ntfy; silence is the
danger with backups.

Durability is the one part still needing a human: the container FS is
ephemeral, so a dump dies on redeploy unless BACKUP_REMOTE (off-box
rsync) or BACKUP_DIR (persistent volume) is set. The scheduler detects
that and pages a WARNING at boot rather than letting an undurable backup
read as "backed up". Runbook rewritten to lead with the code path.

MANUAL REGRADE TRIGGER — scripts/run-snapshot.js, runnable via
docker exec with no VYNDR_INTERNAL_KEY and no new HTTP surface. Runs the
SAME snapshotService.runSnapshot the cron runs (including the team-stats
refresh that powers opp_rank_stat), supports `all` and `--settle`, and
prints the grade/confidence distribution plus p_win/ev_pct presence —
which is the thing you actually want when verifying a grading change.

ACCESS BLOCKER, logged honestly in specs/model-train.md: there is no
VYNDR_INTERNAL_KEY in the local .env and SSH to the box times out from
WSL2, so I can neither curl the internal endpoints (which already exist
from S45) nor docker exec. The trigger is built and correct but only Kev
can run it until a key or SSH access exists. This is the highest-leverage
unblock for phases 2 and 3, which both need on-demand regrade+settle to
verify anything.

Also logged the standing cautions: CLV ledger stays private until
backtest-proven; "self-improving model" is unsupported marketing until
the loop closes; the engine is MLB/WNBA-calibrated and NFL/NBA/soccer
need their own calibration before the hub grades them (scaling gate).

Suite 277/3300 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
This commit is contained in:
Kev
2026-07-19 19:28:15 -04:00
parent bf8ecb45ad
commit b742230d94
6 changed files with 387 additions and 0 deletions
+33
View File
@@ -19,6 +19,39 @@ session). Already set. Optional: `BACKUP_DIR` (mount a Coolify **persistent
volume** here so dumps survive redeploys — e.g. `/var/backups/vyndr`),
`BACKUP_REMOTE` (off-box rsync target, below).
## ✅ THE CRON NOW SHIPS AS CODE (Session 64) — no install step
**Read this before following the manual instructions below; they are now the
FALLBACK, not the primary path.**
`src/backupScheduler.js` runs the nightly backup **inside the API container**,
armed from `server.js` at boot. The container already holds `SUPABASE_DB_URL`,
`pg_dump` and the Supabase network route, so **deploy == installed**. Nothing to
add to crontab, nothing to click in Coolify.
- **Arming is opt-OUT:** armed whenever `SUPABASE_DB_URL` is set. The S62 design
was opt-in (a host cron someone had to add) and nobody ever added it — the
database went unbacked every night for weeks. That failure mode is now
impossible.
- **Kill switch:** `BACKUP_CRON=0`.
- **Schedule:** `BACKUP_HOUR_UTC` (default 3) / `BACKUP_MINUTE_UTC` (default 10).
- **Failure pages high-priority ntfy.** Silence is the danger with backups.
- Boot log line: `[backupScheduler] armed — nightly 03:10 UTC ...`
### ⚠️ DURABILITY — the one thing still requiring a human
The container filesystem is **ephemeral**: a dump written inside it is LOST on the
next redeploy. The scheduler detects this and pages a warning at boot when
neither is configured. Set ONE of:
1. **`BACKUP_REMOTE`** — off-box rsync target (Hetzner Storage Box, ~€3/mo). Best.
2. **`BACKUP_DIR`** pointed at a **Coolify persistent volume** (e.g. `/var/backups/vyndr`).
Until one is set, backups run but do not survive a deploy. An undurable backup
that reads as "backed up" is worse than a loud gap — hence the boot-time page.
---
## Install the cron (host → docker exec into the API container)
Find the API container name (`docker ps | grep vyndr`), then a host cron: