Files
vyndr/scripts/settle-model-snapshots.js
T
builtbykev f976df47b8 Settle model_snapshots + four-stat calibration: works, deploys nowhere
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.

PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.

PHASE 1, two integrity findings both caught by the gate:

1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').

2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.

THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.

THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.

PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):

  hits  n=1140  bias +0.0868  brier 0.2626 -> 0.2511  d -0.0115  CI [-0.0139,-0.0097]
  TB    n=1050  bias +0.0834  brier 0.2490 -> 0.2438  d -0.0052  CI [-0.0061,-0.0045]
  rbi   n= 630  bias +0.0164  brier 0.2011 -> 0.1965  d -0.0046  CI [-0.0092,-0.0010]
  runs  n= 597  bias +0.0410  no map fittable (173 fit rows < 200)

ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.

A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.

PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.

PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.

PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.

p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 15:30:30 -04:00

250 lines
10 KiB
JavaScript

#!/usr/bin/env node
'use strict';
/**
* settle-model-snapshots — pay the standing debt.
*
* 71,192 snapshot rows have never carried an outcome. They are the retention
* table built for exactly this kind of replay, and until they are settled every
* measurement in this programme runs on the far smaller ledger slice.
*
* ── OUTCOME IS SIDE-ALIGNED, NOT RAW ─────────────────────────────────────
* The order specifies `outcome = 1[realized > line]`. That is the OVER
* perspective, and it would be backwards for every under-side prop — `p_win` is
* side-aligned (verified: TB mean p_win 0.5698 against a 0.5074 side-won rate),
* so a raw over-indicator would silently invert the target on the under rows and
* make calibration measure the wrong thing.
*
* So: `actual_value` stores the realized stat (raw, unopinionated) and `outcome`
* stores whether the GRADED SIDE won. Deviation from the literal order, stated
* because it changes the number.
*
* ── INTEGRITY (hard-fail) ────────────────────────────────────────────────
* conservation settled + unresolvable + orphaned == candidates
* no dupes one write per snapshot id
* no orphans a settled row must have matched a real box score
* prediction-time logging captured_at must PRECEDE the game date; a row
* logged after the fact is not a prediction and is refused
*
* node scripts/settle-model-snapshots.js # dry run, verifies only
* SETTLE_WRITE=1 node scripts/settle-model-snapshots.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const axios = require('axios');
const { createClient } = require('@supabase/supabase-js');
const { nameKey } = require('../src/utils/playerName');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const WRITE = process.env.SETTLE_WRITE === '1';
const BOX_CACHE = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** Realized value per stat, from the box-score batting line. */
const FIELD = Object.freeze({
hits: (b) => knownNumber(b.hits),
total_bases: (b) => knownNumber(b.totalBases),
rbi: (b) => knownNumber(b.rbi),
runs: (b) => knownNumber(b.runs),
});
const get = async (url) => (await axios.get(url, { timeout: 45_000 })).data;
/** Eastern first pitch, conservatively. Anything at or after this is in-game. */
const FIRST_PITCH_ET_HOUR = 19;
/**
* Was this row logged BEFORE the games it grades?
*
* The pipeline runs on UTC cron hours, so a 01:00-UTC cycle is 21:00 the
* PREVIOUS evening in Eastern -- same game date, three hours into the slate.
*/
function isPreGame(capturedAt, gameDate) {
if (!capturedAt || !gameDate) return false;
const cap = new Date(capturedAt);
if (Number.isNaN(cap.getTime())) return false;
const et = new Date(cap.getTime() - 4 * 3600 * 1000); // EDT
const etDate = et.toISOString().slice(0, 10);
if (etDate < String(gameDate)) return true; // day before, fine
if (etDate > String(gameDate)) return false; // day after, post-game
return et.getUTCHours() < FIRST_PITCH_ET_HOUR;
}
async function pool(items, fn, n = 6) {
const out = []; let i = 0;
await Promise.all(Array.from({ length: n }, async () => {
while (i < items.length) {
const idx = i; i += 1;
try { out[idx] = await fn(items[idx]); } catch { out[idx] = null; }
}
}));
return out.filter(Boolean);
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
// STABLE ORDER. model_snapshots is a LIVE table -- the snapshot cron writes
// to it at 14/19/22/1/3 UTC -- and an unordered .range() walk over a table
// being appended to returns overlapping pages. The integrity gate caught
// exactly that on the first run.
const { data, error } = await apply(sb.from(table).select(select))
.order('id', { ascending: true })
.range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/** Box-score batting lines for a date range, cached. */
async function battingLines(dates) {
if (fs.existsSync(BOX_CACHE)) {
const c = JSON.parse(fs.readFileSync(BOX_CACHE, 'utf8'));
if (dates.every((d) => c.dates.includes(d))) return c.lines;
}
const games = [];
for (const d of dates) {
try {
const s = await get(`https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}`);
for (const day of s.dates || []) {
for (const g of day.games || []) {
if (String(g.status && g.status.detailedState) === 'Final') {
games.push({ pk: g.gamePk, date: g.officialDate || d });
}
}
}
} catch { /* absent day */ }
}
console.error(`[settle] ${games.length} final games across ${dates.length} dates`);
const lines = {};
const loaded = await pool(games, async (g) => {
const box = await get(`https://statsapi.mlb.com/api/v1/game/${g.pk}/boxscore`);
const out = [];
for (const side of ['home', 'away']) {
const t = box.teams[side];
if (!t) continue;
for (const id of t.batters || []) {
const pl = t.players[`ID${id}`];
const b = pl && pl.stats && pl.stats.batting;
if (!b || b.atBats == null) continue; // did not bat -> absent, not zero
out.push({
date: g.date,
key: nameKey(pl.person && pl.person.fullName),
name: pl.person && pl.person.fullName,
gamePk: g.pk,
hits: b.hits, totalBases: b.totalBases, rbi: b.rbi, runs: b.runs, atBats: b.atBats,
});
}
}
return out;
});
for (const arr of loaded) for (const r of arr) {
const k = `${r.date}|${r.key}`;
// A doubleheader gives two lines; sum them — the prop covers the day.
if (!lines[k]) lines[k] = { ...r, games: 1 };
else {
lines[k].hits += r.hits; lines[k].totalBases += r.totalBases;
lines[k].rbi += r.rbi; lines[k].runs += r.runs; lines[k].atBats += r.atBats;
lines[k].games += 1;
}
}
fs.mkdirSync(path.dirname(BOX_CACHE), { recursive: true });
fs.writeFileSync(BOX_CACHE, JSON.stringify({ dates, lines }));
return lines;
}
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, player_name, line, side, p_win, refused, outcome',
(q) => q.eq('sport', 'mlb').in('stat', STATS).is('outcome', null));
console.error(`[settle] ${snaps.length} unsettled snapshot rows`);
const dates = [...new Set(snaps.map((r) => r.game_date))].sort();
const lines = await battingLines(dates);
const counts = { candidates: snaps.length, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 };
const updates = [];
const seenIds = new Set();
for (const s of snaps) {
// Belt and braces: ordered pagination should make this impossible, and a
// duplicate would double-count a prediction in every downstream measurement.
if (seenIds.has(s.id)) throw new Error(`INTEGRITY: duplicate snapshot id ${s.id}`);
seenIds.add(s.id);
// A row logged after first pitch is not a prediction.
//
// Measured: cycles at ET 21:00/22:00/23:00 on the game date (10,738 rows)
// were captured DURING or AFTER the games they grade, and a further 664 the
// following morning. Games start ~19:05 ET, so the honest cutoff is ET
// first pitch on the game date -- not a UTC date compare, which both keeps
// post-game 01:00-UTC rows and discards legitimate pre-dawn ones.
if (!isPreGame(s.captured_at, s.game_date)) {
counts.post_hoc_logged += 1; counts.unresolvable += 1; continue;
}
const line = knownNumber(s.line);
if (line === null || !s.side) { counts.unresolvable += 1; continue; }
const b = lines[`${s.game_date}|${s.player_key}`];
if (!b) { counts.orphaned += 1; continue; }
const realized = FIELD[s.stat](b);
if (realized === null) { counts.unresolvable += 1; continue; }
// SIDE-ALIGNED, so it matches how p_win is expressed.
const over = realized > line;
const won = String(s.side).toLowerCase() === 'under' ? !over : over;
updates.push({ id: s.id, outcome: won ? 'hit' : 'miss', actual_value: realized });
counts.settled += 1;
}
// CONSERVATION — hard fail.
const acc = counts.settled + counts.unresolvable + counts.orphaned;
if (acc !== counts.candidates) {
throw new Error(`INTEGRITY: conservation violated ${acc} != ${counts.candidates}`);
}
// Hand-verifiable sample.
const sample = updates.slice(0, 12).map((u) => {
const s = snaps.find((x) => x.id === u.id);
return { player: s.player_name, date: s.game_date, stat: s.stat, line: s.line, side: s.side,
realized: u.actual_value, outcome: u.outcome };
});
if (WRITE) {
let written = 0;
for (let i = 0; i < updates.length; i += 500) {
const batch = updates.slice(i, i + 500);
const results = await Promise.all(batch.map((u) => sb.from('model_snapshots')
.update({ outcome: u.outcome, actual_value: u.actual_value, settled_at: new Date().toISOString(), settlement_source: 'statsapi_boxscore' })
.eq('id', u.id).is('outcome', null)));
written += results.filter((r) => !r.error).length;
}
counts.written = written;
}
console.log(JSON.stringify({
mode: WRITE ? 'WRITE' : 'DRY RUN',
counts,
dates_before: 'ledger-only slice',
snapshot_dates: dates.length,
date_span: [dates[0], dates[dates.length - 1]],
hand_verify_sample: sample,
note: 'outcome is SIDE-ALIGNED (matches p_win); actual_value holds the raw realized stat',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });