f976df47b8
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.
PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.
PHASE 1, two integrity findings both caught by the gate:
1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').
2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.
THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.
THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.
PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):
hits n=1140 bias +0.0868 brier 0.2626 -> 0.2511 d -0.0115 CI [-0.0139,-0.0097]
TB n=1050 bias +0.0834 brier 0.2490 -> 0.2438 d -0.0052 CI [-0.0061,-0.0045]
rbi n= 630 bias +0.0164 brier 0.2011 -> 0.1965 d -0.0046 CI [-0.0092,-0.0010]
runs n= 597 bias +0.0410 no map fittable (173 fit rows < 200)
ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.
A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.
PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.
PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.
PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.
p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
268 lines
11 KiB
JavaScript
268 lines
11 KiB
JavaScript
#!/usr/bin/env node
|
|
'use strict';
|
|
|
|
/**
|
|
* calibrate-four-stats — Phases 2, 3, 4 and 7.
|
|
*
|
|
* ── ONE SIDE PER PROP, OR THE MEASUREMENT IS MEANINGLESS ─────────────────
|
|
* 97.6% of snapshot props carry BOTH the over and the under. Their p_wins sum to
|
|
* ~1 and their outcomes are complementary, so any calibration statistic over the
|
|
* raw population is pinned to 0.5 by symmetry. Measured that way the counter
|
|
* looks perfectly calibrated (+0.0002 on hits); deduped to the model-PICKED side
|
|
* it is +0.0868. Same rows, opposite conclusion.
|
|
*
|
|
* ── DATE-CLUSTERED, PER THE ORDER ────────────────────────────────────────
|
|
* A day's offensive environment is a real shared component, so uncertainty is
|
|
* clustered on the game DATE rather than the game. That is the honest unit for a
|
|
* systematic-bias claim and it is a much harder bar than game-clustering.
|
|
*
|
|
* SUPABASE_URL=... node scripts/calibrate-four-stats.js
|
|
*/
|
|
|
|
require('dotenv').config();
|
|
const fs = require('fs');
|
|
const path = require('path');
|
|
const { createClient } = require('@supabase/supabase-js');
|
|
const cal = require('../src/services/model/calibration');
|
|
const { knownNumber } = require('../src/utils/known');
|
|
|
|
const SB_URL = process.env.SUPABASE_URL;
|
|
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
|
|
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
|
|
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
|
|
const PAGE = 1000;
|
|
/** The order's deploy floor: dates, not games. */
|
|
const MIN_DATE_CLUSTERS = 40;
|
|
|
|
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
|
|
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
|
|
const brier = (ps, ys) => mean(ps.map((p, i) => (p - ys[i]) ** 2));
|
|
|
|
async function page(sb, table, select, apply) {
|
|
const out = [];
|
|
for (let from = 0; ; from += PAGE) {
|
|
const { data, error } = await apply(sb.from(table).select(select))
|
|
.order('id', { ascending: true }).range(from, from + PAGE - 1);
|
|
if (error) throw error;
|
|
if (!data || data.length === 0) break;
|
|
out.push(...data);
|
|
if (data.length < PAGE) break;
|
|
}
|
|
return out;
|
|
}
|
|
|
|
const isPreGame = (capturedAt, gameDate) => {
|
|
const et = new Date(new Date(capturedAt).getTime() - 4 * 3600 * 1000);
|
|
const d = et.toISOString().slice(0, 10);
|
|
return d < gameDate || (d === gameDate && et.getUTCHours() < 19);
|
|
};
|
|
|
|
function makeRnd(seed) {
|
|
let s = seed >>> 0;
|
|
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
|
|
}
|
|
|
|
/** Paired bootstrap on the Brier difference, resampling DATES. */
|
|
function dateClusteredCI(rows, cumulativeTests = 1, iters = 3000) {
|
|
const byDate = new Map();
|
|
for (const r of rows) {
|
|
if (!byDate.has(r.date)) byDate.set(r.date, []);
|
|
byDate.get(r.date).push(r);
|
|
}
|
|
const keys = [...byDate.keys()];
|
|
const rnd = makeRnd(20260807);
|
|
const diffs = [];
|
|
for (let it = 0; it < iters; it += 1) {
|
|
const raw = []; const adj = []; const ys = [];
|
|
for (let i = 0; i < keys.length; i += 1) {
|
|
for (const r of byDate.get(keys[Math.floor(rnd() * keys.length)])) {
|
|
raw.push(r.p); adj.push(r.pc); ys.push(r.won);
|
|
}
|
|
}
|
|
diffs.push(brier(adj, ys) - brier(raw, ys));
|
|
}
|
|
diffs.sort((a, b) => a - b);
|
|
const tests = Math.max(1, Math.round(cumulativeTests));
|
|
const alpha = 0.05 / tests;
|
|
const q = (x) => diffs[Math.floor(Math.min(diffs.length - 1, Math.max(0, x * (diffs.length - 1))))];
|
|
return { ci: [round4(q(alpha / 2)), round4(q(1 - alpha / 2))], date_clusters: keys.length, ci_level: round4(1 - alpha) };
|
|
}
|
|
|
|
async function main() {
|
|
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
|
|
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
|
|
|
|
const snaps = await page(sb, 'model_snapshots',
|
|
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused, grade',
|
|
(q) => q.eq('sport', 'mlb').in('stat', STATS));
|
|
|
|
// ── ONE SIDE PER PROP: the side the model picked (its higher p_win). ──
|
|
const picked = new Map();
|
|
const refusedProps = new Map();
|
|
for (const r of snaps) {
|
|
if (!isPreGame(r.captured_at, r.game_date)) continue;
|
|
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
|
|
if (r.refused || knownNumber(r.p_win) === null) {
|
|
if (!refusedProps.has(k)) refusedProps.set(k, r);
|
|
continue;
|
|
}
|
|
const prev = picked.get(k);
|
|
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
|
|
}
|
|
|
|
const resolve = (r) => {
|
|
const b = lines[`${r.game_date}|${r.player_key}`];
|
|
const L = knownNumber(r.line);
|
|
if (!b || L === null || !r.side) return null;
|
|
const v = knownNumber(FIELD[r.stat](b));
|
|
if (v === null) return null;
|
|
const over = v > L;
|
|
return { over, won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0, realized: v };
|
|
};
|
|
|
|
const out = { deploy_floor_date_clusters: MIN_DATE_CLUSTERS, per_stat: {}, refusal_accuracy: {} };
|
|
|
|
for (const stat of STATS) {
|
|
const rows = [];
|
|
for (const r of picked.values()) {
|
|
if (r.stat !== stat) continue;
|
|
const res = resolve(r);
|
|
if (!res) continue;
|
|
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: res.won });
|
|
}
|
|
rows.sort((a, b) => String(a.date).localeCompare(String(b.date)));
|
|
const dates = [...new Set(rows.map((r) => r.date))].sort();
|
|
|
|
if (rows.length < 100 || dates.length < 3) {
|
|
out.per_stat[stat] = { n: rows.length, date_clusters: dates.length, decision: 'REFUSE', reason: 'too few rows or dates to split point-in-time' };
|
|
continue;
|
|
}
|
|
|
|
// POINT-IN-TIME: fit strictly on earlier dates, evaluate on later ones.
|
|
//
|
|
// The cut is placed by ROW COUNT rather than by date index. Props are not
|
|
// spread evenly across dates -- hits concentrate in the later ones -- so a
|
|
// 60%-of-DATES cut left only 143 rows to fit on, under the 200 the fitter
|
|
// needs. Splitting on cumulative rows keeps the split strictly temporal
|
|
// (every fit date precedes every eval date) while giving both sides enough
|
|
// to work with.
|
|
const perDate = new Map();
|
|
for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
|
|
let acc = 0; let cut = dates[dates.length - 1];
|
|
for (const d of dates) {
|
|
acc += perDate.get(d) || 0;
|
|
if (acc >= rows.length * 0.45) { cut = d; break; }
|
|
}
|
|
const fit = rows.filter((r) => r.date < cut);
|
|
const ev = rows.filter((r) => r.date >= cut);
|
|
if (fit.length < 50 || ev.length < 50) {
|
|
out.per_stat[stat] = { n: rows.length, date_clusters: dates.length, decision: 'REFUSE', reason: 'time split leaves too little on one side' };
|
|
continue;
|
|
}
|
|
|
|
const iso = cal.fitIsotonic(fit.map((r) => ({ p: r.p, won: r.won })));
|
|
// NULL IS NOT A PREDICTION. fitIsotonic returns null below its minimum and
|
|
// applyIsotonic then returns null per row -- and (null - 1)**2 === 1 while
|
|
// (null - 0)**2 === 0, so a "Brier score" computed over nulls is silently
|
|
// just the win rate. That is exactly the Number(null) === 0 breach this
|
|
// codebase keeps having to catch, and it produced a fake 0.5567 for hits.
|
|
if (!iso) {
|
|
out.per_stat[stat] = {
|
|
n: rows.length, date_clusters: dates.length, fit_n: fit.length, eval_n: ev.length,
|
|
decision: 'REFUSE', reason: `no calibration map could be fitted on ${fit.length} fit rows`,
|
|
};
|
|
continue;
|
|
}
|
|
const scored = ev.map((r) => ({ ...r, pc: cal.applyIsotonic(iso, r.p) }))
|
|
.filter((r) => knownNumber(r.pc) !== null);
|
|
if (scored.length < 50) {
|
|
out.per_stat[stat] = {
|
|
n: rows.length, date_clusters: dates.length,
|
|
decision: 'REFUSE', reason: `only ${scored.length} eval rows could be mapped`,
|
|
};
|
|
continue;
|
|
}
|
|
const ys = scored.map((r) => r.won);
|
|
const bRaw = brier(scored.map((r) => r.p), ys);
|
|
const bCal = brier(scored.map((r) => r.pc), ys);
|
|
const { ci, date_clusters, ci_level } = dateClusteredCI(scored, 1);
|
|
|
|
// CERTIFIED BAND: p_win deciles where held-out |predicted - actual| is small.
|
|
const bands = [];
|
|
for (let lo = 0.3; lo < 0.95; lo += 0.1) {
|
|
const slice = scored.filter((r) => r.p >= lo && r.p < lo + 0.1);
|
|
if (slice.length < 25) continue;
|
|
const pred = mean(slice.map((r) => r.pc));
|
|
const act = mean(slice.map((r) => r.won));
|
|
bands.push({ range: [round2(lo), round2(lo + 0.1)], n: slice.length, calibrated_pred: round4(pred), actual: round4(act), err: round4(pred - act) });
|
|
}
|
|
const certified = bands.filter((b) => Math.abs(b.err) <= 0.05).map((b) => b.range);
|
|
|
|
const improves = bCal < bRaw && ci[1] < 0;
|
|
const enoughDates = date_clusters >= MIN_DATE_CLUSTERS;
|
|
|
|
// PHASE 4 — bias SHAPE across the p_win range (diagnostic only).
|
|
const shape = [];
|
|
for (let lo = 0.3; lo < 0.95; lo += 0.1) {
|
|
const slice = rows.filter((r) => r.p >= lo && r.p < lo + 0.1);
|
|
if (slice.length < 25) continue;
|
|
shape.push({ range: [round2(lo), round2(lo + 0.1)], n: slice.length, bias: round4(mean(slice.map((r) => r.p)) - mean(slice.map((r) => r.won))) });
|
|
}
|
|
|
|
out.per_stat[stat] = {
|
|
n: rows.length,
|
|
date_clusters: dates.length,
|
|
bias_pre: round4(mean(rows.map((r) => r.p)) - mean(rows.map((r) => r.won))),
|
|
fit_n: fit.length, eval_n: ev.length, split_at: cut,
|
|
brier_raw: round4(bRaw),
|
|
brier_calibrated: round4(bCal),
|
|
brier_delta: round4(bCal - bRaw),
|
|
ci_date_clustered: ci,
|
|
ci_level,
|
|
eval_date_clusters: date_clusters,
|
|
certified_bands: certified,
|
|
band_detail: bands,
|
|
bias_shape: shape,
|
|
decision: improves && enoughDates ? 'DEPLOY' : 'REFUSE',
|
|
reason: improves && enoughDates ? 'held-out Brier improves, date-clustered, and the date floor is met'
|
|
: (!enoughDates
|
|
? `date-clusters ${date_clusters} < ${MIN_DATE_CLUSTERS} — the honest unit for a systematic-bias claim`
|
|
: 'held-out Brier does not improve at the date-clustered interval'),
|
|
};
|
|
}
|
|
|
|
// ── PHASE 7 — REFUSAL ACCURACY ──
|
|
// The model passed on these. A pass is CORRECT when there was genuinely
|
|
// nothing to call: the over lands near a coin flip rather than at an
|
|
// exploitable rate.
|
|
for (const stat of STATS) {
|
|
const refs = [];
|
|
for (const r of refusedProps.values()) {
|
|
if (r.stat !== stat) continue;
|
|
const res = resolve({ ...r, side: 'over' });
|
|
if (res) refs.push(res.over ? 1 : 0);
|
|
}
|
|
const graded = [];
|
|
for (const r of picked.values()) {
|
|
if (r.stat !== stat) continue;
|
|
const res = resolve({ ...r, side: 'over' });
|
|
if (res) graded.push(res.over ? 1 : 0);
|
|
}
|
|
out.refusal_accuracy[stat] = {
|
|
refused_n: refs.length,
|
|
refused_over_rate: refs.length ? round4(mean(refs)) : null,
|
|
graded_over_rate: graded.length ? round4(mean(graded)) : null,
|
|
refused_distance_from_coinflip: refs.length ? round4(Math.abs(mean(refs) - 0.5)) : null,
|
|
graded_distance_from_coinflip: graded.length ? round4(Math.abs(mean(graded) - 0.5)) : null,
|
|
};
|
|
}
|
|
|
|
console.log(JSON.stringify(out, null, 2));
|
|
process.exit(0);
|
|
}
|
|
|
|
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
|
|
const round2 = (v) => Math.round(v * 100) / 100;
|
|
|
|
main().catch((e) => { console.error(e); process.exit(1); });
|