Instrument the calibration duel forward; diagnose the resolution ceiling
— the proven factors were never wired in
PHASE 0 — two truths recorded. The swap is a BET, not an OOS win:
isotonic beat low-param on identical held-out rows (hits +0.0028, rbi
+0.0042, TB tied) and we serve low-param anyway on an untestable prior
about shared daily structure. At 19 dates nothing here can test it. And
the MIN_SLOPE catch is preserved as standing rationale: a near-zero or
negative slope collapses toward base-rate-for-everything, which LOWERS
Brier while destroying all resolution -- a metric win that guts the
product.
PHASE 1 — the duel is now falsifiable. Both corrections computed on every
hits/TB prop; p_win_lowparam served, p_win_isotonic_shadow logged in its
own try so it can never break serving. calibrationDuel.adjudicate encodes
the rule IN CODE before any forward date exists: >=10 forward dates and
isotonic winning with a date-block CI excluding zero => REFUTED, revert;
otherwise UPHELD; under 10 dates PENDING regardless of the numbers. A
date counts as forward only if NEITHER map was fitted on it -- otherwise
we would be scoring which map memorised better. Nothing swaps now.
PHASE 2 — the ceiling, quantified via Murphy decomposition:
stat reliability RESOLUTION uncertainty variance explained
hits 0.01353 0.00252 0.24532 1.03%
TB 0.01419 0.00442 0.24329 1.82%
rbi 0.00654 0.03268 0.22531 14.51%
runs 0.00788 0.00130 0.23182 0.56%
Calibration did exactly what theory says and nothing more: hits
reliability 0.01353 -> 0.00233 (-0.0112, 83% of the error removed) while
resolution moved -0.0002. Unexpected: rbi has 13x the resolution of hits
and is the one stat we do NOT serve corrected -- it needs calibration
least and discriminates most.
PHASE 2 DIAGNOSIS — NOT-TRANSMITTED, and not weak, ABSENT. Traced in code:
sprayDefense.js and platoonSeverity.js are required by NOTHING in src/,
only by analysis scripts and their own tests. The served p_win
(intelligence/probabilityEstimator.js:54) reads exactly four inputs --
game-log frequency, opp_rank_stat +/-0.03, home_away +/-0.015, and a cv
pull -- with zero occurrences of spray, platoon, hard-hit or
contact-profile. And snapshotService grades at line 454 while computing
challenger/context at 640+, so everything proven is computed DOWNSTREAM of
the grade it would inform. The three proven hits factors have never once
moved a served number.
That reframes the recent nulls: "calibrated p_win does not separate within
archetype" was never a statement about factors. The factors were not in
the forecast.
PHASE 3 — bands rebuilt on SERVED values (hits/TB low-param, rbi/runs
raw): 28 archetype slots across four stats, ZERO show lift. No longer an
open shrug -- it is the arithmetic of resolution 0.0013-0.0327 against
uncertainty ~0.23. A forecast explaining 1% of variance cannot produce
separating bands, and no correction to its numbers will change that.
HEADLINE: calibration is complete, delivered honest numbers on two stats
and zero grade separation, because the counter has no resolution -- and
the proven factors are not wired into the forecast at all. The second is
the reason for the first, and it is plumbing rather than a modelling wall.
Per-archetype grades need proven factors that actually reach p_win. Last
calibration order.
Serving unchanged from 74cf1ce. p_win never mutated. No Bonferroni slot.
Counter and frozen clusters verified file-by-file (15 modules).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
@@ -734,46 +734,65 @@ async function runSnapshot(sport, opts = {}) {
|
||||
console.warn(`[challenger] ${sp} skipped:`, e.message);
|
||||
}
|
||||
|
||||
// ── FORWARD CALIBRATION (LODO-gated, per stat) ────────────────────────
|
||||
// ── FORWARD CALIBRATION + THE SHADOW DUEL ─────────────────────────────
|
||||
// Fitted on games that are OVER, applied to tonight's props. `p_win` is NOT
|
||||
// touched — the counter stays byte-identical and the calibrated value rides
|
||||
// touched — the counter stays byte-identical and the corrected value rides
|
||||
// beside it, because a calibration map is a correction TO a forecast, not a
|
||||
// different forecast.
|
||||
//
|
||||
// WHICH STATS SERVE IS MEASURED, NOT ASSUMED. The deploy bar is leave-one-
|
||||
// date-out stability: refit dropping each settled date in turn, and the
|
||||
// improvement must never reverse. That is the right instrument for a monotone
|
||||
// shrink-to-observed layer — the factor gate's >=40 date-cluster interval
|
||||
// floor was built for a CAUSAL claim and does not bind here.
|
||||
// WHAT IS SERVED, AND WHY IT IS A BET RATHER THAN A RESULT. On identical
|
||||
// held-out rows the ISOTONIC map scored BETTER than the low-parameter one
|
||||
// (hits +0.0028, rbi +0.0042, total_bases tied). We serve the low-parameter
|
||||
// map anyway, on the argument that isotonic's in-window edge is daily
|
||||
// structure shared between the fit and evaluation windows and will not
|
||||
// transmit forward. At 19 dates no instrument here can test that argument —
|
||||
// LODO has 1.4-9.3% power — so it is a BET, not evidence.
|
||||
//
|
||||
// Measured 2026-08-07: total_bases passes at every held-size threshold. hits
|
||||
// FAILS (reverses on 2026-07-22 and 2026-07-26), so it is no longer served
|
||||
// calibrated even though it was — a stat that cannot survive dropping one day
|
||||
// was never calibrated, it was fitted to that day. rbi and runs also fail.
|
||||
// So both are computed on every prop and the shadow is logged. Real
|
||||
// out-of-window dates adjudicate it:
|
||||
//
|
||||
// `calibrated` is true only inside a band certified out-of-sample, and it is
|
||||
// what `chain.chainAcross` requires before it will compound anything. Removing
|
||||
// hits here makes hits props unstackable again, which is the honest
|
||||
// consequence of the measurement rather than a regression to work around.
|
||||
// PRE-REGISTERED: once >=10 forward dates have settled that NEITHER map was
|
||||
// fitted on, if isotonic beats low-param with a date-block bootstrap CI
|
||||
// excluding zero, the capacity argument is REFUTED and hits/TB revert to
|
||||
// isotonic. If low-param wins or ties, the bet was right. The season
|
||||
// decides, not the argument.
|
||||
//
|
||||
// Serving is unchanged until that bar is met. `calibrated` is true only inside
|
||||
// a band certified out-of-sample, and it is what `chain.chainAcross` requires
|
||||
// before it will compound anything.
|
||||
if (sp === 'mlb') {
|
||||
for (const stat of CALIBRATION_DEPLOYED) {
|
||||
try {
|
||||
const calSvc = deps.calibrationService || require('./model/lowParamService');
|
||||
const shadowSvc = deps.shadowCalibrationService || require('./model/calibrationService');
|
||||
const sbc = require('../utils/supabase').getSupabaseServiceClient();
|
||||
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat }) : null;
|
||||
// The shadow must never break serving: its own try, and a null shadow
|
||||
// simply means the duel has no entry for tonight.
|
||||
let shadow = null;
|
||||
try { shadow = sbc ? await shadowSvc.fromLedger(sbc, { sport: 'mlb', stat }) : null; } catch { shadow = null; }
|
||||
|
||||
if (calibrator) {
|
||||
let marked = 0;
|
||||
let marked = 0; let shadowed = 0;
|
||||
for (const g of enriched) {
|
||||
if (String(g.stat_type || g.stat || '').toLowerCase() !== stat) continue;
|
||||
const out = calibrator.calibrate(g.p_win);
|
||||
g.p_win_calibrated = out.p_calibrated;
|
||||
g.p_win_lowparam = out.p_calibrated; // named, so the duel is legible
|
||||
g.calibrated = out.calibrated;
|
||||
g.calibration_reason = out.reason;
|
||||
g.calibration_status = 'provisional';
|
||||
g.calibration_basis = CALIBRATION_BASIS[stat] || null;
|
||||
g.calibration_basis = CALIBRATION_BASIS[stat] || null;
|
||||
if (out.calibrated) marked += 1;
|
||||
if (shadow) {
|
||||
const sh = shadow.calibrate(g.p_win);
|
||||
// SHADOW ONLY. Never read by serving, never by chainAcross.
|
||||
g.p_win_isotonic_shadow = sh.p_calibrated;
|
||||
g.calibration_duel_fitted_through = shadow.fitted_through || null;
|
||||
if (sh.p_calibrated != null) shadowed += 1;
|
||||
}
|
||||
}
|
||||
console.log(`[calibration] ${sp} ${stat} (low-param, PROVISIONAL) — ${marked} stackable; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}, a=${calibrator.model.a} shrink=${calibrator.shrinkage}`);
|
||||
console.log(`[calibration] ${sp} ${stat} (low-param, PROVISIONAL) — ${marked} stackable; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}, a=${calibrator.model.a} shrink=${calibrator.shrinkage}; shadow logged on ${shadowed}`);
|
||||
} else {
|
||||
console.log(`[calibration] ${sp} ${stat} — no calibrator (thin history); nothing is stackable`);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user