LODO-gated provisional calibration: total_bases deploys, hits withdrawn

PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration
layer without challenging the binding. That floor is a cluster-robust
interval bar for a CAUSAL claim. Calibration makes no causal claim, has a
bounded failure mode (it can only over- or under-shrink) and consumes no
Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and
leave-one-date-out tests that directly -- a STRICTER bar, since a cluster
count cannot detect a single day carrying the effect. The >=40 floor is
retained, correctly scoped as the PROMOTION bar.

PHASE 1 — both guards codified, 11 tests, green before Phase 2.
Demonstrated on live data: raw population violated=true, mean_p 0.4962,
both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The
null guard's test demonstrates the trap explicitly, since (null-1)**2 is
1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate.

PHASE 2 — LODO:

  hits         n=1140 dates=17  2 reversals (07-22 n=20, 07-26 n=25)  FAIL
  total_bases  n=1050 dates=7   0 reversals, 0 sign flips             PASS
  rbi          n= 630 dates=5   1 reversal  (08-01 n=99)              FAIL
  runs         n= 597 dates=5   2 reversals (08-01 n=86, 08-05 n=244) FAIL

Threshold sensitivity reported because the verdict moves: total_bases
passes at every held-size threshold, runs fails at every one, and hits
fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20
before seeing which stats passed and did not move it afterwards to
preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a
standard error several times the effect, so the instrument is
underpowered per-drop -- an argument for pre-registering a higher
threshold, which is a Roundtable call, not one to make while holding the
results.

PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and
runs REFUSE.

HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService
hardcoded it since S91; it fails LODO, so it is out. A stat that cannot
survive dropping one day was never calibrated, it was fitted to that day.
The consequence is real -- hits props become unstackable for
chain.chainAcross -- and it errs toward withdrawing a claim rather than
preserving one on a fragile verdict. Deployment is now driven by a frozen,
tested CALIBRATION_DEPLOYED set, not a hardcoded stat name.

PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither
waivable. reverify auto-demotes on the first breach (CI stops excluding
zero, or the favourite bias flips sign) and logs the breaking date.
Promotion needs the original >=40 bar. A provisional deploy that cannot be
taken away is just a deploy.

PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The
two-bar rule still bites: calibrated YES, proven NO, so they stay a
base-rate read, now honestly numbered. Every archetype still collapses to
one band -- calibrated p_win separates within archetype no better than raw.

PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI,
and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did
not survive); the refused set is a map of missing inputs; a low-parameter
calibrator is queued unbuilt.

p_win never mutated; calibration rides as p_win_calibrated with
calibration_status provisional. No Bonferroni slot consumed. Counter and
frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-06 18:31:19 -04:00
parent f976df47b8
commit 6ae11f1193
10 changed files with 1064 additions and 23 deletions
+47 -22
View File
@@ -275,6 +275,16 @@ async function loadPitcherArsenals(sport) {
} catch { return out; }
}
/**
* Stats whose calibration passed leave-one-date-out and may serve a calibrated
* number. PROVISIONAL: auto-demoted the first time the held-out interval stops
* excluding zero or the favourite over-prediction flips sign.
*
* hits / rbi / runs are deliberately ABSENT — each fails LODO. See
* specs/lodo-provisional-calibration.md.
*/
const CALIBRATION_DEPLOYED = Object.freeze(['total_bases']);
async function runSnapshot(sport, opts = {}) {
const sp = String(sport || '').toLowerCase();
const deps = {
@@ -700,37 +710,51 @@ async function runSnapshot(sport, opts = {}) {
console.warn(`[challenger] ${sp} skipped:`, e.message);
}
// ── FORWARD CALIBRATION (hits) ────────────────────────────────────────
// ── FORWARD CALIBRATION (LODO-gated, per stat) ────────────────────────
// Fitted on games that are OVER, applied to tonight's props. `p_win` is NOT
// touched — the counter stays byte-identical and the calibrated value rides
// beside it, because a calibration map is a correction TO a forecast, not a
// different forecast.
//
// WHICH STATS SERVE IS MEASURED, NOT ASSUMED. The deploy bar is leave-one-
// date-out stability: refit dropping each settled date in turn, and the
// improvement must never reverse. That is the right instrument for a monotone
// shrink-to-observed layer — the factor gate's >=40 date-cluster interval
// floor was built for a CAUSAL claim and does not bind here.
//
// Measured 2026-08-07: total_bases passes at every held-size threshold. hits
// FAILS (reverses on 2026-07-22 and 2026-07-26), so it is no longer served
// calibrated even though it was — a stat that cannot survive dropping one day
// was never calibrated, it was fitted to that day. rbi and runs also fail.
//
// `calibrated` is true only inside a band certified out-of-sample, and it is
// what `chain.chainAcross` requires before it will compound anything. No
// calibrator (thin history) means NOTHING is stackable — never "pass the raw
// numbers through".
// what `chain.chainAcross` requires before it will compound anything. Removing
// hits here makes hits props unstackable again, which is the honest
// consequence of the measurement rather than a regression to work around.
if (sp === 'mlb') {
try {
const calSvc = deps.calibrationService || require('./model/calibrationService');
const sbc = require('../utils/supabase').getSupabaseServiceClient();
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat: 'hits' }) : null;
if (calibrator) {
let marked = 0;
for (const g of enriched) {
if (String(g.stat_type || g.stat || '').toLowerCase() !== 'hits') continue;
const out = calibrator.calibrate(g.p_win);
g.p_win_calibrated = out.p_calibrated;
g.calibrated = out.calibrated;
g.calibration_reason = out.reason;
if (out.calibrated) marked += 1;
for (const stat of CALIBRATION_DEPLOYED) {
try {
const calSvc = deps.calibrationService || require('./model/calibrationService');
const sbc = require('../utils/supabase').getSupabaseServiceClient();
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat }) : null;
if (calibrator) {
let marked = 0;
for (const g of enriched) {
if (String(g.stat_type || g.stat || '').toLowerCase() !== stat) continue;
const out = calibrator.calibrate(g.p_win);
g.p_win_calibrated = out.p_calibrated;
g.calibrated = out.calibrated;
g.calibration_reason = out.reason;
g.calibration_status = 'provisional';
if (out.calibrated) marked += 1;
}
console.log(`[calibration] ${sp} ${stat} (PROVISIONAL) — ${marked} stackable; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}`);
} else {
console.log(`[calibration] ${sp} ${stat} — no calibrator (thin history); nothing is stackable`);
}
console.log(`[calibration] ${sp} hits — ${marked} stackable of ${enriched.filter((g) => String(g.stat_type || g.stat || '').toLowerCase() === 'hits').length}; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}, bands ${JSON.stringify(calibrator.bands.map((b) => [b.lo, b.hi]))}`);
} else {
console.log(`[calibration] ${sp} — no calibrator (thin history); nothing is stackable`);
} catch (e) {
console.warn(`[calibration] ${stat} skipped:`, e.message);
}
} catch (e) {
console.warn('[calibration] skipped:', e.message);
}
}
@@ -843,5 +867,6 @@ module.exports = {
generateTickerEvents,
pushTickerItems,
ACTIVE_SPORTS,
CALIBRATION_DEPLOYED,
__internals: { propKey, gradedAtFor, indexOdds, lastName, isTopGrade, DELTA_NOISE, DELTA_MOVE, TICKER_CAP },
};