LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration layer without challenging the binding. That floor is a cluster-robust interval bar for a CAUSAL claim. Calibration makes no causal claim, has a bounded failure mode (it can only over- or under-shrink) and consumes no Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and leave-one-date-out tests that directly -- a STRICTER bar, since a cluster count cannot detect a single day carrying the effect. The >=40 floor is retained, correctly scoped as the PROMOTION bar. PHASE 1 — both guards codified, 11 tests, green before Phase 2. Demonstrated on live data: raw population violated=true, mean_p 0.4962, both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The null guard's test demonstrates the trap explicitly, since (null-1)**2 is 1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate. PHASE 2 — LODO: hits n=1140 dates=17 2 reversals (07-22 n=20, 07-26 n=25) FAIL total_bases n=1050 dates=7 0 reversals, 0 sign flips PASS rbi n= 630 dates=5 1 reversal (08-01 n=99) FAIL runs n= 597 dates=5 2 reversals (08-01 n=86, 08-05 n=244) FAIL Threshold sensitivity reported because the verdict moves: total_bases passes at every held-size threshold, runs fails at every one, and hits fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20 before seeing which stats passed and did not move it afterwards to preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a standard error several times the effect, so the instrument is underpowered per-drop -- an argument for pre-registering a higher threshold, which is a Roundtable call, not one to make while holding the results. PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and runs REFUSE. HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService hardcoded it since S91; it fails LODO, so it is out. A stat that cannot survive dropping one day was never calibrated, it was fitted to that day. The consequence is real -- hits props become unstackable for chain.chainAcross -- and it errs toward withdrawing a claim rather than preserving one on a fragile verdict. Deployment is now driven by a frozen, tested CALIBRATION_DEPLOYED set, not a hardcoded stat name. PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither waivable. reverify auto-demotes on the first breach (CI stops excluding zero, or the favourite bias flips sign) and logs the breaking date. Promotion needs the original >=40 bar. A provisional deploy that cannot be taken away is just a deploy. PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The two-bar rule still bites: calibrated YES, proven NO, so they stay a base-rate read, now honestly numbered. Every archetype still collapses to one band -- calibrated p_win separates within archetype no better than raw. PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI, and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did not survive); the refused set is a map of missing inputs; a low-parameter calibrator is queued unbuilt. p_win never mutated; calibration rides as p_win_calibrated with calibration_status provisional. No Bonferroni slot consumed. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
@@ -275,6 +275,16 @@ async function loadPitcherArsenals(sport) {
|
||||
} catch { return out; }
|
||||
}
|
||||
|
||||
/**
|
||||
* Stats whose calibration passed leave-one-date-out and may serve a calibrated
|
||||
* number. PROVISIONAL: auto-demoted the first time the held-out interval stops
|
||||
* excluding zero or the favourite over-prediction flips sign.
|
||||
*
|
||||
* hits / rbi / runs are deliberately ABSENT — each fails LODO. See
|
||||
* specs/lodo-provisional-calibration.md.
|
||||
*/
|
||||
const CALIBRATION_DEPLOYED = Object.freeze(['total_bases']);
|
||||
|
||||
async function runSnapshot(sport, opts = {}) {
|
||||
const sp = String(sport || '').toLowerCase();
|
||||
const deps = {
|
||||
@@ -700,37 +710,51 @@ async function runSnapshot(sport, opts = {}) {
|
||||
console.warn(`[challenger] ${sp} skipped:`, e.message);
|
||||
}
|
||||
|
||||
// ── FORWARD CALIBRATION (hits) ────────────────────────────────────────
|
||||
// ── FORWARD CALIBRATION (LODO-gated, per stat) ────────────────────────
|
||||
// Fitted on games that are OVER, applied to tonight's props. `p_win` is NOT
|
||||
// touched — the counter stays byte-identical and the calibrated value rides
|
||||
// beside it, because a calibration map is a correction TO a forecast, not a
|
||||
// different forecast.
|
||||
//
|
||||
// WHICH STATS SERVE IS MEASURED, NOT ASSUMED. The deploy bar is leave-one-
|
||||
// date-out stability: refit dropping each settled date in turn, and the
|
||||
// improvement must never reverse. That is the right instrument for a monotone
|
||||
// shrink-to-observed layer — the factor gate's >=40 date-cluster interval
|
||||
// floor was built for a CAUSAL claim and does not bind here.
|
||||
//
|
||||
// Measured 2026-08-07: total_bases passes at every held-size threshold. hits
|
||||
// FAILS (reverses on 2026-07-22 and 2026-07-26), so it is no longer served
|
||||
// calibrated even though it was — a stat that cannot survive dropping one day
|
||||
// was never calibrated, it was fitted to that day. rbi and runs also fail.
|
||||
//
|
||||
// `calibrated` is true only inside a band certified out-of-sample, and it is
|
||||
// what `chain.chainAcross` requires before it will compound anything. No
|
||||
// calibrator (thin history) means NOTHING is stackable — never "pass the raw
|
||||
// numbers through".
|
||||
// what `chain.chainAcross` requires before it will compound anything. Removing
|
||||
// hits here makes hits props unstackable again, which is the honest
|
||||
// consequence of the measurement rather than a regression to work around.
|
||||
if (sp === 'mlb') {
|
||||
try {
|
||||
const calSvc = deps.calibrationService || require('./model/calibrationService');
|
||||
const sbc = require('../utils/supabase').getSupabaseServiceClient();
|
||||
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat: 'hits' }) : null;
|
||||
if (calibrator) {
|
||||
let marked = 0;
|
||||
for (const g of enriched) {
|
||||
if (String(g.stat_type || g.stat || '').toLowerCase() !== 'hits') continue;
|
||||
const out = calibrator.calibrate(g.p_win);
|
||||
g.p_win_calibrated = out.p_calibrated;
|
||||
g.calibrated = out.calibrated;
|
||||
g.calibration_reason = out.reason;
|
||||
if (out.calibrated) marked += 1;
|
||||
for (const stat of CALIBRATION_DEPLOYED) {
|
||||
try {
|
||||
const calSvc = deps.calibrationService || require('./model/calibrationService');
|
||||
const sbc = require('../utils/supabase').getSupabaseServiceClient();
|
||||
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat }) : null;
|
||||
if (calibrator) {
|
||||
let marked = 0;
|
||||
for (const g of enriched) {
|
||||
if (String(g.stat_type || g.stat || '').toLowerCase() !== stat) continue;
|
||||
const out = calibrator.calibrate(g.p_win);
|
||||
g.p_win_calibrated = out.p_calibrated;
|
||||
g.calibrated = out.calibrated;
|
||||
g.calibration_reason = out.reason;
|
||||
g.calibration_status = 'provisional';
|
||||
if (out.calibrated) marked += 1;
|
||||
}
|
||||
console.log(`[calibration] ${sp} ${stat} (PROVISIONAL) — ${marked} stackable; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}`);
|
||||
} else {
|
||||
console.log(`[calibration] ${sp} ${stat} — no calibrator (thin history); nothing is stackable`);
|
||||
}
|
||||
console.log(`[calibration] ${sp} hits — ${marked} stackable of ${enriched.filter((g) => String(g.stat_type || g.stat || '').toLowerCase() === 'hits').length}; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}, bands ${JSON.stringify(calibrator.bands.map((b) => [b.lo, b.hi]))}`);
|
||||
} else {
|
||||
console.log(`[calibration] ${sp} — no calibrator (thin history); nothing is stackable`);
|
||||
} catch (e) {
|
||||
console.warn(`[calibration] ${stat} skipped:`, e.message);
|
||||
}
|
||||
} catch (e) {
|
||||
console.warn('[calibration] skipped:', e.message);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -843,5 +867,6 @@ module.exports = {
|
||||
generateTickerEvents,
|
||||
pushTickerItems,
|
||||
ACTIVE_SPORTS,
|
||||
CALIBRATION_DEPLOYED,
|
||||
__internals: { propKey, gradedAtFor, indexOdds, lastName, isTopGrade, DELTA_NOISE, DELTA_MOVE, TICKER_CAP },
|
||||
};
|
||||
|
||||
Reference in New Issue
Block a user