Robust bias established; low-parameter correction replaces isotonic

PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.

PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.

PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:

  hits         a=0.406 shrink 0.565  0.2626 -> 0.2540  CI [-0.0112,-0.0069]  DEPLOY
  total_bases  a=0.472 shrink 0.333  0.2490 -> 0.2429  CI [-0.0062,-0.0059]  DEPLOY
  rbi          a=0.775 shrink 0.231  0.2011 -> 0.2007  CI [-0.0007, 0]       REFUSE
  runs         a=-0.032                                                      REFUSE

A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.

STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.

PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.

PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.

p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-06 23:20:05 -04:00
parent ced40421ed
commit 74cf1ce974
8 changed files with 882 additions and 47 deletions
+27 -26
View File
@@ -276,36 +276,37 @@ async function loadPitcherArsenals(sport) {
}
/**
* Stats that may serve a calibrated number, and ON WHAT BASIS.
* Stats that may serve a calibrated number, and on what basis.
*
* ── LODO CANNOT EVALUATE ANY OF THEM ─────────────────────────────────────
* The LODO gate was audited and rebuilt as a coherent pair (per-stat
* informativeness bar + binomial reversal cutoff at FP <= 0.05). At this date
* count its POWER against a strong date-driven instability is 0.093 / 0.093 /
* 0.045 / 0.014 — it would miss a real failure more than nine times in ten. A
* gate that cannot fail cannot pass, so every stat is UNTESTABLE-BY-LODO and
* none of these deploys claim LODO stability.
* ── THE BIAS IS ROBUST; THE MAP WAS NOT CERTIFIABLE ──────────────────────
* Tested model-free and map-free on the picked-side population: the model
* over-predicts its own favourites, and the sign survives 99.5% of date-block
* resamples pooled and replicates in 4 of 4 stats. Over-prediction rises
* monotonically from -0.008 near p=0.55 to +0.245 above 0.9.
*
* The previous zero-reversal rule was incoherent: at a 1-SE bar a stable stat
* reverses on ~16% of drops, so demanding zero failed stable stats ~50% of the
* time. Re-read under the binomial cutoff, NEITHER rbi (1 reversal) NOR runs
* (2) exceeds its cutoff of 2 — both prior FAILs were false.
* What could NOT be certified on 19 dates is the stability of a specific
* isotonic MAP — LODO has 1.4-9.3% power there. So isotonic is retired and the
* correction is a TWO-PARAMETER Platt curve, which has no capacity to encode a
* single odd day, shrunk toward the raw forecast by fit-date count.
*
* ── SO THE DEPLOY BASIS IS THE DATE-CLUSTERED CI ALONE ───────────────────
* hits, total_bases and rbi each have a point-in-time held-out interval
* excluding zero. That is the ONLY support they have, and it is thin — the
* interval rests on 4, 2 and 2 date clusters respectively. Auto-demotion is
* therefore the sole stability guard, not a backstop to a passed test.
* Validated as a NEW estimator against RAW, date-block bootstrap:
* hits a=0.406 shrink 0.565 0.2626 -> 0.2540 CI [-0.0112,-0.0069]
* total_bases a=0.472 shrink 0.333 0.2490 -> 0.2429 CI [-0.0062,-0.0059]
*
* runs is absent: no isotonic map was fittable at its point-in-time split, so it
* has no CI support to stand on either.
* rbi is WITHDRAWN (deployed last order on isotonic): the low-parameter fit does
* not beat raw, CI [-0.0007, 0] touching zero. runs is refused by the slope
* guard — it fitted a = -0.032, which would invert the forecast rather than
* flatten it. Both now serve RAW.
*/
const CALIBRATION_DEPLOYED = Object.freeze(['hits', 'total_bases']);
/**
* The DIRECTION of the correction is bootstrap-robust; its MAGNITUDE is fitted
* on few dates and deliberately shrunk toward identity. The customer-facing
* letter is unchanged; this is what the internal record says.
*/
const CALIBRATION_DEPLOYED = Object.freeze(['hits', 'total_bases', 'rbi']);
/** Why each deployed stat is allowed to serve. Not one of them is LODO-stable. */
const CALIBRATION_BASIS = Object.freeze({
hits: 'ci_only_lodo_untestable',
total_bases: 'ci_only_lodo_untestable',
rbi: 'ci_only_lodo_untestable',
hits: 'direction_robust_magnitude_provisional',
total_bases: 'direction_robust_magnitude_provisional',
});
async function runSnapshot(sport, opts = {}) {
@@ -757,7 +758,7 @@ async function runSnapshot(sport, opts = {}) {
if (sp === 'mlb') {
for (const stat of CALIBRATION_DEPLOYED) {
try {
const calSvc = deps.calibrationService || require('./model/calibrationService');
const calSvc = deps.calibrationService || require('./model/lowParamService');
const sbc = require('../utils/supabase').getSupabaseServiceClient();
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat }) : null;
if (calibrator) {
@@ -772,7 +773,7 @@ async function runSnapshot(sport, opts = {}) {
g.calibration_basis = CALIBRATION_BASIS[stat] || null;
if (out.calibrated) marked += 1;
}
console.log(`[calibration] ${sp} ${stat} (PROVISIONAL) — ${marked} stackable; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}`);
console.log(`[calibration] ${sp} ${stat} (low-param, PROVISIONAL) — ${marked} stackable; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}, a=${calibrator.model.a} shrink=${calibrator.shrinkage}`);
} else {
console.log(`[calibration] ${sp} ${stat} — no calibrator (thin history); nothing is stackable`);
}