Robust bias established; low-parameter correction replaces isotonic

PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.

PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.

PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:

  hits         a=0.406 shrink 0.565  0.2626 -> 0.2540  CI [-0.0112,-0.0069]  DEPLOY
  total_bases  a=0.472 shrink 0.333  0.2490 -> 0.2429  CI [-0.0062,-0.0059]  DEPLOY
  rbi          a=0.775 shrink 0.231  0.2011 -> 0.2007  CI [-0.0007, 0]       REFUSE
  runs         a=-0.032                                                      REFUSE

A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.

STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.

PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.

PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.

p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-06 23:20:05 -04:00
parent ced40421ed
commit 74cf1ce974
8 changed files with 882 additions and 47 deletions
+102
View File
@@ -0,0 +1,102 @@
'use strict';
/**
* lowParamService — the production side of the two-parameter correction.
*
* Mirrors calibrationService's interface so the serving path swaps cleanly, but
* fits a Platt curve instead of an isotonic map. The reason for the swap is
* capacity, not score: on 19 dates we cannot certify the stability of a map with
* one free parameter per prediction level, and LODO turned out to have 1.4-9.3%
* power to tell us otherwise. Two parameters cannot encode "this Tuesday was
* odd", which is exactly the failure we cannot rule out for isotonic.
*
* Stated plainly because it is a judgement rather than a measurement: on the
* held-out window isotonic scored BETTER than this on hits (+0.0028) and rbi
* (+0.0042) and tied on total_bases. That window spans 2-4 date blocks, so it is
* weak evidence either way, and it is consistent with a flexible map having
* captured structure shared by fit and evaluation periods.
*
* Same point-in-time cut as before: fitted ONLY on games that are already over.
*/
const lp = require('./lowParamCalibrator');
const cal = require('./calibration');
const { knownNumber } = require('../../utils/known');
const MIN_FIT = 200;
const HOLDOUT_FRACTION = 0.35;
/** Build from settled rows: fit on the older part, certify bands on the newer. */
function build(rows, opts = {}) {
const clean = (rows || [])
.map((r) => ({ p: knownNumber(r.p), won: knownNumber(r.won), date: String(r.date || '') }))
.filter((r) => r.p !== null && (r.won === 0 || r.won === 1))
.sort((a, b) => a.date.localeCompare(b.date));
if (clean.length < (opts.minFit ?? MIN_FIT)) return null;
const cut = Math.floor(clean.length * (1 - (opts.holdout ?? HOLDOUT_FRACTION)));
const fitRows = clean.slice(0, cut);
const certRows = clean.slice(cut);
if (fitRows.length < (opts.minFit ?? MIN_FIT) || certRows.length < 50) return null;
const model = lp.fitPlatt(fitRows, opts);
// A refused fit (inverting or collapsed slope) yields no calibrator at all.
if (!model || model.refused) return null;
const corrected = certRows
.map((r) => ({ ...r, p: lp.applyPlatt(model, r.p) }))
.filter((r) => knownNumber(r.p) !== null);
const bands = cal.certifyBands(corrected, {
tolerance: opts.tolerance ?? 0.05,
minBin: opts.minBin ?? 40,
});
return {
model,
bands,
fit_n: fitRows.length,
certify_n: certRows.length,
fitted_through: fitRows[fitRows.length - 1].date,
shrinkage: model.shrinkage,
calibrate(p) {
const raw = knownNumber(p);
if (raw === null) return { p_raw: null, p_calibrated: null, calibrated: false, reason: 'absent' };
const c = lp.applyPlatt(model, raw);
if (c === null) return { p_raw: raw, p_calibrated: null, calibrated: false, reason: 'no_model_value' };
const inBand = cal.inCertifiedBand(bands, c);
return {
p_raw: raw,
p_calibrated: Math.round(c * 1000) / 1000,
calibrated: inBand,
reason: inBand ? null : 'outside_certified_band',
};
},
};
}
/** Load settled history and build, POINT-IN-TIME (strictly before today). */
async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) {
if (!sb) return null;
const cutoff = before || new Intl.DateTimeFormat('en-CA', {
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
}).format(new Date());
const rows = [];
for (let from = 0; ; from += 1000) {
const { data, error } = await sb.from('ledger_entries')
.select('p_win, outcome, game_date, quarantine_reason')
.eq('sport', sport).is('user_id', null).eq('stat', stat)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null)
.lt('game_date', cutoff)
.range(from, from + 999);
if (error || !data || data.length === 0) break;
rows.push(...data);
if (data.length < 1000) break;
}
const clean = rows
.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'))
.map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, date: String(r.game_date) }));
const built = build(clean, opts);
return built ? { ...built, cutoff } : null;
}
module.exports = { build, fromLedger, MIN_FIT, HOLDOUT_FRACTION };