Files
vyndr/specs/robust-bias-lowparam.md
builtbykev 74cf1ce974 Robust bias established; low-parameter correction replaces isotonic
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.

PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.

PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:

  hits         a=0.406 shrink 0.565  0.2626 -> 0.2540  CI [-0.0112,-0.0069]  DEPLOY
  total_bases  a=0.472 shrink 0.333  0.2490 -> 0.2429  CI [-0.0062,-0.0059]  DEPLOY
  rbi          a=0.775 shrink 0.231  0.2011 -> 0.2007  CI [-0.0007, 0]       REFUSE
  runs         a=-0.032                                                      REFUSE

A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.

STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.

PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.

PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.

p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 23:20:05 -04:00

6.2 KiB
Raw Permalink Blame History

The bias is robust; the map was not. Low-parameter correction deployed.

PHASE 0 — the sample-limit truth, on record

On 19 dates, BOTH stability instruments are underpowered. This is the SAMPLE, not a fixable instrument. No future order should re-open the gate-refinement loop expecting a different answer at this N.

  • LODO power 0.0140.093 against a strong date-driven instability. Across every k from 1.0 to 2.0, the best any stat reaches is 0.337.
  • Deploy CIs rest on 24 date clusters. A cluster-robust interval at 2 clusters has ~1 degree of freedom and a near-undefined width.

Neither certifies forward stability of a specific map. Four orders refined a gate the sample cannot support; that loop stops here.

Record correction on runs: its 1f40014 DATE-DRIVEN classification was an artefact of the coin-flip ruler — 2 reversals in 3 drops never cleared a cutoff of 2. runs is an ordinary "no fittable map" refusal. Not date-driven.


PHASE 1 — the robust claim: ROBUST

Model-free, map-free, on the picked-side deduped population. Date-block bootstrap (whole dates resampled, 5,000 draws).

Pooled — the shape is textbook favourite-longshot

bin n predicted realized over-prediction
0.50.6 1,021 0.5477 0.5553 0.0076
0.60.7 961 0.6432 0.6004 +0.0428
0.70.8 745 0.7420 0.6456 +0.0963
0.80.9 459 0.8430 0.6841 +0.1589
0.91.0 157 0.9075 0.6624 +0.2451

Pooled sign stability 0.9946 over 17 date blocks, 90% CI [+0.136, +0.300], zero indeterminate resamples.

Per stat — 4 of 4 replicate

stat n >0.9 bias date blocks sign stability 90% CI replicates
hits 1,140 +0.2435 17 0.994 [0.120, 0.321] yes
total_bases 1,050 +0.2816 7 1.000 [0.154, 0.394] yes
rbi 630 +0.2107 5 1.000 [0.156, 0.245] yes
runs 597 +0.2367 5 0.998 [0.082, 0.314] yes

VERDICT: ROBUST — pooled ≥95% and 4/4 stats (bar was 3/4).

Worth noting alongside it: realized rate plateaus at ~0.650.68 from p=0.7 upward. The 0.9+ bucket (0.6624) performs no better than the 0.80.9 bucket (0.6841). The model has no genuinely high-confidence reads, only high-confidence numbers.


PHASE 3 — low-parameter correction, validated as a new estimator

Platt: p_cal = sigmoid(a·logit(p) + b). Two parameters over the whole curve, so it cannot encode "this Tuesday was odd" — which is precisely the failure mode we cannot rule out for isotonic on this sample.

Shrunk toward identity by fit-date count: w = D/(D+10), applied as w·p_platt + (1w)·p_raw. A thin fit is therefore applied at reduced strength.

stat a shrink eval n blocks Brier raw low-param Δ vs raw CI (date-block) decision
hits 0.406 0.565 765 4 0.2626 0.2540 0.0086 [0.0112, 0.0069] DEPLOY
total_bases 0.472 0.333 625 2 0.2490 0.2429 0.0061 [0.0062, 0.0059] DEPLOY
rbi 0.775 0.231 425 2 0.2011 0.2007 0.0004 [0.0007, 0] REFUSE
runs 0.032 REFUSE (slope)

A guard the first run needed

runs fitted a = 0.032. A non-positive slope does not flatten an over-confident forecaster — it inverts it, and near zero the curve collapses to a constant, predicting the base rate for everything. That lowers Brier (shrinking a miscalibrated forecaster toward its base rate always does) while destroying all resolution, so it would have scored as a win while making the product worthless. MIN_SLOPE now refuses it by name, with a test.

Stated plainly: isotonic scored better, and we are not using it

On the identical held-out rows, isotonic beat the low-parameter fit on hits (+0.0028, CI [0.0013, 0.0045]) and rbi (+0.0042, CI [0.0003, 0.0092]), and tied on total_bases (0.0009, CI spanning zero).

The swap is a capacity judgement, not a measurement. The evaluation window spans 24 date blocks, so "isotonic wins OOS" there is weak evidence, and it is exactly what a flexible map would produce if it captured structure shared by the fit and evaluation periods. That reasoning is a judgement and is labelled as one.


PHASE 4 — deploy and labelling

stat served basis
hits low-parameter correction direction_robust_magnitude_provisional
total_bases low-parameter correction direction_robust_magnitude_provisional
rbi RAW — withdrawn deployed on isotonic at ced4042; low-param does not beat raw
runs RAW slope refused

The direction is bootstrap-robust; the magnitude is thin-sample and conservatively shrunk (0.565 hits, 0.333 TB). Customer-facing letter unchanged.

Auto-demotion remains armed via calibrationRegistry.reverify: a sign flip in the >0.9 bucket or a CI crossing zero demotes to raw and logs the breaking date. Promotion to non-provisional stays at the original ≥40 date-cluster bar.


PHASE 5 — the standing finding, stated hard

Across 18 archetype slots on three stats, calibrated p_win separates within archetype NO BETTER than raw. Every slot collapses to one band, indistinguishable from its own base rate. Zero slots show lift.

Per-archetype grade separation is not coming from calibration. It comes from proven factors or it does not exist.

This reframes the roadmap. Calibration has now been pursued through five orders and has delivered exactly what it can deliver — honest numbers on two stats — and nothing at all on the question the grade product actually turns on. The next real lever is factors on the stats that lack them.


Invariants

p_win never mutated — the correction rides as p_win_calibrated. Calibration consumed no Bonferroni slot. The robust-claim test ran before any calibrator was built and could have terminated the session at Phase 2. Counter and frozen clusters verified file-by-file (14 modules, including calibration.js and calibrationService.js, both untouched and simply no longer on the serving path).