Robust bias established; low-parameter correction replaces isotonic

PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.

PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.

PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:

  hits         a=0.406 shrink 0.565  0.2626 -> 0.2540  CI [-0.0112,-0.0069]  DEPLOY
  total_bases  a=0.472 shrink 0.333  0.2490 -> 0.2429  CI [-0.0062,-0.0059]  DEPLOY
  rbi          a=0.775 shrink 0.231  0.2011 -> 0.2007  CI [-0.0007, 0]       REFUSE
  runs         a=-0.032                                                      REFUSE

A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.

STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.

PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.

PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.

p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-06 23:20:05 -04:00
parent ced40421ed
commit 74cf1ce974
8 changed files with 882 additions and 47 deletions
+137
View File
@@ -0,0 +1,137 @@
# The bias is robust; the map was not. Low-parameter correction deployed.
## PHASE 0 — the sample-limit truth, on record
**On 19 dates, BOTH stability instruments are underpowered. This is the SAMPLE,
not a fixable instrument.** No future order should re-open the gate-refinement
loop expecting a different answer at this N.
- **LODO power 0.0140.093** against a strong date-driven instability. Across
every k from 1.0 to 2.0, the best any stat reaches is 0.337.
- **Deploy CIs rest on 24 date clusters.** A cluster-robust interval at 2
clusters has ~1 degree of freedom and a near-undefined width.
Neither certifies forward stability of a specific map. Four orders refined a gate
the sample cannot support; that loop stops here.
**Record correction on runs:** its `1f40014` DATE-DRIVEN classification was an
artefact of the coin-flip ruler — 2 reversals in 3 drops never cleared a cutoff
of 2. runs is an ordinary "no fittable map" refusal. **Not date-driven.**
---
## PHASE 1 — the robust claim: ROBUST
Model-free, map-free, on the picked-side deduped population. Date-block bootstrap
(whole dates resampled, 5,000 draws).
### Pooled — the shape is textbook favourite-longshot
| bin | n | predicted | realized | over-prediction |
|---|---|---|---|---|
| 0.50.6 | 1,021 | 0.5477 | 0.5553 | 0.0076 |
| 0.60.7 | 961 | 0.6432 | 0.6004 | +0.0428 |
| 0.70.8 | 745 | 0.7420 | 0.6456 | +0.0963 |
| 0.80.9 | 459 | 0.8430 | 0.6841 | +0.1589 |
| **0.91.0** | 157 | 0.9075 | **0.6624** | **+0.2451** |
Pooled sign stability **0.9946** over 17 date blocks, 90% CI [+0.136, +0.300],
zero indeterminate resamples.
### Per stat — 4 of 4 replicate
| stat | n | >0.9 bias | date blocks | sign stability | 90% CI | replicates |
|---|---|---|---|---|---|---|
| hits | 1,140 | +0.2435 | 17 | 0.994 | [0.120, 0.321] | yes |
| total_bases | 1,050 | +0.2816 | 7 | 1.000 | [0.154, 0.394] | yes |
| rbi | 630 | +0.2107 | 5 | 1.000 | [0.156, 0.245] | yes |
| runs | 597 | +0.2367 | 5 | 0.998 | [0.082, 0.314] | yes |
**VERDICT: ROBUST** — pooled ≥95% and 4/4 stats (bar was 3/4).
Worth noting alongside it: **realized rate plateaus at ~0.650.68 from p=0.7
upward.** The 0.9+ bucket (0.6624) performs no better than the 0.80.9 bucket
(0.6841). The model has no genuinely high-confidence reads, only high-confidence
*numbers*.
---
## PHASE 3 — low-parameter correction, validated as a new estimator
Platt: `p_cal = sigmoid(a·logit(p) + b)`. Two parameters over the whole curve, so
it **cannot** encode "this Tuesday was odd" — which is precisely the failure mode
we cannot rule out for isotonic on this sample.
Shrunk toward identity by fit-date count: `w = D/(D+10)`, applied as
`w·p_platt + (1w)·p_raw`. A thin fit is therefore applied at reduced strength.
| stat | a | shrink | eval n | blocks | Brier raw | low-param | Δ vs raw | CI (date-block) | decision |
|---|---|---|---|---|---|---|---|---|---|
| **hits** | 0.406 | 0.565 | 765 | 4 | 0.2626 | 0.2540 | **0.0086** | [0.0112, 0.0069] | **DEPLOY** |
| **total_bases** | 0.472 | 0.333 | 625 | 2 | 0.2490 | 0.2429 | **0.0061** | [0.0062, 0.0059] | **DEPLOY** |
| rbi | 0.775 | 0.231 | 425 | 2 | 0.2011 | 0.2007 | 0.0004 | [0.0007, **0**] | REFUSE |
| runs | 0.032 | — | — | — | — | — | — | — | **REFUSE (slope)** |
### A guard the first run needed
runs fitted **a = 0.032**. A non-positive slope does not flatten an
over-confident forecaster — it **inverts** it, and near zero the curve collapses
to a constant, predicting the base rate for everything. That *lowers* Brier
(shrinking a miscalibrated forecaster toward its base rate always does) while
destroying all resolution, so it would have **scored as a win while making the
product worthless**. `MIN_SLOPE` now refuses it by name, with a test.
### Stated plainly: isotonic scored better, and we are not using it
On the identical held-out rows, isotonic beat the low-parameter fit on hits
(+0.0028, CI [0.0013, 0.0045]) and rbi (+0.0042, CI [0.0003, 0.0092]), and tied
on total_bases (0.0009, CI spanning zero).
**The swap is a capacity judgement, not a measurement.** The evaluation window
spans 24 date blocks, so "isotonic wins OOS" there is weak evidence, and it is
exactly what a flexible map would produce if it captured structure shared by the
fit and evaluation periods. That reasoning is a judgement and is labelled as one.
---
## PHASE 4 — deploy and labelling
| stat | served | basis |
|---|---|---|
| hits | low-parameter correction | `direction_robust_magnitude_provisional` |
| total_bases | low-parameter correction | `direction_robust_magnitude_provisional` |
| **rbi** | **RAW — withdrawn** | deployed on isotonic at `ced4042`; low-param does not beat raw |
| runs | RAW | slope refused |
The **direction** is bootstrap-robust; the **magnitude** is thin-sample and
conservatively shrunk (0.565 hits, 0.333 TB). Customer-facing letter unchanged.
Auto-demotion remains armed via `calibrationRegistry.reverify`: a sign flip in
the >0.9 bucket or a CI crossing zero demotes to raw and logs the breaking date.
Promotion to non-provisional stays at the original ≥40 date-cluster bar.
---
## PHASE 5 — the standing finding, stated hard
**Across 18 archetype slots on three stats, calibrated `p_win` separates within
archetype NO BETTER than raw. Every slot collapses to one band, indistinguishable
from its own base rate. Zero slots show lift.**
Per-archetype grade separation is **not coming from calibration**. It comes from
**proven factors or it does not exist.**
This reframes the roadmap. Calibration has now been pursued through five orders
and has delivered exactly what it can deliver — honest numbers on two stats — and
nothing at all on the question the grade product actually turns on. The next real
lever is factors on the stats that lack them.
---
## Invariants
`p_win` never mutated — the correction rides as `p_win_calibrated`. Calibration
consumed no Bonferroni slot. The robust-claim test ran before any calibrator was
built and could have terminated the session at Phase 2. Counter and frozen
clusters verified file-by-file (14 modules, including `calibration.js` and
`calibrationService.js`, both untouched and simply no longer on the serving path).