PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.
PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.
PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:
hits a=0.406 shrink 0.565 0.2626 -> 0.2540 CI [-0.0112,-0.0069] DEPLOY
total_bases a=0.472 shrink 0.333 0.2490 -> 0.2429 CI [-0.0062,-0.0059] DEPLOY
rbi a=0.775 shrink 0.231 0.2011 -> 0.2007 CI [-0.0007, 0] REFUSE
runs a=-0.032 REFUSE
A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.
STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.
PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.
PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.
p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
6.2 KiB
The bias is robust; the map was not. Low-parameter correction deployed.
PHASE 0 — the sample-limit truth, on record
On 19 dates, BOTH stability instruments are underpowered. This is the SAMPLE, not a fixable instrument. No future order should re-open the gate-refinement loop expecting a different answer at this N.
- LODO power 0.014–0.093 against a strong date-driven instability. Across every k from 1.0 to 2.0, the best any stat reaches is 0.337.
- Deploy CIs rest on 2–4 date clusters. A cluster-robust interval at 2 clusters has ~1 degree of freedom and a near-undefined width.
Neither certifies forward stability of a specific map. Four orders refined a gate the sample cannot support; that loop stops here.
Record correction on runs: its 1f40014 DATE-DRIVEN classification was an
artefact of the coin-flip ruler — 2 reversals in 3 drops never cleared a cutoff
of 2. runs is an ordinary "no fittable map" refusal. Not date-driven.
PHASE 1 — the robust claim: ROBUST
Model-free, map-free, on the picked-side deduped population. Date-block bootstrap (whole dates resampled, 5,000 draws).
Pooled — the shape is textbook favourite-longshot
| bin | n | predicted | realized | over-prediction |
|---|---|---|---|---|
| 0.5–0.6 | 1,021 | 0.5477 | 0.5553 | −0.0076 |
| 0.6–0.7 | 961 | 0.6432 | 0.6004 | +0.0428 |
| 0.7–0.8 | 745 | 0.7420 | 0.6456 | +0.0963 |
| 0.8–0.9 | 459 | 0.8430 | 0.6841 | +0.1589 |
| 0.9–1.0 | 157 | 0.9075 | 0.6624 | +0.2451 |
Pooled sign stability 0.9946 over 17 date blocks, 90% CI [+0.136, +0.300], zero indeterminate resamples.
Per stat — 4 of 4 replicate
| stat | n | >0.9 bias | date blocks | sign stability | 90% CI | replicates |
|---|---|---|---|---|---|---|
| hits | 1,140 | +0.2435 | 17 | 0.994 | [0.120, 0.321] | yes |
| total_bases | 1,050 | +0.2816 | 7 | 1.000 | [0.154, 0.394] | yes |
| rbi | 630 | +0.2107 | 5 | 1.000 | [0.156, 0.245] | yes |
| runs | 597 | +0.2367 | 5 | 0.998 | [0.082, 0.314] | yes |
VERDICT: ROBUST — pooled ≥95% and 4/4 stats (bar was 3/4).
Worth noting alongside it: realized rate plateaus at ~0.65–0.68 from p=0.7 upward. The 0.9+ bucket (0.6624) performs no better than the 0.8–0.9 bucket (0.6841). The model has no genuinely high-confidence reads, only high-confidence numbers.
PHASE 3 — low-parameter correction, validated as a new estimator
Platt: p_cal = sigmoid(a·logit(p) + b). Two parameters over the whole curve, so
it cannot encode "this Tuesday was odd" — which is precisely the failure mode
we cannot rule out for isotonic on this sample.
Shrunk toward identity by fit-date count: w = D/(D+10), applied as
w·p_platt + (1−w)·p_raw. A thin fit is therefore applied at reduced strength.
| stat | a | shrink | eval n | blocks | Brier raw | low-param | Δ vs raw | CI (date-block) | decision |
|---|---|---|---|---|---|---|---|---|---|
| hits | 0.406 | 0.565 | 765 | 4 | 0.2626 | 0.2540 | −0.0086 | [−0.0112, −0.0069] | DEPLOY |
| total_bases | 0.472 | 0.333 | 625 | 2 | 0.2490 | 0.2429 | −0.0061 | [−0.0062, −0.0059] | DEPLOY |
| rbi | 0.775 | 0.231 | 425 | 2 | 0.2011 | 0.2007 | −0.0004 | [−0.0007, 0] | REFUSE |
| runs | −0.032 | — | — | — | — | — | — | — | REFUSE (slope) |
A guard the first run needed
runs fitted a = −0.032. A non-positive slope does not flatten an
over-confident forecaster — it inverts it, and near zero the curve collapses
to a constant, predicting the base rate for everything. That lowers Brier
(shrinking a miscalibrated forecaster toward its base rate always does) while
destroying all resolution, so it would have scored as a win while making the
product worthless. MIN_SLOPE now refuses it by name, with a test.
Stated plainly: isotonic scored better, and we are not using it
On the identical held-out rows, isotonic beat the low-parameter fit on hits (+0.0028, CI [0.0013, 0.0045]) and rbi (+0.0042, CI [0.0003, 0.0092]), and tied on total_bases (−0.0009, CI spanning zero).
The swap is a capacity judgement, not a measurement. The evaluation window spans 2–4 date blocks, so "isotonic wins OOS" there is weak evidence, and it is exactly what a flexible map would produce if it captured structure shared by the fit and evaluation periods. That reasoning is a judgement and is labelled as one.
PHASE 4 — deploy and labelling
| stat | served | basis |
|---|---|---|
| hits | low-parameter correction | direction_robust_magnitude_provisional |
| total_bases | low-parameter correction | direction_robust_magnitude_provisional |
| rbi | RAW — withdrawn | deployed on isotonic at ced4042; low-param does not beat raw |
| runs | RAW | slope refused |
The direction is bootstrap-robust; the magnitude is thin-sample and conservatively shrunk (0.565 hits, 0.333 TB). Customer-facing letter unchanged.
Auto-demotion remains armed via calibrationRegistry.reverify: a sign flip in
the >0.9 bucket or a CI crossing zero demotes to raw and logs the breaking date.
Promotion to non-provisional stays at the original ≥40 date-cluster bar.
PHASE 5 — the standing finding, stated hard
Across 18 archetype slots on three stats, calibrated p_win separates within
archetype NO BETTER than raw. Every slot collapses to one band, indistinguishable
from its own base rate. Zero slots show lift.
Per-archetype grade separation is not coming from calibration. It comes from proven factors or it does not exist.
This reframes the roadmap. Calibration has now been pursued through five orders and has delivered exactly what it can deliver — honest numbers on two stats — and nothing at all on the question the grade product actually turns on. The next real lever is factors on the stats that lack them.
Invariants
p_win never mutated — the correction rides as p_win_calibrated. Calibration
consumed no Bonferroni slot. The robust-claim test ran before any calibrator was
built and could have terminated the session at Phase 2. Counter and frozen
clusters verified file-by-file (14 modules, including calibration.js and
calibrationService.js, both untouched and simply no longer on the serving path).