Files
vyndr/specs/robust-bias-lowparam.md
builtbykev 74cf1ce974 Robust bias established; low-parameter correction replaces isotonic
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.

PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.

PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:

  hits         a=0.406 shrink 0.565  0.2626 -> 0.2540  CI [-0.0112,-0.0069]  DEPLOY
  total_bases  a=0.472 shrink 0.333  0.2490 -> 0.2429  CI [-0.0062,-0.0059]  DEPLOY
  rbi          a=0.775 shrink 0.231  0.2011 -> 0.2007  CI [-0.0007, 0]       REFUSE
  runs         a=-0.032                                                      REFUSE

A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.

STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.

PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.

PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.

p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 23:20:05 -04:00

138 lines
6.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The bias is robust; the map was not. Low-parameter correction deployed.
## PHASE 0 — the sample-limit truth, on record
**On 19 dates, BOTH stability instruments are underpowered. This is the SAMPLE,
not a fixable instrument.** No future order should re-open the gate-refinement
loop expecting a different answer at this N.
- **LODO power 0.0140.093** against a strong date-driven instability. Across
every k from 1.0 to 2.0, the best any stat reaches is 0.337.
- **Deploy CIs rest on 24 date clusters.** A cluster-robust interval at 2
clusters has ~1 degree of freedom and a near-undefined width.
Neither certifies forward stability of a specific map. Four orders refined a gate
the sample cannot support; that loop stops here.
**Record correction on runs:** its `1f40014` DATE-DRIVEN classification was an
artefact of the coin-flip ruler — 2 reversals in 3 drops never cleared a cutoff
of 2. runs is an ordinary "no fittable map" refusal. **Not date-driven.**
---
## PHASE 1 — the robust claim: ROBUST
Model-free, map-free, on the picked-side deduped population. Date-block bootstrap
(whole dates resampled, 5,000 draws).
### Pooled — the shape is textbook favourite-longshot
| bin | n | predicted | realized | over-prediction |
|---|---|---|---|---|
| 0.50.6 | 1,021 | 0.5477 | 0.5553 | 0.0076 |
| 0.60.7 | 961 | 0.6432 | 0.6004 | +0.0428 |
| 0.70.8 | 745 | 0.7420 | 0.6456 | +0.0963 |
| 0.80.9 | 459 | 0.8430 | 0.6841 | +0.1589 |
| **0.91.0** | 157 | 0.9075 | **0.6624** | **+0.2451** |
Pooled sign stability **0.9946** over 17 date blocks, 90% CI [+0.136, +0.300],
zero indeterminate resamples.
### Per stat — 4 of 4 replicate
| stat | n | >0.9 bias | date blocks | sign stability | 90% CI | replicates |
|---|---|---|---|---|---|---|
| hits | 1,140 | +0.2435 | 17 | 0.994 | [0.120, 0.321] | yes |
| total_bases | 1,050 | +0.2816 | 7 | 1.000 | [0.154, 0.394] | yes |
| rbi | 630 | +0.2107 | 5 | 1.000 | [0.156, 0.245] | yes |
| runs | 597 | +0.2367 | 5 | 0.998 | [0.082, 0.314] | yes |
**VERDICT: ROBUST** — pooled ≥95% and 4/4 stats (bar was 3/4).
Worth noting alongside it: **realized rate plateaus at ~0.650.68 from p=0.7
upward.** The 0.9+ bucket (0.6624) performs no better than the 0.80.9 bucket
(0.6841). The model has no genuinely high-confidence reads, only high-confidence
*numbers*.
---
## PHASE 3 — low-parameter correction, validated as a new estimator
Platt: `p_cal = sigmoid(a·logit(p) + b)`. Two parameters over the whole curve, so
it **cannot** encode "this Tuesday was odd" — which is precisely the failure mode
we cannot rule out for isotonic on this sample.
Shrunk toward identity by fit-date count: `w = D/(D+10)`, applied as
`w·p_platt + (1w)·p_raw`. A thin fit is therefore applied at reduced strength.
| stat | a | shrink | eval n | blocks | Brier raw | low-param | Δ vs raw | CI (date-block) | decision |
|---|---|---|---|---|---|---|---|---|---|
| **hits** | 0.406 | 0.565 | 765 | 4 | 0.2626 | 0.2540 | **0.0086** | [0.0112, 0.0069] | **DEPLOY** |
| **total_bases** | 0.472 | 0.333 | 625 | 2 | 0.2490 | 0.2429 | **0.0061** | [0.0062, 0.0059] | **DEPLOY** |
| rbi | 0.775 | 0.231 | 425 | 2 | 0.2011 | 0.2007 | 0.0004 | [0.0007, **0**] | REFUSE |
| runs | 0.032 | — | — | — | — | — | — | — | **REFUSE (slope)** |
### A guard the first run needed
runs fitted **a = 0.032**. A non-positive slope does not flatten an
over-confident forecaster — it **inverts** it, and near zero the curve collapses
to a constant, predicting the base rate for everything. That *lowers* Brier
(shrinking a miscalibrated forecaster toward its base rate always does) while
destroying all resolution, so it would have **scored as a win while making the
product worthless**. `MIN_SLOPE` now refuses it by name, with a test.
### Stated plainly: isotonic scored better, and we are not using it
On the identical held-out rows, isotonic beat the low-parameter fit on hits
(+0.0028, CI [0.0013, 0.0045]) and rbi (+0.0042, CI [0.0003, 0.0092]), and tied
on total_bases (0.0009, CI spanning zero).
**The swap is a capacity judgement, not a measurement.** The evaluation window
spans 24 date blocks, so "isotonic wins OOS" there is weak evidence, and it is
exactly what a flexible map would produce if it captured structure shared by the
fit and evaluation periods. That reasoning is a judgement and is labelled as one.
---
## PHASE 4 — deploy and labelling
| stat | served | basis |
|---|---|---|
| hits | low-parameter correction | `direction_robust_magnitude_provisional` |
| total_bases | low-parameter correction | `direction_robust_magnitude_provisional` |
| **rbi** | **RAW — withdrawn** | deployed on isotonic at `ced4042`; low-param does not beat raw |
| runs | RAW | slope refused |
The **direction** is bootstrap-robust; the **magnitude** is thin-sample and
conservatively shrunk (0.565 hits, 0.333 TB). Customer-facing letter unchanged.
Auto-demotion remains armed via `calibrationRegistry.reverify`: a sign flip in
the >0.9 bucket or a CI crossing zero demotes to raw and logs the breaking date.
Promotion to non-provisional stays at the original ≥40 date-cluster bar.
---
## PHASE 5 — the standing finding, stated hard
**Across 18 archetype slots on three stats, calibrated `p_win` separates within
archetype NO BETTER than raw. Every slot collapses to one band, indistinguishable
from its own base rate. Zero slots show lift.**
Per-archetype grade separation is **not coming from calibration**. It comes from
**proven factors or it does not exist.**
This reframes the roadmap. Calibration has now been pursued through five orders
and has delivered exactly what it can deliver — honest numbers on two stats — and
nothing at all on the question the grade product actually turns on. The next real
lever is factors on the stats that lack them.
---
## Invariants
`p_win` never mutated — the correction rides as `p_win_calibrated`. Calibration
consumed no Bonferroni slot. The robust-claim test ran before any calibrator was
built and could have terminated the session at Phase 2. Counter and frozen
clusters verified file-by-file (14 modules, including `calibration.js` and
`calibrationService.js`, both untouched and simply no longer on the serving path).