1f40014256
routed as date-driven PHASE 0 — threshold derived BLIND, before any stat was re-read. A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count. Per-row Brier difference d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single stat's verdict could shape the threshold deciding it: pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175 n* = (0.09816/0.01175)^2 = 69.8 -> 70 The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect this removes, and why the previous verdict moved with the number. Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it equals the constant, so it cannot drift from its own justification. The derivation script prints no stat verdict, no date and no reversal. PHASE 1 — LODO at n*, applied cold: hits 5 informative drops, 0 reversals PASS total_bases 4 informative drops, 0 reversals PASS rbi reverses 2026-08-01 (n=99) FAIL runs reverses 08-01 (n=86), 08-05 (244) FAIL hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across 123-272 row dates, favourite sign holding on every testable drop. THIS IS THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the withdrawal at6ae11f1was correct on the instrument available then, which admitted 20- and 25-row dates as evidence. Nothing about hits changed; the threshold stopped being chosen. PHASE 2 — both failures are DATE-DRIVEN, not underpowered. Every reversal sits above n*=70 (99, 86, 244), so no threshold and no further accrual rescues either: isotonic is fitting day-structure. Routed to the low-parameter calibrator queue (Platt/beta), not built here. PHASE 3 — CALIBRATION_DEPLOYED is now ['hits','total_bases'], frozen and tested, both PROVISIONAL with auto-demotion armed and the >=40 date-cluster promotion bar unchanged. hits stackability for chain.chainAcross is RESTORED, and the record shows it returned through the powered gate rather than by fiat. hits bands rebuilt on p_win_calibrated (765 eval rows): every archetype still one band, still base_rate -- calibrated YES, proven-per-archetype NO. PHASE 4 logged: the deploy set is now set by a power-derived, pre-committed, tested constant rather than an operator-chosen number. At6ae11f1that rule moved the live path AGAINST the operator; it has now moved it back on the same evidence because the instrument changed. Both directions are the rule working. And calibrated p_win separates within archetype no better than raw across 13 archetype slots on two deployed stats -- per-archetype separation will come from proven factors or not at all. p_win never mutated; no Bonferroni slot consumed; counter and frozen clusters verified byte-identical file by file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
147 lines
5.8 KiB
Markdown
147 lines
5.8 KiB
Markdown
# The LODO threshold, derived from power — hits restored, rbi/runs routed
|
||
|
||
## PHASE 0 — the threshold, derived blind
|
||
|
||
**Estimand:** does dropping date D reverse the SIGN of the out-of-sample Brier
|
||
improvement on D's held-out rows? A reversal is informative only if that date's
|
||
Brier delta is distinguishable from zero at its row count.
|
||
|
||
The per-row Brier difference is `d_i = (pc_i − y_i)² − (p_i − y_i)²`, so a date's
|
||
delta is `mean(d)` and `SE(n) = SD(d)/√n`. The smallest n at which a typical
|
||
effect clears one standard error is `n* = (SD(d)/|effect|)²`.
|
||
|
||
Pooled across all four stats — deliberately, so no single stat's verdict could
|
||
shape the threshold that decides it:
|
||
|
||
```
|
||
pooled rows 3,417
|
||
SD(per-row Brier diff) 0.09816
|
||
|pooled effect| 0.01175
|
||
n* = (0.09816 / 0.01175)^2 = 69.8 -> 70
|
||
```
|
||
|
||
### SE-vs-n
|
||
|
||
| n | SE | effect / SE | informative |
|
||
|---|---|---|---|
|
||
| 10 | 0.0310 | 0.38 | no |
|
||
| **20** (previously chosen by hand) | 0.0220 | **0.54** | **no** |
|
||
| 25 | 0.0196 | 0.60 | no |
|
||
| 30 | 0.0179 | 0.66 | no |
|
||
| 50 | 0.0139 | 0.85 | no |
|
||
| **70 (n\*)** | 0.0117 | **1.00** | **yes** |
|
||
| 100 | 0.0098 | 1.20 | yes |
|
||
| 244 | 0.0063 | 1.87 | yes |
|
||
|
||
The hand-chosen 20 sat at 0.54 SE — a coin flip. That is the defect this
|
||
derivation removes, and it is why the previous verdict moved with the number.
|
||
|
||
**Committed as `calibrationRegistry.LODO_MIN_HELD_ROWS = 70`** with
|
||
`LODO_THRESHOLD_BASIS` recording the inputs. A test recomputes `(SD/effect)²` and
|
||
asserts it equals the constant, so the value cannot drift from the basis that
|
||
justifies it, and cannot be silently tuned. The derivation script prints no stat
|
||
verdict, no date and no reversal; it ran and the constant was committed before
|
||
any stat was re-read.
|
||
|
||
---
|
||
|
||
## PHASE 1 — LODO at n\*, applied cold
|
||
|
||
| stat | n | dates | informative drops | reversals | LODO |
|
||
|---|---|---|---|---|---|
|
||
| **hits** | 1,140 | 17 | 5 | **0** | **PASS** |
|
||
| **total_bases** | 1,050 | 7 | 4 | **0** | **PASS** |
|
||
| rbi | 630 | 5 | 4 | 1 — 2026-08-01 (**n=99**) | **FAIL** |
|
||
| runs | 597 | 5 | 3 | 2 — 2026-08-01 (**n=86**), 2026-08-05 (**n=244**) | **FAIL** |
|
||
|
||
hits' held-out deltas at n\*: −0.0041 / −0.0080 / −0.0192 / −0.0140 / −0.0139
|
||
across 123–272 row dates. Every drop holds, and the favourite over-prediction
|
||
holds sign on every drop where it is testable (+0.52 / +0.318 / +0.272).
|
||
|
||
**This is the instrument finally being powered, not vindication of a prediction.**
|
||
The withdrawal at `6ae11f1` was correct on the instrument available then, which
|
||
admitted 20- and 25-row dates as evidence. Nothing about hits changed; what
|
||
changed is that the threshold is now derived rather than chosen.
|
||
|
||
---
|
||
|
||
## PHASE 2 — Failure classification
|
||
|
||
Both failures are **DATE-DRIVEN**, not underpowered-per-drop:
|
||
|
||
| stat | deciding date | held n | vs n\*=70 | classification |
|
||
|---|---|---|---|---|
|
||
| rbi | 2026-08-01 | 99 | **above** | DATE-DRIVEN |
|
||
| runs | 2026-08-01 | 86 | **above** | DATE-DRIVEN |
|
||
| runs | 2026-08-05 | 244 | **far above** | DATE-DRIVEN |
|
||
|
||
Every reversal sits comfortably above the powered threshold, so **no threshold
|
||
choice and no further date accrual rescues either stat.** Isotonic is fitting
|
||
day-structure on both.
|
||
|
||
**Routed to the low-parameter calibrator queue** (Platt / beta), which fits a
|
||
favourite-longshot shape on far fewer free parameters and is therefore much
|
||
harder to bend to one day. Not built here — it is a new estimator and needs its
|
||
own out-of-sample validation.
|
||
|
||
---
|
||
|
||
## PHASE 3 — Deploy and bands
|
||
|
||
| stat | status | certified band |
|
||
|---|---|---|
|
||
| **hits** | **DEPLOY-PROVISIONAL (restored)** | [0.5–0.7] |
|
||
| **total_bases** | DEPLOY-PROVISIONAL (unchanged from 6ae11f1) | [0.6–0.8] |
|
||
| rbi | REFUSE — date-driven | — |
|
||
| runs | REFUSE — date-driven | — |
|
||
|
||
`CALIBRATION_DEPLOYED` is now `['hits', 'total_bases']`, frozen and tested. Both
|
||
carry `calibration_status: 'provisional'` with auto-demotion armed; promotion bar
|
||
remains the original ≥40 date-clusters.
|
||
|
||
**hits stackability is RESTORED.** It was withdrawn at `6ae11f1`, which removed
|
||
hits props from `chain.chainAcross`. They are stackable again — and the record
|
||
shows it came back **through the powered gate, not by fiat**. A test asserts the
|
||
restoration alongside the threshold's provenance.
|
||
|
||
### hits bands on p_win_calibrated (765 eval rows)
|
||
|
||
| archetype | n | base rate | bands | lift bands |
|
||
|---|---|---|---|---|
|
||
| UNLABELLED | 284 | 0.5211 | 1 | 0 |
|
||
| BOMBER | 271 | 0.5351 | 1 | 0 |
|
||
| GHOST | 110 | 0.6182 | 1 | 0 |
|
||
| BRUSH | 35 | 0.5714 | 1 | 0 |
|
||
| DRIVER | 34 | 0.6765 | 1 | 0 |
|
||
| CATALYST | 21 | 0.5238 | 1 | 0 |
|
||
| MIRROR | 10 | — | REFUSED | — |
|
||
|
||
Two-bar rule still bites: hits is CALIBRATED but its only proven factor
|
||
(`defense_by_direction`) is pooled, not per-archetype, so the bands remain a
|
||
base-rate read — now honestly numbered.
|
||
|
||
---
|
||
|
||
## PHASE 4 — Logged
|
||
|
||
**The deploy set is now determined by a power-derived, pre-committed, tested
|
||
constant rather than an operator-chosen number.** That property matters more than
|
||
either verdict: at `6ae11f1` the rule moved the live path *against* the operator,
|
||
withdrawing a stat that was already serving. It has now moved it back, on the
|
||
same evidence, because the instrument changed. Both directions are the rule
|
||
working. Keep it.
|
||
|
||
**Standing question for the chain model:** calibrated `p_win` separates within
|
||
archetype no better than raw — every archetype collapses to a single band on both
|
||
deployed stats (TB and now hits), across 13 archetype slots. Per-archetype grade
|
||
separation is **not** going to come from calibration. It comes from proven
|
||
per-archetype factors or it does not exist. Descriptive only; no action here.
|
||
|
||
---
|
||
|
||
## Invariants
|
||
|
||
`p_win` never mutated — calibration rides as `p_win_calibrated`. No Bonferroni
|
||
slot consumed. Counter and frozen clusters byte-identical. Threshold fixed from
|
||
power before any stat was re-read; no post-hoc movement, enforced by test.
|