routed as date-driven PHASE 0 — threshold derived BLIND, before any stat was re-read. A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count. Per-row Brier difference d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single stat's verdict could shape the threshold deciding it: pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175 n* = (0.09816/0.01175)^2 = 69.8 -> 70 The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect this removes, and why the previous verdict moved with the number. Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it equals the constant, so it cannot drift from its own justification. The derivation script prints no stat verdict, no date and no reversal. PHASE 1 — LODO at n*, applied cold: hits 5 informative drops, 0 reversals PASS total_bases 4 informative drops, 0 reversals PASS rbi reverses 2026-08-01 (n=99) FAIL runs reverses 08-01 (n=86), 08-05 (244) FAIL hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across 123-272 row dates, favourite sign holding on every testable drop. THIS IS THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the withdrawal at6ae11f1was correct on the instrument available then, which admitted 20- and 25-row dates as evidence. Nothing about hits changed; the threshold stopped being chosen. PHASE 2 — both failures are DATE-DRIVEN, not underpowered. Every reversal sits above n*=70 (99, 86, 244), so no threshold and no further accrual rescues either: isotonic is fitting day-structure. Routed to the low-parameter calibrator queue (Platt/beta), not built here. PHASE 3 — CALIBRATION_DEPLOYED is now ['hits','total_bases'], frozen and tested, both PROVISIONAL with auto-demotion armed and the >=40 date-cluster promotion bar unchanged. hits stackability for chain.chainAcross is RESTORED, and the record shows it returned through the powered gate rather than by fiat. hits bands rebuilt on p_win_calibrated (765 eval rows): every archetype still one band, still base_rate -- calibrated YES, proven-per-archetype NO. PHASE 4 logged: the deploy set is now set by a power-derived, pre-committed, tested constant rather than an operator-chosen number. At6ae11f1that rule moved the live path AGAINST the operator; it has now moved it back on the same evidence because the instrument changed. Both directions are the rule working. And calibrated p_win separates within archetype no better than raw across 13 archetype slots on two deployed stats -- per-archetype separation will come from proven factors or not at all. p_win never mutated; no Bonferroni slot consumed; counter and frozen clusters verified byte-identical file by file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
5.8 KiB
The LODO threshold, derived from power — hits restored, rbi/runs routed
PHASE 0 — the threshold, derived blind
Estimand: does dropping date D reverse the SIGN of the out-of-sample Brier improvement on D's held-out rows? A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count.
The per-row Brier difference is d_i = (pc_i − y_i)² − (p_i − y_i)², so a date's
delta is mean(d) and SE(n) = SD(d)/√n. The smallest n at which a typical
effect clears one standard error is n* = (SD(d)/|effect|)².
Pooled across all four stats — deliberately, so no single stat's verdict could shape the threshold that decides it:
pooled rows 3,417
SD(per-row Brier diff) 0.09816
|pooled effect| 0.01175
n* = (0.09816 / 0.01175)^2 = 69.8 -> 70
SE-vs-n
| n | SE | effect / SE | informative |
|---|---|---|---|
| 10 | 0.0310 | 0.38 | no |
| 20 (previously chosen by hand) | 0.0220 | 0.54 | no |
| 25 | 0.0196 | 0.60 | no |
| 30 | 0.0179 | 0.66 | no |
| 50 | 0.0139 | 0.85 | no |
| 70 (n*) | 0.0117 | 1.00 | yes |
| 100 | 0.0098 | 1.20 | yes |
| 244 | 0.0063 | 1.87 | yes |
The hand-chosen 20 sat at 0.54 SE — a coin flip. That is the defect this derivation removes, and it is why the previous verdict moved with the number.
Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with
LODO_THRESHOLD_BASIS recording the inputs. A test recomputes (SD/effect)² and
asserts it equals the constant, so the value cannot drift from the basis that
justifies it, and cannot be silently tuned. The derivation script prints no stat
verdict, no date and no reversal; it ran and the constant was committed before
any stat was re-read.
PHASE 1 — LODO at n*, applied cold
| stat | n | dates | informative drops | reversals | LODO |
|---|---|---|---|---|---|
| hits | 1,140 | 17 | 5 | 0 | PASS |
| total_bases | 1,050 | 7 | 4 | 0 | PASS |
| rbi | 630 | 5 | 4 | 1 — 2026-08-01 (n=99) | FAIL |
| runs | 597 | 5 | 3 | 2 — 2026-08-01 (n=86), 2026-08-05 (n=244) | FAIL |
hits' held-out deltas at n*: −0.0041 / −0.0080 / −0.0192 / −0.0140 / −0.0139 across 123–272 row dates. Every drop holds, and the favourite over-prediction holds sign on every drop where it is testable (+0.52 / +0.318 / +0.272).
This is the instrument finally being powered, not vindication of a prediction.
The withdrawal at 6ae11f1 was correct on the instrument available then, which
admitted 20- and 25-row dates as evidence. Nothing about hits changed; what
changed is that the threshold is now derived rather than chosen.
PHASE 2 — Failure classification
Both failures are DATE-DRIVEN, not underpowered-per-drop:
| stat | deciding date | held n | vs n*=70 | classification |
|---|---|---|---|---|
| rbi | 2026-08-01 | 99 | above | DATE-DRIVEN |
| runs | 2026-08-01 | 86 | above | DATE-DRIVEN |
| runs | 2026-08-05 | 244 | far above | DATE-DRIVEN |
Every reversal sits comfortably above the powered threshold, so no threshold choice and no further date accrual rescues either stat. Isotonic is fitting day-structure on both.
Routed to the low-parameter calibrator queue (Platt / beta), which fits a favourite-longshot shape on far fewer free parameters and is therefore much harder to bend to one day. Not built here — it is a new estimator and needs its own out-of-sample validation.
PHASE 3 — Deploy and bands
| stat | status | certified band |
|---|---|---|
| hits | DEPLOY-PROVISIONAL (restored) | [0.5–0.7] |
| total_bases | DEPLOY-PROVISIONAL (unchanged from 6ae11f1) |
[0.6–0.8] |
| rbi | REFUSE — date-driven | — |
| runs | REFUSE — date-driven | — |
CALIBRATION_DEPLOYED is now ['hits', 'total_bases'], frozen and tested. Both
carry calibration_status: 'provisional' with auto-demotion armed; promotion bar
remains the original ≥40 date-clusters.
hits stackability is RESTORED. It was withdrawn at 6ae11f1, which removed
hits props from chain.chainAcross. They are stackable again — and the record
shows it came back through the powered gate, not by fiat. A test asserts the
restoration alongside the threshold's provenance.
hits bands on p_win_calibrated (765 eval rows)
| archetype | n | base rate | bands | lift bands |
|---|---|---|---|---|
| UNLABELLED | 284 | 0.5211 | 1 | 0 |
| BOMBER | 271 | 0.5351 | 1 | 0 |
| GHOST | 110 | 0.6182 | 1 | 0 |
| BRUSH | 35 | 0.5714 | 1 | 0 |
| DRIVER | 34 | 0.6765 | 1 | 0 |
| CATALYST | 21 | 0.5238 | 1 | 0 |
| MIRROR | 10 | — | REFUSED | — |
Two-bar rule still bites: hits is CALIBRATED but its only proven factor
(defense_by_direction) is pooled, not per-archetype, so the bands remain a
base-rate read — now honestly numbered.
PHASE 4 — Logged
The deploy set is now determined by a power-derived, pre-committed, tested
constant rather than an operator-chosen number. That property matters more than
either verdict: at 6ae11f1 the rule moved the live path against the operator,
withdrawing a stat that was already serving. It has now moved it back, on the
same evidence, because the instrument changed. Both directions are the rule
working. Keep it.
Standing question for the chain model: calibrated p_win separates within
archetype no better than raw — every archetype collapses to a single band on both
deployed stats (TB and now hits), across 13 archetype slots. Per-archetype grade
separation is not going to come from calibration. It comes from proven
per-archetype factors or it does not exist. Descriptive only; no action here.
Invariants
p_win never mutated — calibration rides as p_win_calibrated. No Bonferroni
slot consumed. Counter and frozen clusters byte-identical. Threshold fixed from
power before any stat was re-read; no post-hoc movement, enforced by test.