# The LODO threshold, derived from power — hits restored, rbi/runs routed ## PHASE 0 — the threshold, derived blind **Estimand:** does dropping date D reverse the SIGN of the out-of-sample Brier improvement on D's held-out rows? A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count. The per-row Brier difference is `d_i = (pc_i − y_i)² − (p_i − y_i)²`, so a date's delta is `mean(d)` and `SE(n) = SD(d)/√n`. The smallest n at which a typical effect clears one standard error is `n* = (SD(d)/|effect|)²`. Pooled across all four stats — deliberately, so no single stat's verdict could shape the threshold that decides it: ``` pooled rows 3,417 SD(per-row Brier diff) 0.09816 |pooled effect| 0.01175 n* = (0.09816 / 0.01175)^2 = 69.8 -> 70 ``` ### SE-vs-n | n | SE | effect / SE | informative | |---|---|---|---| | 10 | 0.0310 | 0.38 | no | | **20** (previously chosen by hand) | 0.0220 | **0.54** | **no** | | 25 | 0.0196 | 0.60 | no | | 30 | 0.0179 | 0.66 | no | | 50 | 0.0139 | 0.85 | no | | **70 (n\*)** | 0.0117 | **1.00** | **yes** | | 100 | 0.0098 | 1.20 | yes | | 244 | 0.0063 | 1.87 | yes | The hand-chosen 20 sat at 0.54 SE — a coin flip. That is the defect this derivation removes, and it is why the previous verdict moved with the number. **Committed as `calibrationRegistry.LODO_MIN_HELD_ROWS = 70`** with `LODO_THRESHOLD_BASIS` recording the inputs. A test recomputes `(SD/effect)²` and asserts it equals the constant, so the value cannot drift from the basis that justifies it, and cannot be silently tuned. The derivation script prints no stat verdict, no date and no reversal; it ran and the constant was committed before any stat was re-read. --- ## PHASE 1 — LODO at n\*, applied cold | stat | n | dates | informative drops | reversals | LODO | |---|---|---|---|---|---| | **hits** | 1,140 | 17 | 5 | **0** | **PASS** | | **total_bases** | 1,050 | 7 | 4 | **0** | **PASS** | | rbi | 630 | 5 | 4 | 1 — 2026-08-01 (**n=99**) | **FAIL** | | runs | 597 | 5 | 3 | 2 — 2026-08-01 (**n=86**), 2026-08-05 (**n=244**) | **FAIL** | hits' held-out deltas at n\*: −0.0041 / −0.0080 / −0.0192 / −0.0140 / −0.0139 across 123–272 row dates. Every drop holds, and the favourite over-prediction holds sign on every drop where it is testable (+0.52 / +0.318 / +0.272). **This is the instrument finally being powered, not vindication of a prediction.** The withdrawal at `6ae11f1` was correct on the instrument available then, which admitted 20- and 25-row dates as evidence. Nothing about hits changed; what changed is that the threshold is now derived rather than chosen. --- ## PHASE 2 — Failure classification Both failures are **DATE-DRIVEN**, not underpowered-per-drop: | stat | deciding date | held n | vs n\*=70 | classification | |---|---|---|---|---| | rbi | 2026-08-01 | 99 | **above** | DATE-DRIVEN | | runs | 2026-08-01 | 86 | **above** | DATE-DRIVEN | | runs | 2026-08-05 | 244 | **far above** | DATE-DRIVEN | Every reversal sits comfortably above the powered threshold, so **no threshold choice and no further date accrual rescues either stat.** Isotonic is fitting day-structure on both. **Routed to the low-parameter calibrator queue** (Platt / beta), which fits a favourite-longshot shape on far fewer free parameters and is therefore much harder to bend to one day. Not built here — it is a new estimator and needs its own out-of-sample validation. --- ## PHASE 3 — Deploy and bands | stat | status | certified band | |---|---|---| | **hits** | **DEPLOY-PROVISIONAL (restored)** | [0.5–0.7] | | **total_bases** | DEPLOY-PROVISIONAL (unchanged from 6ae11f1) | [0.6–0.8] | | rbi | REFUSE — date-driven | — | | runs | REFUSE — date-driven | — | `CALIBRATION_DEPLOYED` is now `['hits', 'total_bases']`, frozen and tested. Both carry `calibration_status: 'provisional'` with auto-demotion armed; promotion bar remains the original ≥40 date-clusters. **hits stackability is RESTORED.** It was withdrawn at `6ae11f1`, which removed hits props from `chain.chainAcross`. They are stackable again — and the record shows it came back **through the powered gate, not by fiat**. A test asserts the restoration alongside the threshold's provenance. ### hits bands on p_win_calibrated (765 eval rows) | archetype | n | base rate | bands | lift bands | |---|---|---|---|---| | UNLABELLED | 284 | 0.5211 | 1 | 0 | | BOMBER | 271 | 0.5351 | 1 | 0 | | GHOST | 110 | 0.6182 | 1 | 0 | | BRUSH | 35 | 0.5714 | 1 | 0 | | DRIVER | 34 | 0.6765 | 1 | 0 | | CATALYST | 21 | 0.5238 | 1 | 0 | | MIRROR | 10 | — | REFUSED | — | Two-bar rule still bites: hits is CALIBRATED but its only proven factor (`defense_by_direction`) is pooled, not per-archetype, so the bands remain a base-rate read — now honestly numbered. --- ## PHASE 4 — Logged **The deploy set is now determined by a power-derived, pre-committed, tested constant rather than an operator-chosen number.** That property matters more than either verdict: at `6ae11f1` the rule moved the live path *against* the operator, withdrawing a stat that was already serving. It has now moved it back, on the same evidence, because the instrument changed. Both directions are the rule working. Keep it. **Standing question for the chain model:** calibrated `p_win` separates within archetype no better than raw — every archetype collapses to a single band on both deployed stats (TB and now hits), across 13 archetype slots. Per-archetype grade separation is **not** going to come from calibration. It comes from proven per-archetype factors or it does not exist. Descriptive only; no action here. --- ## Invariants `p_win` never mutated — calibration rides as `p_win_calibrated`. No Bonferroni slot consumed. Counter and frozen clusters byte-identical. Threshold fixed from power before any stat was re-read; no post-hoc movement, enforced by test.