PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration layer without challenging the binding. That floor is a cluster-robust interval bar for a CAUSAL claim. Calibration makes no causal claim, has a bounded failure mode (it can only over- or under-shrink) and consumes no Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and leave-one-date-out tests that directly -- a STRICTER bar, since a cluster count cannot detect a single day carrying the effect. The >=40 floor is retained, correctly scoped as the PROMOTION bar. PHASE 1 — both guards codified, 11 tests, green before Phase 2. Demonstrated on live data: raw population violated=true, mean_p 0.4962, both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The null guard's test demonstrates the trap explicitly, since (null-1)**2 is 1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate. PHASE 2 — LODO: hits n=1140 dates=17 2 reversals (07-22 n=20, 07-26 n=25) FAIL total_bases n=1050 dates=7 0 reversals, 0 sign flips PASS rbi n= 630 dates=5 1 reversal (08-01 n=99) FAIL runs n= 597 dates=5 2 reversals (08-01 n=86, 08-05 n=244) FAIL Threshold sensitivity reported because the verdict moves: total_bases passes at every held-size threshold, runs fails at every one, and hits fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20 before seeing which stats passed and did not move it afterwards to preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a standard error several times the effect, so the instrument is underpowered per-drop -- an argument for pre-registering a higher threshold, which is a Roundtable call, not one to make while holding the results. PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and runs REFUSE. HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService hardcoded it since S91; it fails LODO, so it is out. A stat that cannot survive dropping one day was never calibrated, it was fitted to that day. The consequence is real -- hits props become unstackable for chain.chainAcross -- and it errs toward withdrawing a claim rather than preserving one on a fragile verdict. Deployment is now driven by a frozen, tested CALIBRATION_DEPLOYED set, not a hardcoded stat name. PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither waivable. reverify auto-demotes on the first breach (CI stops excluding zero, or the favourite bias flips sign) and logs the breaking date. Promotion needs the original >=40 bar. A provisional deploy that cannot be taken away is just a deploy. PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The two-bar rule still bites: calibrated YES, proven NO, so they stay a base-rate read, now honestly numbered. Every archetype still collapses to one band -- calibrated p_win separates within archetype no better than raw. PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI, and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did not survive); the refused set is a map of missing inputs; a low-parameter calibrator is queued unbuilt. p_win never mutated; calibration rides as p_win_calibrated with calibration_status provisional. No Bonferroni slot consumed. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
7.9 KiB
LODO-gated provisional calibration — total_bases deploys, three stats refuse
PHASE 0 — Record correction
The ≥40 date-cluster deploy floor applied to CALIBRATION was the wrong instrument, and I applied it without challenging the binding.
It is factorGate's cluster-robust interval floor, built for a factor making a
CAUSAL claim, where the risk is a false positive dressed as mechanism. A
calibration layer is different in kind:
- it makes no causal claim — it is a monotone shrink toward observed
- its failure mode is bounded: it can only over- or under-shrink
- it consumes no Bonferroni slot
Its real risk is that the correction is date-driven, and leave-one-date-out tests that directly. The replacement bar is stricter on stability, not looser on standard: LODO fails a stat if removing any single day reverses the improvement, which a cluster count cannot detect at all.
The prior order's date premise was also wrong (05-01→08-04, "~90 dates"); the snapshots span 07-19→08-06 = 19 dates. That was corrected in the settlement session. The mis-bound instrument is mine.
The ≥40 floor is retained, correctly scoped as the PROMOTION bar — the point at which a stat leaves provisional status.
PHASE 1 — Both guards codified (11 tests)
src/services/model/calibrationGuards.js
GUARD 1 — the both-sides tell. Picked-side dedup is now mandatory preprocessing, asserted. The guard fires on the CONJUNCTION of both sides being present AND mean p_win pinned near 0.5 — either alone is unremarkable, and flagging a genuinely balanced one-sided book would be a false alarm.
Demonstrated on live data in this run:
raw population violated=true mean_p 0.4962 both_sides_share 0.9763
after dedup violated=false mean_p 0.6694
GUARD 2 — a null that scores itself. safeBrier refuses when any prediction
is null; applyOrRefuse drops unmappable rows rather than passing nulls
downstream. A test demonstrates the trap explicitly — (null−1)² === 1 and
(null−0)² === 0, so a Brier over nulls silently equals the win rate.
PHASE 2 — LODO table
Refit dropping each date; measure held-out Brier delta and the sign of the >0.9 favourite bias. Drops with fewer than 20 held rows are marked UNINFORMATIVE rather than counted either way.
A note on what LODO is: refitting on all-but-one date uses dates that follow the held-out one, so this is a STABILITY test, not a point-in-time backtest. The point-in-time result is separate and already established. Both are required.
| stat | n | dates | informative drops | reversals | sign flips | LODO |
|---|---|---|---|---|---|---|
| hits | 1,140 | 17 | 7 | 2 (07-22 n=20, 07-26 n=25) | 0 | FAIL |
| total_bases | 1,050 | 7 | 5 | 0 | 0 | PASS |
| rbi | 630 | 5 | 5 | 1 (08-01 n=99) | 0 | FAIL |
| runs | 597 | 5 | 5 | 2 (08-01 n=86, 08-05 n=244) | 0 | FAIL |
Threshold sensitivity — reported because the verdict moves
| min held rows | hits | total_bases | rbi | runs |
|---|---|---|---|---|
| 20 (applied) | FAIL | PASS | FAIL | FAIL |
| 30 | PASS | PASS | FAIL | FAIL |
| 50 / 75 | PASS | PASS | FAIL | FAIL |
| 100 | PASS | PASS | PASS | FAIL |
- total_bases passes at every threshold — the only unambiguous result.
- runs fails at every threshold, reversing on a 244-row date.
- hits' failure is threshold-fragile: it fails only when 20- and 25-row dates are admitted, and those are the two smallest informative drops in the set.
I chose MIN_HELD_ROWS = 20 before seeing which stats passed, and did not move
it afterwards to preserve a deploy. The honest caveat: a per-date Brier delta on
20 rows has a standard error several times the effect being tested, so the LODO
instrument is underpowered per-drop at this sample size. That argues for
pre-registering a higher threshold — a Roundtable decision, not one to make while
holding the results.
PHASE 3 — Deploy decisions
| stat | LODO | point-in-time CI | decision |
|---|---|---|---|
| total_bases | PASS | [−0.0061, −0.0045] | DEPLOY-PROVISIONAL |
| hits | FAIL | [−0.0139, −0.0097] | REFUSE — improvement reverses on 07-22 / 07-26 |
| rbi | FAIL | [−0.0092, −0.0010] | REFUSE — improvement reverses on 08-01 |
| runs | FAIL | no fittable map at the point-in-time split | REFUSE — honest null |
Certified band for total_bases: [0.6–0.8]. Outside it → refuse, fall to base rate.
hits was being served calibrated, and is not any more
snapshotService hardcoded hits calibration since S91. hits fails LODO, so it
has been removed from the deployed set. A stat that cannot survive dropping one
day was never calibrated — it was fitted to that day. The consequence is real:
hits props become unstackable again for chain.chainAcross. That is the honest
result of measuring it, not a regression to route around, and it errs toward
withdrawing a claim rather than preserving one on a fragile verdict.
Deployment is now driven by CALIBRATION_DEPLOYED (frozen, tested), not a
hardcoded stat name.
PHASE 4 — Auto-demotion (14 tests)
src/services/model/calibrationRegistry.js
- Deploy needs BOTH gates — LODO pass AND a point-in-time CI excluding zero. Neither is waivable.
reverifydemotes on the first breach: the CI ceasing to exclude zero, or the favourite over-prediction flipping sign (which would mean the correction is now pushing the wrong way). The breaking date is logged.- Promotion to non-provisional requires the original ≥40 date-cluster bar, with the interval still holding.
A provisional deploy that cannot be taken away is just a deploy; reverify is
what makes the label mean something.
PHASE 5 — Bands rebuilt on p_win_calibrated (total_bases only)
625 eval rows on calibrated values. The two-bar rule still bites: TB is now CALIBRATED but no factor is PROVEN for it (barrel, exit velo and hard-contact-allowed were all THEATER), so bands remain a base-rate read — now an honestly-numbered one.
| archetype | n | base rate | bands | separation |
|---|---|---|---|---|
| UNLABELLED | 275 | 0.6255 | 1 | indistinguishable from base rate |
| BOMBER | 200 | 0.6100 | 1 | indistinguishable |
| GHOST | 87 | 0.5747 | 1 | indistinguishable |
| DRIVER | 24 | 0.7917 | 1 (PROVISIONAL) | indistinguishable |
| BRUSH | 19 | 0.4737 | 1 (PROVISIONAL) | indistinguishable |
| MIRROR | 6 | — | REFUSED | insufficient outcomes |
Calibration compressed the served range to 0.4286–1.0. Every archetype still collapses to a single band — calibrated p_win does not separate within archetype any better than raw p_win did. Refused stats keep base-rate bands on raw p_win.
PHASE 6 — Logged, not acted on
The dead gradient is buried. Over-prediction ordering on the fuller settled set is hits ≈ TB > runs > RBI, not TB > RBI > runs. The skill-driven-gradient mechanism did not survive — RBI has the SMALLEST bias (+0.0164). Descriptive only; no mechanism claimed.
Refusal coverage. Refused props are predictable-but-input-less rather than genuinely uncertain (3.20 vs 3.39 AB rules out playing time). The refused set is a MAP OF MISSING INPUTS and feeds the input-coverage roadmap. Not this order.
Queued candidate: a low-parameter calibrator (Platt / beta) fits a favourite-longshot shape on far fewer points than isotonic needs, which is exactly the constraint that refused runs. It is a NEW estimator requiring its own out-of-sample validation. Not built here.
Programme-level finding: calibration beats every factor tried on TB / RBI / runs, and the defect is systematic over-prediction concentrated in favourites (+0.21 to +0.28 above p_win 0.9 on all four stats) rather than a uniform shift.
Invariants
p_win never mutated — calibration rides as p_win_calibrated with
calibration_status: 'provisional'. No Bonferroni slot consumed; testLedger
factor count untouched. Counter and frozen clusters byte-identical.