Files
vyndr/specs/lodo-provisional-calibration.md
builtbykev 6ae11f1193 LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration
layer without challenging the binding. That floor is a cluster-robust
interval bar for a CAUSAL claim. Calibration makes no causal claim, has a
bounded failure mode (it can only over- or under-shrink) and consumes no
Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and
leave-one-date-out tests that directly -- a STRICTER bar, since a cluster
count cannot detect a single day carrying the effect. The >=40 floor is
retained, correctly scoped as the PROMOTION bar.

PHASE 1 — both guards codified, 11 tests, green before Phase 2.
Demonstrated on live data: raw population violated=true, mean_p 0.4962,
both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The
null guard's test demonstrates the trap explicitly, since (null-1)**2 is
1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate.

PHASE 2 — LODO:

  hits         n=1140 dates=17  2 reversals (07-22 n=20, 07-26 n=25)  FAIL
  total_bases  n=1050 dates=7   0 reversals, 0 sign flips             PASS
  rbi          n= 630 dates=5   1 reversal  (08-01 n=99)              FAIL
  runs         n= 597 dates=5   2 reversals (08-01 n=86, 08-05 n=244) FAIL

Threshold sensitivity reported because the verdict moves: total_bases
passes at every held-size threshold, runs fails at every one, and hits
fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20
before seeing which stats passed and did not move it afterwards to
preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a
standard error several times the effect, so the instrument is
underpowered per-drop -- an argument for pre-registering a higher
threshold, which is a Roundtable call, not one to make while holding the
results.

PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and
runs REFUSE.

HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService
hardcoded it since S91; it fails LODO, so it is out. A stat that cannot
survive dropping one day was never calibrated, it was fitted to that day.
The consequence is real -- hits props become unstackable for
chain.chainAcross -- and it errs toward withdrawing a claim rather than
preserving one on a fragile verdict. Deployment is now driven by a frozen,
tested CALIBRATION_DEPLOYED set, not a hardcoded stat name.

PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither
waivable. reverify auto-demotes on the first breach (CI stops excluding
zero, or the favourite bias flips sign) and logs the breaking date.
Promotion needs the original >=40 bar. A provisional deploy that cannot be
taken away is just a deploy.

PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The
two-bar rule still bites: calibrated YES, proven NO, so they stay a
base-rate read, now honestly numbered. Every archetype still collapses to
one band -- calibrated p_win separates within archetype no better than raw.

PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI,
and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did
not survive); the refused set is a map of missing inputs; a low-parameter
calibrator is queued unbuilt.

p_win never mutated; calibration rides as p_win_calibrated with
calibration_status provisional. No Bonferroni slot consumed. Counter and
frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 18:31:19 -04:00

7.9 KiB
Raw Permalink Blame History

LODO-gated provisional calibration — total_bases deploys, three stats refuse

PHASE 0 — Record correction

The ≥40 date-cluster deploy floor applied to CALIBRATION was the wrong instrument, and I applied it without challenging the binding.

It is factorGate's cluster-robust interval floor, built for a factor making a CAUSAL claim, where the risk is a false positive dressed as mechanism. A calibration layer is different in kind:

  • it makes no causal claim — it is a monotone shrink toward observed
  • its failure mode is bounded: it can only over- or under-shrink
  • it consumes no Bonferroni slot

Its real risk is that the correction is date-driven, and leave-one-date-out tests that directly. The replacement bar is stricter on stability, not looser on standard: LODO fails a stat if removing any single day reverses the improvement, which a cluster count cannot detect at all.

The prior order's date premise was also wrong (05-01→08-04, "~90 dates"); the snapshots span 07-19→08-06 = 19 dates. That was corrected in the settlement session. The mis-bound instrument is mine.

The ≥40 floor is retained, correctly scoped as the PROMOTION bar — the point at which a stat leaves provisional status.


PHASE 1 — Both guards codified (11 tests)

src/services/model/calibrationGuards.js

GUARD 1 — the both-sides tell. Picked-side dedup is now mandatory preprocessing, asserted. The guard fires on the CONJUNCTION of both sides being present AND mean p_win pinned near 0.5 — either alone is unremarkable, and flagging a genuinely balanced one-sided book would be a false alarm.

Demonstrated on live data in this run:

raw population    violated=true   mean_p 0.4962   both_sides_share 0.9763
after dedup       violated=false  mean_p 0.6694

GUARD 2 — a null that scores itself. safeBrier refuses when any prediction is null; applyOrRefuse drops unmappable rows rather than passing nulls downstream. A test demonstrates the trap explicitly — (null1)² === 1 and (null0)² === 0, so a Brier over nulls silently equals the win rate.


PHASE 2 — LODO table

Refit dropping each date; measure held-out Brier delta and the sign of the >0.9 favourite bias. Drops with fewer than 20 held rows are marked UNINFORMATIVE rather than counted either way.

A note on what LODO is: refitting on all-but-one date uses dates that follow the held-out one, so this is a STABILITY test, not a point-in-time backtest. The point-in-time result is separate and already established. Both are required.

stat n dates informative drops reversals sign flips LODO
hits 1,140 17 7 2 (07-22 n=20, 07-26 n=25) 0 FAIL
total_bases 1,050 7 5 0 0 PASS
rbi 630 5 5 1 (08-01 n=99) 0 FAIL
runs 597 5 5 2 (08-01 n=86, 08-05 n=244) 0 FAIL

Threshold sensitivity — reported because the verdict moves

min held rows hits total_bases rbi runs
20 (applied) FAIL PASS FAIL FAIL
30 PASS PASS FAIL FAIL
50 / 75 PASS PASS FAIL FAIL
100 PASS PASS PASS FAIL
  • total_bases passes at every threshold — the only unambiguous result.
  • runs fails at every threshold, reversing on a 244-row date.
  • hits' failure is threshold-fragile: it fails only when 20- and 25-row dates are admitted, and those are the two smallest informative drops in the set.

I chose MIN_HELD_ROWS = 20 before seeing which stats passed, and did not move it afterwards to preserve a deploy. The honest caveat: a per-date Brier delta on 20 rows has a standard error several times the effect being tested, so the LODO instrument is underpowered per-drop at this sample size. That argues for pre-registering a higher threshold — a Roundtable decision, not one to make while holding the results.


PHASE 3 — Deploy decisions

stat LODO point-in-time CI decision
total_bases PASS [0.0061, 0.0045] DEPLOY-PROVISIONAL
hits FAIL [0.0139, 0.0097] REFUSE — improvement reverses on 07-22 / 07-26
rbi FAIL [0.0092, 0.0010] REFUSE — improvement reverses on 08-01
runs FAIL no fittable map at the point-in-time split REFUSE — honest null

Certified band for total_bases: [0.60.8]. Outside it → refuse, fall to base rate.

hits was being served calibrated, and is not any more

snapshotService hardcoded hits calibration since S91. hits fails LODO, so it has been removed from the deployed set. A stat that cannot survive dropping one day was never calibrated — it was fitted to that day. The consequence is real: hits props become unstackable again for chain.chainAcross. That is the honest result of measuring it, not a regression to route around, and it errs toward withdrawing a claim rather than preserving one on a fragile verdict.

Deployment is now driven by CALIBRATION_DEPLOYED (frozen, tested), not a hardcoded stat name.


PHASE 4 — Auto-demotion (14 tests)

src/services/model/calibrationRegistry.js

  • Deploy needs BOTH gates — LODO pass AND a point-in-time CI excluding zero. Neither is waivable.
  • reverify demotes on the first breach: the CI ceasing to exclude zero, or the favourite over-prediction flipping sign (which would mean the correction is now pushing the wrong way). The breaking date is logged.
  • Promotion to non-provisional requires the original ≥40 date-cluster bar, with the interval still holding.

A provisional deploy that cannot be taken away is just a deploy; reverify is what makes the label mean something.


PHASE 5 — Bands rebuilt on p_win_calibrated (total_bases only)

625 eval rows on calibrated values. The two-bar rule still bites: TB is now CALIBRATED but no factor is PROVEN for it (barrel, exit velo and hard-contact-allowed were all THEATER), so bands remain a base-rate read — now an honestly-numbered one.

archetype n base rate bands separation
UNLABELLED 275 0.6255 1 indistinguishable from base rate
BOMBER 200 0.6100 1 indistinguishable
GHOST 87 0.5747 1 indistinguishable
DRIVER 24 0.7917 1 (PROVISIONAL) indistinguishable
BRUSH 19 0.4737 1 (PROVISIONAL) indistinguishable
MIRROR 6 REFUSED insufficient outcomes

Calibration compressed the served range to 0.42861.0. Every archetype still collapses to a single band — calibrated p_win does not separate within archetype any better than raw p_win did. Refused stats keep base-rate bands on raw p_win.


PHASE 6 — Logged, not acted on

The dead gradient is buried. Over-prediction ordering on the fuller settled set is hits ≈ TB > runs > RBI, not TB > RBI > runs. The skill-driven-gradient mechanism did not survive — RBI has the SMALLEST bias (+0.0164). Descriptive only; no mechanism claimed.

Refusal coverage. Refused props are predictable-but-input-less rather than genuinely uncertain (3.20 vs 3.39 AB rules out playing time). The refused set is a MAP OF MISSING INPUTS and feeds the input-coverage roadmap. Not this order.

Queued candidate: a low-parameter calibrator (Platt / beta) fits a favourite-longshot shape on far fewer points than isotonic needs, which is exactly the constraint that refused runs. It is a NEW estimator requiring its own out-of-sample validation. Not built here.

Programme-level finding: calibration beats every factor tried on TB / RBI / runs, and the defect is systematic over-prediction concentrated in favourites (+0.21 to +0.28 above p_win 0.9 on all four stats) rather than a uniform shift.


Invariants

p_win never mutated — calibration rides as p_win_calibrated with calibration_status: 'provisional'. No Bonferroni slot consumed; testLedger factor count untouched. Counter and frozen clusters byte-identical.