Files
vyndr/specs/lodo-provisional-calibration.md
T
builtbykev 6ae11f1193 LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration
layer without challenging the binding. That floor is a cluster-robust
interval bar for a CAUSAL claim. Calibration makes no causal claim, has a
bounded failure mode (it can only over- or under-shrink) and consumes no
Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and
leave-one-date-out tests that directly -- a STRICTER bar, since a cluster
count cannot detect a single day carrying the effect. The >=40 floor is
retained, correctly scoped as the PROMOTION bar.

PHASE 1 — both guards codified, 11 tests, green before Phase 2.
Demonstrated on live data: raw population violated=true, mean_p 0.4962,
both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The
null guard's test demonstrates the trap explicitly, since (null-1)**2 is
1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate.

PHASE 2 — LODO:

  hits         n=1140 dates=17  2 reversals (07-22 n=20, 07-26 n=25)  FAIL
  total_bases  n=1050 dates=7   0 reversals, 0 sign flips             PASS
  rbi          n= 630 dates=5   1 reversal  (08-01 n=99)              FAIL
  runs         n= 597 dates=5   2 reversals (08-01 n=86, 08-05 n=244) FAIL

Threshold sensitivity reported because the verdict moves: total_bases
passes at every held-size threshold, runs fails at every one, and hits
fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20
before seeing which stats passed and did not move it afterwards to
preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a
standard error several times the effect, so the instrument is
underpowered per-drop -- an argument for pre-registering a higher
threshold, which is a Roundtable call, not one to make while holding the
results.

PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and
runs REFUSE.

HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService
hardcoded it since S91; it fails LODO, so it is out. A stat that cannot
survive dropping one day was never calibrated, it was fitted to that day.
The consequence is real -- hits props become unstackable for
chain.chainAcross -- and it errs toward withdrawing a claim rather than
preserving one on a fragile verdict. Deployment is now driven by a frozen,
tested CALIBRATION_DEPLOYED set, not a hardcoded stat name.

PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither
waivable. reverify auto-demotes on the first breach (CI stops excluding
zero, or the favourite bias flips sign) and logs the breaking date.
Promotion needs the original >=40 bar. A provisional deploy that cannot be
taken away is just a deploy.

PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The
two-bar rule still bites: calibrated YES, proven NO, so they stay a
base-rate read, now honestly numbered. Every archetype still collapses to
one band -- calibrated p_win separates within archetype no better than raw.

PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI,
and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did
not survive); the refused set is a map of missing inputs; a low-parameter
calibrator is queued unbuilt.

p_win never mutated; calibration rides as p_win_calibrated with
calibration_status provisional. No Bonferroni slot consumed. Counter and
frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 18:31:19 -04:00

185 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# LODO-gated provisional calibration — total_bases deploys, three stats refuse
## PHASE 0 — Record correction
**The ≥40 date-cluster deploy floor applied to CALIBRATION was the wrong
instrument, and I applied it without challenging the binding.**
It is `factorGate`'s cluster-robust interval floor, built for a factor making a
CAUSAL claim, where the risk is a false positive dressed as mechanism. A
calibration layer is different in kind:
- it makes **no causal claim** — it is a monotone shrink toward observed
- its failure mode is **bounded**: it can only over- or under-shrink
- it consumes **no Bonferroni slot**
Its real risk is that the correction is **date-driven**, and leave-one-date-out
tests that directly. The replacement bar is **stricter on stability**, not looser
on standard: LODO fails a stat if removing any single day reverses the
improvement, which a cluster count cannot detect at all.
The prior order's date premise was also wrong (05-01→08-04, "~90 dates"); the
snapshots span 07-19→08-06 = 19 dates. That was corrected in the settlement
session. The mis-bound instrument is mine.
**The ≥40 floor is retained, correctly scoped as the PROMOTION bar** — the point
at which a stat leaves provisional status.
---
## PHASE 1 — Both guards codified (11 tests)
`src/services/model/calibrationGuards.js`
**GUARD 1 — the both-sides tell.** Picked-side dedup is now mandatory
preprocessing, asserted. The guard fires on the CONJUNCTION of both sides being
present AND mean p_win pinned near 0.5 — either alone is unremarkable, and
flagging a genuinely balanced one-sided book would be a false alarm.
Demonstrated on live data in this run:
```
raw population violated=true mean_p 0.4962 both_sides_share 0.9763
after dedup violated=false mean_p 0.6694
```
**GUARD 2 — a null that scores itself.** `safeBrier` refuses when any prediction
is null; `applyOrRefuse` drops unmappable rows rather than passing nulls
downstream. A test demonstrates the trap explicitly — `(null1)² === 1` and
`(null0)² === 0`, so a Brier over nulls silently equals the win rate.
---
## PHASE 2 — LODO table
Refit dropping each date; measure held-out Brier delta and the sign of the >0.9
favourite bias. Drops with fewer than 20 held rows are marked UNINFORMATIVE
rather than counted either way.
**A note on what LODO is:** refitting on all-but-one date uses dates that follow
the held-out one, so this is a STABILITY test, not a point-in-time backtest. The
point-in-time result is separate and already established. Both are required.
| stat | n | dates | informative drops | reversals | sign flips | LODO |
|---|---|---|---|---|---|---|
| hits | 1,140 | 17 | 7 | **2** (07-22 n=20, 07-26 n=25) | 0 | **FAIL** |
| **total_bases** | 1,050 | 7 | 5 | 0 | 0 | **PASS** |
| rbi | 630 | 5 | 5 | **1** (08-01 n=99) | 0 | **FAIL** |
| runs | 597 | 5 | 5 | **2** (08-01 n=86, 08-05 n=244) | 0 | **FAIL** |
### Threshold sensitivity — reported because the verdict moves
| min held rows | hits | total_bases | rbi | runs |
|---|---|---|---|---|
| **20** (applied) | FAIL | **PASS** | FAIL | FAIL |
| 30 | PASS | **PASS** | FAIL | FAIL |
| 50 / 75 | PASS | **PASS** | FAIL | FAIL |
| 100 | PASS | **PASS** | PASS | FAIL |
- **total_bases passes at every threshold** — the only unambiguous result.
- **runs fails at every threshold**, reversing on a 244-row date.
- **hits' failure is threshold-fragile**: it fails only when 20- and 25-row dates
are admitted, and those are the two smallest informative drops in the set.
I chose `MIN_HELD_ROWS = 20` before seeing which stats passed, and did not move
it afterwards to preserve a deploy. The honest caveat: a per-date Brier delta on
20 rows has a standard error several times the effect being tested, so the LODO
instrument is underpowered per-drop at this sample size. That argues for
pre-registering a higher threshold — a Roundtable decision, not one to make while
holding the results.
---
## PHASE 3 — Deploy decisions
| stat | LODO | point-in-time CI | decision |
|---|---|---|---|
| **total_bases** | PASS | [0.0061, 0.0045] | **DEPLOY-PROVISIONAL** |
| hits | FAIL | [0.0139, 0.0097] | REFUSE — improvement reverses on 07-22 / 07-26 |
| rbi | FAIL | [0.0092, 0.0010] | REFUSE — improvement reverses on 08-01 |
| runs | FAIL | no fittable map at the point-in-time split | REFUSE — honest null |
Certified band for total_bases: **[0.60.8]**. Outside it → refuse, fall to base
rate.
### hits was being served calibrated, and is not any more
`snapshotService` hardcoded hits calibration since S91. hits fails LODO, so it
has been removed from the deployed set. **A stat that cannot survive dropping one
day was never calibrated — it was fitted to that day.** The consequence is real:
hits props become unstackable again for `chain.chainAcross`. That is the honest
result of measuring it, not a regression to route around, and it errs toward
withdrawing a claim rather than preserving one on a fragile verdict.
Deployment is now driven by `CALIBRATION_DEPLOYED` (frozen, tested), not a
hardcoded stat name.
---
## PHASE 4 — Auto-demotion (14 tests)
`src/services/model/calibrationRegistry.js`
- **Deploy needs BOTH gates** — LODO pass AND a point-in-time CI excluding zero.
Neither is waivable.
- **`reverify` demotes on the first breach**: the CI ceasing to exclude zero, or
the favourite over-prediction flipping sign (which would mean the correction is
now pushing the wrong way). The breaking date is logged.
- **Promotion to non-provisional** requires the original ≥40 date-cluster bar,
with the interval still holding.
A provisional deploy that cannot be taken away is just a deploy; `reverify` is
what makes the label mean something.
---
## PHASE 5 — Bands rebuilt on p_win_calibrated (total_bases only)
625 eval rows on calibrated values. **The two-bar rule still bites**: TB is now
CALIBRATED but no factor is PROVEN for it (barrel, exit velo and
hard-contact-allowed were all THEATER), so bands remain a base-rate read — now an
honestly-numbered one.
| archetype | n | base rate | bands | separation |
|---|---|---|---|---|
| UNLABELLED | 275 | 0.6255 | 1 | indistinguishable from base rate |
| BOMBER | 200 | 0.6100 | 1 | indistinguishable |
| GHOST | 87 | 0.5747 | 1 | indistinguishable |
| DRIVER | 24 | 0.7917 | 1 (PROVISIONAL) | indistinguishable |
| BRUSH | 19 | 0.4737 | 1 (PROVISIONAL) | indistinguishable |
| MIRROR | 6 | — | REFUSED | insufficient outcomes |
Calibration compressed the served range to 0.42861.0. Every archetype still
collapses to a single band — calibrated p_win does not separate within archetype
any better than raw p_win did. Refused stats keep base-rate bands on raw p_win.
---
## PHASE 6 — Logged, not acted on
**The dead gradient is buried.** Over-prediction ordering on the fuller settled
set is **hits ≈ TB > runs > RBI**, not TB > RBI > runs. The skill-driven-gradient
mechanism did not survive — **RBI has the SMALLEST bias** (+0.0164). Descriptive
only; no mechanism claimed.
**Refusal coverage.** Refused props are predictable-but-input-less rather than
genuinely uncertain (3.20 vs 3.39 AB rules out playing time). The refused set is
a MAP OF MISSING INPUTS and feeds the input-coverage roadmap. Not this order.
**Queued candidate:** a low-parameter calibrator (Platt / beta) fits a
favourite-longshot shape on far fewer points than isotonic needs, which is
exactly the constraint that refused runs. It is a NEW estimator requiring its own
out-of-sample validation. Not built here.
**Programme-level finding:** calibration beats every factor tried on TB / RBI /
runs, and the defect is systematic over-prediction **concentrated in favourites**
(+0.21 to +0.28 above p_win 0.9 on all four stats) rather than a uniform shift.
---
## Invariants
`p_win` never mutated — calibration rides as `p_win_calibrated` with
`calibration_status: 'provisional'`. No Bonferroni slot consumed; testLedger
factor count untouched. Counter and frozen clusters byte-identical.