Settle model_snapshots + four-stat calibration: works, deploys nowhere

Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.

PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.

PHASE 1, two integrity findings both caught by the gate:

1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').

2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.

THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.

THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.

PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):

  hits  n=1140  bias +0.0868  brier 0.2626 -> 0.2511  d -0.0115  CI [-0.0139,-0.0097]
  TB    n=1050  bias +0.0834  brier 0.2490 -> 0.2438  d -0.0052  CI [-0.0061,-0.0045]
  rbi   n= 630  bias +0.0164  brier 0.2011 -> 0.1965  d -0.0046  CI [-0.0092,-0.0010]
  runs  n= 597  bias +0.0410  no map fittable (173 fit rows < 200)

ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.

A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.

PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.

PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.

PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.

p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-06 15:30:30 -04:00
parent 23d1b13176
commit f976df47b8
4 changed files with 696 additions and 1 deletions
@@ -0,0 +1,179 @@
# Settlement + four-stat calibration — the enabling move did not enable
**Settlement is done (15,484 rows written). Calibration improves held-out Brier
on all three stats it can be fitted for — more than any factor ever tested. And
no stat can deploy, because the date-cluster ceiling is 17, not 90.**
---
## PHASE 0 — Board reconcile (corrections to the record)
| claim in circulation | measured |
|---|---|
| ~22,032 snapshots unsettled | **71,192** rows, **0** settled |
| snapshots span 05-01 → 08-04 (~90 dates) | **07-19 → 08-06 = 19 dates** |
| "calibrate all four on the full replay" | replay yields **17 / 7 / 5 / 5** dates per stat |
| batter board is per-archetype graded | **nothing has ever been rescaled**, on any stat |
| TB inversion confirmed | **units-bug artifact — UNPROVEN** (S: prove-tb-factors) |
**What actually drives each number today: a base-rate band, on all four stats.**
No stat is per-archetype graded. Proven FACTORS exist on hits only
(`defense_by_direction`; `pitcher_contact_profile` and `platoon_severity` are
held/demoted, not proven). `gradeBands` remains built, gated and unwired.
---
## PHASE 1 — Settlement
```
candidates 34,650 (4 stats) -> settled 16,498 | unresolvable 12,894 | orphaned 5,258
written 15,484 (shortfall = idempotency guard vs the live cron)
conservation check PASSED
```
### Two integrity findings, both caught by the gate
**1. Unordered pagination over a LIVE table.** The dupe check hard-failed on
snapshot id 33875. `model_snapshots` is written by the cron at 14/19/22/1/3 UTC,
and an unordered `.range()` walk over a table being appended to returns
overlapping pages. Fixed with `.order('id')`. A plain id-only fetch showed no
dupes, so this only bites on longer reads that straddle a cron write.
**2. 12,894 rows were logged AFTER first pitch.** Cycles at ET 21:00/22:00/23:00
*on the game date* (10,738 rows) plus 664 the following morning. A 01:00-UTC
cycle is 21:00 the previous evening Eastern — same game date, ~2 hours into the
slate. These are not predictions and are excluded.
Tested whether they were outcome-contaminated: bias +0.0058 in-game vs +0.0008
pre-game — **not sharper, just late.** Excluded for provenance, not because they
cheated.
### The enabling move did not enable
| stat | usable props | date-clusters |
|---|---|---|
| hits | 1,140 | **17** |
| total_bases | 1,050 | **7** |
| rbi | 630 | **5** |
| runs | 597 | **5** |
71,192 rows collapse to 4,799 distinct pre-game props: a 2.5× cycle fan-out,
then 97.6% both-sides duplication, then the pre-game filter. **Hits ends with
1,140 rows against the ledger's existing 1,312.** Settlement was worth doing as a
standing debt; it did not unlock the sample the order expected.
---
## THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY
**97.6% of props carry BOTH sides.** Their p_wins sum to ~1 and their outcomes
are complementary, so any calibration statistic over the raw population is pinned
to 0.5 by construction:
| | hits | TB | rbi | runs |
|---|---|---|---|---|
| both-sides population | +0.0002 | +0.0012 | +0.0019 | +0.0000 |
| **model-picked side only** | **+0.0868** | **+0.0834** | **+0.0164** | **+0.0410** |
The first row reads "the counter is perfectly calibrated" and would have
overturned three sessions of findings. Same rows, opposite conclusion, and the
tell was mean p_win sitting at 0.4998 on every stat.
---
## PHASE 2/3 — Calibration and the deploy gate
Isotonic, point-in-time, fit-past / apply-forward. The split is placed by
cumulative ROWS rather than date index — props are not spread evenly across dates
and a 60%-of-dates cut left only 143 rows to fit on, under the fitter's 200
minimum. Still strictly temporal: every fit date precedes every eval date.
| stat | n | dates | bias | fit / eval | Brier raw | Brier cal | Δ | CI (date-clustered) | eval dates | decision |
|---|---|---|---|---|---|---|---|---|---|---|
| hits | 1,140 | 17 | +0.0868 | 375 / 765 | 0.2626 | 0.2511 | **0.0115** | [0.0139, 0.0097] | 4 | **REFUSE** |
| total_bases | 1,050 | 7 | +0.0834 | 425 / 625 | 0.2490 | 0.2438 | **0.0052** | [0.0061, 0.0045] | 2 | **REFUSE** |
| rbi | 630 | 5 | +0.0164 | 205 / 425 | 0.2011 | 0.1965 | **0.0046** | [0.0092, 0.0010] | 2 | **REFUSE** |
| runs | 597 | 5 | +0.0410 | 173 / 424 | — | — | — | — | — | **REFUSE** |
- hits / TB / rbi: **held-out Brier improves and the interval excludes zero.**
Every one of these beats every factor ever tested on any stat.
- runs: **no map could be fitted** — 173 fit rows under the 200 minimum.
- **All four refuse on the date floor: 24 eval date-clusters against 40.**
Certified bands (held-out |err| ≤ 0.05): hits [0.50.7], TB [0.60.8],
rbi [0.50.9]. Outside band → refuse, fall to base rate.
### A null that scored itself
The first run reported hits at Brier **0.5567** — worse than predicting 0.5 for
everything. `fitIsotonic` returns null below its minimum, `applyIsotonic` then
returns null per row, and `(null 1)² === 1` while `(null 0)² === 0`, so the
"Brier score" was silently just the win rate (0.5684). **This project's signature
`Number(null) === 0` breach, in my own measurement code.** Now a hard refuse.
---
## PHASE 4 — Bias shape (diagnostic only)
| stat | 0.50.6 | 0.60.7 | 0.70.8 | 0.80.9 | 0.91.0 |
|---|---|---|---|---|---|
| hits | +0.024 | +0.054 | +0.145 | +0.244 | +0.244 |
| total_bases | 0.019 | +0.076 | +0.156 | +0.160 | +0.282 |
| rbi | 0.026 | 0.045 | 0.039 | +0.095 | +0.211 |
| runs | 0.050 | +0.031 | +0.064 | +0.155 | +0.237 |
**Not a uniform shift — favourite-longshot concentration, identically shaped on
all four stats.** Near zero or slightly negative at the bottom, then rising
sharply. The counter is over-confident specifically about its favourites, which
is the population a user acts on.
Cross-stat gradient measured: hits +0.0868 ≈ TB +0.0834 > runs +0.0410 > rbi
+0.0164 — not the TB > RBI > runs the order anticipated.
---
## PHASE 5/6 — Not run, honestly
Both are gated on a Phase 3 deploy. Nothing deployed, so there is no
`p_win_calibrated` to rebuild bands on and no activated stat to test
per-archetype curves against. Running them would be building on a gate that did
not open.
---
## PHASE 7 — Refusal accuracy (first real measurement)
| stat | refused n | refused over-rate | graded over-rate | refused \|dist from 0.5\| | graded |
|---|---|---|---|---|---|
| hits | 59 | 0.4237 | 0.5368 | 0.076 | 0.037 |
| total_bases | 74 | 0.2162 | 0.4029 | **0.284** | 0.097 |
| rbi | 141 | 0.1773 | 0.2429 | **0.323** | 0.257 |
| runs | 12 | 0.6667 | 0.3015 | 0.167 | 0.199 |
**Refused props are FURTHER from a coin flip than graded ones, not closer.** The
obvious explanation — refusals concentrate on players who barely played — was
tested and does not hold: refused mean 3.20 AB vs graded 3.39, and 6.6% vs 6.2%
with ≤1 AB.
So "we pass on what we can't call" is not quite what happens. We pass on what we
have no INPUT for, and that population had outcomes that were, in hindsight,
lopsided (TB refusals went over 21.6% of the time). Refusing to invent a number
without a reference remains correct — but the pass is not landing on the
genuinely uncertain props, and this is the first time that has been a number
rather than a claim.
---
## Verdict
- **Calibration works.** Three stats improve held-out, all beating every factor
ever tested. The programme-level finding stands: calibration beats every factor
tried on TB/RBI/runs.
- **Nothing deploys.** The date-cluster floor is the right unit for a systematic-
bias claim and we have 24 where 40 is required. Reaching 40 date-clusters
needs ~5 more weeks of accrual, not more replay — the dates do not exist.
- **The floor is the order's own** and it was not relaxed to force a pass.
`p_win` never mutated; no `p_win_calibrated` written since nothing deployed.
Calibration consumed no Bonferroni slot. Counter and frozen clusters
byte-identical.