Settle model_snapshots + four-stat calibration: works, deploys nowhere
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.
PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.
PHASE 1, two integrity findings both caught by the gate:
1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').
2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.
THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.
THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.
PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):
hits n=1140 bias +0.0868 brier 0.2626 -> 0.2511 d -0.0115 CI [-0.0139,-0.0097]
TB n=1050 bias +0.0834 brier 0.2490 -> 0.2438 d -0.0052 CI [-0.0061,-0.0045]
rbi n= 630 bias +0.0164 brier 0.2011 -> 0.1965 d -0.0046 CI [-0.0092,-0.0010]
runs n= 597 bias +0.0410 no map fittable (173 fit rows < 200)
ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.
A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.
PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.
PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.
PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.
p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
@@ -0,0 +1,179 @@
|
||||
# Settlement + four-stat calibration — the enabling move did not enable
|
||||
|
||||
**Settlement is done (15,484 rows written). Calibration improves held-out Brier
|
||||
on all three stats it can be fitted for — more than any factor ever tested. And
|
||||
no stat can deploy, because the date-cluster ceiling is 17, not 90.**
|
||||
|
||||
---
|
||||
|
||||
## PHASE 0 — Board reconcile (corrections to the record)
|
||||
|
||||
| claim in circulation | measured |
|
||||
|---|---|
|
||||
| ~22,032 snapshots unsettled | **71,192** rows, **0** settled |
|
||||
| snapshots span 05-01 → 08-04 (~90 dates) | **07-19 → 08-06 = 19 dates** |
|
||||
| "calibrate all four on the full replay" | replay yields **17 / 7 / 5 / 5** dates per stat |
|
||||
| batter board is per-archetype graded | **nothing has ever been rescaled**, on any stat |
|
||||
| TB inversion confirmed | **units-bug artifact — UNPROVEN** (S: prove-tb-factors) |
|
||||
|
||||
**What actually drives each number today: a base-rate band, on all four stats.**
|
||||
No stat is per-archetype graded. Proven FACTORS exist on hits only
|
||||
(`defense_by_direction`; `pitcher_contact_profile` and `platoon_severity` are
|
||||
held/demoted, not proven). `gradeBands` remains built, gated and unwired.
|
||||
|
||||
---
|
||||
|
||||
## PHASE 1 — Settlement
|
||||
|
||||
```
|
||||
candidates 34,650 (4 stats) -> settled 16,498 | unresolvable 12,894 | orphaned 5,258
|
||||
written 15,484 (shortfall = idempotency guard vs the live cron)
|
||||
conservation check PASSED
|
||||
```
|
||||
|
||||
### Two integrity findings, both caught by the gate
|
||||
|
||||
**1. Unordered pagination over a LIVE table.** The dupe check hard-failed on
|
||||
snapshot id 33875. `model_snapshots` is written by the cron at 14/19/22/1/3 UTC,
|
||||
and an unordered `.range()` walk over a table being appended to returns
|
||||
overlapping pages. Fixed with `.order('id')`. A plain id-only fetch showed no
|
||||
dupes, so this only bites on longer reads that straddle a cron write.
|
||||
|
||||
**2. 12,894 rows were logged AFTER first pitch.** Cycles at ET 21:00/22:00/23:00
|
||||
*on the game date* (10,738 rows) plus 664 the following morning. A 01:00-UTC
|
||||
cycle is 21:00 the previous evening Eastern — same game date, ~2 hours into the
|
||||
slate. These are not predictions and are excluded.
|
||||
|
||||
Tested whether they were outcome-contaminated: bias +0.0058 in-game vs +0.0008
|
||||
pre-game — **not sharper, just late.** Excluded for provenance, not because they
|
||||
cheated.
|
||||
|
||||
### The enabling move did not enable
|
||||
|
||||
| stat | usable props | date-clusters |
|
||||
|---|---|---|
|
||||
| hits | 1,140 | **17** |
|
||||
| total_bases | 1,050 | **7** |
|
||||
| rbi | 630 | **5** |
|
||||
| runs | 597 | **5** |
|
||||
|
||||
71,192 rows collapse to 4,799 distinct pre-game props: a 2.5× cycle fan-out,
|
||||
then 97.6% both-sides duplication, then the pre-game filter. **Hits ends with
|
||||
1,140 rows against the ledger's existing 1,312.** Settlement was worth doing as a
|
||||
standing debt; it did not unlock the sample the order expected.
|
||||
|
||||
---
|
||||
|
||||
## THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY
|
||||
|
||||
**97.6% of props carry BOTH sides.** Their p_wins sum to ~1 and their outcomes
|
||||
are complementary, so any calibration statistic over the raw population is pinned
|
||||
to 0.5 by construction:
|
||||
|
||||
| | hits | TB | rbi | runs |
|
||||
|---|---|---|---|---|
|
||||
| both-sides population | +0.0002 | +0.0012 | +0.0019 | +0.0000 |
|
||||
| **model-picked side only** | **+0.0868** | **+0.0834** | **+0.0164** | **+0.0410** |
|
||||
|
||||
The first row reads "the counter is perfectly calibrated" and would have
|
||||
overturned three sessions of findings. Same rows, opposite conclusion, and the
|
||||
tell was mean p_win sitting at 0.4998 on every stat.
|
||||
|
||||
---
|
||||
|
||||
## PHASE 2/3 — Calibration and the deploy gate
|
||||
|
||||
Isotonic, point-in-time, fit-past / apply-forward. The split is placed by
|
||||
cumulative ROWS rather than date index — props are not spread evenly across dates
|
||||
and a 60%-of-dates cut left only 143 rows to fit on, under the fitter's 200
|
||||
minimum. Still strictly temporal: every fit date precedes every eval date.
|
||||
|
||||
| stat | n | dates | bias | fit / eval | Brier raw | Brier cal | Δ | CI (date-clustered) | eval dates | decision |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| hits | 1,140 | 17 | +0.0868 | 375 / 765 | 0.2626 | 0.2511 | **−0.0115** | [−0.0139, −0.0097] | 4 | **REFUSE** |
|
||||
| total_bases | 1,050 | 7 | +0.0834 | 425 / 625 | 0.2490 | 0.2438 | **−0.0052** | [−0.0061, −0.0045] | 2 | **REFUSE** |
|
||||
| rbi | 630 | 5 | +0.0164 | 205 / 425 | 0.2011 | 0.1965 | **−0.0046** | [−0.0092, −0.0010] | 2 | **REFUSE** |
|
||||
| runs | 597 | 5 | +0.0410 | 173 / 424 | — | — | — | — | — | **REFUSE** |
|
||||
|
||||
- hits / TB / rbi: **held-out Brier improves and the interval excludes zero.**
|
||||
Every one of these beats every factor ever tested on any stat.
|
||||
- runs: **no map could be fitted** — 173 fit rows under the 200 minimum.
|
||||
- **All four refuse on the date floor: 2–4 eval date-clusters against 40.**
|
||||
|
||||
Certified bands (held-out |err| ≤ 0.05): hits [0.5–0.7], TB [0.6–0.8],
|
||||
rbi [0.5–0.9]. Outside band → refuse, fall to base rate.
|
||||
|
||||
### A null that scored itself
|
||||
|
||||
The first run reported hits at Brier **0.5567** — worse than predicting 0.5 for
|
||||
everything. `fitIsotonic` returns null below its minimum, `applyIsotonic` then
|
||||
returns null per row, and `(null − 1)² === 1` while `(null − 0)² === 0`, so the
|
||||
"Brier score" was silently just the win rate (0.5684). **This project's signature
|
||||
`Number(null) === 0` breach, in my own measurement code.** Now a hard refuse.
|
||||
|
||||
---
|
||||
|
||||
## PHASE 4 — Bias shape (diagnostic only)
|
||||
|
||||
| stat | 0.5–0.6 | 0.6–0.7 | 0.7–0.8 | 0.8–0.9 | 0.9–1.0 |
|
||||
|---|---|---|---|---|---|
|
||||
| hits | +0.024 | +0.054 | +0.145 | +0.244 | +0.244 |
|
||||
| total_bases | −0.019 | +0.076 | +0.156 | +0.160 | +0.282 |
|
||||
| rbi | −0.026 | −0.045 | −0.039 | +0.095 | +0.211 |
|
||||
| runs | −0.050 | +0.031 | +0.064 | +0.155 | +0.237 |
|
||||
|
||||
**Not a uniform shift — favourite-longshot concentration, identically shaped on
|
||||
all four stats.** Near zero or slightly negative at the bottom, then rising
|
||||
sharply. The counter is over-confident specifically about its favourites, which
|
||||
is the population a user acts on.
|
||||
|
||||
Cross-stat gradient measured: hits +0.0868 ≈ TB +0.0834 > runs +0.0410 > rbi
|
||||
+0.0164 — not the TB > RBI > runs the order anticipated.
|
||||
|
||||
---
|
||||
|
||||
## PHASE 5/6 — Not run, honestly
|
||||
|
||||
Both are gated on a Phase 3 deploy. Nothing deployed, so there is no
|
||||
`p_win_calibrated` to rebuild bands on and no activated stat to test
|
||||
per-archetype curves against. Running them would be building on a gate that did
|
||||
not open.
|
||||
|
||||
---
|
||||
|
||||
## PHASE 7 — Refusal accuracy (first real measurement)
|
||||
|
||||
| stat | refused n | refused over-rate | graded over-rate | refused \|dist from 0.5\| | graded |
|
||||
|---|---|---|---|---|---|
|
||||
| hits | 59 | 0.4237 | 0.5368 | 0.076 | 0.037 |
|
||||
| total_bases | 74 | 0.2162 | 0.4029 | **0.284** | 0.097 |
|
||||
| rbi | 141 | 0.1773 | 0.2429 | **0.323** | 0.257 |
|
||||
| runs | 12 | 0.6667 | 0.3015 | 0.167 | 0.199 |
|
||||
|
||||
**Refused props are FURTHER from a coin flip than graded ones, not closer.** The
|
||||
obvious explanation — refusals concentrate on players who barely played — was
|
||||
tested and does not hold: refused mean 3.20 AB vs graded 3.39, and 6.6% vs 6.2%
|
||||
with ≤1 AB.
|
||||
|
||||
So "we pass on what we can't call" is not quite what happens. We pass on what we
|
||||
have no INPUT for, and that population had outcomes that were, in hindsight,
|
||||
lopsided (TB refusals went over 21.6% of the time). Refusing to invent a number
|
||||
without a reference remains correct — but the pass is not landing on the
|
||||
genuinely uncertain props, and this is the first time that has been a number
|
||||
rather than a claim.
|
||||
|
||||
---
|
||||
|
||||
## Verdict
|
||||
|
||||
- **Calibration works.** Three stats improve held-out, all beating every factor
|
||||
ever tested. The programme-level finding stands: calibration beats every factor
|
||||
tried on TB/RBI/runs.
|
||||
- **Nothing deploys.** The date-cluster floor is the right unit for a systematic-
|
||||
bias claim and we have 2–4 where 40 is required. Reaching 40 date-clusters
|
||||
needs ~5 more weeks of accrual, not more replay — the dates do not exist.
|
||||
- **The floor is the order's own** and it was not relaxed to force a pass.
|
||||
|
||||
`p_win` never mutated; no `p_win_calibrated` written since nothing deployed.
|
||||
Calibration consumed no Bonferroni slot. Counter and frozen clusters
|
||||
byte-identical.
|
||||
Reference in New Issue
Block a user