runs + RBI: mostly-base-rate confirmed, and one level deeper than expected

Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest
grade is the POOLED base rate.

PREMISE NOTE: the order's closing line says the batter board is
per-archetype-graded after this. Nothing has been rescaled for hits or
total_bases either -- no archetype slot has ever reached sample and
gradeBands remains built, gated and unwired. This is the fourth stat
measured, not the completion of three.

AUDIT: RBI 935 clean / 43 games; RUNS 617 clean / 33 games. Zero
quarantined. No archetype slot reaches 500 -- and the signature
archetypes the order names are the two SMALLEST slots on the board,
RBI->DRIVER at n=24 and runs->CATALYST at n=9. RUNS is refused
structurally before any factor is tested: 33 game clusters against a 40
floor.

INPUTS RECONSTRUCTED rather than declared missing. lineup_context only
covers 08-04 onward while settled rows start 07-31, so 187/617 runs rows
joined. But the play-by-play cache runs from 05-01 and the batting order
IS the order batters first appear -- slot, power-behind and reach-base
all rebuilt point-in-time, coverage 187 -> 574.

RBI, all THEATER: risp_opportunity +0.0047, extra_base_skill +0.0010,
risp x extra_base +0.0056. RUNS, all refused on clusters and all pointing
the wrong way: +0.0043 / +0.0008 / +0.0054.

THE COMPOUND IS THE WORST VERSION IN BOTH STATS. The causally-correct
compound was the most promising factor on the sheet and is the most
harmful in each. Two multipliers that individually carry nothing do not
cancel -- they compound each other's noise. Distinct from the
collapsed-sequence lesson: there the product of two REAL effects was too
small to use; here the product of two NULL effects is worse than either.

THE ARCHETYPE DOES NOT RESCUE IT, and this is where the session nearly
went wrong. The base rates look strongly differentiated (RBI DRIVER 0.609
vs BOMBER 0.413; runs GHOST 0.716 vs BOMBER 0.460). Gated directly
against the pooled base rate: RBI +0.0010 CI [-0.0034,+0.0050] THEATER;
runs -0.0028 CI [-0.0147,+0.0108] candidate at k=33. DRIVER's 0.609 is
n=23 -- small-slot noise wearing a decimal point. Read off the table
instead of gated, this would have shipped as "archetype differentiation
is real and large". It is not.

THE CROSS-STAT PATTERN THAT IS REAL -- the counter over-predicts every
batter counting stat measured:

  total_bases  p_win 0.5698 vs actual 0.5074  bias +0.0624
  rbi          p_win 0.4860 vs actual 0.4313  bias +0.0547
  runs         p_win 0.5949 vs actual 0.5749  bias +0.0200

Across four stats and three sessions, calibration is the systematic
defect and factor scarcity is not. TB's held-out isotonic fix (-0.0039)
still outperforms every factor tried on any stat, all null or theatre.

NO RESCALE. Nothing proved, nothing certified calibrated, no slot at
sample, and for RBI the archetype split is itself theatre -- so the
honest band is the pooled base rate, which gradeBands returns by
construction.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-06 14:42:33 -04:00
parent 6a327d9114
commit 23d1b13176
3 changed files with 445 additions and 1 deletions
+143
View File
@@ -0,0 +1,143 @@
# runs + RBI — mostly-base-rate confirmed, and stronger than expected
**Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest
grade is the POOLED base rate.** The order anticipated that mostly-base-rate
would be the correct answer for context-heavy stats. It is, and one level deeper
than predicted.
## Premise note
The order's closing line says the batter board is "per-archetype-graded" after
this. It is not: **nothing has been rescaled for hits or total_bases either**
no archetype slot has ever reached sample, and `gradeBands` remains built, gated
and unwired. This is the fourth stat measured, not the completion of three.
---
## STEP 1 — Full-history audit
| | clean | with p_win | players | dates | **games** | lines |
|---|---|---|---|---|---|---|
| RBI | 935 | 931 | 344 | 5 | **43** | 0.5 on 885 |
| RUNS | 617 | 614 | 334 | 5 | **33** | 0.5 on all |
Zero quarantined in either. **No archetype slot reaches 500.**
| archetype | RBI n | RUNS n |
|---|---|---|
| UNLABELLED | 546 | 341 |
| BOMBER | 207 | 148 |
| GHOST | 92 | 67 |
| BRUSH | 43 | 26 |
| **DRIVER** | **24** | 20 |
| **CATALYST** | 13 | **9** |
**The signature archetypes the order names are the two smallest slots on the
board** — RBI→DRIVER at n=24, runs→CATALYST at n=9. The hypothesis is reasonable
and we are three orders of magnitude from being able to test it.
**RUNS is refused structurally before any factor is tested: 33 game clusters
against a 40 floor.** Reported as such rather than dressed up as a factor result.
## Inputs reconstructed rather than declared missing
`lineup_context` only covers 2026-08-04/05/06 (ingest began last week) while
settled rows start 07-31, so just 187 of 617 runs rows join to a batting order.
That reads as input-blocked — but the play-by-play cache runs from 05-01, and
**the batting order IS the order batters first appear.** Batting slot, power
behind (mean barrel of the next three slots) and reach-base rate were all
reconstructed point-in-time from it: coverage went 187 → 574.
---
## STEP 2/3 — The gate (153 / 168 cumulative tests)
### RBI — all THEATER
| factor | n | clusters | shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|
| `risp_opportunity` | 803 | 43 | 0.0585 | **+0.0047** | [0.0023, +0.0122] | **THEATER** |
| `extra_base_skill` | 881 | 43 | 0.0271 | **+0.0010** | [0.0023, +0.0042] | **THEATER** |
| `risp × extra_base` | 803 | 43 | 0.0624 | **+0.0056** | [0.0042, +0.0135] | **THEATER** |
### RUNS — all refused on cluster count, all pointing the wrong way
| factor | n | clusters | shift | Brier Δ | verdict |
|---|---|---|---|---|---|
| `reach_base` | 525 | 32 | 0.0383 | +0.0043 | PENDING — k<40 |
| `lineup_power_behind` | 574 | 32 | 0.0232 | +0.0008 | PENDING — k<40 |
| `reach × power_behind` | 525 | 32 | 0.0482 | +0.0054 | PENDING — k<40 |
### The compound is the WORST version, in both stats
RBI: 0.0047 / 0.0010 → **0.0056** compounded. RUNS: 0.0043 / 0.0008 →
**0.0054** compounded.
The causally-correct compound was the most promising factor on the sheet and is
the most harmful in both. Two multipliers that individually carry nothing do not
cancel — they compound each other's noise. Related to but distinct from the
collapsed-sequence lesson: there the product of two REAL effects was too small to
use; here the product of two NULL effects is actively worse than either.
---
## The archetype itself does not rescue it
The archetype base rates look strongly differentiated, and that appearance is
most of the trap:
| | RBI | RUNS |
|---|---|---|
| DRIVER | **0.609** (n=23) | 0.500 (n=20) |
| GHOST | 0.467 | **0.716** (n=67) |
| BOMBER | 0.413 | 0.460 |
| pooled | 0.4305 | 0.5721 |
Tested directly — is the archetype's leave-one-out base rate better than the
pooled one?
| | shift | Brier Δ | CI | verdict |
|---|---|---|---|---|
| **RBI** | 0.0231 | +0.0010 | [0.0034, +0.0050] | **THEATER** |
| **RUNS** | 0.0642 | 0.0028 | [0.0147, +0.0108] | CANDIDATE — k=33 |
**For RBI, knowing the archetype makes the forecast worse.** DRIVER's 0.609 is
n=23 — the spread is small-slot noise wearing a decimal point. Runs is at least
directionally favourable, and unproven.
Had this been read off the base-rate table instead of gated, the session would
have shipped "archetype differentiation is real and large" as a finding. It is
not one.
---
## The cross-stat pattern that IS real
| stat | mean p_win | actual | counter bias |
|---|---|---|---|
| total_bases | 0.5698 | 0.5074 | **+0.0624** |
| RBI | 0.4860 | 0.4313 | **+0.0547** |
| runs | 0.5949 | 0.5749 | **+0.0200** |
**The counter over-predicts every batter counting stat measured.** Across four
stats and three sessions, calibration is the systematic defect and factor
scarcity is not — TB's held-out isotonic fix (0.0039 Brier) still outperforms
every factor tried on any stat, all of which have been null or theatre.
## STEP 4 — No rescale
Two-bar rule: nothing proved, nothing certified calibrated, no slot at sample —
and for RBI the archetype split is itself theatre, so the honest band is the
POOLED base rate rather than a per-archetype one. `gradeBands` returns exactly
that by construction.
## Next, by value
1. **Calibrate the counter across all four stats.** One systematic bias,
measured four times, larger than anything else on the board.
2. **Stop adding factors to context stats.** Six tested across runs/RBI, six
null-or-worse, and the compounds worst of all.
3. Runs needs game-date accrual to clear the cluster floor before it can be
gated at all.
Counter and frozen clusters byte-identical.