PREMISE CORRECTION: the per-archetype rescale is not "proven and live on hits". gradeBands was built, gated and explicitly NOT wired two orders ago -- no hits archetype slot reached sample, every band came back base-rate, and only defense_by_direction proved pooled. This applies an unvalidated-at-archetype-level method to a second stat. FULL-HISTORY AUDIT: 988 clean settled TB rows (101 quarantined, 948 with p_win), 341 players, and only 9 DISTINCT GAME DATES. No archetype slot reaches 500 -- BOMBER 340, GHOST 147, BRUSH 55. Confirmed short on full history, not a windowed artifact. The 9-date figure matters more than the row count: ~49 games means any game- or venue-borne factor has almost no replication here. THE BASELINE HAD TO CHANGE, to a harder null. TB lines vary (1.5 on 559 rows, 0.5 on 345), so a per-line personal base rate would rest on ~2 rows per player-line and would have to be invented. The null is the counter's own p_win, which already prices the line -- beating the champion, not beating "he's due". THE UNITS BUG, caught, and it had produced the best result in the programme. The first run reported barrel_rate at Brier -0.0095, the largest improvement ever measured here. fromStatcastRow returns barrel_pct as a FRACTION (0.06) while the raw table stores 0-100, so (0.06 - 7.8) * 0.018 clamped EVERY row to the maximum negative shift. That uniform downward push "improved" Brier purely by leaning on the counter's over-prediction and contained no barrel information at all. Same family as the S80 trap, inverted. exit_velo was a second bug -- the column is avg_exit_velo, so it read null on every row and reported n=0. A zero is a wiring bug until proven an honest absence. GATE with units fixed, 138 cumulative tests: barrel_rate n=707 shift 0.0364 brier +0.0036 THEATER exit_velo n=707 shift 0.0229 brier +0.0022 THEATER hard_contact_allowed n=707 shift 0.0260 brier +0.0033 THEATER park_weather_hit_type n=651 36 entities PENDING (k<40) platoon_severity n=481 PENDING (n<500) THE PREDICTED INVERSION WENT THE OTHER WAY. BOMBER x barrel_rate is +0.0114, the single most harmful cell in the table, exactly where the strongest proof was predicted. GHOST +0.0012. All sample-blocked so not a verdict, but recorded so it is not claimed later. AND IT IS NOT DOUBLE-COUNTING -- tested and refuted: corr(barrel, p_win) = -0.061, the counter is not pricing barrel at all. The duller answer is corr(barrel, counter RESIDUAL) = -0.012. Barrel is a real skill that carries no information about what the counter gets wrong at this line. That also closes the S81 lead: hard_hit r=0.153 at n=295 drifted to 0.135 at n=383 and is THEATER at n=707. THE REAL FINDING: TB is miscalibrated, not under-factored. mean p_win 0.5698 vs actual 0.5074, bias +0.0624. Held out on a strict time split (fit < 2026-08-02, eval 651 unseen rows): raw 0.25007, constant de-bias 0.24740 (-0.00267), isotonic 0.24621 (-0.00386). Worth more than any factor tested and the only intervention pointing the right way -- and still refused at the corrected bar on 32 clusters. A CANDIDATE, not a result. It also explains the units bug's fake success exactly: a blanket downward shift is a crude de-bias. NO RESCALE. Nothing proved, nothing certified calibrated, no slot at sample -- every band would be the honest base-rate band gradeBands already returns by construction. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
6.2 KiB
total_bases — every power factor is THEATER, and the units bug nearly hid it
Nothing proved. The predicted inversion did not appear. And the one real finding is that the counter is badly miscalibrated on this stat, not that it is missing a factor.
Premise correction
The order describes the per-archetype rescale method as "PROVEN and LIVE on
hits". It is neither. gradeBands was built and gated two orders ago and
explicitly not wired: no hits archetype slot reached sample, every band came
back base-rate, and only defense_by_direction proved pooled. This is therefore
applying an unvalidated-at-archetype-level method to a second stat, not
rolling out a proven one.
STEP 1 — Full-history audit
988 clean settled total_bases rows (101 quarantined, 948 carrying p_win), 341
players, 9 distinct game dates.
| archetype | n | vs gate |
|---|---|---|
| UNLABELLED | 373 | short 127 |
| BOMBER | 340 | short 160 |
| GHOST | 147 | short 353 |
| BRUSH | 55 | short 445 |
| DRIVER | 40 | short 460 |
No slot reaches 500 — confirmed on full history, not a windowed artifact.
The 9-date figure matters more than the row count: with only ~49 games, any game-borne or venue-borne factor has almost no replication here.
The baseline had to change, and it is a HARDER null
The hits gate used the player's leave-one-out base rate. That cannot be reproduced: TB lines vary (1.5 on 559 rows, 0.5 on 345, 2.5 on 45), and a player's rate of clearing 1.5 bases is a different quantity from 0.5. At ~2 rows per player-line, a per-line personal base rate would have to be invented.
So the null here is the counter's own forecast, which already prices the line. A factor must beat the champion, not "he's due" — strictly harder.
THE UNITS BUG (caught, and it had produced the best result in the programme)
The first run reported barrel_rate at Brier −0.0095, the largest
improvement ever measured here. It was an artifact.
fromStatcastRow returns barrel_pct as a FRACTION (0.06); the raw table
stores 0–100. The factor was written against the percentage scale, so
(0.06 − 7.8) × 0.018 clamped every row to the maximum negative shift. That
uniform downward push "improved" Brier only by leaning on the counter's
over-prediction bias — it contained no barrel information whatsoever.
Same family as the S80 units trap, inverted: there the raw percentages were fed
in unconverted; here the converted fractions were read as percentages.
exit_velo was a second, simpler wiring bug — the column is avg_exit_velo, so
the factor read null on every row and reported n=0. A zero is a wiring bug
until proven an honest absence.
STEP 2/3 — The gate, units fixed (138 cumulative tests)
| factor | n | entities | clustered on | k | shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|---|---|
barrel_rate |
707 | 179 | game | 49 | 0.0364 | +0.0036 | [−0.0025, +0.0101] | THEATER |
exit_velo |
707 | 179 | game | 49 | 0.0229 | +0.0022 | [−0.0030, +0.0072] | THEATER |
hard_contact_allowed |
707 | 93 | game | 49 | 0.0260 | +0.0033 | [−0.0019, +0.0082] | THEATER |
park_weather_hit_type |
651 | 36 | treatment entity | 36 | 0.0070 | +0.0016 | [+0.0001, +0.0038] | PENDING — k<40 |
platoon_severity |
481 | 120 | game | 49 | 0.0243 | +0.0006 | [−0.0045, +0.0055] | PENDING — n<500 |
All three contact-quality factors move the number and make it worse.
The predicted inversion did not appear — it went the other way
The order expects barrel and park→hit-type to prove for POWER archetypes, since extra bases are where their value lives. Measured:
| BOMBER | GHOST | |
|---|---|---|
barrel_rate |
+0.0114 (most harmful cell in the table) | +0.0012 |
exit_velo |
+0.0067 | +0.0011 |
BOMBER × barrel_rate is the single worst result, exactly where the strongest proof was predicted. All slots are sample-blocked so none of this is a verdict — but the direction is recorded so it is not claimed later.
Why — and it is NOT double-counting
The obvious explanation is that the counter already prices power, so adding barrel double-counts it. Tested and refuted: corr(barrel_pct, p_win) = −0.061. The counter is not pricing barrel at all.
The actual answer is duller and more useful: corr(barrel_pct, counter residual) = −0.012. Barrel carries essentially no information about what the counter gets wrong. It is a real skill that does not help at this line.
This also closes the S81 lead: hard_hit_pct marginal r = 0.153 at n=295, which
drifted to 0.135 at n=383 (S82) and is now THEATER at n=707. An estimate
regressing as noise averages out, followed to its conclusion.
The real finding: TB is miscalibrated, not under-factored
mean p_win 0.5698 actual hit rate 0.5074 counter bias +0.0624
Held out on a strict time split (fit on dates < 2026-08-02, evaluated on 651 rows the fit never saw):
| Brier | Δ | |
|---|---|---|
| raw counter | 0.25007 | — |
| constant de-bias | 0.24740 | −0.00267 |
| isotonic | 0.24621 | −0.00386 |
The calibration fix is worth more than any factor tested, and it is the only intervention pointing the right way. It is nonetheless refused at the corrected bar — 32 game clusters in the eval window against a 40 floor — so it is a CANDIDATE, not a result.
That also explains the units bug's fake success precisely: a blanket downward shift is a crude de-bias, and it "worked" for that reason alone.
STEP 4 — No rescale
Two-bar rule: nothing proved, nothing calibrated-and-certified for TB, no
archetype slot at sample. Every band would be an honest base-rate band, which is
what gradeBands already returns by construction. Running it would add nothing.
Next, in order of value
- Calibrate total_bases — largest measured effect, needs game-date accrual to clear the cluster floor, not new inputs.
- Stop adding contact-quality factors to TB. Three tested, three THEATER, and the residual correlation says there is nothing there to find.
- Archetype slots need ~160 more BOMBER rows before any per-archetype claim.
Counter and frozen clusters byte-identical.