READ-ONLY. Nothing built or fixed; the four challengers untouched. 41% OF THE REPORTED GAP WAS A MEASUREMENT ARTIFACT. p_win is P(graded side); proj_p_over_line is P(over); 31.4% of settled rows are UNDER-graded, so comparing them raw measures the ladder backwards on a third of the sample. Matched + direction-aligned (n=437): 0.252 vs champion 0.352, not 0.108 vs 0.331. The PRODUCT is not making this mistake -- I checked; projectionChallenger normalises both to the over basis deliberately. The error was in the measurement. THE LOSS IS CONCENTRATED. hits (n=245, res 0.060) and total_bases (n=49, res 0.009) are 67% of rows and carry essentially no signal. Everything else is fine or better: walks 0.519 vs champion 0.544, runs mean 0.345 vs 0.392, and on DOUBLES the ladder's mean BEATS the champion's (0.207 vs -0.062). IT IS THE MEAN, NOT THE SHAPE. On the two failing families the mean itself carries no signal (0.052, -0.019) against the champion's 0.158 and 0.085. Where the mean is good the probability is good -- shape follows mean. A HYPOTHESIS I TESTED AND DISPROVED: prediction compression. I expected P(>=1 hit) to sit in a narrow band and fail to rank. It does not -- spread ratio 0.94 overall, 0.80 for hits, 0.94 for total_bases. The ladder has comparable spread; it is spread in a direction uncorrelated with outcomes. Recorded because it was a plausible story the data refused. PRIORS AND PLUMBING CLEAN. proj_factors carries form_rate, combined_multiplier and breakdown on every row; proj_point 100% populated with sane centres (hits 0.830 vs line 0.578). Not the environment-style silent-null failure. NAMED CAUSE (structural, flagged as hypothesis not finding): the count model mismatches those two stats. total_bases is a WEIGHTED SUM (1B..HR = 1..4), so an NB treats one home run as four events and mis-states variance -- and TB has the worst result in the table. hits is BOUNDED BY AT-BATS and mostly traded at 0.5, so almost everything rides on P(0), the region where the wrong family hurts most. walks/runs/doubles ARE genuine low-rate counts and are exactly the ones that work. FIX BRANCH: targeted per-stat fix for hits and total_bases. THIS REMOVES THE MLB SIMILARITY BUILD FROM THE CRITICAL PATH -- that branch assumed a GLOBAL mean weakness, and the mean is fine or better on three of six stat families. Similarity may be worth building later, on evidence, not on this. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
6.2 KiB
WHY proj-v1.1 LOSES — DIAGNOSIS
Date: 2026-08-02 · READ-ONLY — nothing built, fixed or wired. The four accruing challengers were not touched.
HEADLINE: THE GAP IS REAL BUT WAS OVERSTATED, AND IT IS NOT A SIMILARITY PROBLEM
CAUSE: a bad MEAN on TWO stat families —
hitsandtotal_bases— which are 67% of the sample. Everything else works.FIX BRANCH: targeted per-stat model fix. NOT an MLB similarity build.
The diagnosis removes the biggest remaining build, which is what it was for.
STEP 1 — MATCHED ROWS: 41% OF THE "GAP" WAS A MEASUREMENT ARTIFACT
p_win is P(graded side). proj_p_over_line is P(over). 31.4% of settled
rows are UNDER-graded, so comparing raw P(over) against a graded-under outcome
measures the ladder backwards on a third of the sample.
Matched rows, n=437:
| resolution | |
|---|---|
champion p_win |
0.3523 |
| proj-v1.1, unaligned (as previously reported) | 0.1494 |
| proj-v1.1, direction-aligned | 0.2521 |
Aligning direction recovers 41% of the apparent gap. The previously reported 0.108 vs 0.331 substantially overstated the loss.
The product itself is NOT making this mistake — I checked. projectionChallenger
deliberately normalises both quantities to the over basis (proj_book_implied =
fairOver, converting the graded-side fair when direction is under). The error
was in the measurement, not the model. The real gap is 0.252 vs 0.352.
STEP 2 — PER-STAT: THE LOSS IS CONCENTRATED, NOT SYSTEMIC
| stat | n | base | res champ | res proj | champ MEAN | proj MEAN |
|---|---|---|---|---|---|---|
| hits | 245 | .588 | 0.204 | 0.060 | 0.158 | 0.052 |
| total_bases | 49 | .612 | 0.273 | 0.009 | 0.085 | −0.019 |
| rbi | 42 | .333 | 0.412 | 0.218 | 0.349 | 0.159 |
| runs | 38 | .579 | 0.467 | 0.270 | 0.392 | 0.345 |
| walks | 29 | .517 | 0.544 | 0.519 | 0.297 | 0.519 |
| doubles | 28 | .179 | 0.315 | 0.213 | −0.062 | 0.207 |
| ALL | 437 | .533 | 0.352 | 0.252 | 0.199 | 0.166 |
hits + total_bases = 294 of 437 rows (67%), and on both the ladder has
essentially NO signal. They drag the aggregate on their own.
Where the ladder works, it works well:
walks: 0.519 vs the champion's 0.544 — competitive.doubles: the ladder's MEAN (0.207) BEATS the champion's (−0.062).runs: mean 0.345 vs 0.392 — close.
The machinery is not broken. Two stat families are.
STEP 3 — MEAN vs SHAPE: IT IS THE MEAN
For the two failing families the mean itself carries no signal — hits 0.052,
total_bases −0.019 — while the champion's mean on the same rows carries 0.158
and 0.085. The distribution cannot rescue a mean that does not rank.
Conversely, where the mean is good (walks 0.519, runs 0.345, doubles 0.207)
the ladder's probability is good. Shape follows mean, cleanly.
A hypothesis I tested and DISPROVED
I expected prediction compression — that P(≥1 hit) would land in a narrow band across players and so could not rank. Wrong:
| stat | sd champion | sd proj | ratio |
|---|---|---|---|
| ALL | 0.2048 | 0.1932 | 0.94 |
| hits | 0.1703 | 0.1369 | 0.80 |
| total_bases | 0.1876 | 0.1758 | 0.94 |
The ladder has comparable spread. It is not compressed — it is spread in a direction uncorrelated with outcomes. Recording this because it was a plausible story that the data refused.
STEP 4 — PRIOR / PLUMBING INTEGRITY: CLEAN
No silent default found. On a real snapshot, proj_factors carries form_rate,
combined_multiplier, breakdown, book_implied_basis on every row, and
proj_point is populated 100%. The central values are sane:
| stat | mean proj_point |
mean line |
|---|---|---|
| hits | 0.830 | 0.578 |
| total_bases | 1.639 | 1.500 |
| strikeouts | 4.972 | 5.204 |
The priors are reaching the posterior. This is not the environment-style silent-null failure.
THE CAUSE, AND WHY THOSE TWO STATS
Named cause: the count model does not match the generative structure of hits
and total_bases.
Stated as a hypothesis — it follows from the structure, and this diagnosis did not test it directly:
total_basesis not a count of events — it is a WEIGHTED SUM (1B=1, 2B=2, 3B=3, HR=4). A negative binomial fitted to TB treats one home run as "four events", which mis-states the variance badly. This is a structural mismatch, not a tuning error — and TB has the worst result in the table (−0.019).hitsis bounded by at-bats (~4/game). It is closer to binomial(AB, avg) than to an unbounded Poisson/NB, and most lines are 0.5 (mean line 0.578), so almost everything rides on P(0) — the exact region where the wrong family hurts most.walks,runs,doublesARE genuine low-rate event counts — and they are precisely the ones that work.
FIX BRANCH — and what it rules OUT
RECOMMENDED NEXT ORDER: a targeted per-stat model fix for hits and
total_bases. Small, contained, and aimed at 67% of the sample.
This diagnosis REMOVES the MLB similarity build from the critical path. The
order's "bad MEAN → build similarity (large)" branch assumed a global mean
weakness. It is not global: the mean is fine or better than the champion's on
walks, runs and doubles. A similarity layer would not fix a
family-mismatched count model, and building one now would be a large project aimed
at the wrong defect.
Similarity may still be worth building later — but on evidence, not on this.
TAGS
VERIFIED: matched-row aligned gap 0.252 vs 0.352 (n=437) · 31.4% under-graded ·
projectionChallenger normalises direction correctly, so the misalignment was
measurement-only · per-stat table above · mean-vs-shape isolation · priors and
plumbing clean.
DISPROVED: prediction compression (spread ratio 0.80–0.94).
HYPOTHESIS, NOT VERIFIED: that the NB family mismatch is why hits and TB fail. Structurally motivated; the fix order should test it before committing to a family change.
CORRECTED: "proj-v1.1 loses 0.108 vs 0.331" — that comparison was unaligned on direction and across different row sets.