Files
vyndr/specs/proj-v11-diagnosis.md
T
builtbykev 48706210fe Diagnose proj-v1.1: concentrated mean failure, NOT a similarity problem
READ-ONLY. Nothing built or fixed; the four challengers untouched.

41% OF THE REPORTED GAP WAS A MEASUREMENT ARTIFACT. p_win is P(graded
side); proj_p_over_line is P(over); 31.4% of settled rows are UNDER-graded,
so comparing them raw measures the ladder backwards on a third of the
sample. Matched + direction-aligned (n=437): 0.252 vs champion 0.352, not
0.108 vs 0.331. The PRODUCT is not making this mistake -- I checked;
projectionChallenger normalises both to the over basis deliberately. The
error was in the measurement.

THE LOSS IS CONCENTRATED. hits (n=245, res 0.060) and total_bases (n=49,
res 0.009) are 67% of rows and carry essentially no signal. Everything else
is fine or better: walks 0.519 vs champion 0.544, runs mean 0.345 vs 0.392,
and on DOUBLES the ladder's mean BEATS the champion's (0.207 vs -0.062).

IT IS THE MEAN, NOT THE SHAPE. On the two failing families the mean itself
carries no signal (0.052, -0.019) against the champion's 0.158 and 0.085.
Where the mean is good the probability is good -- shape follows mean.

A HYPOTHESIS I TESTED AND DISPROVED: prediction compression. I expected
P(>=1 hit) to sit in a narrow band and fail to rank. It does not -- spread
ratio 0.94 overall, 0.80 for hits, 0.94 for total_bases. The ladder has
comparable spread; it is spread in a direction uncorrelated with outcomes.
Recorded because it was a plausible story the data refused.

PRIORS AND PLUMBING CLEAN. proj_factors carries form_rate,
combined_multiplier and breakdown on every row; proj_point 100% populated
with sane centres (hits 0.830 vs line 0.578). Not the environment-style
silent-null failure.

NAMED CAUSE (structural, flagged as hypothesis not finding): the count
model mismatches those two stats. total_bases is a WEIGHTED SUM (1B..HR =
1..4), so an NB treats one home run as four events and mis-states variance
-- and TB has the worst result in the table. hits is BOUNDED BY AT-BATS and
mostly traded at 0.5, so almost everything rides on P(0), the region where
the wrong family hurts most. walks/runs/doubles ARE genuine low-rate counts
and are exactly the ones that work.

FIX BRANCH: targeted per-stat fix for hits and total_bases. THIS REMOVES
THE MLB SIMILARITY BUILD FROM THE CRITICAL PATH -- that branch assumed a
GLOBAL mean weakness, and the mean is fine or better on three of six stat
families. Similarity may be worth building later, on evidence, not on this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 02:50:13 -04:00

6.2 KiB
Raw Blame History

WHY proj-v1.1 LOSES — DIAGNOSIS

Date: 2026-08-02 · READ-ONLY — nothing built, fixed or wired. The four accruing challengers were not touched.


HEADLINE: THE GAP IS REAL BUT WAS OVERSTATED, AND IT IS NOT A SIMILARITY PROBLEM

CAUSE: a bad MEAN on TWO stat families — hits and total_bases — which are 67% of the sample. Everything else works.

FIX BRANCH: targeted per-stat model fix. NOT an MLB similarity build.

The diagnosis removes the biggest remaining build, which is what it was for.


STEP 1 — MATCHED ROWS: 41% OF THE "GAP" WAS A MEASUREMENT ARTIFACT

p_win is P(graded side). proj_p_over_line is P(over). 31.4% of settled rows are UNDER-graded, so comparing raw P(over) against a graded-under outcome measures the ladder backwards on a third of the sample.

Matched rows, n=437:

resolution
champion p_win 0.3523
proj-v1.1, unaligned (as previously reported) 0.1494
proj-v1.1, direction-aligned 0.2521

Aligning direction recovers 41% of the apparent gap. The previously reported 0.108 vs 0.331 substantially overstated the loss.

The product itself is NOT making this mistake — I checked. projectionChallenger deliberately normalises both quantities to the over basis (proj_book_implied = fairOver, converting the graded-side fair when direction is under). The error was in the measurement, not the model. The real gap is 0.252 vs 0.352.


STEP 2 — PER-STAT: THE LOSS IS CONCENTRATED, NOT SYSTEMIC

stat n base res champ res proj champ MEAN proj MEAN
hits 245 .588 0.204 0.060 0.158 0.052
total_bases 49 .612 0.273 0.009 0.085 −0.019
rbi 42 .333 0.412 0.218 0.349 0.159
runs 38 .579 0.467 0.270 0.392 0.345
walks 29 .517 0.544 0.519 0.297 0.519
doubles 28 .179 0.315 0.213 −0.062 0.207
ALL 437 .533 0.352 0.252 0.199 0.166

hits + total_bases = 294 of 437 rows (67%), and on both the ladder has essentially NO signal. They drag the aggregate on their own.

Where the ladder works, it works well:

  • walks: 0.519 vs the champion's 0.544 — competitive.
  • doubles: the ladder's MEAN (0.207) BEATS the champion's (−0.062).
  • runs: mean 0.345 vs 0.392 — close.

The machinery is not broken. Two stat families are.


STEP 3 — MEAN vs SHAPE: IT IS THE MEAN

For the two failing families the mean itself carries no signal — hits 0.052, total_bases −0.019 — while the champion's mean on the same rows carries 0.158 and 0.085. The distribution cannot rescue a mean that does not rank.

Conversely, where the mean is good (walks 0.519, runs 0.345, doubles 0.207) the ladder's probability is good. Shape follows mean, cleanly.

A hypothesis I tested and DISPROVED

I expected prediction compression — that P(≥1 hit) would land in a narrow band across players and so could not rank. Wrong:

stat sd champion sd proj ratio
ALL 0.2048 0.1932 0.94
hits 0.1703 0.1369 0.80
total_bases 0.1876 0.1758 0.94

The ladder has comparable spread. It is not compressed — it is spread in a direction uncorrelated with outcomes. Recording this because it was a plausible story that the data refused.


STEP 4 — PRIOR / PLUMBING INTEGRITY: CLEAN

No silent default found. On a real snapshot, proj_factors carries form_rate, combined_multiplier, breakdown, book_implied_basis on every row, and proj_point is populated 100%. The central values are sane:

stat mean proj_point mean line
hits 0.830 0.578
total_bases 1.639 1.500
strikeouts 4.972 5.204

The priors are reaching the posterior. This is not the environment-style silent-null failure.


THE CAUSE, AND WHY THOSE TWO STATS

Named cause: the count model does not match the generative structure of hits and total_bases.

Stated as a hypothesis — it follows from the structure, and this diagnosis did not test it directly:

  • total_bases is not a count of events — it is a WEIGHTED SUM (1B=1, 2B=2, 3B=3, HR=4). A negative binomial fitted to TB treats one home run as "four events", which mis-states the variance badly. This is a structural mismatch, not a tuning error — and TB has the worst result in the table (−0.019).
  • hits is bounded by at-bats (~4/game). It is closer to binomial(AB, avg) than to an unbounded Poisson/NB, and most lines are 0.5 (mean line 0.578), so almost everything rides on P(0) — the exact region where the wrong family hurts most.
  • walks, runs, doubles ARE genuine low-rate event counts — and they are precisely the ones that work.

FIX BRANCH — and what it rules OUT

RECOMMENDED NEXT ORDER: a targeted per-stat model fix for hits and total_bases. Small, contained, and aimed at 67% of the sample.

This diagnosis REMOVES the MLB similarity build from the critical path. The order's "bad MEAN → build similarity (large)" branch assumed a global mean weakness. It is not global: the mean is fine or better than the champion's on walks, runs and doubles. A similarity layer would not fix a family-mismatched count model, and building one now would be a large project aimed at the wrong defect.

Similarity may still be worth building later — but on evidence, not on this.

TAGS

VERIFIED: matched-row aligned gap 0.252 vs 0.352 (n=437) · 31.4% under-graded · projectionChallenger normalises direction correctly, so the misalignment was measurement-only · per-stat table above · mean-vs-shape isolation · priors and plumbing clean.

DISPROVED: prediction compression (spread ratio 0.80–0.94).

HYPOTHESIS, NOT VERIFIED: that the NB family mismatch is why hits and TB fail. Structurally motivated; the fix order should test it before committing to a family change.

CORRECTED: "proj-v1.1 loses 0.108 vs 0.331" — that comparison was unaligned on direction and across different row sets.