Files
vyndr/specs/proj-v11-diagnosis.md
T
builtbykev 48706210fe Diagnose proj-v1.1: concentrated mean failure, NOT a similarity problem
READ-ONLY. Nothing built or fixed; the four challengers untouched.

41% OF THE REPORTED GAP WAS A MEASUREMENT ARTIFACT. p_win is P(graded
side); proj_p_over_line is P(over); 31.4% of settled rows are UNDER-graded,
so comparing them raw measures the ladder backwards on a third of the
sample. Matched + direction-aligned (n=437): 0.252 vs champion 0.352, not
0.108 vs 0.331. The PRODUCT is not making this mistake -- I checked;
projectionChallenger normalises both to the over basis deliberately. The
error was in the measurement.

THE LOSS IS CONCENTRATED. hits (n=245, res 0.060) and total_bases (n=49,
res 0.009) are 67% of rows and carry essentially no signal. Everything else
is fine or better: walks 0.519 vs champion 0.544, runs mean 0.345 vs 0.392,
and on DOUBLES the ladder's mean BEATS the champion's (0.207 vs -0.062).

IT IS THE MEAN, NOT THE SHAPE. On the two failing families the mean itself
carries no signal (0.052, -0.019) against the champion's 0.158 and 0.085.
Where the mean is good the probability is good -- shape follows mean.

A HYPOTHESIS I TESTED AND DISPROVED: prediction compression. I expected
P(>=1 hit) to sit in a narrow band and fail to rank. It does not -- spread
ratio 0.94 overall, 0.80 for hits, 0.94 for total_bases. The ladder has
comparable spread; it is spread in a direction uncorrelated with outcomes.
Recorded because it was a plausible story the data refused.

PRIORS AND PLUMBING CLEAN. proj_factors carries form_rate,
combined_multiplier and breakdown on every row; proj_point 100% populated
with sane centres (hits 0.830 vs line 0.578). Not the environment-style
silent-null failure.

NAMED CAUSE (structural, flagged as hypothesis not finding): the count
model mismatches those two stats. total_bases is a WEIGHTED SUM (1B..HR =
1..4), so an NB treats one home run as four events and mis-states variance
-- and TB has the worst result in the table. hits is BOUNDED BY AT-BATS and
mostly traded at 0.5, so almost everything rides on P(0), the region where
the wrong family hurts most. walks/runs/doubles ARE genuine low-rate counts
and are exactly the ones that work.

FIX BRANCH: targeted per-stat fix for hits and total_bases. THIS REMOVES
THE MLB SIMILARITY BUILD FROM THE CRITICAL PATH -- that branch assumed a
GLOBAL mean weakness, and the mean is fine or better on three of six stat
families. Similarity may be worth building later, on evidence, not on this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 02:50:13 -04:00

6.2 KiB
Raw Blame History

WHY proj-v1.1 LOSES — DIAGNOSIS

Date: 2026-08-02 · READ-ONLY — nothing built, fixed or wired. The four accruing challengers were not touched.


HEADLINE: THE GAP IS REAL BUT WAS OVERSTATED, AND IT IS NOT A SIMILARITY PROBLEM

CAUSE: a bad MEAN on TWO stat families — hits and total_bases — which are 67% of the sample. Everything else works.

FIX BRANCH: targeted per-stat model fix. NOT an MLB similarity build.

The diagnosis removes the biggest remaining build, which is what it was for.


STEP 1 — MATCHED ROWS: 41% OF THE "GAP" WAS A MEASUREMENT ARTIFACT

p_win is P(graded side). proj_p_over_line is P(over). 31.4% of settled rows are UNDER-graded, so comparing raw P(over) against a graded-under outcome measures the ladder backwards on a third of the sample.

Matched rows, n=437:

resolution
champion p_win 0.3523
proj-v1.1, unaligned (as previously reported) 0.1494
proj-v1.1, direction-aligned 0.2521

Aligning direction recovers 41% of the apparent gap. The previously reported 0.108 vs 0.331 substantially overstated the loss.

The product itself is NOT making this mistake — I checked. projectionChallenger deliberately normalises both quantities to the over basis (proj_book_implied = fairOver, converting the graded-side fair when direction is under). The error was in the measurement, not the model. The real gap is 0.252 vs 0.352.


STEP 2 — PER-STAT: THE LOSS IS CONCENTRATED, NOT SYSTEMIC

stat n base res champ res proj champ MEAN proj MEAN
hits 245 .588 0.204 0.060 0.158 0.052
total_bases 49 .612 0.273 0.009 0.085 0.019
rbi 42 .333 0.412 0.218 0.349 0.159
runs 38 .579 0.467 0.270 0.392 0.345
walks 29 .517 0.544 0.519 0.297 0.519
doubles 28 .179 0.315 0.213 0.062 0.207
ALL 437 .533 0.352 0.252 0.199 0.166

hits + total_bases = 294 of 437 rows (67%), and on both the ladder has essentially NO signal. They drag the aggregate on their own.

Where the ladder works, it works well:

  • walks: 0.519 vs the champion's 0.544 — competitive.
  • doubles: the ladder's MEAN (0.207) BEATS the champion's (0.062).
  • runs: mean 0.345 vs 0.392 — close.

The machinery is not broken. Two stat families are.


STEP 3 — MEAN vs SHAPE: IT IS THE MEAN

For the two failing families the mean itself carries no signalhits 0.052, total_bases 0.019 — while the champion's mean on the same rows carries 0.158 and 0.085. The distribution cannot rescue a mean that does not rank.

Conversely, where the mean is good (walks 0.519, runs 0.345, doubles 0.207) the ladder's probability is good. Shape follows mean, cleanly.

A hypothesis I tested and DISPROVED

I expected prediction compression — that P(≥1 hit) would land in a narrow band across players and so could not rank. Wrong:

stat sd champion sd proj ratio
ALL 0.2048 0.1932 0.94
hits 0.1703 0.1369 0.80
total_bases 0.1876 0.1758 0.94

The ladder has comparable spread. It is not compressed — it is spread in a direction uncorrelated with outcomes. Recording this because it was a plausible story that the data refused.


STEP 4 — PRIOR / PLUMBING INTEGRITY: CLEAN

No silent default found. On a real snapshot, proj_factors carries form_rate, combined_multiplier, breakdown, book_implied_basis on every row, and proj_point is populated 100%. The central values are sane:

stat mean proj_point mean line
hits 0.830 0.578
total_bases 1.639 1.500
strikeouts 4.972 5.204

The priors are reaching the posterior. This is not the environment-style silent-null failure.


THE CAUSE, AND WHY THOSE TWO STATS

Named cause: the count model does not match the generative structure of hits and total_bases.

Stated as a hypothesis — it follows from the structure, and this diagnosis did not test it directly:

  • total_bases is not a count of events — it is a WEIGHTED SUM (1B=1, 2B=2, 3B=3, HR=4). A negative binomial fitted to TB treats one home run as "four events", which mis-states the variance badly. This is a structural mismatch, not a tuning error — and TB has the worst result in the table (0.019).
  • hits is bounded by at-bats (~4/game). It is closer to binomial(AB, avg) than to an unbounded Poisson/NB, and most lines are 0.5 (mean line 0.578), so almost everything rides on P(0) — the exact region where the wrong family hurts most.
  • walks, runs, doubles ARE genuine low-rate event counts — and they are precisely the ones that work.

FIX BRANCH — and what it rules OUT

RECOMMENDED NEXT ORDER: a targeted per-stat model fix for hits and total_bases. Small, contained, and aimed at 67% of the sample.

This diagnosis REMOVES the MLB similarity build from the critical path. The order's "bad MEAN → build similarity (large)" branch assumed a global mean weakness. It is not global: the mean is fine or better than the champion's on walks, runs and doubles. A similarity layer would not fix a family-mismatched count model, and building one now would be a large project aimed at the wrong defect.

Similarity may still be worth building later — but on evidence, not on this.

TAGS

VERIFIED: matched-row aligned gap 0.252 vs 0.352 (n=437) · 31.4% under-graded · projectionChallenger normalises direction correctly, so the misalignment was measurement-only · per-stat table above · mean-vs-shape isolation · priors and plumbing clean.

DISPROVED: prediction compression (spread ratio 0.800.94).

HYPOTHESIS, NOT VERIFIED: that the NB family mismatch is why hits and TB fail. Structurally motivated; the fix order should test it before committing to a family change.

CORRECTED: "proj-v1.1 loses 0.108 vs 0.331" — that comparison was unaligned on direction and across different row sets.