Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.
FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.
TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.
GATE on the concentrated subset, 114 cumulative tests:
early-exit x WEAK pen n=1931 brier -0.0001 CI [-0.0014,+0.0010] NOT_PROVEN
early-exit x STRONG pen n=2574 brier 0.0000 CI [-0.0011,+0.0010] THEATER
all early-exit later ABs n=6869 brier -0.0001 CI [-0.0007,+0.0005] NOT_PROVEN
pooled all later ABs n=17891 brier 0.0000 CI [-0.0004,+0.0003] THEATER
Not pooled-diluted -- the concentrated subset was gated alone and is no
better.
THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925
hit-rate swing across pen quality 0.0394
MAX JUSTIFIABLE ADJUSTMENT 0.00365
actually applied 0.01930 -> 5.3x over-movement
A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.
A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.
NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.
THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.
Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
5.3 KiB
The collapsed sequence edge — two proven links whose product is too small to use
Both components are real. Their product is 0.37pp, and detecting it would take 52 seasons. This is the most instructive negative in the programme so far, because nothing in it failed: link-by-link proof did not produce a usable edge.
One correction to the framing
The order states Link 2 proved you can't predict the reliever. Half true, and the other half matters: the individual grain failed (17.2% accuracy), but the quality grain PROVED — predicted pen quality separates 2.70pp of realized hit rate. Pen-season-quality here is a measured predictor, not a fallback after a failure. That strengthened the plan going in.
Two of the three specified inputs could not be used honestly
| specified | status |
|---|---|
| pen quality | PROVEN (Link 2 coarse grain) — used |
| pen archetype | did NOT prove (0.5669 vs 0.5309 modal baseline, corrected interval spanning zero) — excluded, building it in would chain on an unproven link |
| hitter approach identity ("fastball-hunter", "finesse-vulnerable") | does not exist in this registry. MLB batter archetypes are BOMBER / GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID / CATALYST. Inventing an identity to condition on is the fabrication the gate exists to catch |
A hitter power/contact split derived from the sequence data itself was tested as a separate gated addition rather than assumed into the main effect. Neither half proved (power −0.0002, contact −0.0001, both intervals spanning zero).
The gate — two-part, on the concentrated subset, 114 cumulative tests
| subset | n | games | mean shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|
| concentrated (early-exit × WEAK pen) | 1,931 | 141 | 0.0193 | −0.0001 | [−0.0014, +0.0010] | NOT_PROVEN |
| mirror (early-exit × STRONG pen) | 2,574 | 189 | 0.0186 | 0.0000 | [−0.0011, +0.0010] | THEATER |
| all early-exit later ABs | 6,869 | 451 | 0.0140 | −0.0001 | [−0.0007, +0.0005] | NOT_PROVEN |
| pooled all later ABs | 17,891 | 803 | 0.0141 | 0.0000 | [−0.0004, +0.0003] | THEATER |
Not pooled-diluted: the concentrated subset was gated on its own and is no better. Two subsets are THEATER by the gate's own definition — the adjustment moves the number ~1.9pp and improves accuracy by essentially nothing.
Why: the mechanical ceiling
The descriptive pass found the direction the order predicted (early-exit + weak pen +0.74pp, early-exit + strong pen −0.79pp vs a deep-starter baseline). The signs are right. The magnitude is the problem, and it is structural:
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925 <- NOT a switch to the pen
hit-rate swing across pen quality 0.0394 (weak 0.2491 vs strong 0.2038)
MAX JUSTIFIABLE ADJUSTMENT = 0.0925 x 0.0394 = 0.00365 (0.37pp)
adjustment actually applied (mean |shift|) = 0.01930 (1.93pp)
OVER-MOVEMENT FACTOR = 5.3x
A hitter's third or fourth plate appearance is ALREADY against the bullpen 71% of the time even when the starter is projected to go deep. Link 1 lifts that to 81%. It buys nine points of extra pen exposure, not a change of opponent — so any adjustment riding on it is capped at about a tenth of the pen-quality swing.
The 5.3× over-movement is exactly why the mirror subset reads as THEATER rather than as a small true effect: the adjustment asserts five times more than the mechanism can support.
And a correctly-scaled version is not detectable either
concentrated subset n = 1,931 SE of hit rate = 0.00985
max justifiable effect 0.00365 = 0.37 SE
to detect at the corrected bar (z~3.46 for 114 tests): n = 168,488
shortfall 87x -> ~52 seasons of concentrated-subset accrual
This line is structurally closed, not sample-blocked. Waiting does not fix it.
Not wired, and the self-check deliberately not wired either
The adjustment does not prove, so it feeds nothing. The order also asks for a self-check flagging where our sequence read diverges from the line's starter-script, as an opportunity signal. That is not wired, because flagging divergence on an adjustment measured as absent would advertise an edge we have just shown does not exist — the same failure as fabricated reasoning, one layer up.
The lesson worth keeping
Link 1 proved (MAE 3.22 → 2.80 batters faced). Link 2's quality grain proved (2.70pp of realized separation). Both are real, both are point-in-time, both survived cumulative correction. Their product is still too small to use.
Link-by-link validation guarantees each link is real. It does not guarantee the chain transmits anything. The multiplicative structure has to be sized BEFORE building — one exposure term of 0.09 is enough to reduce a genuine 3.94pp signal to noise, and no amount of downstream care recovers it.
Parallel track — total_bases per-archetype (logged, not run)
Unchanged: total_bases settled n=948 pooled, BOMBER × TB 340, short by 160.
Sample-readiness only, not a verdict. The specs/per-archetype-grade-bands.md
blocker still stands — the grade does not yet separate within any archetype.
Link 3 confirmed SKIPPED. Counter and frozen clusters byte-identical.