Collapsed sequence edge: two proven links whose product is too small to use
Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.
FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.
TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.
GATE on the concentrated subset, 114 cumulative tests:
early-exit x WEAK pen n=1931 brier -0.0001 CI [-0.0014,+0.0010] NOT_PROVEN
early-exit x STRONG pen n=2574 brier 0.0000 CI [-0.0011,+0.0010] THEATER
all early-exit later ABs n=6869 brier -0.0001 CI [-0.0007,+0.0005] NOT_PROVEN
pooled all later ABs n=17891 brier 0.0000 CI [-0.0004,+0.0003] THEATER
Not pooled-diluted -- the concentrated subset was gated alone and is no
better.
THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925
hit-rate swing across pen quality 0.0394
MAX JUSTIFIABLE ADJUSTMENT 0.00365
actually applied 0.01930 -> 5.3x over-movement
A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.
A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.
NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.
THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.
Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
@@ -0,0 +1,111 @@
|
||||
# The collapsed sequence edge — two proven links whose product is too small to use
|
||||
|
||||
**Both components are real. Their product is 0.37pp, and detecting it would take
|
||||
52 seasons.** This is the most instructive negative in the programme so far,
|
||||
because nothing in it failed: link-by-link proof did not produce a usable edge.
|
||||
|
||||
---
|
||||
|
||||
## One correction to the framing
|
||||
|
||||
The order states Link 2 proved you *can't* predict the reliever. Half true, and
|
||||
the other half matters: the **individual** grain failed (17.2% accuracy), but the
|
||||
**quality** grain PROVED — predicted pen quality separates 2.70pp of realized hit
|
||||
rate. Pen-season-quality here is a measured predictor, not a fallback after a
|
||||
failure. That strengthened the plan going in.
|
||||
|
||||
## Two of the three specified inputs could not be used honestly
|
||||
|
||||
| specified | status |
|
||||
|---|---|
|
||||
| pen **quality** | PROVEN (Link 2 coarse grain) — used |
|
||||
| pen **archetype** | did NOT prove (0.5669 vs 0.5309 modal baseline, corrected interval spanning zero) — **excluded**, building it in would chain on an unproven link |
|
||||
| hitter **approach identity** ("fastball-hunter", "finesse-vulnerable") | **does not exist** in this registry. MLB batter archetypes are BOMBER / GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID / CATALYST. Inventing an identity to condition on is the fabrication the gate exists to catch |
|
||||
|
||||
A hitter power/contact split derived from the sequence data itself was tested as a
|
||||
**separate gated addition** rather than assumed into the main effect. Neither half
|
||||
proved (power −0.0002, contact −0.0001, both intervals spanning zero).
|
||||
|
||||
---
|
||||
|
||||
## The gate — two-part, on the concentrated subset, 114 cumulative tests
|
||||
|
||||
| subset | n | games | mean shift | Brier Δ | CI | verdict |
|
||||
|---|---|---|---|---|---|---|
|
||||
| concentrated (early-exit × WEAK pen) | 1,931 | 141 | 0.0193 | −0.0001 | [−0.0014, +0.0010] | NOT_PROVEN |
|
||||
| mirror (early-exit × STRONG pen) | 2,574 | 189 | 0.0186 | 0.0000 | [−0.0011, +0.0010] | **THEATER** |
|
||||
| all early-exit later ABs | 6,869 | 451 | 0.0140 | −0.0001 | [−0.0007, +0.0005] | NOT_PROVEN |
|
||||
| pooled all later ABs | 17,891 | 803 | 0.0141 | 0.0000 | [−0.0004, +0.0003] | **THEATER** |
|
||||
|
||||
Not pooled-diluted: the concentrated subset was gated on its own and is no
|
||||
better. Two subsets are THEATER by the gate's own definition — the adjustment
|
||||
moves the number ~1.9pp and improves accuracy by essentially nothing.
|
||||
|
||||
---
|
||||
|
||||
## Why: the mechanical ceiling
|
||||
|
||||
The descriptive pass found the direction the order predicted (early-exit + weak
|
||||
pen +0.74pp, early-exit + strong pen −0.79pp vs a deep-starter baseline). The
|
||||
signs are right. The magnitude is the problem, and it is structural:
|
||||
|
||||
```
|
||||
P(faces pen | early-exit flagged) 0.8075
|
||||
P(faces pen | starter goes deep) 0.7149
|
||||
exposure the flag actually buys 0.0925 <- NOT a switch to the pen
|
||||
|
||||
hit-rate swing across pen quality 0.0394 (weak 0.2491 vs strong 0.2038)
|
||||
|
||||
MAX JUSTIFIABLE ADJUSTMENT = 0.0925 x 0.0394 = 0.00365 (0.37pp)
|
||||
adjustment actually applied (mean |shift|) = 0.01930 (1.93pp)
|
||||
OVER-MOVEMENT FACTOR = 5.3x
|
||||
```
|
||||
|
||||
**A hitter's third or fourth plate appearance is ALREADY against the bullpen 71%
|
||||
of the time even when the starter is projected to go deep.** Link 1 lifts that to
|
||||
81%. It buys nine points of extra pen exposure, not a change of opponent — so any
|
||||
adjustment riding on it is capped at about a tenth of the pen-quality swing.
|
||||
|
||||
The 5.3× over-movement is exactly why the mirror subset reads as THEATER rather
|
||||
than as a small true effect: the adjustment asserts five times more than the
|
||||
mechanism can support.
|
||||
|
||||
### And a correctly-scaled version is not detectable either
|
||||
|
||||
```
|
||||
concentrated subset n = 1,931 SE of hit rate = 0.00985
|
||||
max justifiable effect 0.00365 = 0.37 SE
|
||||
to detect at the corrected bar (z~3.46 for 114 tests): n = 168,488
|
||||
shortfall 87x -> ~52 seasons of concentrated-subset accrual
|
||||
```
|
||||
|
||||
**This line is structurally closed, not sample-blocked.** Waiting does not fix it.
|
||||
|
||||
---
|
||||
|
||||
## Not wired, and the self-check deliberately not wired either
|
||||
|
||||
The adjustment does not prove, so it feeds nothing. The order also asks for a
|
||||
self-check flagging where our sequence read diverges from the line's
|
||||
starter-script, as an opportunity signal. **That is not wired**, because flagging
|
||||
divergence on an adjustment measured as absent would advertise an edge we have
|
||||
just shown does not exist — the same failure as fabricated reasoning, one layer up.
|
||||
|
||||
## The lesson worth keeping
|
||||
|
||||
Link 1 proved (MAE 3.22 → 2.80 batters faced). Link 2's quality grain proved
|
||||
(2.70pp of realized separation). Both are real, both are point-in-time, both
|
||||
survived cumulative correction. **Their product is still too small to use.**
|
||||
|
||||
Link-by-link validation guarantees each link is real. It does not guarantee the
|
||||
chain transmits anything. The multiplicative structure has to be sized BEFORE
|
||||
building — one exposure term of 0.09 is enough to reduce a genuine 3.94pp signal
|
||||
to noise, and no amount of downstream care recovers it.
|
||||
|
||||
## Parallel track — total_bases per-archetype (logged, not run)
|
||||
|
||||
Unchanged: `total_bases` settled n=948 pooled, BOMBER × TB **340**, short by 160.
|
||||
Sample-readiness only, not a verdict. The `specs/per-archetype-grade-bands.md`
|
||||
blocker still stands — the grade does not yet separate within any archetype.
|
||||
|
||||
Link 3 confirmed SKIPPED. Counter and frozen clusters byte-identical.
|
||||
Reference in New Issue
Block a user