Collapsed sequence edge: two proven links whose product is too small to use

Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.

FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.

TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.

GATE on the concentrated subset, 114 cumulative tests:

  early-exit x WEAK pen    n=1931  brier -0.0001  CI [-0.0014,+0.0010]  NOT_PROVEN
  early-exit x STRONG pen  n=2574  brier  0.0000  CI [-0.0011,+0.0010]  THEATER
  all early-exit later ABs n=6869  brier -0.0001  CI [-0.0007,+0.0005]  NOT_PROVEN
  pooled all later ABs    n=17891  brier  0.0000  CI [-0.0004,+0.0003]  THEATER

Not pooled-diluted -- the concentrated subset was gated alone and is no
better.

THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:

  P(faces pen | early-exit flagged)  0.8075
  P(faces pen | starter goes deep)   0.7149
    exposure the flag actually buys  0.0925
  hit-rate swing across pen quality  0.0394
  MAX JUSTIFIABLE ADJUSTMENT         0.00365
  actually applied                   0.01930   -> 5.3x over-movement

A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.

A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.

NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.

THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.

Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-06 02:43:23 -04:00
parent b2e4c6c4fb
commit 8ab6557faa
3 changed files with 322 additions and 1 deletions
+111
View File
@@ -0,0 +1,111 @@
# The collapsed sequence edge — two proven links whose product is too small to use
**Both components are real. Their product is 0.37pp, and detecting it would take
52 seasons.** This is the most instructive negative in the programme so far,
because nothing in it failed: link-by-link proof did not produce a usable edge.
---
## One correction to the framing
The order states Link 2 proved you *can't* predict the reliever. Half true, and
the other half matters: the **individual** grain failed (17.2% accuracy), but the
**quality** grain PROVED — predicted pen quality separates 2.70pp of realized hit
rate. Pen-season-quality here is a measured predictor, not a fallback after a
failure. That strengthened the plan going in.
## Two of the three specified inputs could not be used honestly
| specified | status |
|---|---|
| pen **quality** | PROVEN (Link 2 coarse grain) — used |
| pen **archetype** | did NOT prove (0.5669 vs 0.5309 modal baseline, corrected interval spanning zero) — **excluded**, building it in would chain on an unproven link |
| hitter **approach identity** ("fastball-hunter", "finesse-vulnerable") | **does not exist** in this registry. MLB batter archetypes are BOMBER / GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID / CATALYST. Inventing an identity to condition on is the fabrication the gate exists to catch |
A hitter power/contact split derived from the sequence data itself was tested as a
**separate gated addition** rather than assumed into the main effect. Neither half
proved (power 0.0002, contact 0.0001, both intervals spanning zero).
---
## The gate — two-part, on the concentrated subset, 114 cumulative tests
| subset | n | games | mean shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|
| concentrated (early-exit × WEAK pen) | 1,931 | 141 | 0.0193 | 0.0001 | [0.0014, +0.0010] | NOT_PROVEN |
| mirror (early-exit × STRONG pen) | 2,574 | 189 | 0.0186 | 0.0000 | [0.0011, +0.0010] | **THEATER** |
| all early-exit later ABs | 6,869 | 451 | 0.0140 | 0.0001 | [0.0007, +0.0005] | NOT_PROVEN |
| pooled all later ABs | 17,891 | 803 | 0.0141 | 0.0000 | [0.0004, +0.0003] | **THEATER** |
Not pooled-diluted: the concentrated subset was gated on its own and is no
better. Two subsets are THEATER by the gate's own definition — the adjustment
moves the number ~1.9pp and improves accuracy by essentially nothing.
---
## Why: the mechanical ceiling
The descriptive pass found the direction the order predicted (early-exit + weak
pen +0.74pp, early-exit + strong pen 0.79pp vs a deep-starter baseline). The
signs are right. The magnitude is the problem, and it is structural:
```
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925 <- NOT a switch to the pen
hit-rate swing across pen quality 0.0394 (weak 0.2491 vs strong 0.2038)
MAX JUSTIFIABLE ADJUSTMENT = 0.0925 x 0.0394 = 0.00365 (0.37pp)
adjustment actually applied (mean |shift|) = 0.01930 (1.93pp)
OVER-MOVEMENT FACTOR = 5.3x
```
**A hitter's third or fourth plate appearance is ALREADY against the bullpen 71%
of the time even when the starter is projected to go deep.** Link 1 lifts that to
81%. It buys nine points of extra pen exposure, not a change of opponent — so any
adjustment riding on it is capped at about a tenth of the pen-quality swing.
The 5.3× over-movement is exactly why the mirror subset reads as THEATER rather
than as a small true effect: the adjustment asserts five times more than the
mechanism can support.
### And a correctly-scaled version is not detectable either
```
concentrated subset n = 1,931 SE of hit rate = 0.00985
max justifiable effect 0.00365 = 0.37 SE
to detect at the corrected bar (z~3.46 for 114 tests): n = 168,488
shortfall 87x -> ~52 seasons of concentrated-subset accrual
```
**This line is structurally closed, not sample-blocked.** Waiting does not fix it.
---
## Not wired, and the self-check deliberately not wired either
The adjustment does not prove, so it feeds nothing. The order also asks for a
self-check flagging where our sequence read diverges from the line's
starter-script, as an opportunity signal. **That is not wired**, because flagging
divergence on an adjustment measured as absent would advertise an edge we have
just shown does not exist — the same failure as fabricated reasoning, one layer up.
## The lesson worth keeping
Link 1 proved (MAE 3.22 → 2.80 batters faced). Link 2's quality grain proved
(2.70pp of realized separation). Both are real, both are point-in-time, both
survived cumulative correction. **Their product is still too small to use.**
Link-by-link validation guarantees each link is real. It does not guarantee the
chain transmits anything. The multiplicative structure has to be sized BEFORE
building — one exposure term of 0.09 is enough to reduce a genuine 3.94pp signal
to noise, and no amount of downstream care recovers it.
## Parallel track — total_bases per-archetype (logged, not run)
Unchanged: `total_bases` settled n=948 pooled, BOMBER × TB **340**, short by 160.
Sample-readiness only, not a verdict. The `specs/per-archetype-grade-bands.md`
blocker still stands — the grade does not yet separate within any archetype.
Link 3 confirmed SKIPPED. Counter and frozen clusters byte-identical.