Files
vyndr/specs/collapsed-sequence-edge.md
T
builtbykev 8ab6557faa Collapsed sequence edge: two proven links whose product is too small to use
Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.

FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.

TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.

GATE on the concentrated subset, 114 cumulative tests:

  early-exit x WEAK pen    n=1931  brier -0.0001  CI [-0.0014,+0.0010]  NOT_PROVEN
  early-exit x STRONG pen  n=2574  brier  0.0000  CI [-0.0011,+0.0010]  THEATER
  all early-exit later ABs n=6869  brier -0.0001  CI [-0.0007,+0.0005]  NOT_PROVEN
  pooled all later ABs    n=17891  brier  0.0000  CI [-0.0004,+0.0003]  THEATER

Not pooled-diluted -- the concentrated subset was gated alone and is no
better.

THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:

  P(faces pen | early-exit flagged)  0.8075
  P(faces pen | starter goes deep)   0.7149
    exposure the flag actually buys  0.0925
  hit-rate swing across pen quality  0.0394
  MAX JUSTIFIABLE ADJUSTMENT         0.00365
  actually applied                   0.01930   -> 5.3x over-movement

A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.

A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.

NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.

THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.

Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 02:43:23 -04:00

112 lines
5.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# The collapsed sequence edge — two proven links whose product is too small to use
**Both components are real. Their product is 0.37pp, and detecting it would take
52 seasons.** This is the most instructive negative in the programme so far,
because nothing in it failed: link-by-link proof did not produce a usable edge.
---
## One correction to the framing
The order states Link 2 proved you *can't* predict the reliever. Half true, and
the other half matters: the **individual** grain failed (17.2% accuracy), but the
**quality** grain PROVED — predicted pen quality separates 2.70pp of realized hit
rate. Pen-season-quality here is a measured predictor, not a fallback after a
failure. That strengthened the plan going in.
## Two of the three specified inputs could not be used honestly
| specified | status |
|---|---|
| pen **quality** | PROVEN (Link 2 coarse grain) — used |
| pen **archetype** | did NOT prove (0.5669 vs 0.5309 modal baseline, corrected interval spanning zero) — **excluded**, building it in would chain on an unproven link |
| hitter **approach identity** ("fastball-hunter", "finesse-vulnerable") | **does not exist** in this registry. MLB batter archetypes are BOMBER / GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID / CATALYST. Inventing an identity to condition on is the fabrication the gate exists to catch |
A hitter power/contact split derived from the sequence data itself was tested as a
**separate gated addition** rather than assumed into the main effect. Neither half
proved (power 0.0002, contact 0.0001, both intervals spanning zero).
---
## The gate — two-part, on the concentrated subset, 114 cumulative tests
| subset | n | games | mean shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|
| concentrated (early-exit × WEAK pen) | 1,931 | 141 | 0.0193 | 0.0001 | [0.0014, +0.0010] | NOT_PROVEN |
| mirror (early-exit × STRONG pen) | 2,574 | 189 | 0.0186 | 0.0000 | [0.0011, +0.0010] | **THEATER** |
| all early-exit later ABs | 6,869 | 451 | 0.0140 | 0.0001 | [0.0007, +0.0005] | NOT_PROVEN |
| pooled all later ABs | 17,891 | 803 | 0.0141 | 0.0000 | [0.0004, +0.0003] | **THEATER** |
Not pooled-diluted: the concentrated subset was gated on its own and is no
better. Two subsets are THEATER by the gate's own definition — the adjustment
moves the number ~1.9pp and improves accuracy by essentially nothing.
---
## Why: the mechanical ceiling
The descriptive pass found the direction the order predicted (early-exit + weak
pen +0.74pp, early-exit + strong pen 0.79pp vs a deep-starter baseline). The
signs are right. The magnitude is the problem, and it is structural:
```
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925 <- NOT a switch to the pen
hit-rate swing across pen quality 0.0394 (weak 0.2491 vs strong 0.2038)
MAX JUSTIFIABLE ADJUSTMENT = 0.0925 x 0.0394 = 0.00365 (0.37pp)
adjustment actually applied (mean |shift|) = 0.01930 (1.93pp)
OVER-MOVEMENT FACTOR = 5.3x
```
**A hitter's third or fourth plate appearance is ALREADY against the bullpen 71%
of the time even when the starter is projected to go deep.** Link 1 lifts that to
81%. It buys nine points of extra pen exposure, not a change of opponent — so any
adjustment riding on it is capped at about a tenth of the pen-quality swing.
The 5.3× over-movement is exactly why the mirror subset reads as THEATER rather
than as a small true effect: the adjustment asserts five times more than the
mechanism can support.
### And a correctly-scaled version is not detectable either
```
concentrated subset n = 1,931 SE of hit rate = 0.00985
max justifiable effect 0.00365 = 0.37 SE
to detect at the corrected bar (z~3.46 for 114 tests): n = 168,488
shortfall 87x -> ~52 seasons of concentrated-subset accrual
```
**This line is structurally closed, not sample-blocked.** Waiting does not fix it.
---
## Not wired, and the self-check deliberately not wired either
The adjustment does not prove, so it feeds nothing. The order also asks for a
self-check flagging where our sequence read diverges from the line's
starter-script, as an opportunity signal. **That is not wired**, because flagging
divergence on an adjustment measured as absent would advertise an edge we have
just shown does not exist — the same failure as fabricated reasoning, one layer up.
## The lesson worth keeping
Link 1 proved (MAE 3.22 → 2.80 batters faced). Link 2's quality grain proved
(2.70pp of realized separation). Both are real, both are point-in-time, both
survived cumulative correction. **Their product is still too small to use.**
Link-by-link validation guarantees each link is real. It does not guarantee the
chain transmits anything. The multiplicative structure has to be sized BEFORE
building — one exposure term of 0.09 is enough to reduce a genuine 3.94pp signal
to noise, and no amount of downstream care recovers it.
## Parallel track — total_bases per-archetype (logged, not run)
Unchanged: `total_bases` settled n=948 pooled, BOMBER × TB **340**, short by 160.
Sample-readiness only, not a verdict. The `specs/per-archetype-grade-bands.md`
blocker still stands — the grade does not yet separate within any archetype.
Link 3 confirmed SKIPPED. Counter and frozen clusters byte-identical.