Files
vyndr/specs/chain-v1.md
T
builtbykev 6c34af3414 checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion
recon (Tony resurrection, STEP 0). This work existed only on the
laptop disk.

- chain shadow accrual + probe script (038_chain_shadow.sql)
- WNBA possession feed: ESPN adapter, usage service, verify script
  (039_wnba_player_game.sql)
- baseball chain
- retention/snapshot service updates, tableKeys, matchupKeys
- specs: chain-v1, wnba-possession-feed, wnba-source-survey
- unit tests for the above

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
2026-08-14 16:53:37 -04:00

504 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# chain-v1 / v2 — the portable engine, made whole, fed, and shadowed on MLB
> **v2 (2026-08-12) is appended at §8.** It corrects the order's premise (the
> chain was already firing on 99.6%, `paOutcome` reads no handedness, and there
> is no fallback path anywhere) and plumbs the hand split as a matchup
> conditioner on the RATE. Read §8 for the current numbers; §4 below is the v1
> measurement and is superseded.
**Status:** SHADOW. Served by nothing. Nothing promoted.
**Date:** 2026-08-12
**Predecessors:** chaining-v1 (Session 90 — the design, which never had a spec
file), A5 (factor freeze), A6 (the join keys, shadow), A7 (shadow accrual).
---
## 0. The finding this order started from
The extraction pass found `chain.js` was a **shell of its own header**:
| the header promised | the code had |
|---|---|
| `atoms` | a positional array ✅ |
| `context` | passed to `redistribute` only, partially |
| **`chainFn`** — atom → per-entity probability | **ABSENT. No parameter, no call site, no export.** |
| `aggregator` — ACROSS / UP | two functions ✅ |
| `redistribute` | **on `chainUp` only** |
And it was inert for a second, undocumented reason: `chainAcross` requires
`calibrated: true`, and the only writer of that flag is a loop over
`CALIBRATION_DEPLOYED`, which is `Object.freeze([])`. **No grade in the system
carries the flag**, so `chainAcross` had zero callers *and* would have refused
any caller it had.
The missing `chainFn` is why the "portable core" was not portable: with no slot
for the atom→probability stage, every sport's real work had to live somewhere
else. For MLB it lived in `scripts/`, reachable from no pipeline.
---
## 1. The three core fixes (`src/services/model/chain.js`)
### 1.1 `chainFn(atom, context) → per-entity probability`
- `applyChainFn` runs before usability filtering. **Default is identity-on-`p`**,
so every pre-existing caller is byte-identical.
- Returning a **number** sets `p`; returning an **object** merges (so a chainFn
can attach its own trace); returning **null or throwing** makes the atom
UNREADABLE ⇒ **DROPPED, counted, never `p=0`**. A zero leg would zero an entire
ticket, and "we could not read him" is not "he cannot do it".
- The refusals are surfaced as `chain_fn_refused` rather than swallowed, so a
chainFn quietly failing across the board is visible instead of looking like a
thin slate.
### 1.2 Correlation is SIGNED, and the direction follows the sign
Was clamped `[0, 1]` with the joint always shifted toward the weakest leg. That
is baseball's shape — same-game legs share the pitcher, the park and the weather
— and it made basketball's case **inexpressible**: teammates compete for finite
possessions, so one player's shot is another's non-shot and their props are
NEGATIVELY correlated.
Correlation is now clamped `[-1, 1]` and `|corr|` interpolates from independence
toward the **Fréchet–Hoeffding bound its sign selects**:
```
corr = +1 → joint = min(p_i) (upper bound, co-monotone)
corr = 0 → joint = Π p_i (independent)
corr = −1 → joint = max(0, Σp − (n−1)) (lower bound, counter-monotone)
```
This is why the direction is principled rather than chosen. The positive branch
is arithmetically unchanged — the weakest leg **is** the upper bound — so every
previously-correct number stays exactly what it was.
### 1.3 `redistribute` reaches BOTH readings
Previously on `chainUp` only. A redistribution that reached the team read and not
the across read would leave `selfCheck` comparing post-redistribution to
pre-redistribution atoms and flagging an INTERNAL_INCONSISTENCY **the model had
itself just manufactured**.
`prepareAtoms(atoms, opts)` is exported so a caller prepares ONCE — chainFn, then
usability, then redistribution — and hands the identical legs to both readings.
---
## 2. Baseball's chainFn (`src/services/model/baseballChain.js`)
A chained forecast is always `rate × opportunity`. Baseball is the clean case
because **opportunity is fixed**: the batting order is set before first pitch, a
nine-run lead does not change who bats next, and PA/game varies over a narrow
range set almost entirely by lineup slot.
```
p_hit_per_PA ← skillProjection.paOutcome the RATE (modelled)
expected PA ← lineup slot the OPPORTUNITY (a lookup)
P(hits ≥ k) ← Binomial(PA, p) mixed over paDistribution
```
**Atoms routed** — all pre-existing, all previously reachable only from
`scripts/`:
| atom | source | role |
|---|---|---|
| `fromStatcastRow` | skillProjection:161 | the ONE legal units conversion (statcast stores PERCENTAGES 0–100) |
| `paOutcome` | skillProjection:268 | K / BB via log5 odds-ratio vs league; remainder = balls in play |
| `hitOnContact` | skillProjection:213 | archetype-selected barrel / hard-hit / exit-velo / GB-speed × pitcher contact allowed × park |
| `paDistribution`, `binomialPmf`, `atLeast` | skillProjection:327/308/339 | the chain over opportunity |
`PA_BY_SLOT` is a **lookup, not a fit** — 4.65 (leadoff) down to 3.85 (nine hole),
the documented ~0.1-PA-per-slot decline. Nothing was tuned on settled rows; a
tuned opportunity term on 1,741 rows is curve-fitting dressed as physics.
**Refusals:** no batter profile ⇒ null ⇒ dropped (never a league hitter). A
missing pitcher is different and handled inside `paOutcome` — the batter's own
rate stands rather than being pulled toward average.
**Scope: hits only.** `projectSkill` also routes `total_bases`, but TB's
head-to-head is on record as INCONCLUSIVE-under-contamination. Widening the first
shadow to a stat whose verdict is already muddy buys noise.
**`redistribute` is dormant** and returns the legs unchanged — the honest dormant
behaviour. Returning null would read as "the hook failed".
---
## 3. The shadow (`src/services/model/chainShadow.js`)
Runs at the post-enriched / pre-persist point in `snapshotService` — the A6
position, where the slate exists as a set. Reuses the statcast rows and resolved
opposing starter the challenger pass already fetched: **zero new I/O**.
### The unit of evidence is the TRIPLE
```
(chain_p, counter_p, outcome)
```
all three on one row, all side-aligned. `counter_p` is the served
`estimateProbability` value for that exact side; `outcome` arrives from the
ordinary settle pass; `chain_p` is the chain's. **Evidence you cannot adjudicate
is not evidence** — a chain probability stored without the number it must beat,
or without the result, can only ever be compared to itself.
**Side alignment is load-bearing.** The chain computes P(over the line); `p_win`
is expressed for the graded SIDE. An under row stores `1 − p_over`. Storing the
raw over-probability against an under row would invert every later comparison,
silently.
### UN-SERVABLE, said in the data
`requireCalibrated: false` is legitimate **only** because nothing downstream
reads the result. Every stored block carries `status: 'UN-SERVABLE'` and
`servable: false` in its own payload, not merely in a comment — a caveat that
lives only in a comment is not attached to the data once something else queries
it. Flipping that requires passing the gate, not editing a file.
### The self-check is VACUOUS today, and says so
`selfCheck` earns its keep by comparing per-entity reads to an **independent**
team read. None exists — the game-script projection was deliberately not built
(no atom has passed the gate). So the up-read is assembled from the SAME atoms as
the across-read and agreement between them is arithmetic. Every block carries
`self_check.vacuous: true` with its reason, so nobody later mistakes a tautology
for a passing consistency test.
### Storage
`model_snapshots.chain_shadow jsonb` (migration 038). A separate column, not
extra keys inside `features` — `champion-ablation.js` iterates every `features`
key for its residual scan, so widening it would silently enlarge that
multiple-comparisons denominator. Same reasoning as 034.
---
## 4. First measurement (`scripts/chain-shadow-probe.js`)
Real board, 2026-08-07 → 2026-08-12, MLB hits, **9,376 graded rows**:
```
CHAIN FIRE — 9,340/9,376 atoms read (36 refused) across 82/82 games [99.6%]
min p25 median p75 max mean
chain_p 0.013 0.382 0.500 0.618 0.987 0.500
counter_p 0.050 0.396 0.500 0.604 0.950 0.500
divergence (signed) −0.556 −0.082 0.000 0.082 0.556 −0.000
divergence (abs) 0.000 0.038 0.082 0.142 0.556 0.100
|divergence| < 0.02 1,203 (12.9%) agrees with the counter
0.02 – 0.05 1,760 (18.8%)
0.05 – 0.10 2,577 (27.6%)
0.10 – 0.20 2,835 (30.4%)
≥ 0.20 965 (10.3%) a different read entirely
OVER SIDE ONLY (n=4,673)
divergence (signed) −0.526 −0.045 0.031 0.104 0.556 0.029
chain HIGHER on 2,837 (60.7%)
```
**Reading it honestly:**
- The chain **fires**, on 99.6% of the board. It is not the A5 case (built,
correct, never invoked).
- It is **not a relabelled counter**: 40.7% of rows differ by ≥0.10, and 10.3%
by ≥0.20. It is also not noise — 12.9% agree inside 0.02.
- The **symmetry of the two-sided signed distribution is arithmetic**, not a
finding: both sides of every prop are in the sample, so each pair contributes
`+d` and `−d`. Reading `mean −0.000` as "unbiased" would be reading the
sampling scheme. The over-side slice is the one that can lean, and it does:
**+2.9pp mean, higher on 60.7%**.
- **Divergence is not merit.** A challenger that disagrees is interesting, not
right. Which of the two is closer to what happened is the settle pass's
question, and it is exactly why the triple is stored.
- **The probe is contaminated by construction** and reports only a divergence
(a property of two forecasts) rather than a resolution (a property of a
forecast against an outcome): `statcast_aggregates` is upserted in place and
keeps one as-of date, so profiles read for a row graded three days ago are
today's. The forward accrual on the cron does not have this problem.
- The probe resolves **no opposing pitcher** (offline), so it measures the
batter-side read. The live shadow does resolve it.
---
## 5. What is NOT claimed
- **Nothing is promoted.** The proven set remains empty.
- **The chain is not calibrated** and has passed no gate.
- **No head-to-head has been run.** That needs settled outcomes against the
stored triples and must go through `factorGate` / the cumulative Bonferroni
denominator, as a NEW hypothesis.
- **`CALIBRATION_DEPLOYED` stays `[]`.** Nothing is served calibrated.
- **WNBA is untouched.** Its contested-possession chainFn and its possession feed
are later orders. The core changes (signed correlation, redistribute on both
readings) were built now because they are engine honesty, not because MLB needs
them — MLB exercises neither.
---
## 6. The pre-registered next step
Once settled outcomes accrue against `chain_shadow`:
1. Score `chain_p` vs `counter_p` on the SAME rows (paired bootstrap — comparing
independent SEs overstates uncertainty and has previously read a reliable
−0.022 as noise).
2. Per stat, never pooled (pooled resolution is inflated by base-rate structure).
3. Through the cumulative test ledger; the CI widens to `1 − 0.05/tests`.
4. **Pre-registered fallback, stated before the answer is known:** if the chain
moves ~87% of the board and does not improve Brier, it is THEATER by the
`factorGate` definition and the correct action is to leave it off — not to
re-tune the opportunity term until it passes.
---
## 7. Incidental defect found and fixed
`runSnapshot`'s `deps` object was an **allowlist of 17 keys**, but fourteen call
sites read `deps.challenger`, `deps.loadStatcast`, `deps.environmentContext`,
`deps.lineupContext`, `deps.hitsFactorContext`, `deps.matchupKeys`,
`deps.gameBinder`, `deps.archetypeAxes`, `deps.contactChallenger`,
`deps.projectionChallenger`, `deps.loadArsenals`, `deps.mlbAdapter` — each
documented as injectable, each **permanently `undefined`**, each always falling
through to the real module. The seam existed in the comment and not in the code.
Fixed by spreading `...opts` FIRST in the literal: every explicit key is declared
after and already reads `opts.X`, so no existing behaviour moves, while an
unlisted dep now actually arrives. Found because the chain shadow's own test
could not inject a statcast map — the test would have passed while measuring
nothing, which is the failure this whole line of work exists to stop repeating.
---
# §8 — chain v2: FEED THE ENGINE (2026-08-12)
## 8.1 The order's premise, checked before building
The v2 order opened from four numbers. Each was checked against the code and the
board before anything was written:
| claim | measured |
|---|---|
| "`paOutcome` returned null on 596/725 rows" | **False.** Fire rate is **9,752/9,792 (99.6%)**. The only refusal reason on the whole board is `no_batter_profile: 40`. `pa_outcome_refused` never fires. |
| "it needs the batter-vs-pitcher-hand split" *to run* | **False.** `paOutcome` and `hitOnContact` read `k_pct`, `bb_pct`, `barrel_pct`, `hard_hit_pct`, `avg_exit_velo`, `avg_launch_angle` and the pitcher's `k_pct`/`bb_pct`/`hard_hit_pct`. Neither reads `bats` or `throws` at all. A hand split cannot change whether they run. |
| "82% FELL BACK to the seasonal rate, i.e. became the counter" | **No fallback path exists.** `baseballChain.chainFn` returns null on any refusal and `chain.applyChainFn` DROPS the atom. Nothing in the module reads a season rate as a substitute. |
| "425/425 served-identical" | That figure is from commit `7c8ef8b` (the A1–A7 deploy verification), not from the chain. v1 measured **2/2 served-identical** in the harness and byte-identical served payloads. |
Chain v1 was also never committed or deployed and migration 038 was never
applied, so no chain-shadow rows exist in production — the premise numbers cannot
have come from a chain-shadow run.
**The premise was wrong about the mechanism. It was right about the thing that
matters:** the chain was reading a hitter's SEASON rates, which already average
his platoon split over whichever hands he happened to face. That is a season read
wearing a matchup read's clothes, and un-averaging it is real work. So v2 plumbs
the hand split — as a **conditioner on the rate**, not as a fix to the fire rate.
## 8.2 What was built
**The split enters at the per-PA hit rate, not at the output probability.**
Multiplying `P(hits ≥ 1)` by a platoon factor would scale a number that has
already been through the opportunity term — a different and wrong claim. So
`baseballChain.chainFn` now writes the chain out explicitly:
```
paOutcome → p_hit_per_pa (season)
→ × platoonRead multiplier ← THE MATCHUP CONDITIONER
→ Binomial(PA, p_hit) over paDistribution
→ P(hits ≥ k)
```
A test asserts that with no split supplied this is **arithmetically identical to
`projectSkill`** across three archetypes × three PA values, so the restructuring
cannot quietly become a second model.
**As-of-correct (A4).** The hitter's own split comes from `hitsFactorContext`
(whose reads are `lte('as_of_date', asOf)`), the opposing starter's hand through
the A6 `matchupKeys` resolve. The probe bounds every read at the row's own
`game_date`. Nothing later than the grade can enter.
**Refuse, never substitute.** `platoonSeverity` already refuses below 60 PA on
the smaller side and declares switch hitters unreadable. On a refusal the season
rate stands **exactly** untouched (asserted to 6 dp) and the reason is recorded.
**Refusals are counted by reason** (`context.onRefusal`), and the count is taken
off the **stored blocks**, not the legs — a prop with both an over and an under
row produces two legs sharing one block, and the first draft reported 48.1% and
55.1% for the same fact over two different denominators.
## 8.3 Measured — real board, 2026-08-07 → 08-12, 9,792 graded hits rows
```
CHAIN FIRE — 9,752/9,792 atoms read (40 refused) across 83/83 games
refusal reasons: { no_batter_profile: 40 }
HAND SPLIT — fired on 288/553 unique props (52.1%)
season-rate reasons:
insufficient_split_sample 130 the honest 60-PA refusal
no_pitcher_hand 71 the only FIXABLE gap
switch_hitter_side_value_unknown 60 genuinely unreadable
no_splits / missing_split 4
CHAIN vs COUNTER — n=9,752
min p25 median p75 max mean
chain_p 0.016 0.359 0.498 0.641 0.984 0.500
counter_p 0.050 0.397 0.500 0.603 0.950 0.500
|divergence| 0.000 0.044 0.095 0.164 0.509 0.113
|div| < 0.02 1,142 (11.7%) 0.10–0.20 3,129 (32.1%)
0.02–0.05 1,608 (16.5%) >= 0.20 1,538 (15.8%)
0.05–0.10 2,335 (23.9%)
OVER SIDE ONLY (n=4,879): mean +0.045, chain HIGHER on 66.3%
MATCHUP READ vs SEASON READ
split FIRED (n=5,372 rows) median |div| 0.091 · >= 0.10 on 46.0% · within 0.02 on 13.0%
split REFUSED (n=4,380 rows) median |div| 0.100 · >= 0.10 on 50.1% · within 0.02 on 10.2%
```
## 8.4 Reading it honestly
- **The hand split fires on 52.1% of props**, up from 0. Of the 47.9% that do
not, **190 of 265 (72%) are principled refusals** — a thin split or a switch
hitter. Only `no_pitcher_hand` (71) is a plumbing gap, and it is the A5 shape
exactly: the hitter's split is sitting right there and the pitcher hand is
missing because that player had no lineup row.
- **Feeding the split widened the divergence**: median |div| 0.082 → **0.095**,
and the over-side lean +2.9pp → **+4.5pp**. The chain moved further from the
counter, which is what a matchup conditioner should do and is *not* evidence
it moved in the right direction.
- **The rows where the split FIRED disagree with the counter slightly LESS**
(median 0.091 vs 0.100) than the rows where it refused. **This comparison is
confounded and must not be read as an effect**: a hitter with 60+ PA on both
sides is an established regular, and the counter has more game log on him too.
It is two different populations, not two treatments.
- **Divergence is still not merit.** Nothing here says the chain is closer to
what happened. That needs settled outcomes against the stored triple, and it
goes through the gate as a new hypothesis against the cumulative denominator.
- **The probe remains contaminated** for the batter profile
(`statcast_aggregates` keeps one as-of date) and so reports a divergence, never
a resolution. The forward cron accrual does not have this problem.
## 8.5 Still not claimed
Unchanged from §5: nothing promoted, no head-to-head run, `CALIBRATION_DEPLOYED`
still `[]`, A8 untouched, `hitsFactors` untouched, the counter still serves,
WNBA untouched.
## 8.6 Second incidental defect found and fixed
`hitsFactorContext.build` and `matchupKeys.build` were both gated on
`require('../utils/supabase').getSupabaseServiceClient()` called inline, so the
entire factor and hand-split path was **unreachable from any test**. "It is
wired" could only ever have rested on reading the code — which is precisely how
A5 shipped three factors that never fired. Both now read `deps.supabase ||` the
real client (additive; `undefined` gives identical behaviour).
---
# §9 — chain v3: the opportunity term, and the artifact/signal diagnostic (2026-08-12)
## 9.1 The premise, again checked first — one half right, one half wrong
**RIGHT, and a real defect:** the v2 shadow passed **no `lineupSlotFor`**, so
every hitter fell to `skillProjection.DEFAULT_PA = 4.1`. "A regular" was asserted
about the leadoff man and the nine hole alike. The chain is `rate × opportunity`
and the opportunity half was a constant across the entire lineup — a free,
known, pre-game fact thrown away. Fixed.
**WRONG about the conversion, and about the direction:**
- *"mis-handles the multiple-chances structure"* — it does not.
`paDistribution` is a mean-preserving two-point mixture, and
`atLeast(Binomial(n, p), 1)` **is** `1 − (1−p)^n` averaged over n. The order's
proposed formula is what the code already computes. The defect was the INPUT
`E[PA]`, not the conversion.
- *"5.7pts BELOW the counter on 92% of rows"* — **measured the other way.** On
the over side the chain runs **+4.5pp ABOVE** the counter and is below on
**33.6%**. It was never 92%-below, at any point, in any measurement here.
## 9.2 Phase 1 — the E[PA] fix, A/B on identical rows
`matchup_keys` now carries `batting_order` (one extra column on a read it
already performs — no new query), and the shadow feeds it as the opportunity
term. Absent ⇒ `DEFAULT_PA` and the block records `default_regular`, so a
league-shaped opportunity term is never mistaken for a posted one.
```
OPPORTUNITY — posted lineup slot on 491/553 props (88.8%)
fell to a default regular: 62
n mean median chain BELOW counter
REAL E[PA] (posted slot) 4,879 +0.0450 +0.0499 33.6%
CONSTANT 4.1 (the v2 shadow) 4,879 +0.0322 +0.0363 39.6%
```
**The uniform bias did not collapse, because it was never there to collapse.**
The fix moved the chain *further above* the counter, not toward it — which is
correct behaviour, not a regression: real slots raise E[PA] for the top of the
order and lower it for the bottom, and top-of-order hitters are over-represented
in the prop board.
## 9.3 Phase 2 — THE DIAGNOSTIC: artifact or signal?
Over side, n=3,705 rows carrying an opposing-pitcher profile, quintiles of
opposing-pitcher K% (low = soft matchup):
| bucket | opp K% | n | mean divergence | chain below counter |
|---|---|---|---|---|
| Q1 | 10.8–18.4 | 741 | **+0.0749** | 27.8% |
| Q2 | 18.4–20.3 | 741 | +0.0560 | 30.8% |
| Q3 | 20.3–23.1 | 741 | +0.0287 | 34.1% |
| Q4 | 23.1–26.6 | 741 | +0.0396 | 35.0% |
| Q5 | 26.6–40.6 | 741 | **+0.0256** | 38.6% |
**Q1 − Q5 spread +0.0494 · Pearson r(oppK, divergence) = −0.119**
### The answer is BOTH, and the order's binary framing does not fit
The order asked for *uniform (broken)* **or** *difficulty-correlated (signal)*.
The data is a **mixture of the two, and both components should be named**:
- **A difficulty-correlated component, ~+0.049 across the range.** The chain
reads soft matchups higher and hard matchups lower than the counter does. The
`below %` is **monotone across all five quintiles** (27.8 → 30.8 → 34.1 → 35.0
→ 38.6), which is a cleaner signature than the means (Q4 breaks order).
- **A uniform positive offset of ~+0.026.** Even in the HARDEST quintile the
chain sits +2.6pp above the counter. That floor does not move with the matchup
and is therefore not conditioning — it is exactly the shape a residual
mechanical bias makes. It is smaller than the conditioning component but it
has not been explained, and calling the whole result "signal" would bury it.
### The caveat that must travel with the correlation
**This is close to mechanically guaranteed and is NOT evidence of correctness.**
The chain reads the opposing pitcher's K rate directly (log5 odds-ratio in
`paOutcome`) and the counter reads nothing about the pitcher at all. So a
monotone relationship between opposing-pitcher K% and chain-minus-counter is
approximately a proof that *the wiring works* — that the pitcher input reaches
the number. It says nothing about whether the adjustment is the right size, the
right direction on any individual row, or better than ignoring the pitcher.
**Does the chain see the game, or just miscompute it?** It demonstrably
CONDITIONS on the game — the pitcher input reaches the forecast and moves it in
the theorised direction. Whether that conditioning is *right* is unanswerable
from a divergence and needs settled outcomes against the stored triple. There is
also a residual ~2.6pp offset that conditioning does not explain and that should
be chased before anyone reads the correlation as a win.
## 9.4 Open item created by this measurement
The ~+0.026 floor. Candidates not yet tested: the chain is unclamped where the
counter clamps to `[0.10, 0.95]` (chain min 0.016 vs counter min 0.050); the
`LEAGUE.babip = 0.291` anchor in `hitOnContact`; the ±35% BABIP bound. Naming it
as unexplained is the honest state — it is not yet an artifact and not yet
signal.
## 9.5 Unchanged
Nothing promoted, no head-to-head, `CALIBRATION_DEPLOYED` still `[]`, A8 and
`hitsFactors` untouched, the counter still serves, WNBA untouched. Migration 038
is still an unapplied blocking precondition of deploy.