# chain-v1 / v2 — the portable engine, made whole, fed, and shadowed on MLB > **v2 (2026-08-12) is appended at §8.** It corrects the order's premise (the > chain was already firing on 99.6%, `paOutcome` reads no handedness, and there > is no fallback path anywhere) and plumbs the hand split as a matchup > conditioner on the RATE. Read §8 for the current numbers; §4 below is the v1 > measurement and is superseded. **Status:** SHADOW. Served by nothing. Nothing promoted. **Date:** 2026-08-12 **Predecessors:** chaining-v1 (Session 90 — the design, which never had a spec file), A5 (factor freeze), A6 (the join keys, shadow), A7 (shadow accrual). --- ## 0. The finding this order started from The extraction pass found `chain.js` was a **shell of its own header**: | the header promised | the code had | |---|---| | `atoms` | a positional array ✅ | | `context` | passed to `redistribute` only, partially | | **`chainFn`** — atom → per-entity probability | **ABSENT. No parameter, no call site, no export.** | | `aggregator` — ACROSS / UP | two functions ✅ | | `redistribute` | **on `chainUp` only** | And it was inert for a second, undocumented reason: `chainAcross` requires `calibrated: true`, and the only writer of that flag is a loop over `CALIBRATION_DEPLOYED`, which is `Object.freeze([])`. **No grade in the system carries the flag**, so `chainAcross` had zero callers *and* would have refused any caller it had. The missing `chainFn` is why the "portable core" was not portable: with no slot for the atom→probability stage, every sport's real work had to live somewhere else. For MLB it lived in `scripts/`, reachable from no pipeline. --- ## 1. The three core fixes (`src/services/model/chain.js`) ### 1.1 `chainFn(atom, context) → per-entity probability` - `applyChainFn` runs before usability filtering. **Default is identity-on-`p`**, so every pre-existing caller is byte-identical. - Returning a **number** sets `p`; returning an **object** merges (so a chainFn can attach its own trace); returning **null or throwing** makes the atom UNREADABLE ⇒ **DROPPED, counted, never `p=0`**. A zero leg would zero an entire ticket, and "we could not read him" is not "he cannot do it". - The refusals are surfaced as `chain_fn_refused` rather than swallowed, so a chainFn quietly failing across the board is visible instead of looking like a thin slate. ### 1.2 Correlation is SIGNED, and the direction follows the sign Was clamped `[0, 1]` with the joint always shifted toward the weakest leg. That is baseball's shape — same-game legs share the pitcher, the park and the weather — and it made basketball's case **inexpressible**: teammates compete for finite possessions, so one player's shot is another's non-shot and their props are NEGATIVELY correlated. Correlation is now clamped `[-1, 1]` and `|corr|` interpolates from independence toward the **Fréchet–Hoeffding bound its sign selects**: ``` corr = +1 → joint = min(p_i) (upper bound, co-monotone) corr = 0 → joint = Π p_i (independent) corr = −1 → joint = max(0, Σp − (n−1)) (lower bound, counter-monotone) ``` This is why the direction is principled rather than chosen. The positive branch is arithmetically unchanged — the weakest leg **is** the upper bound — so every previously-correct number stays exactly what it was. ### 1.3 `redistribute` reaches BOTH readings Previously on `chainUp` only. A redistribution that reached the team read and not the across read would leave `selfCheck` comparing post-redistribution to pre-redistribution atoms and flagging an INTERNAL_INCONSISTENCY **the model had itself just manufactured**. `prepareAtoms(atoms, opts)` is exported so a caller prepares ONCE — chainFn, then usability, then redistribution — and hands the identical legs to both readings. --- ## 2. Baseball's chainFn (`src/services/model/baseballChain.js`) A chained forecast is always `rate × opportunity`. Baseball is the clean case because **opportunity is fixed**: the batting order is set before first pitch, a nine-run lead does not change who bats next, and PA/game varies over a narrow range set almost entirely by lineup slot. ``` p_hit_per_PA ← skillProjection.paOutcome the RATE (modelled) expected PA ← lineup slot the OPPORTUNITY (a lookup) P(hits ≥ k) ← Binomial(PA, p) mixed over paDistribution ``` **Atoms routed** — all pre-existing, all previously reachable only from `scripts/`: | atom | source | role | |---|---|---| | `fromStatcastRow` | skillProjection:161 | the ONE legal units conversion (statcast stores PERCENTAGES 0–100) | | `paOutcome` | skillProjection:268 | K / BB via log5 odds-ratio vs league; remainder = balls in play | | `hitOnContact` | skillProjection:213 | archetype-selected barrel / hard-hit / exit-velo / GB-speed × pitcher contact allowed × park | | `paDistribution`, `binomialPmf`, `atLeast` | skillProjection:327/308/339 | the chain over opportunity | `PA_BY_SLOT` is a **lookup, not a fit** — 4.65 (leadoff) down to 3.85 (nine hole), the documented ~0.1-PA-per-slot decline. Nothing was tuned on settled rows; a tuned opportunity term on 1,741 rows is curve-fitting dressed as physics. **Refusals:** no batter profile ⇒ null ⇒ dropped (never a league hitter). A missing pitcher is different and handled inside `paOutcome` — the batter's own rate stands rather than being pulled toward average. **Scope: hits only.** `projectSkill` also routes `total_bases`, but TB's head-to-head is on record as INCONCLUSIVE-under-contamination. Widening the first shadow to a stat whose verdict is already muddy buys noise. **`redistribute` is dormant** and returns the legs unchanged — the honest dormant behaviour. Returning null would read as "the hook failed". --- ## 3. The shadow (`src/services/model/chainShadow.js`) Runs at the post-enriched / pre-persist point in `snapshotService` — the A6 position, where the slate exists as a set. Reuses the statcast rows and resolved opposing starter the challenger pass already fetched: **zero new I/O**. ### The unit of evidence is the TRIPLE ``` (chain_p, counter_p, outcome) ``` all three on one row, all side-aligned. `counter_p` is the served `estimateProbability` value for that exact side; `outcome` arrives from the ordinary settle pass; `chain_p` is the chain's. **Evidence you cannot adjudicate is not evidence** — a chain probability stored without the number it must beat, or without the result, can only ever be compared to itself. **Side alignment is load-bearing.** The chain computes P(over the line); `p_win` is expressed for the graded SIDE. An under row stores `1 − p_over`. Storing the raw over-probability against an under row would invert every later comparison, silently. ### UN-SERVABLE, said in the data `requireCalibrated: false` is legitimate **only** because nothing downstream reads the result. Every stored block carries `status: 'UN-SERVABLE'` and `servable: false` in its own payload, not merely in a comment — a caveat that lives only in a comment is not attached to the data once something else queries it. Flipping that requires passing the gate, not editing a file. ### The self-check is VACUOUS today, and says so `selfCheck` earns its keep by comparing per-entity reads to an **independent** team read. None exists — the game-script projection was deliberately not built (no atom has passed the gate). So the up-read is assembled from the SAME atoms as the across-read and agreement between them is arithmetic. Every block carries `self_check.vacuous: true` with its reason, so nobody later mistakes a tautology for a passing consistency test. ### Storage `model_snapshots.chain_shadow jsonb` (migration 038). A separate column, not extra keys inside `features` — `champion-ablation.js` iterates every `features` key for its residual scan, so widening it would silently enlarge that multiple-comparisons denominator. Same reasoning as 034. --- ## 4. First measurement (`scripts/chain-shadow-probe.js`) Real board, 2026-08-07 → 2026-08-12, MLB hits, **9,376 graded rows**: ``` CHAIN FIRE — 9,340/9,376 atoms read (36 refused) across 82/82 games [99.6%] min p25 median p75 max mean chain_p 0.013 0.382 0.500 0.618 0.987 0.500 counter_p 0.050 0.396 0.500 0.604 0.950 0.500 divergence (signed) −0.556 −0.082 0.000 0.082 0.556 −0.000 divergence (abs) 0.000 0.038 0.082 0.142 0.556 0.100 |divergence| < 0.02 1,203 (12.9%) agrees with the counter 0.02 – 0.05 1,760 (18.8%) 0.05 – 0.10 2,577 (27.6%) 0.10 – 0.20 2,835 (30.4%) ≥ 0.20 965 (10.3%) a different read entirely OVER SIDE ONLY (n=4,673) divergence (signed) −0.526 −0.045 0.031 0.104 0.556 0.029 chain HIGHER on 2,837 (60.7%) ``` **Reading it honestly:** - The chain **fires**, on 99.6% of the board. It is not the A5 case (built, correct, never invoked). - It is **not a relabelled counter**: 40.7% of rows differ by ≥0.10, and 10.3% by ≥0.20. It is also not noise — 12.9% agree inside 0.02. - The **symmetry of the two-sided signed distribution is arithmetic**, not a finding: both sides of every prop are in the sample, so each pair contributes `+d` and `−d`. Reading `mean −0.000` as "unbiased" would be reading the sampling scheme. The over-side slice is the one that can lean, and it does: **+2.9pp mean, higher on 60.7%**. - **Divergence is not merit.** A challenger that disagrees is interesting, not right. Which of the two is closer to what happened is the settle pass's question, and it is exactly why the triple is stored. - **The probe is contaminated by construction** and reports only a divergence (a property of two forecasts) rather than a resolution (a property of a forecast against an outcome): `statcast_aggregates` is upserted in place and keeps one as-of date, so profiles read for a row graded three days ago are today's. The forward accrual on the cron does not have this problem. - The probe resolves **no opposing pitcher** (offline), so it measures the batter-side read. The live shadow does resolve it. --- ## 5. What is NOT claimed - **Nothing is promoted.** The proven set remains empty. - **The chain is not calibrated** and has passed no gate. - **No head-to-head has been run.** That needs settled outcomes against the stored triples and must go through `factorGate` / the cumulative Bonferroni denominator, as a NEW hypothesis. - **`CALIBRATION_DEPLOYED` stays `[]`.** Nothing is served calibrated. - **WNBA is untouched.** Its contested-possession chainFn and its possession feed are later orders. The core changes (signed correlation, redistribute on both readings) were built now because they are engine honesty, not because MLB needs them — MLB exercises neither. --- ## 6. The pre-registered next step Once settled outcomes accrue against `chain_shadow`: 1. Score `chain_p` vs `counter_p` on the SAME rows (paired bootstrap — comparing independent SEs overstates uncertainty and has previously read a reliable −0.022 as noise). 2. Per stat, never pooled (pooled resolution is inflated by base-rate structure). 3. Through the cumulative test ledger; the CI widens to `1 − 0.05/tests`. 4. **Pre-registered fallback, stated before the answer is known:** if the chain moves ~87% of the board and does not improve Brier, it is THEATER by the `factorGate` definition and the correct action is to leave it off — not to re-tune the opportunity term until it passes. --- ## 7. Incidental defect found and fixed `runSnapshot`'s `deps` object was an **allowlist of 17 keys**, but fourteen call sites read `deps.challenger`, `deps.loadStatcast`, `deps.environmentContext`, `deps.lineupContext`, `deps.hitsFactorContext`, `deps.matchupKeys`, `deps.gameBinder`, `deps.archetypeAxes`, `deps.contactChallenger`, `deps.projectionChallenger`, `deps.loadArsenals`, `deps.mlbAdapter` — each documented as injectable, each **permanently `undefined`**, each always falling through to the real module. The seam existed in the comment and not in the code. Fixed by spreading `...opts` FIRST in the literal: every explicit key is declared after and already reads `opts.X`, so no existing behaviour moves, while an unlisted dep now actually arrives. Found because the chain shadow's own test could not inject a statcast map — the test would have passed while measuring nothing, which is the failure this whole line of work exists to stop repeating. --- # §8 — chain v2: FEED THE ENGINE (2026-08-12) ## 8.1 The order's premise, checked before building The v2 order opened from four numbers. Each was checked against the code and the board before anything was written: | claim | measured | |---|---| | "`paOutcome` returned null on 596/725 rows" | **False.** Fire rate is **9,752/9,792 (99.6%)**. The only refusal reason on the whole board is `no_batter_profile: 40`. `pa_outcome_refused` never fires. | | "it needs the batter-vs-pitcher-hand split" *to run* | **False.** `paOutcome` and `hitOnContact` read `k_pct`, `bb_pct`, `barrel_pct`, `hard_hit_pct`, `avg_exit_velo`, `avg_launch_angle` and the pitcher's `k_pct`/`bb_pct`/`hard_hit_pct`. Neither reads `bats` or `throws` at all. A hand split cannot change whether they run. | | "82% FELL BACK to the seasonal rate, i.e. became the counter" | **No fallback path exists.** `baseballChain.chainFn` returns null on any refusal and `chain.applyChainFn` DROPS the atom. Nothing in the module reads a season rate as a substitute. | | "425/425 served-identical" | That figure is from commit `7c8ef8b` (the A1–A7 deploy verification), not from the chain. v1 measured **2/2 served-identical** in the harness and byte-identical served payloads. | Chain v1 was also never committed or deployed and migration 038 was never applied, so no chain-shadow rows exist in production — the premise numbers cannot have come from a chain-shadow run. **The premise was wrong about the mechanism. It was right about the thing that matters:** the chain was reading a hitter's SEASON rates, which already average his platoon split over whichever hands he happened to face. That is a season read wearing a matchup read's clothes, and un-averaging it is real work. So v2 plumbs the hand split — as a **conditioner on the rate**, not as a fix to the fire rate. ## 8.2 What was built **The split enters at the per-PA hit rate, not at the output probability.** Multiplying `P(hits ≥ 1)` by a platoon factor would scale a number that has already been through the opportunity term — a different and wrong claim. So `baseballChain.chainFn` now writes the chain out explicitly: ``` paOutcome → p_hit_per_pa (season) → × platoonRead multiplier ← THE MATCHUP CONDITIONER → Binomial(PA, p_hit) over paDistribution → P(hits ≥ k) ``` A test asserts that with no split supplied this is **arithmetically identical to `projectSkill`** across three archetypes × three PA values, so the restructuring cannot quietly become a second model. **As-of-correct (A4).** The hitter's own split comes from `hitsFactorContext` (whose reads are `lte('as_of_date', asOf)`), the opposing starter's hand through the A6 `matchupKeys` resolve. The probe bounds every read at the row's own `game_date`. Nothing later than the grade can enter. **Refuse, never substitute.** `platoonSeverity` already refuses below 60 PA on the smaller side and declares switch hitters unreadable. On a refusal the season rate stands **exactly** untouched (asserted to 6 dp) and the reason is recorded. **Refusals are counted by reason** (`context.onRefusal`), and the count is taken off the **stored blocks**, not the legs — a prop with both an over and an under row produces two legs sharing one block, and the first draft reported 48.1% and 55.1% for the same fact over two different denominators. ## 8.3 Measured — real board, 2026-08-07 → 08-12, 9,792 graded hits rows ``` CHAIN FIRE — 9,752/9,792 atoms read (40 refused) across 83/83 games refusal reasons: { no_batter_profile: 40 } HAND SPLIT — fired on 288/553 unique props (52.1%) season-rate reasons: insufficient_split_sample 130 the honest 60-PA refusal no_pitcher_hand 71 the only FIXABLE gap switch_hitter_side_value_unknown 60 genuinely unreadable no_splits / missing_split 4 CHAIN vs COUNTER — n=9,752 min p25 median p75 max mean chain_p 0.016 0.359 0.498 0.641 0.984 0.500 counter_p 0.050 0.397 0.500 0.603 0.950 0.500 |divergence| 0.000 0.044 0.095 0.164 0.509 0.113 |div| < 0.02 1,142 (11.7%) 0.10–0.20 3,129 (32.1%) 0.02–0.05 1,608 (16.5%) >= 0.20 1,538 (15.8%) 0.05–0.10 2,335 (23.9%) OVER SIDE ONLY (n=4,879): mean +0.045, chain HIGHER on 66.3% MATCHUP READ vs SEASON READ split FIRED (n=5,372 rows) median |div| 0.091 · >= 0.10 on 46.0% · within 0.02 on 13.0% split REFUSED (n=4,380 rows) median |div| 0.100 · >= 0.10 on 50.1% · within 0.02 on 10.2% ``` ## 8.4 Reading it honestly - **The hand split fires on 52.1% of props**, up from 0. Of the 47.9% that do not, **190 of 265 (72%) are principled refusals** — a thin split or a switch hitter. Only `no_pitcher_hand` (71) is a plumbing gap, and it is the A5 shape exactly: the hitter's split is sitting right there and the pitcher hand is missing because that player had no lineup row. - **Feeding the split widened the divergence**: median |div| 0.082 → **0.095**, and the over-side lean +2.9pp → **+4.5pp**. The chain moved further from the counter, which is what a matchup conditioner should do and is *not* evidence it moved in the right direction. - **The rows where the split FIRED disagree with the counter slightly LESS** (median 0.091 vs 0.100) than the rows where it refused. **This comparison is confounded and must not be read as an effect**: a hitter with 60+ PA on both sides is an established regular, and the counter has more game log on him too. It is two different populations, not two treatments. - **Divergence is still not merit.** Nothing here says the chain is closer to what happened. That needs settled outcomes against the stored triple, and it goes through the gate as a new hypothesis against the cumulative denominator. - **The probe remains contaminated** for the batter profile (`statcast_aggregates` keeps one as-of date) and so reports a divergence, never a resolution. The forward cron accrual does not have this problem. ## 8.5 Still not claimed Unchanged from §5: nothing promoted, no head-to-head run, `CALIBRATION_DEPLOYED` still `[]`, A8 untouched, `hitsFactors` untouched, the counter still serves, WNBA untouched. ## 8.6 Second incidental defect found and fixed `hitsFactorContext.build` and `matchupKeys.build` were both gated on `require('../utils/supabase').getSupabaseServiceClient()` called inline, so the entire factor and hand-split path was **unreachable from any test**. "It is wired" could only ever have rested on reading the code — which is precisely how A5 shipped three factors that never fired. Both now read `deps.supabase ||` the real client (additive; `undefined` gives identical behaviour). --- # §9 — chain v3: the opportunity term, and the artifact/signal diagnostic (2026-08-12) ## 9.1 The premise, again checked first — one half right, one half wrong **RIGHT, and a real defect:** the v2 shadow passed **no `lineupSlotFor`**, so every hitter fell to `skillProjection.DEFAULT_PA = 4.1`. "A regular" was asserted about the leadoff man and the nine hole alike. The chain is `rate × opportunity` and the opportunity half was a constant across the entire lineup — a free, known, pre-game fact thrown away. Fixed. **WRONG about the conversion, and about the direction:** - *"mis-handles the multiple-chances structure"* — it does not. `paDistribution` is a mean-preserving two-point mixture, and `atLeast(Binomial(n, p), 1)` **is** `1 − (1−p)^n` averaged over n. The order's proposed formula is what the code already computes. The defect was the INPUT `E[PA]`, not the conversion. - *"5.7pts BELOW the counter on 92% of rows"* — **measured the other way.** On the over side the chain runs **+4.5pp ABOVE** the counter and is below on **33.6%**. It was never 92%-below, at any point, in any measurement here. ## 9.2 Phase 1 — the E[PA] fix, A/B on identical rows `matchup_keys` now carries `batting_order` (one extra column on a read it already performs — no new query), and the shadow feeds it as the opportunity term. Absent ⇒ `DEFAULT_PA` and the block records `default_regular`, so a league-shaped opportunity term is never mistaken for a posted one. ``` OPPORTUNITY — posted lineup slot on 491/553 props (88.8%) fell to a default regular: 62 n mean median chain BELOW counter REAL E[PA] (posted slot) 4,879 +0.0450 +0.0499 33.6% CONSTANT 4.1 (the v2 shadow) 4,879 +0.0322 +0.0363 39.6% ``` **The uniform bias did not collapse, because it was never there to collapse.** The fix moved the chain *further above* the counter, not toward it — which is correct behaviour, not a regression: real slots raise E[PA] for the top of the order and lower it for the bottom, and top-of-order hitters are over-represented in the prop board. ## 9.3 Phase 2 — THE DIAGNOSTIC: artifact or signal? Over side, n=3,705 rows carrying an opposing-pitcher profile, quintiles of opposing-pitcher K% (low = soft matchup): | bucket | opp K% | n | mean divergence | chain below counter | |---|---|---|---|---| | Q1 | 10.8–18.4 | 741 | **+0.0749** | 27.8% | | Q2 | 18.4–20.3 | 741 | +0.0560 | 30.8% | | Q3 | 20.3–23.1 | 741 | +0.0287 | 34.1% | | Q4 | 23.1–26.6 | 741 | +0.0396 | 35.0% | | Q5 | 26.6–40.6 | 741 | **+0.0256** | 38.6% | **Q1 − Q5 spread +0.0494 · Pearson r(oppK, divergence) = −0.119** ### The answer is BOTH, and the order's binary framing does not fit The order asked for *uniform (broken)* **or** *difficulty-correlated (signal)*. The data is a **mixture of the two, and both components should be named**: - **A difficulty-correlated component, ~+0.049 across the range.** The chain reads soft matchups higher and hard matchups lower than the counter does. The `below %` is **monotone across all five quintiles** (27.8 → 30.8 → 34.1 → 35.0 → 38.6), which is a cleaner signature than the means (Q4 breaks order). - **A uniform positive offset of ~+0.026.** Even in the HARDEST quintile the chain sits +2.6pp above the counter. That floor does not move with the matchup and is therefore not conditioning — it is exactly the shape a residual mechanical bias makes. It is smaller than the conditioning component but it has not been explained, and calling the whole result "signal" would bury it. ### The caveat that must travel with the correlation **This is close to mechanically guaranteed and is NOT evidence of correctness.** The chain reads the opposing pitcher's K rate directly (log5 odds-ratio in `paOutcome`) and the counter reads nothing about the pitcher at all. So a monotone relationship between opposing-pitcher K% and chain-minus-counter is approximately a proof that *the wiring works* — that the pitcher input reaches the number. It says nothing about whether the adjustment is the right size, the right direction on any individual row, or better than ignoring the pitcher. **Does the chain see the game, or just miscompute it?** It demonstrably CONDITIONS on the game — the pitcher input reaches the forecast and moves it in the theorised direction. Whether that conditioning is *right* is unanswerable from a divergence and needs settled outcomes against the stored triple. There is also a residual ~2.6pp offset that conditioning does not explain and that should be chased before anyone reads the correlation as a win. ## 9.4 Open item created by this measurement The ~+0.026 floor. Candidates not yet tested: the chain is unclamped where the counter clamps to `[0.10, 0.95]` (chain min 0.016 vs counter min 0.050); the `LEAGUE.babip = 0.291` anchor in `hitOnContact`; the ±35% BABIP bound. Naming it as unexplained is the honest state — it is not yet an artifact and not yet signal. ## 9.5 Unchanged Nothing promoted, no head-to-head, `CALIBRATION_DEPLOYED` still `[]`, A8 and `hitsFactors` untouched, the counter still serves, WNBA untouched. Migration 038 is still an unapplied blocking precondition of deploy.