Files
builtbykev 6c34af3414 checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion
recon (Tony resurrection, STEP 0). This work existed only on the
laptop disk.

- chain shadow accrual + probe script (038_chain_shadow.sql)
- WNBA possession feed: ESPN adapter, usage service, verify script
  (039_wnba_player_game.sql)
- baseball chain
- retention/snapshot service updates, tableKeys, matchupKeys
- specs: chain-v1, wnba-possession-feed, wnba-source-survey
- unit tests for the above

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
2026-08-14 16:53:37 -04:00

24 KiB
Raw Permalink Blame History

chain-v1 / v2 — the portable engine, made whole, fed, and shadowed on MLB

v2 (2026-08-12) is appended at §8. It corrects the order's premise (the chain was already firing on 99.6%, paOutcome reads no handedness, and there is no fallback path anywhere) and plumbs the hand split as a matchup conditioner on the RATE. Read §8 for the current numbers; §4 below is the v1 measurement and is superseded.

Status: SHADOW. Served by nothing. Nothing promoted. Date: 2026-08-12 Predecessors: chaining-v1 (Session 90 — the design, which never had a spec file), A5 (factor freeze), A6 (the join keys, shadow), A7 (shadow accrual).


0. The finding this order started from

The extraction pass found chain.js was a shell of its own header:

the header promised the code had
atoms a positional array ✅
context passed to redistribute only, partially
chainFn — atom → per-entity probability ABSENT. No parameter, no call site, no export.
aggregator — ACROSS / UP two functions ✅
redistribute on chainUp only

And it was inert for a second, undocumented reason: chainAcross requires calibrated: true, and the only writer of that flag is a loop over CALIBRATION_DEPLOYED, which is Object.freeze([]). No grade in the system carries the flag, so chainAcross had zero callers and would have refused any caller it had.

The missing chainFn is why the "portable core" was not portable: with no slot for the atom→probability stage, every sport's real work had to live somewhere else. For MLB it lived in scripts/, reachable from no pipeline.


1. The three core fixes (src/services/model/chain.js)

1.1 chainFn(atom, context) → per-entity probability

  • applyChainFn runs before usability filtering. Default is identity-on-p, so every pre-existing caller is byte-identical.
  • Returning a number sets p; returning an object merges (so a chainFn can attach its own trace); returning null or throwing makes the atom UNREADABLE ⇒ DROPPED, counted, never p=0. A zero leg would zero an entire ticket, and "we could not read him" is not "he cannot do it".
  • The refusals are surfaced as chain_fn_refused rather than swallowed, so a chainFn quietly failing across the board is visible instead of looking like a thin slate.

1.2 Correlation is SIGNED, and the direction follows the sign

Was clamped [0, 1] with the joint always shifted toward the weakest leg. That is baseball's shape — same-game legs share the pitcher, the park and the weather — and it made basketball's case inexpressible: teammates compete for finite possessions, so one player's shot is another's non-shot and their props are NEGATIVELY correlated.

Correlation is now clamped [-1, 1] and |corr| interpolates from independence toward the Fréchet–Hoeffding bound its sign selects:

corr = +1  →  joint = min(p_i)                  (upper bound, co-monotone)
corr =  0  →  joint = Π p_i                     (independent)
corr = −1  →  joint = max(0, Σp − (n−1))        (lower bound, counter-monotone)

This is why the direction is principled rather than chosen. The positive branch is arithmetically unchanged — the weakest leg is the upper bound — so every previously-correct number stays exactly what it was.

1.3 redistribute reaches BOTH readings

Previously on chainUp only. A redistribution that reached the team read and not the across read would leave selfCheck comparing post-redistribution to pre-redistribution atoms and flagging an INTERNAL_INCONSISTENCY the model had itself just manufactured.

prepareAtoms(atoms, opts) is exported so a caller prepares ONCE — chainFn, then usability, then redistribution — and hands the identical legs to both readings.


2. Baseball's chainFn (src/services/model/baseballChain.js)

A chained forecast is always rate × opportunity. Baseball is the clean case because opportunity is fixed: the batting order is set before first pitch, a nine-run lead does not change who bats next, and PA/game varies over a narrow range set almost entirely by lineup slot.

p_hit_per_PA   ← skillProjection.paOutcome     the RATE (modelled)
expected PA    ← lineup slot                    the OPPORTUNITY (a lookup)
P(hits ≥ k)    ← Binomial(PA, p) mixed over paDistribution

Atoms routed — all pre-existing, all previously reachable only from scripts/:

atom source role
fromStatcastRow skillProjection:161 the ONE legal units conversion (statcast stores PERCENTAGES 0–100)
paOutcome skillProjection:268 K / BB via log5 odds-ratio vs league; remainder = balls in play
hitOnContact skillProjection:213 archetype-selected barrel / hard-hit / exit-velo / GB-speed × pitcher contact allowed × park
paDistribution, binomialPmf, atLeast skillProjection:327/308/339 the chain over opportunity

PA_BY_SLOT is a lookup, not a fit — 4.65 (leadoff) down to 3.85 (nine hole), the documented ~0.1-PA-per-slot decline. Nothing was tuned on settled rows; a tuned opportunity term on 1,741 rows is curve-fitting dressed as physics.

Refusals: no batter profile ⇒ null ⇒ dropped (never a league hitter). A missing pitcher is different and handled inside paOutcome — the batter's own rate stands rather than being pulled toward average.

Scope: hits only. projectSkill also routes total_bases, but TB's head-to-head is on record as INCONCLUSIVE-under-contamination. Widening the first shadow to a stat whose verdict is already muddy buys noise.

redistribute is dormant and returns the legs unchanged — the honest dormant behaviour. Returning null would read as "the hook failed".


3. The shadow (src/services/model/chainShadow.js)

Runs at the post-enriched / pre-persist point in snapshotService — the A6 position, where the slate exists as a set. Reuses the statcast rows and resolved opposing starter the challenger pass already fetched: zero new I/O.

The unit of evidence is the TRIPLE

(chain_p, counter_p, outcome)

all three on one row, all side-aligned. counter_p is the served estimateProbability value for that exact side; outcome arrives from the ordinary settle pass; chain_p is the chain's. Evidence you cannot adjudicate is not evidence — a chain probability stored without the number it must beat, or without the result, can only ever be compared to itself.

Side alignment is load-bearing. The chain computes P(over the line); p_win is expressed for the graded SIDE. An under row stores 1 − p_over. Storing the raw over-probability against an under row would invert every later comparison, silently.

UN-SERVABLE, said in the data

requireCalibrated: false is legitimate only because nothing downstream reads the result. Every stored block carries status: 'UN-SERVABLE' and servable: false in its own payload, not merely in a comment — a caveat that lives only in a comment is not attached to the data once something else queries it. Flipping that requires passing the gate, not editing a file.

The self-check is VACUOUS today, and says so

selfCheck earns its keep by comparing per-entity reads to an independent team read. None exists — the game-script projection was deliberately not built (no atom has passed the gate). So the up-read is assembled from the SAME atoms as the across-read and agreement between them is arithmetic. Every block carries self_check.vacuous: true with its reason, so nobody later mistakes a tautology for a passing consistency test.

Storage

model_snapshots.chain_shadow jsonb (migration 038). A separate column, not extra keys inside features — champion-ablation.js iterates every features key for its residual scan, so widening it would silently enlarge that multiple-comparisons denominator. Same reasoning as 034.


4. First measurement (scripts/chain-shadow-probe.js)

Real board, 2026-08-07 → 2026-08-12, MLB hits, 9,376 graded rows:

CHAIN FIRE — 9,340/9,376 atoms read (36 refused) across 82/82 games   [99.6%]

                       min     p25    median    p75     max    mean
chain_p              0.013   0.382    0.500   0.618   0.987   0.500
counter_p            0.050   0.396    0.500   0.604   0.950   0.500
divergence (signed) −0.556  −0.082    0.000   0.082   0.556  −0.000
divergence (abs)     0.000   0.038    0.082   0.142   0.556   0.100

|divergence| < 0.02    1,203  (12.9%)   agrees with the counter
       0.02 – 0.05     1,760  (18.8%)
       0.05 – 0.10     2,577  (27.6%)
       0.10 – 0.20     2,835  (30.4%)
            ≥ 0.20       965  (10.3%)   a different read entirely

OVER SIDE ONLY (n=4,673)
divergence (signed) −0.526  −0.045    0.031   0.104   0.556   0.029
chain HIGHER on 2,837 (60.7%)

Reading it honestly:

  • The chain fires, on 99.6% of the board. It is not the A5 case (built, correct, never invoked).
  • It is not a relabelled counter: 40.7% of rows differ by ≥0.10, and 10.3% by ≥0.20. It is also not noise — 12.9% agree inside 0.02.
  • The symmetry of the two-sided signed distribution is arithmetic, not a finding: both sides of every prop are in the sample, so each pair contributes +d and −d. Reading mean −0.000 as "unbiased" would be reading the sampling scheme. The over-side slice is the one that can lean, and it does: +2.9pp mean, higher on 60.7%.
  • Divergence is not merit. A challenger that disagrees is interesting, not right. Which of the two is closer to what happened is the settle pass's question, and it is exactly why the triple is stored.
  • The probe is contaminated by construction and reports only a divergence (a property of two forecasts) rather than a resolution (a property of a forecast against an outcome): statcast_aggregates is upserted in place and keeps one as-of date, so profiles read for a row graded three days ago are today's. The forward accrual on the cron does not have this problem.
  • The probe resolves no opposing pitcher (offline), so it measures the batter-side read. The live shadow does resolve it.

5. What is NOT claimed

  • Nothing is promoted. The proven set remains empty.
  • The chain is not calibrated and has passed no gate.
  • No head-to-head has been run. That needs settled outcomes against the stored triples and must go through factorGate / the cumulative Bonferroni denominator, as a NEW hypothesis.
  • CALIBRATION_DEPLOYED stays []. Nothing is served calibrated.
  • WNBA is untouched. Its contested-possession chainFn and its possession feed are later orders. The core changes (signed correlation, redistribute on both readings) were built now because they are engine honesty, not because MLB needs them — MLB exercises neither.

6. The pre-registered next step

Once settled outcomes accrue against chain_shadow:

  1. Score chain_p vs counter_p on the SAME rows (paired bootstrap — comparing independent SEs overstates uncertainty and has previously read a reliable −0.022 as noise).
  2. Per stat, never pooled (pooled resolution is inflated by base-rate structure).
  3. Through the cumulative test ledger; the CI widens to 1 − 0.05/tests.
  4. Pre-registered fallback, stated before the answer is known: if the chain moves ~87% of the board and does not improve Brier, it is THEATER by the factorGate definition and the correct action is to leave it off — not to re-tune the opportunity term until it passes.

7. Incidental defect found and fixed

runSnapshot's deps object was an allowlist of 17 keys, but fourteen call sites read deps.challenger, deps.loadStatcast, deps.environmentContext, deps.lineupContext, deps.hitsFactorContext, deps.matchupKeys, deps.gameBinder, deps.archetypeAxes, deps.contactChallenger, deps.projectionChallenger, deps.loadArsenals, deps.mlbAdapter — each documented as injectable, each permanently undefined, each always falling through to the real module. The seam existed in the comment and not in the code.

Fixed by spreading ...opts FIRST in the literal: every explicit key is declared after and already reads opts.X, so no existing behaviour moves, while an unlisted dep now actually arrives. Found because the chain shadow's own test could not inject a statcast map — the test would have passed while measuring nothing, which is the failure this whole line of work exists to stop repeating.


§8 — chain v2: FEED THE ENGINE (2026-08-12)

8.1 The order's premise, checked before building

The v2 order opened from four numbers. Each was checked against the code and the board before anything was written:

claim measured
"paOutcome returned null on 596/725 rows" False. Fire rate is 9,752/9,792 (99.6%). The only refusal reason on the whole board is no_batter_profile: 40. pa_outcome_refused never fires.
"it needs the batter-vs-pitcher-hand split" to run False. paOutcome and hitOnContact read k_pct, bb_pct, barrel_pct, hard_hit_pct, avg_exit_velo, avg_launch_angle and the pitcher's k_pct/bb_pct/hard_hit_pct. Neither reads bats or throws at all. A hand split cannot change whether they run.
"82% FELL BACK to the seasonal rate, i.e. became the counter" No fallback path exists. baseballChain.chainFn returns null on any refusal and chain.applyChainFn DROPS the atom. Nothing in the module reads a season rate as a substitute.
"425/425 served-identical" That figure is from commit 7c8ef8b (the A1–A7 deploy verification), not from the chain. v1 measured 2/2 served-identical in the harness and byte-identical served payloads.

Chain v1 was also never committed or deployed and migration 038 was never applied, so no chain-shadow rows exist in production — the premise numbers cannot have come from a chain-shadow run.

The premise was wrong about the mechanism. It was right about the thing that matters: the chain was reading a hitter's SEASON rates, which already average his platoon split over whichever hands he happened to face. That is a season read wearing a matchup read's clothes, and un-averaging it is real work. So v2 plumbs the hand split — as a conditioner on the rate, not as a fix to the fire rate.

8.2 What was built

The split enters at the per-PA hit rate, not at the output probability. Multiplying P(hits ≥ 1) by a platoon factor would scale a number that has already been through the opportunity term — a different and wrong claim. So baseballChain.chainFn now writes the chain out explicitly:

paOutcome  →  p_hit_per_pa (season)
           →  × platoonRead multiplier   ← THE MATCHUP CONDITIONER
           →  Binomial(PA, p_hit) over paDistribution
           →  P(hits ≥ k)

A test asserts that with no split supplied this is arithmetically identical to projectSkill across three archetypes × three PA values, so the restructuring cannot quietly become a second model.

As-of-correct (A4). The hitter's own split comes from hitsFactorContext (whose reads are lte('as_of_date', asOf)), the opposing starter's hand through the A6 matchupKeys resolve. The probe bounds every read at the row's own game_date. Nothing later than the grade can enter.

Refuse, never substitute. platoonSeverity already refuses below 60 PA on the smaller side and declares switch hitters unreadable. On a refusal the season rate stands exactly untouched (asserted to 6 dp) and the reason is recorded.

Refusals are counted by reason (context.onRefusal), and the count is taken off the stored blocks, not the legs — a prop with both an over and an under row produces two legs sharing one block, and the first draft reported 48.1% and 55.1% for the same fact over two different denominators.

8.3 Measured — real board, 2026-08-07 → 08-12, 9,792 graded hits rows

CHAIN FIRE — 9,752/9,792 atoms read (40 refused) across 83/83 games
  refusal reasons: { no_batter_profile: 40 }

HAND SPLIT — fired on 288/553 unique props (52.1%)
  season-rate reasons:
    insufficient_split_sample          130   the honest 60-PA refusal
    no_pitcher_hand                     71   the only FIXABLE gap
    switch_hitter_side_value_unknown    60   genuinely unreadable
    no_splits / missing_split            4

CHAIN vs COUNTER — n=9,752
                       min     p25   median    p75     max    mean
  chain_p            0.016   0.359   0.498   0.641   0.984   0.500
  counter_p          0.050   0.397   0.500   0.603   0.950   0.500
  |divergence|       0.000   0.044   0.095   0.164   0.509   0.113

  |div| < 0.02   1,142 (11.7%)      0.10–0.20   3,129 (32.1%)
  0.02–0.05      1,608 (16.5%)      >= 0.20     1,538 (15.8%)
  0.05–0.10      2,335 (23.9%)

OVER SIDE ONLY (n=4,879): mean +0.045, chain HIGHER on 66.3%

MATCHUP READ vs SEASON READ
  split FIRED   (n=5,372 rows)  median |div| 0.091 · >= 0.10 on 46.0% · within 0.02 on 13.0%
  split REFUSED (n=4,380 rows)  median |div| 0.100 · >= 0.10 on 50.1% · within 0.02 on 10.2%

8.4 Reading it honestly

  • The hand split fires on 52.1% of props, up from 0. Of the 47.9% that do not, 190 of 265 (72%) are principled refusals — a thin split or a switch hitter. Only no_pitcher_hand (71) is a plumbing gap, and it is the A5 shape exactly: the hitter's split is sitting right there and the pitcher hand is missing because that player had no lineup row.
  • Feeding the split widened the divergence: median |div| 0.082 → 0.095, and the over-side lean +2.9pp → +4.5pp. The chain moved further from the counter, which is what a matchup conditioner should do and is not evidence it moved in the right direction.
  • The rows where the split FIRED disagree with the counter slightly LESS (median 0.091 vs 0.100) than the rows where it refused. This comparison is confounded and must not be read as an effect: a hitter with 60+ PA on both sides is an established regular, and the counter has more game log on him too. It is two different populations, not two treatments.
  • Divergence is still not merit. Nothing here says the chain is closer to what happened. That needs settled outcomes against the stored triple, and it goes through the gate as a new hypothesis against the cumulative denominator.
  • The probe remains contaminated for the batter profile (statcast_aggregates keeps one as-of date) and so reports a divergence, never a resolution. The forward cron accrual does not have this problem.

8.5 Still not claimed

Unchanged from §5: nothing promoted, no head-to-head run, CALIBRATION_DEPLOYED still [], A8 untouched, hitsFactors untouched, the counter still serves, WNBA untouched.

8.6 Second incidental defect found and fixed

hitsFactorContext.build and matchupKeys.build were both gated on require('../utils/supabase').getSupabaseServiceClient() called inline, so the entire factor and hand-split path was unreachable from any test. "It is wired" could only ever have rested on reading the code — which is precisely how A5 shipped three factors that never fired. Both now read deps.supabase || the real client (additive; undefined gives identical behaviour).


§9 — chain v3: the opportunity term, and the artifact/signal diagnostic (2026-08-12)

9.1 The premise, again checked first — one half right, one half wrong

RIGHT, and a real defect: the v2 shadow passed no lineupSlotFor, so every hitter fell to skillProjection.DEFAULT_PA = 4.1. "A regular" was asserted about the leadoff man and the nine hole alike. The chain is rate × opportunity and the opportunity half was a constant across the entire lineup — a free, known, pre-game fact thrown away. Fixed.

WRONG about the conversion, and about the direction:

  • "mis-handles the multiple-chances structure" — it does not. paDistribution is a mean-preserving two-point mixture, and atLeast(Binomial(n, p), 1) is 1 − (1−p)^n averaged over n. The order's proposed formula is what the code already computes. The defect was the INPUT E[PA], not the conversion.
  • "5.7pts BELOW the counter on 92% of rows" — measured the other way. On the over side the chain runs +4.5pp ABOVE the counter and is below on 33.6%. It was never 92%-below, at any point, in any measurement here.

9.2 Phase 1 — the E[PA] fix, A/B on identical rows

matchup_keys now carries batting_order (one extra column on a read it already performs — no new query), and the shadow feeds it as the opportunity term. Absent ⇒ DEFAULT_PA and the block records default_regular, so a league-shaped opportunity term is never mistaken for a posted one.

OPPORTUNITY — posted lineup slot on 491/553 props (88.8%)
              fell to a default regular: 62

                                 n      mean      median    chain BELOW counter
  REAL E[PA] (posted slot)     4,879   +0.0450   +0.0499          33.6%
  CONSTANT 4.1 (the v2 shadow) 4,879   +0.0322   +0.0363          39.6%

The uniform bias did not collapse, because it was never there to collapse. The fix moved the chain further above the counter, not toward it — which is correct behaviour, not a regression: real slots raise E[PA] for the top of the order and lower it for the bottom, and top-of-order hitters are over-represented in the prop board.

9.3 Phase 2 — THE DIAGNOSTIC: artifact or signal?

Over side, n=3,705 rows carrying an opposing-pitcher profile, quintiles of opposing-pitcher K% (low = soft matchup):

bucket opp K% n mean divergence chain below counter
Q1 10.8–18.4 741 +0.0749 27.8%
Q2 18.4–20.3 741 +0.0560 30.8%
Q3 20.3–23.1 741 +0.0287 34.1%
Q4 23.1–26.6 741 +0.0396 35.0%
Q5 26.6–40.6 741 +0.0256 38.6%

Q1 − Q5 spread +0.0494 · Pearson r(oppK, divergence) = −0.119

The answer is BOTH, and the order's binary framing does not fit

The order asked for uniform (broken) or difficulty-correlated (signal). The data is a mixture of the two, and both components should be named:

  • A difficulty-correlated component, ~+0.049 across the range. The chain reads soft matchups higher and hard matchups lower than the counter does. The below % is monotone across all five quintiles (27.8 → 30.8 → 34.1 → 35.0 → 38.6), which is a cleaner signature than the means (Q4 breaks order).
  • A uniform positive offset of ~+0.026. Even in the HARDEST quintile the chain sits +2.6pp above the counter. That floor does not move with the matchup and is therefore not conditioning — it is exactly the shape a residual mechanical bias makes. It is smaller than the conditioning component but it has not been explained, and calling the whole result "signal" would bury it.

The caveat that must travel with the correlation

This is close to mechanically guaranteed and is NOT evidence of correctness. The chain reads the opposing pitcher's K rate directly (log5 odds-ratio in paOutcome) and the counter reads nothing about the pitcher at all. So a monotone relationship between opposing-pitcher K% and chain-minus-counter is approximately a proof that the wiring works — that the pitcher input reaches the number. It says nothing about whether the adjustment is the right size, the right direction on any individual row, or better than ignoring the pitcher.

Does the chain see the game, or just miscompute it? It demonstrably CONDITIONS on the game — the pitcher input reaches the forecast and moves it in the theorised direction. Whether that conditioning is right is unanswerable from a divergence and needs settled outcomes against the stored triple. There is also a residual ~2.6pp offset that conditioning does not explain and that should be chased before anyone reads the correlation as a win.

9.4 Open item created by this measurement

The ~+0.026 floor. Candidates not yet tested: the chain is unclamped where the counter clamps to [0.10, 0.95] (chain min 0.016 vs counter min 0.050); the LEAGUE.babip = 0.291 anchor in hitOnContact; the ±35% BABIP bound. Naming it as unexplained is the honest state — it is not yet an artifact and not yet signal.

9.5 Unchanged

Nothing promoted, no head-to-head, CALIBRATION_DEPLOYED still [], A8 and hitsFactors untouched, the counter still serves, WNBA untouched. Migration 038 is still an unapplied blocking precondition of deploy.