checkpoint: chain shadow, WNBA possession feed, baseball chain

Backup commit of uncommitted working-tree state found during Legion
recon (Tony resurrection, STEP 0). This work existed only on the
laptop disk.

- chain shadow accrual + probe script (038_chain_shadow.sql)
- WNBA possession feed: ESPN adapter, usage service, verify script
  (039_wnba_player_game.sql)
- baseball chain
- retention/snapshot service updates, tableKeys, matchupKeys
- specs: chain-v1, wnba-possession-feed, wnba-source-survey
- unit tests for the above

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
This commit is contained in:
Kev
2026-08-14 16:53:37 -04:00
parent 0657b71d18
commit 6c34af3414
22 changed files with 4958 additions and 32 deletions
+183
View File
@@ -2163,6 +2163,189 @@ phased plan in the Session-57 conversation / BUILD-STATE Next section).
- Safety net was exact: the newest pg_dump held **367,595 lock_lines rows —
matching the live count row-for-row**, pg_restore-verified before the drop.
## chain v1 — the slot that was never there (2026-08-12 — non-obvious)
- **`chainFn` was in `chain.js`'s header and NOT IN THE CODE** — no parameter, no
call site, no export, for its whole life. That single absence is why the
"portable core" was not portable: with no stage for atom→probability, every
sport's real work had to live elsewhere, and for MLB it lived in `scripts/`
reachable from no pipeline. Added; **defaults to identity-on-`p`** so every
prior caller is byte-identical. A chainFn returning null or throwing makes the
atom UNREADABLE ⇒ dropped and COUNTED (`chain_fn_refused`), never `p=0` — one
zero leg would zero a whole ticket.
- **`chainAcross` had a SECOND, undocumented reason for zero callers:** it
requires `calibrated: true`, and the only writer of that flag is the loop over
`CALIBRATION_DEPLOYED`, which is `Object.freeze([])`. No grade in the system
carries it. It was uncalled AND would have refused any caller it had.
- **Correlation is now signed `[-1,1]` and interpolates toward the FRÉCHET bound
its sign selects** (+1 → `min(p_i)`, −1 → `max(0, Σp−(n−1))`). The old `[0,1]`
clamp always shifted the joint toward the weakest leg — baseball's shape — and
made basketball's negative usage-competition case *inexpressible*, not merely
mismodelled. Positive is arithmetically unchanged (the weakest leg IS the
upper bound), so no previously-correct number moved.
- **`redistribute` must reach BOTH readings or `selfCheck` lies about itself.**
On `chainUp` only, the across-read was built from pre-redistribution atoms and
the check flagged an INTERNAL_INCONSISTENCY the model had just manufactured.
Use `chain.prepareAtoms(atoms, opts)` once and feed both readings its legs.
- **`baseballChain.chainFn` is a ROUTER, not new modelling** — `paOutcome` →
Binomial(PA, p_hit) over `paDistribution`. **Opportunity is a LOOKUP**
(`PA_BY_SLOT` 4.65→3.85, the documented ~0.1-PA-per-slot decline), NOT a fit:
baseball's opportunity is fixed, which is precisely why it is the clean first
fill. Basketball's is a contested model and is a later order. Hits ONLY — TB's
head-to-head is already inconclusive-under-contamination.
- **`selfCheck` IS VACUOUS TODAY and the data says so** (`vacuous: true`). It
earns its keep only against an INDEPENDENT team read; none exists (the
game-script projection was deliberately never built), so the up-read is built
from the same atoms and agreement is arithmetic.
- **The shadow stores the TRIPLE `(chain_p, counter_p, outcome)`**, side-aligned,
on `model_snapshots.chain_shadow` (migration 038). `chain_p` for an UNDER row
is `1 − p_over` — storing the raw over-probability there inverts every later
comparison silently. Every block carries `servable:false` IN THE PAYLOAD; a
caveat that lives only in a comment is not attached to the data.
- **MEASURED 2026-08-12 on 9,376 graded rows: fires on 99.6%, median
|divergence| 0.082, 40.7% differ by ≥0.10.** The two-sided signed distribution
is symmetric BY ARITHMETIC (both sides of every prop are sampled, contributing
+d and −d) — reading `mean −0.000` as "unbiased" reads the sampling scheme.
The over-side slice is the one that can lean: **+2.9pp, higher on 60.7%**.
**Divergence is not merit** — which forecast is closer is the settle pass's
question, which is why the triple is stored.
- **`runSnapshot`'s `deps` WAS AN ALLOWLIST** of 17 keys while 14 call sites read
`deps.challenger` / `loadStatcast` / `environmentContext` / `lineupContext` /
`hitsFactorContext` / `matchupKeys` / `gameBinder` / … — all permanently
`undefined`, all silently falling through to the real module. Every "injectable"
comment on those was fiction. Fixed by spreading `...opts` FIRST (explicit keys
are declared after and already read `opts.X`, so nothing moved). Found only
because the shadow's test could not inject a statcast map — it would have
passed while measuring nothing.
## chain v2 — the hand split, and a premise that had to be checked first (non-obvious)
- **`paOutcome` AND `hitOnContact` READ NO HANDEDNESS.** `fromStatcastRow` copies
`bats`/`throws` onto the profile at `:169-170` and **nothing downstream reads
them**. So a hand split cannot change whether the PA tree runs — it is a
conditioner on the RATE, never a precondition. Verify with a body-scan
(`grep` the function bodies), not by seeing the fields on the object.
- **THERE IS NO FALLBACK PATH IN `baseballChain`.** A refusal returns null and
`chain.applyChainFn` DROPS the atom. Nothing ever silently becomes a season
rate or the counter — which matters because a silent fallback would make the
shadow agree with the counter for a reason that LOOKS like agreement and is
not. A test locks it.
- **MEASURED fire rate is 99.6% (9,752/9,792), sole refusal `no_batter_profile`.**
If an order quotes a low fire rate, re-measure before building — the numbers in
the v2 order (596/725, 129/725, 425/425) matched nothing in the code or the
board; "425" is from `7c8ef8b`, the A1–A7 deploy verification.
- **THE SPLIT ENTERS AT THE PER-PA RATE, not at the output probability.**
`p_hit_per_pa × platoonRead.multiplier` → THEN Binomial over PA. Multiplying
`P(hits ≥ 1)` would scale a number already through the opportunity term — a
different and wrong claim. The chain is written out in `baseballChain` for
exactly this reason; a test asserts the no-split path is **arithmetically
identical to `projectSkill`** so it cannot drift into a second model.
- **HAND SPLIT FIRES ON 52.1% OF PROPS (288/553), and 72% of the misses are
PRINCIPLED** — `insufficient_split_sample` 130 (the 60-PA floor),
`switch_hitter_side_value_unknown` 60. Only `no_pitcher_hand` (71) is a
plumbing gap. Report the reasons, never a bare rate: "48% didn't fire" and
"48% couldn't honestly be read" are different claims.
- **COUNT THE RATE OFF THE STORED BLOCKS, NOT THE LEGS.** A prop's over and under
legs share ONE `perProp` block, so a per-leg count and a per-block slice give
two rates over two denominators that look comparable — the first draft printed
48.1% and 55.1% for the same fact.
- **`chainAcross`, `chainUp` and `prepareAtoms` EACH apply the chainFn.** Running
raw atoms through all three evaluates every hitter 3x and triple-counts every
refusal. `chainShadow` prepares ONCE and passes `prep.legs` (identity-on-p) to
the two readings.
- **Feeding the split WIDENED divergence** (median |div| 0.082 → 0.095, over-side
lean +2.9 → +4.5pp) — which is what a conditioner should do and is **not**
evidence it moved the right way. And fired rows disagree slightly LESS than
refused rows (0.091 vs 0.100): **confounded, not an effect** — 60+ PA on both
sides means an established regular, whom the counter also has more log on.
- **`hitsFactorContext.build` / `matchupKeys.build` were gated on an INLINE
`getSupabaseServiceClient()`**, making the whole factor + hand-split path
unreachable from any test — "it is wired" could only rest on reading the code,
which is exactly how A5 shipped three factors that never fired. Both now read
`deps.supabase ||` the real client.
- **The hand-split gate is `factorContext` ALONE, not `factorContext && matchupKeys`.**
Without the keys the hitter's own split is still readable and only the pitcher
hand is missing, so the reason must record `no_pitcher_hand` — naming the input
that is actually absent. Gating on both would record `no_splits` and point at
the half that was there all along.
## chain v3 — the opportunity term, and reading a divergence honestly (non-obvious)
- **`paDistribution` IS NOT BROKEN and never was.** It is a mean-preserving
two-point mixture (mean 4.65 → `out[4]=0.35, out[5]=0.65`), and
`atLeast(Binomial(n,p), 1)` **is exactly** `1−(1−p)^n` averaged over n. If an
order proposes "replace the conversion with 1−(1−p)^E[PA]", that is what the
code already computes — check the INPUT `E[PA]` instead.
- **THE REAL v2 DEFECT: the shadow passed no `lineupSlotFor`**, so every hitter
ran on `DEFAULT_PA = 4.1`. `rate × opportunity` with opportunity CONSTANT
across the lineup. `matchupKeys` now carries `batting_order` — one extra column
on the `lineup_context` read it already performs, so zero new I/O. Posted slot
reaches **88.8%** of props. Absent ⇒ `default_regular`, counted separately;
never silently 4.1-as-if-read.
- **REPORT `rate` AND `opportunity` COVERAGE SEPARATELY.** `readable` (did it
produce a number), `platoon_applied` (did it read the matchup), and
`opportunity_posted` (did it read the lineup) are THREE different questions.
v2 shipped with the opportunity half constant and no summary field could show it.
- **The chain runs ABOVE the counter, not below** — over side mean **+4.5pp**,
below on **33.6%**. Fixing E[PA] moved it FURTHER above (+0.032 → +0.045), and
that is correct, not a regression: real slots raise E[PA] at the top of the
order and top-of-order hitters dominate the prop board.
- **A DIVERGENCE DECOMPOSES; DON'T ACCEPT AN EITHER/OR FRAMING.** Bucketed by
opposing-pitcher K%, the result was BOTH: a difficulty-correlated component
(Q1 +0.0749 → Q5 +0.0256, spread +0.049, r = −0.119) AND a **uniform +0.026
offset surviving into the hardest quintile**. A floor that does not move with
the matchup is not conditioning. Reporting only the correlation would have
buried it.
- **The `below %` was monotone across all five quintiles (27.8→38.6) while the
MEANS were not (Q4 breaks order).** When one statistic is monotone and another
is not, lead with the monotone one and say the other isn't.
- **A DIFFICULTY CORRELATION HERE IS NEAR-MECHANICAL AND IS NOT EVIDENCE OF
CORRECTNESS.** The chain reads opposing-pitcher K% directly (log5 in
`paOutcome`); the counter reads nothing about the pitcher. So the correlation
proves the WIRING reaches the forecast — not that the adjustment's size or
per-row direction is right. Only settled outcomes can say that. Say this in the
same breath as the correlation, every time.
- **A/B an input fix on IDENTICAL ROWS** (`runShadow` twice, one arm with
`lineupSlotFor: () => null`). Comparing two slates would measure the slate.
## WNBA possession/usage feed (2026-08-13 — non-obvious)
- **ESPN's WNBA box score serves the COMPONENTS, never the rates.** No usage, no
possessions, no pace fields exist. `usage_rate`/`team_possessions`/`team_pace`/
`ts_pct`/`efg_pct` are DERIVED in `espnWnbaAdapter` from the standard
identities. **Say "derived", not "proxied"** — a proxy stands in for something
unseen; these are the quantity itself, recomputed from counted events. The ONLY
estimated term in the whole feed is the **0.44 free-throw-trip coefficient**.
- **PER-GAME GRAIN MAKES POINT-IN-TIME NATIVE — no history twin.**
`statcast_aggregates` needed `statcast_history` because it upserts a SEASON
AGGREGATE in place. `wnba_player_game` stores per-game rows, and a completed
box score never changes, so as-of is `WHERE game_date < asOf` — a filter, not a
snapshot. **Strictly `<`, never `<=`:** a game ON the as-of date may tip after
grade time (the `isPreGame` rule). If you add another feed, ask whether
per-event grain removes the retention problem before building a dated snapshot.
- **TWO NON-LEAGUE GAMES ARE IN THE ESPN WNBA SCOREBOARD** and will silently
pollute usage profiles: an exhibition vs a national team (2026-05-02, `NIGER`)
and the ALL-STAR game (2026-07-25, `SPO` vs `COOP`). Measured effect: Natasha
Howard 37→36 games, usage 21.897→22.057. The filter reads **ESPN's `/teams`**
rather than a hardcoded fifteen — the WNBA has expanded twice in three years.
**An EMPTY `/teams` response filters NOTHING** (unknown membership ≠ nobody is
in the league) — the `fielding_oaa` lesson: a failed feed degrading to an empty
index looks exactly like an honest absence.
- **Team minutes carry overtime for free.** Pace divides by `teamMinutes / 5`, so
a 225-team-minute OT game needs no branch. Regulation is 40 min in the WNBA
(not 48) — `REGULATION_MINUTES` is the constant, don't copy an NBA one.
- **`totalTurnovers`, not `turnovers`, in the team totals** — a shot-clock
violation belongs to the possession count even though no player committed it.
- **A DNP is OMITTED, never a zero line.** A zero-minute row asserts he was
available and produced nothing; and the usage denominator divides by minutes.
`ts_pct`/`efg_pct` null on zero attempts is a REFUSAL, not a coverage gap —
don't "fix" those numbers.
- **`profileAsOf` usage is MINUTES-WEIGHTED** (the lineup-K-rate lesson: an
unweighted aggregate counts a 6-minute cameo like a 34-minute start, and
unweighted HURT that model), and REFUSES below 3 games rather than returning a
league-average player — usage feeds the chain's opportunity term directly.
- **The local `.env` cannot reach the DB from WSL2:** `SUPABASE_DB_PASSWORD`
fails pooler auth (`aws-1-us-east-1...`) and `db.<ref>.supabase.co` resolves
IPv6-only → ENETUNREACH. Migrations 038 and 039 are written and UNAPPLIED. A
measurement run in memory verifies parse/derivation/as-of but NOT the DB
round-trip — state that distinction rather than implying a feed exists.
## Active Skills
- vyndr-voice (all user-facing output)
- prop-analysis (grading methodology)