checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion recon (Tony resurrection, STEP 0). This work existed only on the laptop disk. - chain shadow accrual + probe script (038_chain_shadow.sql) - WNBA possession feed: ESPN adapter, usage service, verify script (039_wnba_player_game.sql) - baseball chain - retention/snapshot service updates, tableKeys, matchupKeys - specs: chain-v1, wnba-possession-feed, wnba-source-survey - unit tests for the above Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
This commit is contained in:
+163
@@ -3,6 +3,169 @@
|
||||
## Last Updated
|
||||
2026-08-12
|
||||
|
||||
## WNBA v1 (2026-08-13) — the possession/usage feed established ✅
|
||||
Spec: `specs/wnba-possession-feed.md`. 4,905 tests / 377 suites green, web build
|
||||
exit 0. **DATA INFRA ONLY — no chainFn, no archetype wiring, no shadow, nothing
|
||||
served. MLB / chain shadow / A8 / CALIBRATION_DEPLOYED untouched (test-asserted).**
|
||||
- **SOURCE ESTABLISHED BEFORE PLUMBING.** Python `nba_api` rejected as a
|
||||
dependency (offline in prod). **ESPN site API** chosen — free, no auth, already
|
||||
the host for WNBA schedules/box/live. Verified live 2026-08-13.
|
||||
- **chainFn inputs: ALL PRESENT.** minutes SERVED; usage / possessions / pace /
|
||||
TS% / eFG% **DERIVED EXACTLY** from box components (`POSS = FGA−OREB+TOV+
|
||||
0.44·FTA`; the 0.44 free-throw-trip coefficient is the only estimated term).
|
||||
**"Derived" ≠ "proxied"** — these are the quantities themselves, recomputed.
|
||||
Game-state for the `redistribute` hook (starter, final_margin) SERVED.
|
||||
MISSING and named: shot location, on/off lineups, opponent DRtg (the last is
|
||||
derivable later from the same rows).
|
||||
- **POINT-IN-TIME IS NATIVE — no `statcast_history` twin needed.** Stores
|
||||
**PER-GAME** rows, not a season aggregate upserted in place, so an as-of
|
||||
profile is `WHERE game_date < asOf` over immutable completed box scores.
|
||||
**Strictly `<`** — a game ON the date may tip after grade time (`isPreGame`).
|
||||
This is the statcast lesson applied at design time instead of retrofitted.
|
||||
- **`wnba_player_game`** (migration 039), key `(game_id, source_id)`, registered
|
||||
in `tableKeys`. Every derived rate stored **beside its components** so it is
|
||||
re-derivable (`factorFreeze` rule for a feed). Scoped to the chainFn — not a
|
||||
stats dump (the `closing_captures` lesson).
|
||||
- **MEASURED, full live season:** 5,085 rows · 256 games · 241 players · 15 teams ·
|
||||
2026-05-01→08-12 · **0 fetch errors**. minutes/possessions/pace/starter/margin
|
||||
5085/5085; usage 5072/5085. ts_pct/efg_pct gaps are **correct refusals** (no
|
||||
attempts ⇒ no percentage), not missing data.
|
||||
- **TWO CONTAMINATING GAMES FOUND AND EXCLUDED:** an exhibition vs **Nigeria**
|
||||
(05-02) and the **All-Star game** (07-25). Filter reads **ESPN's own `/teams`**,
|
||||
not a hardcoded fifteen — this league has expanded twice in three years. An
|
||||
EMPTY teams response filters NOTHING (a failed feed must not masquerade as an
|
||||
honest absence — the `fielding_oaa` lesson). Concrete effect: Natasha Howard
|
||||
37→36 games, usage 21.897→22.057.
|
||||
- **AS-OF VERIFIED ON REAL DATA through the real function:** asOf 06-24 → 18
|
||||
games (usage 24.826), asOf 08-12 → 35 (22.057), asOf 12-31 → 36 (22.027) — the
|
||||
profile MOVES with the date. Same-day game EXCLUDED; thin history REFUSED
|
||||
(null, never a league-average stand-in). Usage is **minutes-weighted** (the
|
||||
lineup-K-rate lesson).
|
||||
- **⛔ NOT INGESTED TO THE DB.** `SUPABASE_DB_PASSWORD` fails pooler auth and the
|
||||
direct host is IPv6-only (unreachable from WSL2). Migration 039 written and
|
||||
UNAPPLIED; `ingestRange` written and has never written a row. All coverage
|
||||
above was measured by running the real parsers + real as-of read **in memory**
|
||||
over the live season — parse/derivation/as-of verified, **DB round-trip not**.
|
||||
Migration 038 also still outstanding.
|
||||
|
||||
## chain v3 (2026-08-12) — E[PA] fixed; artifact vs signal separated ✅
|
||||
Spec: `specs/chain-v1.md` §9. 4,882 tests / 376 suites green, web build exit 0.
|
||||
**Still SHADOW — served byte-identical, `CALIBRATION_DEPLOYED` `[]`, counter
|
||||
serves, A8 / hitsFactors / WNBA untouched.**
|
||||
- **THE ORDER WAS HALF RIGHT, and the right half was a real defect.** The v2
|
||||
shadow passed **no `lineupSlotFor`**, so every hitter ran on
|
||||
`DEFAULT_PA = 4.1` — "a regular" asserted about the leadoff man and the nine
|
||||
hole alike. `rate × opportunity` with opportunity held constant across the
|
||||
lineup. FIXED: `matchupKeys` now carries `batting_order` (one extra column on
|
||||
a read it already does — no new query); **posted slot on 491/553 props (88.8%)**.
|
||||
- **WRONG about the conversion:** `paDistribution` is a mean-preserving two-point
|
||||
mixture and `atLeast(Binomial(n,p),1)` **IS** `1−(1−p)^n` averaged over n — the
|
||||
order's proposed formula is what the code already computed. The defect was the
|
||||
INPUT E[PA], never the structure.
|
||||
- **WRONG about the direction, and this matters:** claimed "5.7pts BELOW on 92%".
|
||||
MEASURED **+4.5pp ABOVE**, below on **33.6%**. There was never a uniform low
|
||||
bias to collapse. A/B on identical rows: real E[PA] mean **+0.0450** / below
|
||||
33.6%; constant 4.1 **+0.0322** / below 39.6%. The fix moved the chain FURTHER
|
||||
above — correct, not a regression (real slots raise E[PA] for the top of the
|
||||
order, and top-of-order hitters dominate the prop board).
|
||||
- **THE DIAGNOSTIC — the answer is BOTH, not the order's either/or.** Over side,
|
||||
n=3,705, quintiles of opposing-pitcher K%:
|
||||
Q1 (soft) **+0.0749** → Q5 (hard) **+0.0256**; spread **+0.0494**,
|
||||
Pearson r = **−0.119**; `below %` **monotone across all five** (27.8 → 30.8 →
|
||||
34.1 → 35.0 → 38.6).
|
||||
- A **difficulty-correlated component** (~+0.049) — the chain conditions.
|
||||
- A **uniform positive offset** (~+0.026) surviving into the HARDEST quintile —
|
||||
does not move with the matchup, so it is not conditioning. Unexplained.
|
||||
Calling the whole result "signal" would bury it.
|
||||
- **THE CAVEAT THAT TRAVELS WITH IT:** the chain reads opposing-pitcher K%
|
||||
directly (log5 in `paOutcome`) and the counter reads nothing about the pitcher,
|
||||
so the correlation is close to mechanical — it proves the WIRING reaches the
|
||||
number, **not** that the adjustment is right. Correctness needs settled
|
||||
outcomes against the stored triple.
|
||||
- **Open item:** the ~+0.026 floor. Untested candidates — the chain is unclamped
|
||||
where the counter clamps `[0.10,0.95]` (chain min 0.016 vs counter 0.050);
|
||||
`LEAGUE.babip = 0.291`; the ±35% BABIP bound.
|
||||
|
||||
## chain v2 (2026-08-12) — the hand split plumbed; the premise corrected ✅
|
||||
Spec: `specs/chain-v1.md` §8. 4,877 tests / 376 suites green, web build exit 0.
|
||||
**Still SHADOW — served payload byte-identical, `CALIBRATION_DEPLOYED` still `[]`,
|
||||
counter still serves, A8 / hitsFactors / WNBA untouched.**
|
||||
- **THE ORDER'S PREMISE WAS WRONG, and checking it first was the work.** Claimed
|
||||
"paOutcome returned null on 596/725, 82% fell back to the seasonal rate".
|
||||
MEASURED: fire rate **9,752/9,792 (99.6%)**, sole refusal reason
|
||||
`no_batter_profile: 40`. `paOutcome`/`hitOnContact` **read no handedness at
|
||||
all** — a hand split cannot change whether they run. And **there is no fallback
|
||||
path**: a refusal DROPS the atom, it never becomes a season rate. (The "425/425"
|
||||
figure is from `7c8ef8b`, the A1–A7 deploy verification, not the chain.)
|
||||
- **The premise was right about what MATTERS, though:** the chain read a hitter's
|
||||
SEASON rates, which already average his platoon split over whatever hands he
|
||||
faced. That is a season read wearing a matchup read's clothes.
|
||||
- **The split now enters at the RATE** (`p_hit_per_pa × platoonRead`), then the
|
||||
binomial over PA — not at the output probability, which would scale a number
|
||||
already through the opportunity term. A test asserts the no-split path is
|
||||
**arithmetically identical to `projectSkill`** so the restructuring cannot
|
||||
become a second model.
|
||||
- **MEASURED: hand split fires on 288/553 unique props (52.1%)**, from 0. Of the
|
||||
265 that don't, **190 (72%) are principled refusals** — `insufficient_split_
|
||||
sample` 130 (the 60-PA floor), `switch_hitter` 60. Only `no_pitcher_hand` 71 is
|
||||
a plumbing gap, and it is the A5 shape: split present, pitcher hand absent.
|
||||
- **Divergence WIDENED**: median |div| 0.082 → **0.095**, over-side lean +2.9pp →
|
||||
**+4.5pp**. Fired rows disagree slightly LESS than refused rows (0.091 vs
|
||||
0.100) — **confounded, not an effect**: 60+PA-both-sides means an established
|
||||
regular, and the counter has more log on him too.
|
||||
- **Counting bug caught in the first draft:** the hand-split rate was counted per
|
||||
LEG while divergence was sliced per BLOCK (a prop's over+under share one
|
||||
block), reporting 48.1% and 55.1% for the same fact. One denominator now, taken
|
||||
from what is persisted.
|
||||
- **Second unreachable seam found:** `hitsFactorContext.build` /
|
||||
`matchupKeys.build` were gated on an inline `getSupabaseServiceClient()`, so the
|
||||
whole factor + hand-split path was untestable — "it is wired" could only rest on
|
||||
reading the code, which is how A5 shipped three factors that never fired. Both
|
||||
now take `deps.supabase ||`.
|
||||
|
||||
## chain v1 (2026-08-12) — the engine made whole, shadowed on MLB ✅
|
||||
Spec: `specs/chain-v1.md`. 4,859 tests / 376 suites green, web build exit 0.
|
||||
**SERVED BY NOTHING — served payload byte-identical, proven with the shadow on
|
||||
and off. `CALIBRATION_DEPLOYED` still `[]`. A8, hitsFactors and WNBA untouched.**
|
||||
- **THE CORE WAS A SHELL OF ITS OWN HEADER.** `chainFn` — the atom→probability
|
||||
slot — was described in the header and **absent from the code** (no parameter,
|
||||
no call site, no export). That is why the "portable core" was not portable:
|
||||
with no slot for that stage, MLB's real work lived in `scripts/`, called by no
|
||||
pipeline. Added, defaulting to identity-on-`p` so every prior caller is
|
||||
unchanged. An unreadable atom is DROPPED and COUNTED, never `p=0`.
|
||||
- **Correlation un-clamped to [-1,1], direction follows the sign.** `|corr|`
|
||||
interpolates from independence toward the **Fréchet bound its sign selects**
|
||||
(+1 → `min(p)`, −1 → `max(0, Σp−(n−1))`). Positive is arithmetically unchanged.
|
||||
Basketball's negative usage-competition case was previously *inexpressible*.
|
||||
- **`redistribute` now reaches chainAcross too.** Reaching only the team read
|
||||
left `selfCheck` comparing post- to pre-redistribution atoms and flagging an
|
||||
inconsistency the model had just manufactured. `prepareAtoms` is exported so
|
||||
both readings are built from identical legs.
|
||||
- **Baseball's chainFn** (`baseballChain.js`) routes `paOutcome` →
|
||||
Binomial(PA, p_hit) over `paDistribution`. Opportunity is a **LOOKUP**
|
||||
(`PA_BY_SLOT`, 4.65→3.85), not a fit — baseball's opportunity is fixed, which
|
||||
is why it is the clean first fill. Hits only; TB's verdict is already muddy.
|
||||
- **First real measurement** (`scripts/chain-shadow-probe.js`, 9,376 graded rows,
|
||||
08-07→08-12): chain fires on **99.6%**; median |divergence| **0.082**; 40.7%
|
||||
differ by ≥0.10, 10.3% by ≥0.20, 12.9% agree inside 0.02. The two-sided signed
|
||||
symmetry is **arithmetic** (both sides sampled); the over-side slice leans
|
||||
**+2.9pp, higher on 60.7%**. Divergence is not merit.
|
||||
- **The triple `(chain_p, counter_p, outcome)`** lands on `model_snapshots.
|
||||
chain_shadow` (migration 038), side-aligned. `selfCheck` is labelled
|
||||
`vacuous: true` — no independent team read exists, so agreement is arithmetic.
|
||||
- **Found + fixed:** `runSnapshot`'s `deps` was an allowlist of 17 keys while 14
|
||||
call sites read unlisted `deps.X` — every one permanently `undefined`, every
|
||||
documented injection seam a comment only. `...opts` spread first.
|
||||
|
||||
### ⛔ BLOCKING PRECONDITION — migration 038 BEFORE this deploys
|
||||
`retentionService.rowsFromSides` now declares `chain_shadow` on **every** row
|
||||
(it must — PostgREST builds a bulk insert from the FIRST row's shape, so a key
|
||||
present on only some rows is dropped for the whole batch). If the column does
|
||||
not exist, **every retention insert fails**, and retention is best-effort, so it
|
||||
fails SILENTLY — the exact "retention writes nothing" shape the S64 ops alarm
|
||||
exists to catch. Apply `supabase/migrations/038_chain_shadow.sql` first. Not
|
||||
applied yet: this branch is uncommitted, pending Roundtable review.
|
||||
|
||||
## Reclaim 2 (2026-08-12) — lock_lines dropped, writer retired ✅
|
||||
4,803 tests / 374 suites, web build exit 0, app healthy. **Moat untouched
|
||||
(`ledger_entries`, `model_snapshots`); grade path untouched.**
|
||||
|
||||
@@ -2163,6 +2163,189 @@ phased plan in the Session-57 conversation / BUILD-STATE Next section).
|
||||
- Safety net was exact: the newest pg_dump held **367,595 lock_lines rows —
|
||||
matching the live count row-for-row**, pg_restore-verified before the drop.
|
||||
|
||||
## chain v1 — the slot that was never there (2026-08-12 — non-obvious)
|
||||
- **`chainFn` was in `chain.js`'s header and NOT IN THE CODE** — no parameter, no
|
||||
call site, no export, for its whole life. That single absence is why the
|
||||
"portable core" was not portable: with no stage for atom→probability, every
|
||||
sport's real work had to live elsewhere, and for MLB it lived in `scripts/`
|
||||
reachable from no pipeline. Added; **defaults to identity-on-`p`** so every
|
||||
prior caller is byte-identical. A chainFn returning null or throwing makes the
|
||||
atom UNREADABLE ⇒ dropped and COUNTED (`chain_fn_refused`), never `p=0` — one
|
||||
zero leg would zero a whole ticket.
|
||||
- **`chainAcross` had a SECOND, undocumented reason for zero callers:** it
|
||||
requires `calibrated: true`, and the only writer of that flag is the loop over
|
||||
`CALIBRATION_DEPLOYED`, which is `Object.freeze([])`. No grade in the system
|
||||
carries it. It was uncalled AND would have refused any caller it had.
|
||||
- **Correlation is now signed `[-1,1]` and interpolates toward the FRÉCHET bound
|
||||
its sign selects** (+1 → `min(p_i)`, −1 → `max(0, Σp−(n−1))`). The old `[0,1]`
|
||||
clamp always shifted the joint toward the weakest leg — baseball's shape — and
|
||||
made basketball's negative usage-competition case *inexpressible*, not merely
|
||||
mismodelled. Positive is arithmetically unchanged (the weakest leg IS the
|
||||
upper bound), so no previously-correct number moved.
|
||||
- **`redistribute` must reach BOTH readings or `selfCheck` lies about itself.**
|
||||
On `chainUp` only, the across-read was built from pre-redistribution atoms and
|
||||
the check flagged an INTERNAL_INCONSISTENCY the model had just manufactured.
|
||||
Use `chain.prepareAtoms(atoms, opts)` once and feed both readings its legs.
|
||||
- **`baseballChain.chainFn` is a ROUTER, not new modelling** — `paOutcome` →
|
||||
Binomial(PA, p_hit) over `paDistribution`. **Opportunity is a LOOKUP**
|
||||
(`PA_BY_SLOT` 4.65→3.85, the documented ~0.1-PA-per-slot decline), NOT a fit:
|
||||
baseball's opportunity is fixed, which is precisely why it is the clean first
|
||||
fill. Basketball's is a contested model and is a later order. Hits ONLY — TB's
|
||||
head-to-head is already inconclusive-under-contamination.
|
||||
- **`selfCheck` IS VACUOUS TODAY and the data says so** (`vacuous: true`). It
|
||||
earns its keep only against an INDEPENDENT team read; none exists (the
|
||||
game-script projection was deliberately never built), so the up-read is built
|
||||
from the same atoms and agreement is arithmetic.
|
||||
- **The shadow stores the TRIPLE `(chain_p, counter_p, outcome)`**, side-aligned,
|
||||
on `model_snapshots.chain_shadow` (migration 038). `chain_p` for an UNDER row
|
||||
is `1 − p_over` — storing the raw over-probability there inverts every later
|
||||
comparison silently. Every block carries `servable:false` IN THE PAYLOAD; a
|
||||
caveat that lives only in a comment is not attached to the data.
|
||||
- **MEASURED 2026-08-12 on 9,376 graded rows: fires on 99.6%, median
|
||||
|divergence| 0.082, 40.7% differ by ≥0.10.** The two-sided signed distribution
|
||||
is symmetric BY ARITHMETIC (both sides of every prop are sampled, contributing
|
||||
+d and −d) — reading `mean −0.000` as "unbiased" reads the sampling scheme.
|
||||
The over-side slice is the one that can lean: **+2.9pp, higher on 60.7%**.
|
||||
**Divergence is not merit** — which forecast is closer is the settle pass's
|
||||
question, which is why the triple is stored.
|
||||
- **`runSnapshot`'s `deps` WAS AN ALLOWLIST** of 17 keys while 14 call sites read
|
||||
`deps.challenger` / `loadStatcast` / `environmentContext` / `lineupContext` /
|
||||
`hitsFactorContext` / `matchupKeys` / `gameBinder` / … — all permanently
|
||||
`undefined`, all silently falling through to the real module. Every "injectable"
|
||||
comment on those was fiction. Fixed by spreading `...opts` FIRST (explicit keys
|
||||
are declared after and already read `opts.X`, so nothing moved). Found only
|
||||
because the shadow's test could not inject a statcast map — it would have
|
||||
passed while measuring nothing.
|
||||
|
||||
## chain v2 — the hand split, and a premise that had to be checked first (non-obvious)
|
||||
- **`paOutcome` AND `hitOnContact` READ NO HANDEDNESS.** `fromStatcastRow` copies
|
||||
`bats`/`throws` onto the profile at `:169-170` and **nothing downstream reads
|
||||
them**. So a hand split cannot change whether the PA tree runs — it is a
|
||||
conditioner on the RATE, never a precondition. Verify with a body-scan
|
||||
(`grep` the function bodies), not by seeing the fields on the object.
|
||||
- **THERE IS NO FALLBACK PATH IN `baseballChain`.** A refusal returns null and
|
||||
`chain.applyChainFn` DROPS the atom. Nothing ever silently becomes a season
|
||||
rate or the counter — which matters because a silent fallback would make the
|
||||
shadow agree with the counter for a reason that LOOKS like agreement and is
|
||||
not. A test locks it.
|
||||
- **MEASURED fire rate is 99.6% (9,752/9,792), sole refusal `no_batter_profile`.**
|
||||
If an order quotes a low fire rate, re-measure before building — the numbers in
|
||||
the v2 order (596/725, 129/725, 425/425) matched nothing in the code or the
|
||||
board; "425" is from `7c8ef8b`, the A1–A7 deploy verification.
|
||||
- **THE SPLIT ENTERS AT THE PER-PA RATE, not at the output probability.**
|
||||
`p_hit_per_pa × platoonRead.multiplier` → THEN Binomial over PA. Multiplying
|
||||
`P(hits ≥ 1)` would scale a number already through the opportunity term — a
|
||||
different and wrong claim. The chain is written out in `baseballChain` for
|
||||
exactly this reason; a test asserts the no-split path is **arithmetically
|
||||
identical to `projectSkill`** so it cannot drift into a second model.
|
||||
- **HAND SPLIT FIRES ON 52.1% OF PROPS (288/553), and 72% of the misses are
|
||||
PRINCIPLED** — `insufficient_split_sample` 130 (the 60-PA floor),
|
||||
`switch_hitter_side_value_unknown` 60. Only `no_pitcher_hand` (71) is a
|
||||
plumbing gap. Report the reasons, never a bare rate: "48% didn't fire" and
|
||||
"48% couldn't honestly be read" are different claims.
|
||||
- **COUNT THE RATE OFF THE STORED BLOCKS, NOT THE LEGS.** A prop's over and under
|
||||
legs share ONE `perProp` block, so a per-leg count and a per-block slice give
|
||||
two rates over two denominators that look comparable — the first draft printed
|
||||
48.1% and 55.1% for the same fact.
|
||||
- **`chainAcross`, `chainUp` and `prepareAtoms` EACH apply the chainFn.** Running
|
||||
raw atoms through all three evaluates every hitter 3x and triple-counts every
|
||||
refusal. `chainShadow` prepares ONCE and passes `prep.legs` (identity-on-p) to
|
||||
the two readings.
|
||||
- **Feeding the split WIDENED divergence** (median |div| 0.082 → 0.095, over-side
|
||||
lean +2.9 → +4.5pp) — which is what a conditioner should do and is **not**
|
||||
evidence it moved the right way. And fired rows disagree slightly LESS than
|
||||
refused rows (0.091 vs 0.100): **confounded, not an effect** — 60+ PA on both
|
||||
sides means an established regular, whom the counter also has more log on.
|
||||
- **`hitsFactorContext.build` / `matchupKeys.build` were gated on an INLINE
|
||||
`getSupabaseServiceClient()`**, making the whole factor + hand-split path
|
||||
unreachable from any test — "it is wired" could only rest on reading the code,
|
||||
which is exactly how A5 shipped three factors that never fired. Both now read
|
||||
`deps.supabase ||` the real client.
|
||||
- **The hand-split gate is `factorContext` ALONE, not `factorContext && matchupKeys`.**
|
||||
Without the keys the hitter's own split is still readable and only the pitcher
|
||||
hand is missing, so the reason must record `no_pitcher_hand` — naming the input
|
||||
that is actually absent. Gating on both would record `no_splits` and point at
|
||||
the half that was there all along.
|
||||
|
||||
## chain v3 — the opportunity term, and reading a divergence honestly (non-obvious)
|
||||
- **`paDistribution` IS NOT BROKEN and never was.** It is a mean-preserving
|
||||
two-point mixture (mean 4.65 → `out[4]=0.35, out[5]=0.65`), and
|
||||
`atLeast(Binomial(n,p), 1)` **is exactly** `1−(1−p)^n` averaged over n. If an
|
||||
order proposes "replace the conversion with 1−(1−p)^E[PA]", that is what the
|
||||
code already computes — check the INPUT `E[PA]` instead.
|
||||
- **THE REAL v2 DEFECT: the shadow passed no `lineupSlotFor`**, so every hitter
|
||||
ran on `DEFAULT_PA = 4.1`. `rate × opportunity` with opportunity CONSTANT
|
||||
across the lineup. `matchupKeys` now carries `batting_order` — one extra column
|
||||
on the `lineup_context` read it already performs, so zero new I/O. Posted slot
|
||||
reaches **88.8%** of props. Absent ⇒ `default_regular`, counted separately;
|
||||
never silently 4.1-as-if-read.
|
||||
- **REPORT `rate` AND `opportunity` COVERAGE SEPARATELY.** `readable` (did it
|
||||
produce a number), `platoon_applied` (did it read the matchup), and
|
||||
`opportunity_posted` (did it read the lineup) are THREE different questions.
|
||||
v2 shipped with the opportunity half constant and no summary field could show it.
|
||||
- **The chain runs ABOVE the counter, not below** — over side mean **+4.5pp**,
|
||||
below on **33.6%**. Fixing E[PA] moved it FURTHER above (+0.032 → +0.045), and
|
||||
that is correct, not a regression: real slots raise E[PA] at the top of the
|
||||
order and top-of-order hitters dominate the prop board.
|
||||
- **A DIVERGENCE DECOMPOSES; DON'T ACCEPT AN EITHER/OR FRAMING.** Bucketed by
|
||||
opposing-pitcher K%, the result was BOTH: a difficulty-correlated component
|
||||
(Q1 +0.0749 → Q5 +0.0256, spread +0.049, r = −0.119) AND a **uniform +0.026
|
||||
offset surviving into the hardest quintile**. A floor that does not move with
|
||||
the matchup is not conditioning. Reporting only the correlation would have
|
||||
buried it.
|
||||
- **The `below %` was monotone across all five quintiles (27.8→38.6) while the
|
||||
MEANS were not (Q4 breaks order).** When one statistic is monotone and another
|
||||
is not, lead with the monotone one and say the other isn't.
|
||||
- **A DIFFICULTY CORRELATION HERE IS NEAR-MECHANICAL AND IS NOT EVIDENCE OF
|
||||
CORRECTNESS.** The chain reads opposing-pitcher K% directly (log5 in
|
||||
`paOutcome`); the counter reads nothing about the pitcher. So the correlation
|
||||
proves the WIRING reaches the forecast — not that the adjustment's size or
|
||||
per-row direction is right. Only settled outcomes can say that. Say this in the
|
||||
same breath as the correlation, every time.
|
||||
- **A/B an input fix on IDENTICAL ROWS** (`runShadow` twice, one arm with
|
||||
`lineupSlotFor: () => null`). Comparing two slates would measure the slate.
|
||||
|
||||
## WNBA possession/usage feed (2026-08-13 — non-obvious)
|
||||
- **ESPN's WNBA box score serves the COMPONENTS, never the rates.** No usage, no
|
||||
possessions, no pace fields exist. `usage_rate`/`team_possessions`/`team_pace`/
|
||||
`ts_pct`/`efg_pct` are DERIVED in `espnWnbaAdapter` from the standard
|
||||
identities. **Say "derived", not "proxied"** — a proxy stands in for something
|
||||
unseen; these are the quantity itself, recomputed from counted events. The ONLY
|
||||
estimated term in the whole feed is the **0.44 free-throw-trip coefficient**.
|
||||
- **PER-GAME GRAIN MAKES POINT-IN-TIME NATIVE — no history twin.**
|
||||
`statcast_aggregates` needed `statcast_history` because it upserts a SEASON
|
||||
AGGREGATE in place. `wnba_player_game` stores per-game rows, and a completed
|
||||
box score never changes, so as-of is `WHERE game_date < asOf` — a filter, not a
|
||||
snapshot. **Strictly `<`, never `<=`:** a game ON the as-of date may tip after
|
||||
grade time (the `isPreGame` rule). If you add another feed, ask whether
|
||||
per-event grain removes the retention problem before building a dated snapshot.
|
||||
- **TWO NON-LEAGUE GAMES ARE IN THE ESPN WNBA SCOREBOARD** and will silently
|
||||
pollute usage profiles: an exhibition vs a national team (2026-05-02, `NIGER`)
|
||||
and the ALL-STAR game (2026-07-25, `SPO` vs `COOP`). Measured effect: Natasha
|
||||
Howard 37→36 games, usage 21.897→22.057. The filter reads **ESPN's `/teams`**
|
||||
rather than a hardcoded fifteen — the WNBA has expanded twice in three years.
|
||||
**An EMPTY `/teams` response filters NOTHING** (unknown membership ≠ nobody is
|
||||
in the league) — the `fielding_oaa` lesson: a failed feed degrading to an empty
|
||||
index looks exactly like an honest absence.
|
||||
- **Team minutes carry overtime for free.** Pace divides by `teamMinutes / 5`, so
|
||||
a 225-team-minute OT game needs no branch. Regulation is 40 min in the WNBA
|
||||
(not 48) — `REGULATION_MINUTES` is the constant, don't copy an NBA one.
|
||||
- **`totalTurnovers`, not `turnovers`, in the team totals** — a shot-clock
|
||||
violation belongs to the possession count even though no player committed it.
|
||||
- **A DNP is OMITTED, never a zero line.** A zero-minute row asserts he was
|
||||
available and produced nothing; and the usage denominator divides by minutes.
|
||||
`ts_pct`/`efg_pct` null on zero attempts is a REFUSAL, not a coverage gap —
|
||||
don't "fix" those numbers.
|
||||
- **`profileAsOf` usage is MINUTES-WEIGHTED** (the lineup-K-rate lesson: an
|
||||
unweighted aggregate counts a 6-minute cameo like a 34-minute start, and
|
||||
unweighted HURT that model), and REFUSES below 3 games rather than returning a
|
||||
league-average player — usage feeds the chain's opportunity term directly.
|
||||
- **The local `.env` cannot reach the DB from WSL2:** `SUPABASE_DB_PASSWORD`
|
||||
fails pooler auth (`aws-1-us-east-1...`) and `db.<ref>.supabase.co` resolves
|
||||
IPv6-only → ENETUNREACH. Migrations 038 and 039 are written and UNAPPLIED. A
|
||||
measurement run in memory verifies parse/derivation/as-of but NOT the DB
|
||||
round-trip — state that distinction rather than implying a feed exists.
|
||||
|
||||
## Active Skills
|
||||
- vyndr-voice (all user-facing output)
|
||||
- prop-analysis (grading methodology)
|
||||
|
||||
@@ -0,0 +1,392 @@
|
||||
#!/usr/bin/env node
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* chain-shadow-probe — WHAT WOULD THE CHAIN SAY, on the real board, right now?
|
||||
*
|
||||
* The shadow accrues on the cron. This runs the SAME code against the rows
|
||||
* already on record so the first divergence distribution is available before a
|
||||
* single new snapshot fires, and so the wiring is measured rather than assumed.
|
||||
*
|
||||
* IT IS A MEASUREMENT, NOT A VERDICT. The chain has never been through a
|
||||
* calibration gate and its head-to-head against the counter is not run here —
|
||||
* that needs settled outcomes, which is why the shadow stores the triple. What
|
||||
* this answers is narrower and comes first: does the chain fire at all, on how
|
||||
* much of the board, and how far does it sit from the number being served? A
|
||||
* challenger that agrees with the incumbent everywhere is carrying nothing; one
|
||||
* that disagrees everywhere is probably broken. Both are worth knowing before
|
||||
* anyone waits a fortnight for outcomes.
|
||||
*
|
||||
* POINT-IN-TIME CAVEAT, STATED UP FRONT: `statcast_aggregates` is upserted in
|
||||
* place and keeps ONE as-of date, so profiles read here are TODAY's. For rows
|
||||
* graded earlier that is contamination, and it is why this reports a
|
||||
* DIVERGENCE (a property of two forecasts) and never a resolution or a lift (a
|
||||
* property of a forecast against an outcome).
|
||||
*
|
||||
* SUPABASE_URL=... node scripts/chain-shadow-probe.js [--days 3]
|
||||
*/
|
||||
|
||||
require('dotenv').config();
|
||||
const { createClient } = require('@supabase/supabase-js');
|
||||
const { paginate } = require('../src/utils/safePaginate');
|
||||
const cs = require('../src/services/model/chainShadow');
|
||||
const reg = require('../src/services/model/featureRegistry');
|
||||
const mk = require('../src/services/model/matchupKeys');
|
||||
const { nameKey } = require('../src/utils/playerName');
|
||||
|
||||
const SB_URL = process.env.SUPABASE_URL;
|
||||
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
|
||||
const DAYS = Number((process.argv.find((a) => a.startsWith('--days=')) || '').split('=')[1]) || 3;
|
||||
|
||||
function daysAgo(n) {
|
||||
const d = new Date(Date.now() - n * 86_400_000);
|
||||
return d.toISOString().slice(0, 10);
|
||||
}
|
||||
|
||||
const pct = (v) => `${(v * 100).toFixed(1)}%`;
|
||||
|
||||
function quantiles(xs) {
|
||||
if (!xs.length) return null;
|
||||
const s = [...xs].sort((a, b) => a - b);
|
||||
const at = (q) => s[Math.min(s.length - 1, Math.max(0, Math.floor(q * (s.length - 1))))];
|
||||
return {
|
||||
min: s[0], p10: at(0.10), p25: at(0.25), median: at(0.50),
|
||||
p75: at(0.75), p90: at(0.90), max: s[s.length - 1],
|
||||
mean: s.reduce((a, b) => a + b, 0) / s.length,
|
||||
};
|
||||
}
|
||||
|
||||
async function main() {
|
||||
if (!SB_URL || !SB_KEY) {
|
||||
console.error('SUPABASE_URL + SUPABASE_SERVICE_ROLE_KEY required.');
|
||||
process.exit(0);
|
||||
}
|
||||
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
|
||||
const since = daysAgo(DAYS);
|
||||
|
||||
// The graded rows, walked through the SAFE paginator (a `.range()` with no
|
||||
// stable order returns the right COUNT and the wrong ROWS — measured at up to
|
||||
// 33.6% duplication on this very table).
|
||||
const rows = await paginate(
|
||||
() => sb.from('model_snapshots')
|
||||
.select('id,player_key,player_name,stat,line,side,p_win,game_id,game_date,archetype,team,refused')
|
||||
.eq('sport', 'mlb').eq('stat', 'hits').gte('game_date', since),
|
||||
{ key: 'id', label: 'chain-shadow-probe:model_snapshots' },
|
||||
);
|
||||
const graded = rows.filter((r) => !r.refused && r.p_win != null);
|
||||
console.log(`\nBOARD — ${rows.length} hits rows since ${since}, ${graded.length} graded\n`);
|
||||
if (!graded.length) { process.exit(0); }
|
||||
|
||||
// Statcast profiles, keyed the way the shadow keys them.
|
||||
const sc = await paginate(
|
||||
() => sb.from('statcast_aggregates').select('*').eq('sport', 'mlb'),
|
||||
// The real unique key — statcast_aggregates is upserted in place and keeps
|
||||
// no as_of_date (which is exactly why this probe cannot be point-in-time).
|
||||
{ key: ['sport', 'season', 'source_id', 'role'], label: 'chain-shadow-probe:statcast' },
|
||||
);
|
||||
const byKey = new Map();
|
||||
for (const r of sc) {
|
||||
if (!r.player_key) continue;
|
||||
const prev = byKey.get(r.player_key);
|
||||
const size = Number(r.sample_pa || r.sample_ip || 0);
|
||||
if (!prev || size > Number(prev.sample_pa || prev.sample_ip || 0)) byKey.set(r.player_key, r);
|
||||
}
|
||||
console.log(`statcast profiles: ${byKey.size}`);
|
||||
|
||||
// ── THE MATCHUP INPUTS, resolved PER GAME DATE and AS-OF-CORRECT ────────
|
||||
// v1 passed `pitcherRowFor: () => null` and no hand split, so it measured a
|
||||
// batter-side SEASON read and said so. This resolves what the live shadow
|
||||
// resolves: the opposing starter (A6 keys) and the hitter's own platoon split
|
||||
// (A4-dated), each bounded at the row's own game date so nothing later than
|
||||
// the grade can leak in.
|
||||
const mlbAdapter = require('../src/services/adapters/mlbStatsAdapter');
|
||||
const dates = [...new Set(graded.map((r) => r.game_date).filter(Boolean))].sort();
|
||||
|
||||
const platoon = await paginate(
|
||||
() => sb.from('platoon_splits').select('*').eq('sport', 'mlb'),
|
||||
{ key: ['as_of_date', 'sport', 'season', 'player_key'], label: 'chain-shadow-probe:platoon' },
|
||||
);
|
||||
const lineups = await paginate(
|
||||
() => sb.from('lineup_context').select('*').eq('sport', 'mlb').in('game_date', dates),
|
||||
{ key: ['as_of_date', 'sport', 'game_pk', 'team', 'player_key'], label: 'chain-shadow-probe:lineups' },
|
||||
);
|
||||
console.log(`platoon splits: ${platoon.length} rows · lineup rows: ${lineups.length}`);
|
||||
|
||||
/** Latest row at or before `asOf` — refuse rather than reach forward (A4). */
|
||||
const latestAsOf = (rows, asOf) => {
|
||||
let best = null;
|
||||
for (const r of rows) {
|
||||
if (!r.as_of_date || r.as_of_date > asOf) continue;
|
||||
if (!best || r.as_of_date > best.as_of_date) best = r;
|
||||
}
|
||||
return best;
|
||||
};
|
||||
const platoonByPlayer = new Map();
|
||||
for (const r of platoon) {
|
||||
if (!r.player_key) continue;
|
||||
if (!platoonByPlayer.has(r.player_key)) platoonByPlayer.set(r.player_key, []);
|
||||
platoonByPlayer.get(r.player_key).push(r);
|
||||
}
|
||||
const slotByPlayerDate = new Map();
|
||||
for (const r of lineups) {
|
||||
if (!r.player_key || r.batting_order == null) continue;
|
||||
slotByPlayerDate.set(`${r.player_key}|${r.game_date}`, r.batting_order);
|
||||
}
|
||||
|
||||
// Pitcher statcast rows by NAME (the A6 key resolves a pitcher name).
|
||||
const pitcherByKey = new Map();
|
||||
for (const r of sc) if (r.role === 'pitcher' && r.player_key) pitcherByKey.set(r.player_key, r);
|
||||
|
||||
const keysByDate = new Map();
|
||||
for (const d of dates) {
|
||||
try {
|
||||
const resolve = await mk.build({
|
||||
sb, getSchedule: (dd) => mlbAdapter.getScheduleWithPitchers(dd), gameDate: d, asOf: d,
|
||||
});
|
||||
if (resolve) keysByDate.set(d, resolve);
|
||||
} catch (e) { console.warn(` matchupKeys ${d}: ${e.message}`); }
|
||||
}
|
||||
console.log(`matchup key indexes built for ${keysByDate.size}/${dates.length} dates`);
|
||||
|
||||
// One "grade" per row, in the shape runShadow reads. The LOCKED line is the
|
||||
// row's own line, which is what the counter was graded against.
|
||||
const grades = graded.map((r) => ({
|
||||
player: r.player_name || r.player_key,
|
||||
stat_type: r.stat,
|
||||
direction: r.side,
|
||||
p_win: r.p_win,
|
||||
game_id: r.game_id,
|
||||
game_date: r.game_date,
|
||||
team: r.team,
|
||||
archetype: r.archetype,
|
||||
gradedAt: { line: r.line },
|
||||
}));
|
||||
|
||||
const throwsFor = (g) => {
|
||||
const resolve = keysByDate.get(g.game_date);
|
||||
if (!resolve) return null;
|
||||
const k = resolve({ player: g.player });
|
||||
if (!k || !k.opposing_pitcher) return null;
|
||||
const row = pitcherByKey.get(nameKey(k.opposing_pitcher));
|
||||
return row || null;
|
||||
};
|
||||
|
||||
const shadowDeps = (over = {}) => ({
|
||||
statcastByKey: byKey,
|
||||
pitcherRowFor: throwsFor,
|
||||
lineupSlotFor: (g) => slotByPlayerDate.get(`${nameKey(g.player)}|${g.game_date}`) ?? null,
|
||||
handSplitFor: (g) => {
|
||||
const pk = nameKey(g.player);
|
||||
const rows = platoonByPlayer.get(pk) || [];
|
||||
// AS-OF-CORRECT: the split as it stood on the row's own game date.
|
||||
const sp = latestAsOf(rows, g.game_date);
|
||||
const pitcherRow = throwsFor(g);
|
||||
return {
|
||||
bats: sp && sp.bats ? String(sp.bats)[0] : ((byKey.get(pk) || {}).bats || null),
|
||||
throws: pitcherRow && pitcherRow.throws ? String(pitcherRow.throws)[0] : null,
|
||||
platoonSplits: sp ? {
|
||||
vl: { pa: sp.vl_pa, atBats: sp.vl_ab, hits: sp.vl_hits },
|
||||
vr: { pa: sp.vr_pa, atBats: sp.vr_ab, hits: sp.vr_hits },
|
||||
} : null,
|
||||
};
|
||||
},
|
||||
allowed: reg.candidateFeatures('mlb'),
|
||||
...over,
|
||||
});
|
||||
|
||||
const out = cs.runShadow(grades, shadowDeps());
|
||||
// THE A/B THE FIX IS JUDGED ON: identical rows, identical everything, except
|
||||
// the opportunity term is forced back to the constant `DEFAULT_PA` the v2
|
||||
// shadow actually ran on. Anything else would compare two different slates.
|
||||
const outFixedPa = cs.runShadow(grades, shadowDeps({ lineupSlotFor: () => null }));
|
||||
|
||||
const s = out.summary;
|
||||
console.log(`\nCHAIN FIRE — ${s.readable}/${s.atoms} atoms read (${s.refused} refused) `
|
||||
+ `across ${s.games_read}/${s.games} games`);
|
||||
if (s.refused) console.log(` refusal reasons: ${JSON.stringify(s.refusal_reasons)}`);
|
||||
console.log(`\nHAND SPLIT — fired on ${s.platoon_applied}/${s.props} unique props `
|
||||
+ `(${pct(s.platoon_applied / Math.max(1, s.props))})`);
|
||||
console.log(` season-rate reasons: ${JSON.stringify(s.platoon_reasons)}`);
|
||||
console.log(`\nOPPORTUNITY — posted lineup slot on ${s.opportunity_posted}/${s.props} props `
|
||||
+ `(${pct(s.opportunity_posted / Math.max(1, s.props))})`);
|
||||
console.log(` fell to a default regular: ${JSON.stringify(s.opportunity_reasons)}`);
|
||||
|
||||
// Side-align every row against its OWN served p_win — the triple's first two
|
||||
// thirds, exactly as the shadow stores them.
|
||||
const divergences = [];
|
||||
const chainPs = [];
|
||||
const counterPs = [];
|
||||
const firedDiv = []; // hand split APPLIED — a genuine matchup read
|
||||
const seasonDiv = []; // hand split refused — a season read
|
||||
let matched = 0;
|
||||
for (const r of graded) {
|
||||
const block = out.byKey.get(cs.shadowKey(r.player_key, r.stat, r.line));
|
||||
if (!block) continue;
|
||||
const aligned = cs.alignToSide(block, r.side, r.p_win);
|
||||
if (!aligned || aligned.divergence == null) continue;
|
||||
matched += 1;
|
||||
divergences.push(aligned.divergence);
|
||||
chainPs.push(aligned.chain_p);
|
||||
counterPs.push(aligned.counter_p);
|
||||
(block.platoon_applied ? firedDiv : seasonDiv).push(aligned.divergence);
|
||||
}
|
||||
|
||||
if (!matched) { console.log('\nNo comparable rows — nothing to report.'); process.exit(0); }
|
||||
|
||||
const abs = divergences.map(Math.abs);
|
||||
const q = quantiles(divergences);
|
||||
const qa = quantiles(abs);
|
||||
const qc = quantiles(chainPs);
|
||||
const qk = quantiles(counterPs);
|
||||
|
||||
console.log(`\nCHAIN vs COUNTER — n=${matched} (${pct(matched / graded.length)} of graded)\n`);
|
||||
const row = (label, x) => console.log(
|
||||
` ${label.padEnd(22)} min ${x.min.toFixed(3)} p25 ${x.p25.toFixed(3)} median ${x.median.toFixed(3)}`
|
||||
+ ` p75 ${x.p75.toFixed(3)} max ${x.max.toFixed(3)} mean ${x.mean.toFixed(3)}`);
|
||||
row('chain_p', qc);
|
||||
row('counter_p', qk);
|
||||
row('divergence (signed)', q);
|
||||
row('divergence (absolute)', qa);
|
||||
|
||||
const band = (lo, hi) => abs.filter((d) => d >= lo && d < hi).length;
|
||||
console.log(`\n |divergence| < 0.02 ${band(0, 0.02)} (${pct(band(0, 0.02) / matched)}) — agrees with the counter`);
|
||||
console.log(` 0.02 - 0.05 ${band(0.02, 0.05)} (${pct(band(0.02, 0.05) / matched)})`);
|
||||
console.log(` 0.05 - 0.10 ${band(0.05, 0.10)} (${pct(band(0.05, 0.10) / matched)})`);
|
||||
console.log(` 0.10 - 0.20 ${band(0.10, 0.20)} (${pct(band(0.10, 0.20) / matched)})`);
|
||||
console.log(` >= 0.20 ${band(0.20, 99)} (${pct(band(0.20, 99) / matched)}) — a different read entirely`);
|
||||
|
||||
const higher = divergences.filter((d) => d > 0).length;
|
||||
console.log(`\n chain HIGHER than counter on ${higher} (${pct(higher / matched)}), lower on ${matched - higher}`);
|
||||
|
||||
// BOTH SIDES OF EVERY PROP ARE IN THE SAMPLE, so each pair contributes +d and
|
||||
// -d and the signed distribution is forced to be symmetric about zero. That
|
||||
// symmetry is arithmetic, not a finding — reading it as "the chain is unbiased"
|
||||
// would be reading the sampling scheme. The OVER slice is where a directional
|
||||
// lean is visible at all.
|
||||
const overs = [];
|
||||
for (const r of graded) {
|
||||
if (String(r.side).toLowerCase() !== 'over') continue;
|
||||
const block = out.byKey.get(cs.shadowKey(r.player_key, r.stat, r.line));
|
||||
const aligned = block && cs.alignToSide(block, r.side, r.p_win);
|
||||
if (aligned && aligned.divergence != null) overs.push(aligned.divergence);
|
||||
}
|
||||
if (overs.length) {
|
||||
const qo = quantiles(overs);
|
||||
const hi = overs.filter((d) => d > 0).length;
|
||||
console.log(`\n OVER SIDE ONLY (n=${overs.length}) — the signed read that is not forced symmetric`);
|
||||
row(' divergence (signed)', qo);
|
||||
console.log(` chain HIGHER on ${hi} (${pct(hi / overs.length)}) — mean ${qo.mean.toFixed(4)}`);
|
||||
}
|
||||
|
||||
// ── THE SPLIT THAT ACTUALLY MATTERS ─────────────────────────────────────
|
||||
// A row where the hand split refused is a SEASON read; a row where it fired is
|
||||
// a MATCHUP read. Pooling them reports an average of two different models, and
|
||||
// any later adjudication would be measuring the mixture rather than the chain.
|
||||
const slice = (label, arr) => {
|
||||
if (!arr.length) { console.log(`\n ${label}: none`); return; }
|
||||
const q2 = quantiles(arr.map(Math.abs));
|
||||
console.log(`\n ${label} (n=${arr.length})`);
|
||||
row(' |divergence|', q2);
|
||||
const far = arr.filter((d) => Math.abs(d) >= 0.10).length;
|
||||
const near = arr.filter((d) => Math.abs(d) < 0.02).length;
|
||||
console.log(` disagree >= 0.10 on ${far} (${pct(far / arr.length)}) · `
|
||||
+ `agree within 0.02 on ${near} (${pct(near / arr.length)})`);
|
||||
};
|
||||
console.log('\n──────── MATCHUP READ vs SEASON READ ────────');
|
||||
slice('HAND SPLIT FIRED — a genuine matchup read', firedDiv);
|
||||
slice('HAND SPLIT REFUSED — a season read', seasonDiv);
|
||||
|
||||
// ── PHASE 1 CHECK: DID THE UNIFORM BIAS COLLAPSE? ───────────────────────
|
||||
// Signed bias on the OVER side only (the two-sided set is forced symmetric).
|
||||
// Same rows, both arms, so the difference is the opportunity term and nothing
|
||||
// else. `%below` is the shape that matters: a MECHANICAL bias pushes nearly
|
||||
// every row the same way, and a real conditioner does not.
|
||||
const armStats = (shadow) => {
|
||||
const d = [];
|
||||
for (const r of graded) {
|
||||
if (String(r.side).toLowerCase() !== 'over') continue;
|
||||
const b = shadow.byKey.get(cs.shadowKey(r.player_key, r.stat, r.line));
|
||||
const a = b && cs.alignToSide(b, r.side, r.p_win);
|
||||
if (a && a.divergence != null) d.push(a.divergence);
|
||||
}
|
||||
if (!d.length) return null;
|
||||
const q = quantiles(d);
|
||||
return { n: d.length, mean: q.mean, median: q.median, below: d.filter((x) => x < 0).length };
|
||||
};
|
||||
console.log('\n──────── PHASE 1 — E[PA]: REAL LINEUP SLOT vs CONSTANT 4.1 ────────');
|
||||
for (const [label, shadow] of [['REAL E[PA] (posted slot)', out], ['CONSTANT 4.1 (the v2 shadow)', outFixedPa]]) {
|
||||
const a = armStats(shadow);
|
||||
if (!a) { console.log(` ${label}: none`); continue; }
|
||||
console.log(` ${label.padEnd(30)} n=${a.n} mean ${a.mean >= 0 ? '+' : ''}${a.mean.toFixed(4)} `
|
||||
+ `median ${a.median >= 0 ? '+' : ''}${a.median.toFixed(4)} chain BELOW counter on ${pct(a.below / a.n)}`);
|
||||
}
|
||||
|
||||
// ── PHASE 2 — THE DIAGNOSTIC: ARTIFACT or SIGNAL? ───────────────────────
|
||||
// A MECHANICAL bias is flat across matchup difficulty: a broken conversion
|
||||
// does not know who is pitching. A CONDITIONER is not flat — it moves with the
|
||||
// matchup the counter is blind to.
|
||||
//
|
||||
// READ THE CAVEAT WITH THE RESULT: the chain reads the opposing pitcher's K
|
||||
// rate directly and the counter does not, so a monotone relationship here is
|
||||
// close to mechanical proof that the chain CONDITIONS on the matchup. It is
|
||||
// NOT evidence the conditioning is CORRECT. Only settled outcomes can say that.
|
||||
const rowsWithDifficulty = [];
|
||||
for (const r of graded) {
|
||||
if (String(r.side).toLowerCase() !== 'over') continue;
|
||||
const b = out.byKey.get(cs.shadowKey(r.player_key, r.stat, r.line));
|
||||
const a = b && cs.alignToSide(b, r.side, r.p_win);
|
||||
if (!a || a.divergence == null) continue;
|
||||
const g = grades.find((x) => nameKey(x.player) === r.player_key && x.game_date === r.game_date);
|
||||
const pit = g ? throwsFor(g) : null;
|
||||
const k = pit && pit.k_pct != null ? Number(pit.k_pct) : null;
|
||||
rowsWithDifficulty.push({ div: a.divergence, oppK: k, platoon: b.platoon_applied ? b.platoon_multiplier : null });
|
||||
}
|
||||
const withK = rowsWithDifficulty.filter((x) => Number.isFinite(x.oppK));
|
||||
console.log('\n──────── PHASE 2 — DIVERGENCE BY MATCHUP DIFFICULTY ────────');
|
||||
if (withK.length < 50) {
|
||||
console.log(` only ${withK.length} rows carry an opposing-pitcher profile — too few to bucket.`);
|
||||
} else {
|
||||
withK.sort((a, b) => a.oppK - b.oppK);
|
||||
const B = 5;
|
||||
const size = Math.floor(withK.length / B);
|
||||
console.log(` n=${withK.length}, quintiles of OPPOSING PITCHER K% (low = soft matchup)`);
|
||||
const means = [];
|
||||
for (let i = 0; i < B; i += 1) {
|
||||
const chunk = withK.slice(i * size, i === B - 1 ? withK.length : (i + 1) * size);
|
||||
const m = chunk.reduce((acc, x) => acc + x.div, 0) / chunk.length;
|
||||
means.push(m);
|
||||
const kLo = chunk[0].oppK.toFixed(1);
|
||||
const kHi = chunk[chunk.length - 1].oppK.toFixed(1);
|
||||
const below = chunk.filter((x) => x.div < 0).length;
|
||||
console.log(` Q${i + 1} oppK ${kLo}-${kHi}%`.padEnd(28)
|
||||
+ `n=${String(chunk.length).padEnd(6)}mean div ${m >= 0 ? '+' : ''}${m.toFixed(4)} below ${pct(below / chunk.length)}`);
|
||||
}
|
||||
const spread = means[0] - means[means.length - 1];
|
||||
// Pearson r between opposing-pitcher K% and divergence.
|
||||
const xs = withK.map((x) => x.oppK); const ys = withK.map((x) => x.div);
|
||||
const mx = xs.reduce((a, b) => a + b, 0) / xs.length;
|
||||
const my = ys.reduce((a, b) => a + b, 0) / ys.length;
|
||||
let sxy = 0; let sxx = 0; let syy = 0;
|
||||
for (let i = 0; i < xs.length; i += 1) {
|
||||
sxy += (xs[i] - mx) * (ys[i] - my); sxx += (xs[i] - mx) ** 2; syy += (ys[i] - my) ** 2;
|
||||
}
|
||||
const r = sxy / Math.sqrt(sxx * syy);
|
||||
console.log(`\n Q1 - Q5 spread: ${spread >= 0 ? '+' : ''}${spread.toFixed(4)} `
|
||||
+ `Pearson r(oppK, divergence) = ${r.toFixed(4)}`);
|
||||
console.log(` VERDICT: ${Math.abs(spread) < 0.02 && Math.abs(r) < 0.05
|
||||
? 'FLAT across difficulty — the divergence is MECHANICAL, the chain is still miscomputing.'
|
||||
: 'MOVES with difficulty — the chain is CONDITIONING on the matchup the counter cannot see.'}`);
|
||||
console.log(' (Conditioning is not correctness. Only settled outcomes can say which read is right.)');
|
||||
}
|
||||
console.log('\n UN-SERVABLE. No calibration gate has been passed; this is a divergence, not a verdict.\n');
|
||||
|
||||
// Redis runs degraded locally and a reconnect timer holds the process open,
|
||||
// which loses piped output to SIGTERM.
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
if (require.main === module) {
|
||||
main().catch((e) => { console.error(e.message); process.exit(1); });
|
||||
}
|
||||
|
||||
module.exports = { quantiles };
|
||||
@@ -0,0 +1,144 @@
|
||||
#!/usr/bin/env node
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* wnba-feed-verify — prove the tap, the ingest and the as-of read on REAL data.
|
||||
*
|
||||
* Runs the production parsers over the live WNBA season and holds the rows in
|
||||
* memory, then drives `wnbaUsageService.profileAsOf` against them through a stub
|
||||
* client. It exercises the SAME functions the pipeline will call — not a
|
||||
* restatement of them — so a passing run is evidence about the code that ships.
|
||||
*
|
||||
* node scripts/wnba-feed-verify.js [--from=2026-05-01] [--to=2026-08-12]
|
||||
*/
|
||||
|
||||
const adapter = require('../src/services/adapters/espnWnbaAdapter');
|
||||
const usage = require('../src/services/wnbaUsageService');
|
||||
const { nameKey } = require('../src/utils/playerName');
|
||||
|
||||
const arg = (k, d) => {
|
||||
const hit = process.argv.find((a) => a.startsWith(`--${k}=`));
|
||||
return hit ? hit.split('=')[1] : d;
|
||||
};
|
||||
|
||||
/** A stub client that serves in-memory rows through the real read path. */
|
||||
function memClient(rows) {
|
||||
return {
|
||||
from() {
|
||||
const state = { rows: rows.slice(), order: [], range: null };
|
||||
const q = {
|
||||
select: () => q,
|
||||
eq: (col, val) => { state.rows = state.rows.filter((r) => r[col] === val); return q; },
|
||||
lt: (col, val) => { state.rows = state.rows.filter((r) => String(r[col]) < String(val)); return q; },
|
||||
order: (col, o) => { state.order.push([col, o && o.ascending !== false]); return q; },
|
||||
range: async (from, to) => {
|
||||
const sorted = state.rows.slice().sort((a, b) => {
|
||||
for (const [c, asc] of state.order) {
|
||||
const cmp = String(a[c]).localeCompare(String(b[c]));
|
||||
if (cmp) return asc ? cmp : -cmp;
|
||||
}
|
||||
return 0;
|
||||
});
|
||||
return { data: sorted.slice(from, to + 1), error: null };
|
||||
},
|
||||
};
|
||||
return q;
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
async function main() {
|
||||
const to = arg('to', usage.dateET());
|
||||
const from = arg('from', '2026-05-01');
|
||||
console.log(`\nWNBA FEED — pulling completed games ${from} .. ${to}\n`);
|
||||
|
||||
const rows = [];
|
||||
const games = new Set();
|
||||
const dateStats = [];
|
||||
let errors = 0;
|
||||
|
||||
for (const date of usage.eachDate(from, to)) {
|
||||
let ids = [];
|
||||
try { ids = await adapter.getFinalGameIds(date); } catch (e) { errors += 1; continue; }
|
||||
if (!ids.length) continue;
|
||||
let dayRows = 0;
|
||||
for (const id of ids) {
|
||||
try {
|
||||
const parsed = await adapter.getGameRows(id);
|
||||
if (parsed.skipped || !parsed.rows.length) continue;
|
||||
games.add(id);
|
||||
for (const r of parsed.rows) {
|
||||
if (!r.source_id || !r.player_name) continue;
|
||||
rows.push({ ...r, player_key: nameKey(r.player_name) });
|
||||
dayRows += 1;
|
||||
}
|
||||
} catch { errors += 1; }
|
||||
}
|
||||
dateStats.push({ date, games: ids.length, rows: dayRows });
|
||||
process.stdout.write(` ${date}: ${ids.length} games, ${dayRows} rows\r`);
|
||||
}
|
||||
|
||||
const dates = [...new Set(rows.map((r) => r.game_date))].sort();
|
||||
console.log('\n\n──────── COVERAGE ────────');
|
||||
console.log(` rows ${rows.length}`);
|
||||
console.log(` games ${games.size}`);
|
||||
console.log(` players ${new Set(rows.map((r) => r.player_key)).size}`);
|
||||
console.log(` teams ${new Set(rows.map((r) => r.team)).size}`);
|
||||
console.log(` date range ${dates[0]} .. ${dates[dates.length - 1]} (${dates.length} game days)`);
|
||||
console.log(` fetch errors ${errors}`);
|
||||
const cov = (f) => `${rows.filter((r) => r[f] != null).length}/${rows.length}`;
|
||||
console.log('\n FIELD COVERAGE (the chainFn inputs)');
|
||||
for (const f of ['minutes', 'usage_rate', 'team_possessions', 'team_pace', 'ts_pct', 'efg_pct', 'starter', 'final_margin']) {
|
||||
console.log(` ${f.padEnd(20)} ${cov(f)}`);
|
||||
}
|
||||
|
||||
// A worked example, so the derivation is checkable by hand.
|
||||
const sample = rows.find((r) => r.minutes > 25 && r.usage_rate != null);
|
||||
if (sample) {
|
||||
console.log('\n WORKED EXAMPLE (derivation is exact arithmetic, not a proxy)');
|
||||
console.log(` ${sample.player_name} (${sample.team}) ${sample.game_date}`);
|
||||
console.log(` MIN ${sample.minutes} FGA ${sample.fga} FTA ${sample.fta} TOV ${sample.tov} PTS ${sample.points}`);
|
||||
console.log(` team: MIN ${sample.team_minutes} FGA ${sample.team_fga} FTA ${sample.team_fta} TOV ${sample.team_tov} OREB ${sample.team_oreb}`);
|
||||
const poss = sample.team_fga - sample.team_oreb + sample.team_tov + 0.44 * sample.team_fta;
|
||||
const usg = 100 * ((sample.fga + 0.44 * sample.fta + sample.tov) * (sample.team_minutes / 5))
|
||||
/ (sample.minutes * (sample.team_fga + 0.44 * sample.team_fta + sample.team_tov));
|
||||
console.log(` possessions stored ${sample.team_possessions} recomputed ${poss.toFixed(3)}`);
|
||||
console.log(` usage_rate stored ${sample.usage_rate} recomputed ${usg.toFixed(3)}`);
|
||||
console.log(` pace ${sample.team_pace} TS% ${sample.ts_pct} starter ${sample.starter} margin ${sample.final_margin}`);
|
||||
}
|
||||
|
||||
// ── AS-OF VERIFICATION, through the REAL read path ──────────────────────
|
||||
console.log('\n──────── AS-OF READ (through wnbaUsageService.profileAsOf) ────────');
|
||||
const sb = memClient(rows);
|
||||
const counts = new Map();
|
||||
for (const r of rows) counts.set(r.player_key, (counts.get(r.player_key) || 0) + 1);
|
||||
const busiest = [...counts.entries()].sort((a, b) => b[1] - a[1])[0];
|
||||
const pk = busiest[0];
|
||||
const theirs = rows.filter((r) => r.player_key === pk).sort((a, b) => a.game_date.localeCompare(b.game_date));
|
||||
const name = theirs[0].player_name;
|
||||
console.log(` player: ${name} (${pk}) — ${theirs.length} games on record\n`);
|
||||
|
||||
const mid = theirs[Math.floor(theirs.length / 2)].game_date;
|
||||
for (const asOf of [theirs[2].game_date, mid, theirs[theirs.length - 1].game_date, '2026-12-31']) {
|
||||
const p = await usage.profileAsOf(sb, { playerKey: pk, asOf });
|
||||
const expected = theirs.filter((r) => r.game_date < asOf).length;
|
||||
const ok = p ? p.games === expected : expected < usage.MIN_GAMES;
|
||||
console.log(` asOf ${asOf} games=${p ? p.games : 'REFUSED'} (expected ${expected}) ${ok ? 'OK' : 'MISMATCH'}`
|
||||
+ (p ? ` usage ${p.usage_rate} min/g ${p.minutes_per_game} pace ${p.team_pace} TS ${p.ts_pct}` : ''));
|
||||
}
|
||||
|
||||
// THE LEAK TEST: a game ON the as-of date must never be counted.
|
||||
const onDate = theirs[theirs.length - 1].game_date;
|
||||
const p = await usage.profileAsOf(sb, { playerKey: pk, asOf: onDate });
|
||||
const includesSameDay = p && p.last_game_date === onDate;
|
||||
console.log(`\n STRICT CUTOFF — a game ON the as-of date is ${includesSameDay ? 'INCLUDED (LEAK)' : 'EXCLUDED (correct)'}`);
|
||||
console.log(` last game counted: ${p ? p.last_game_date : 'n/a'} vs as-of ${onDate}`);
|
||||
|
||||
// And the refusal.
|
||||
const thin = await usage.profileAsOf(sb, { playerKey: pk, asOf: theirs[0].game_date });
|
||||
console.log(` THIN HISTORY — asOf before his first game returns ${thin === null ? 'null (refused, correct)' : 'A PROFILE (WRONG)'}`);
|
||||
|
||||
process.exit(0);
|
||||
}
|
||||
|
||||
if (require.main === module) main().catch((e) => { console.error(e); process.exit(1); });
|
||||
@@ -0,0 +1,503 @@
|
||||
# chain-v1 / v2 — the portable engine, made whole, fed, and shadowed on MLB
|
||||
|
||||
> **v2 (2026-08-12) is appended at §8.** It corrects the order's premise (the
|
||||
> chain was already firing on 99.6%, `paOutcome` reads no handedness, and there
|
||||
> is no fallback path anywhere) and plumbs the hand split as a matchup
|
||||
> conditioner on the RATE. Read §8 for the current numbers; §4 below is the v1
|
||||
> measurement and is superseded.
|
||||
|
||||
|
||||
|
||||
**Status:** SHADOW. Served by nothing. Nothing promoted.
|
||||
**Date:** 2026-08-12
|
||||
**Predecessors:** chaining-v1 (Session 90 — the design, which never had a spec
|
||||
file), A5 (factor freeze), A6 (the join keys, shadow), A7 (shadow accrual).
|
||||
|
||||
---
|
||||
|
||||
## 0. The finding this order started from
|
||||
|
||||
The extraction pass found `chain.js` was a **shell of its own header**:
|
||||
|
||||
| the header promised | the code had |
|
||||
|---|---|
|
||||
| `atoms` | a positional array ✅ |
|
||||
| `context` | passed to `redistribute` only, partially |
|
||||
| **`chainFn`** — atom → per-entity probability | **ABSENT. No parameter, no call site, no export.** |
|
||||
| `aggregator` — ACROSS / UP | two functions ✅ |
|
||||
| `redistribute` | **on `chainUp` only** |
|
||||
|
||||
And it was inert for a second, undocumented reason: `chainAcross` requires
|
||||
`calibrated: true`, and the only writer of that flag is a loop over
|
||||
`CALIBRATION_DEPLOYED`, which is `Object.freeze([])`. **No grade in the system
|
||||
carries the flag**, so `chainAcross` had zero callers *and* would have refused
|
||||
any caller it had.
|
||||
|
||||
The missing `chainFn` is why the "portable core" was not portable: with no slot
|
||||
for the atom→probability stage, every sport's real work had to live somewhere
|
||||
else. For MLB it lived in `scripts/`, reachable from no pipeline.
|
||||
|
||||
---
|
||||
|
||||
## 1. The three core fixes (`src/services/model/chain.js`)
|
||||
|
||||
### 1.1 `chainFn(atom, context) → per-entity probability`
|
||||
|
||||
- `applyChainFn` runs before usability filtering. **Default is identity-on-`p`**,
|
||||
so every pre-existing caller is byte-identical.
|
||||
- Returning a **number** sets `p`; returning an **object** merges (so a chainFn
|
||||
can attach its own trace); returning **null or throwing** makes the atom
|
||||
UNREADABLE ⇒ **DROPPED, counted, never `p=0`**. A zero leg would zero an entire
|
||||
ticket, and "we could not read him" is not "he cannot do it".
|
||||
- The refusals are surfaced as `chain_fn_refused` rather than swallowed, so a
|
||||
chainFn quietly failing across the board is visible instead of looking like a
|
||||
thin slate.
|
||||
|
||||
### 1.2 Correlation is SIGNED, and the direction follows the sign
|
||||
|
||||
Was clamped `[0, 1]` with the joint always shifted toward the weakest leg. That
|
||||
is baseball's shape — same-game legs share the pitcher, the park and the weather
|
||||
— and it made basketball's case **inexpressible**: teammates compete for finite
|
||||
possessions, so one player's shot is another's non-shot and their props are
|
||||
NEGATIVELY correlated.
|
||||
|
||||
Correlation is now clamped `[-1, 1]` and `|corr|` interpolates from independence
|
||||
toward the **Fréchet–Hoeffding bound its sign selects**:
|
||||
|
||||
```
|
||||
corr = +1 → joint = min(p_i) (upper bound, co-monotone)
|
||||
corr = 0 → joint = Π p_i (independent)
|
||||
corr = −1 → joint = max(0, Σp − (n−1)) (lower bound, counter-monotone)
|
||||
```
|
||||
|
||||
This is why the direction is principled rather than chosen. The positive branch
|
||||
is arithmetically unchanged — the weakest leg **is** the upper bound — so every
|
||||
previously-correct number stays exactly what it was.
|
||||
|
||||
### 1.3 `redistribute` reaches BOTH readings
|
||||
|
||||
Previously on `chainUp` only. A redistribution that reached the team read and not
|
||||
the across read would leave `selfCheck` comparing post-redistribution to
|
||||
pre-redistribution atoms and flagging an INTERNAL_INCONSISTENCY **the model had
|
||||
itself just manufactured**.
|
||||
|
||||
`prepareAtoms(atoms, opts)` is exported so a caller prepares ONCE — chainFn, then
|
||||
usability, then redistribution — and hands the identical legs to both readings.
|
||||
|
||||
---
|
||||
|
||||
## 2. Baseball's chainFn (`src/services/model/baseballChain.js`)
|
||||
|
||||
A chained forecast is always `rate × opportunity`. Baseball is the clean case
|
||||
because **opportunity is fixed**: the batting order is set before first pitch, a
|
||||
nine-run lead does not change who bats next, and PA/game varies over a narrow
|
||||
range set almost entirely by lineup slot.
|
||||
|
||||
```
|
||||
p_hit_per_PA ← skillProjection.paOutcome the RATE (modelled)
|
||||
expected PA ← lineup slot the OPPORTUNITY (a lookup)
|
||||
P(hits ≥ k) ← Binomial(PA, p) mixed over paDistribution
|
||||
```
|
||||
|
||||
**Atoms routed** — all pre-existing, all previously reachable only from
|
||||
`scripts/`:
|
||||
|
||||
| atom | source | role |
|
||||
|---|---|---|
|
||||
| `fromStatcastRow` | skillProjection:161 | the ONE legal units conversion (statcast stores PERCENTAGES 0–100) |
|
||||
| `paOutcome` | skillProjection:268 | K / BB via log5 odds-ratio vs league; remainder = balls in play |
|
||||
| `hitOnContact` | skillProjection:213 | archetype-selected barrel / hard-hit / exit-velo / GB-speed × pitcher contact allowed × park |
|
||||
| `paDistribution`, `binomialPmf`, `atLeast` | skillProjection:327/308/339 | the chain over opportunity |
|
||||
|
||||
`PA_BY_SLOT` is a **lookup, not a fit** — 4.65 (leadoff) down to 3.85 (nine hole),
|
||||
the documented ~0.1-PA-per-slot decline. Nothing was tuned on settled rows; a
|
||||
tuned opportunity term on 1,741 rows is curve-fitting dressed as physics.
|
||||
|
||||
**Refusals:** no batter profile ⇒ null ⇒ dropped (never a league hitter). A
|
||||
missing pitcher is different and handled inside `paOutcome` — the batter's own
|
||||
rate stands rather than being pulled toward average.
|
||||
|
||||
**Scope: hits only.** `projectSkill` also routes `total_bases`, but TB's
|
||||
head-to-head is on record as INCONCLUSIVE-under-contamination. Widening the first
|
||||
shadow to a stat whose verdict is already muddy buys noise.
|
||||
|
||||
**`redistribute` is dormant** and returns the legs unchanged — the honest dormant
|
||||
behaviour. Returning null would read as "the hook failed".
|
||||
|
||||
---
|
||||
|
||||
## 3. The shadow (`src/services/model/chainShadow.js`)
|
||||
|
||||
Runs at the post-enriched / pre-persist point in `snapshotService` — the A6
|
||||
position, where the slate exists as a set. Reuses the statcast rows and resolved
|
||||
opposing starter the challenger pass already fetched: **zero new I/O**.
|
||||
|
||||
### The unit of evidence is the TRIPLE
|
||||
|
||||
```
|
||||
(chain_p, counter_p, outcome)
|
||||
```
|
||||
|
||||
all three on one row, all side-aligned. `counter_p` is the served
|
||||
`estimateProbability` value for that exact side; `outcome` arrives from the
|
||||
ordinary settle pass; `chain_p` is the chain's. **Evidence you cannot adjudicate
|
||||
is not evidence** — a chain probability stored without the number it must beat,
|
||||
or without the result, can only ever be compared to itself.
|
||||
|
||||
**Side alignment is load-bearing.** The chain computes P(over the line); `p_win`
|
||||
is expressed for the graded SIDE. An under row stores `1 − p_over`. Storing the
|
||||
raw over-probability against an under row would invert every later comparison,
|
||||
silently.
|
||||
|
||||
### UN-SERVABLE, said in the data
|
||||
|
||||
`requireCalibrated: false` is legitimate **only** because nothing downstream
|
||||
reads the result. Every stored block carries `status: 'UN-SERVABLE'` and
|
||||
`servable: false` in its own payload, not merely in a comment — a caveat that
|
||||
lives only in a comment is not attached to the data once something else queries
|
||||
it. Flipping that requires passing the gate, not editing a file.
|
||||
|
||||
### The self-check is VACUOUS today, and says so
|
||||
|
||||
`selfCheck` earns its keep by comparing per-entity reads to an **independent**
|
||||
team read. None exists — the game-script projection was deliberately not built
|
||||
(no atom has passed the gate). So the up-read is assembled from the SAME atoms as
|
||||
the across-read and agreement between them is arithmetic. Every block carries
|
||||
`self_check.vacuous: true` with its reason, so nobody later mistakes a tautology
|
||||
for a passing consistency test.
|
||||
|
||||
### Storage
|
||||
|
||||
`model_snapshots.chain_shadow jsonb` (migration 038). A separate column, not
|
||||
extra keys inside `features` — `champion-ablation.js` iterates every `features`
|
||||
key for its residual scan, so widening it would silently enlarge that
|
||||
multiple-comparisons denominator. Same reasoning as 034.
|
||||
|
||||
---
|
||||
|
||||
## 4. First measurement (`scripts/chain-shadow-probe.js`)
|
||||
|
||||
Real board, 2026-08-07 → 2026-08-12, MLB hits, **9,376 graded rows**:
|
||||
|
||||
```
|
||||
CHAIN FIRE — 9,340/9,376 atoms read (36 refused) across 82/82 games [99.6%]
|
||||
|
||||
min p25 median p75 max mean
|
||||
chain_p 0.013 0.382 0.500 0.618 0.987 0.500
|
||||
counter_p 0.050 0.396 0.500 0.604 0.950 0.500
|
||||
divergence (signed) −0.556 −0.082 0.000 0.082 0.556 −0.000
|
||||
divergence (abs) 0.000 0.038 0.082 0.142 0.556 0.100
|
||||
|
||||
|divergence| < 0.02 1,203 (12.9%) agrees with the counter
|
||||
0.02 – 0.05 1,760 (18.8%)
|
||||
0.05 – 0.10 2,577 (27.6%)
|
||||
0.10 – 0.20 2,835 (30.4%)
|
||||
≥ 0.20 965 (10.3%) a different read entirely
|
||||
|
||||
OVER SIDE ONLY (n=4,673)
|
||||
divergence (signed) −0.526 −0.045 0.031 0.104 0.556 0.029
|
||||
chain HIGHER on 2,837 (60.7%)
|
||||
```
|
||||
|
||||
**Reading it honestly:**
|
||||
|
||||
- The chain **fires**, on 99.6% of the board. It is not the A5 case (built,
|
||||
correct, never invoked).
|
||||
- It is **not a relabelled counter**: 40.7% of rows differ by ≥0.10, and 10.3%
|
||||
by ≥0.20. It is also not noise — 12.9% agree inside 0.02.
|
||||
- The **symmetry of the two-sided signed distribution is arithmetic**, not a
|
||||
finding: both sides of every prop are in the sample, so each pair contributes
|
||||
`+d` and `−d`. Reading `mean −0.000` as "unbiased" would be reading the
|
||||
sampling scheme. The over-side slice is the one that can lean, and it does:
|
||||
**+2.9pp mean, higher on 60.7%**.
|
||||
- **Divergence is not merit.** A challenger that disagrees is interesting, not
|
||||
right. Which of the two is closer to what happened is the settle pass's
|
||||
question, and it is exactly why the triple is stored.
|
||||
- **The probe is contaminated by construction** and reports only a divergence
|
||||
(a property of two forecasts) rather than a resolution (a property of a
|
||||
forecast against an outcome): `statcast_aggregates` is upserted in place and
|
||||
keeps one as-of date, so profiles read for a row graded three days ago are
|
||||
today's. The forward accrual on the cron does not have this problem.
|
||||
- The probe resolves **no opposing pitcher** (offline), so it measures the
|
||||
batter-side read. The live shadow does resolve it.
|
||||
|
||||
---
|
||||
|
||||
## 5. What is NOT claimed
|
||||
|
||||
- **Nothing is promoted.** The proven set remains empty.
|
||||
- **The chain is not calibrated** and has passed no gate.
|
||||
- **No head-to-head has been run.** That needs settled outcomes against the
|
||||
stored triples and must go through `factorGate` / the cumulative Bonferroni
|
||||
denominator, as a NEW hypothesis.
|
||||
- **`CALIBRATION_DEPLOYED` stays `[]`.** Nothing is served calibrated.
|
||||
- **WNBA is untouched.** Its contested-possession chainFn and its possession feed
|
||||
are later orders. The core changes (signed correlation, redistribute on both
|
||||
readings) were built now because they are engine honesty, not because MLB needs
|
||||
them — MLB exercises neither.
|
||||
|
||||
---
|
||||
|
||||
## 6. The pre-registered next step
|
||||
|
||||
Once settled outcomes accrue against `chain_shadow`:
|
||||
|
||||
1. Score `chain_p` vs `counter_p` on the SAME rows (paired bootstrap — comparing
|
||||
independent SEs overstates uncertainty and has previously read a reliable
|
||||
−0.022 as noise).
|
||||
2. Per stat, never pooled (pooled resolution is inflated by base-rate structure).
|
||||
3. Through the cumulative test ledger; the CI widens to `1 − 0.05/tests`.
|
||||
4. **Pre-registered fallback, stated before the answer is known:** if the chain
|
||||
moves ~87% of the board and does not improve Brier, it is THEATER by the
|
||||
`factorGate` definition and the correct action is to leave it off — not to
|
||||
re-tune the opportunity term until it passes.
|
||||
|
||||
---
|
||||
|
||||
## 7. Incidental defect found and fixed
|
||||
|
||||
`runSnapshot`'s `deps` object was an **allowlist of 17 keys**, but fourteen call
|
||||
sites read `deps.challenger`, `deps.loadStatcast`, `deps.environmentContext`,
|
||||
`deps.lineupContext`, `deps.hitsFactorContext`, `deps.matchupKeys`,
|
||||
`deps.gameBinder`, `deps.archetypeAxes`, `deps.contactChallenger`,
|
||||
`deps.projectionChallenger`, `deps.loadArsenals`, `deps.mlbAdapter` — each
|
||||
documented as injectable, each **permanently `undefined`**, each always falling
|
||||
through to the real module. The seam existed in the comment and not in the code.
|
||||
|
||||
Fixed by spreading `...opts` FIRST in the literal: every explicit key is declared
|
||||
after and already reads `opts.X`, so no existing behaviour moves, while an
|
||||
unlisted dep now actually arrives. Found because the chain shadow's own test
|
||||
could not inject a statcast map — the test would have passed while measuring
|
||||
nothing, which is the failure this whole line of work exists to stop repeating.
|
||||
|
||||
---
|
||||
|
||||
# §8 — chain v2: FEED THE ENGINE (2026-08-12)
|
||||
|
||||
## 8.1 The order's premise, checked before building
|
||||
|
||||
The v2 order opened from four numbers. Each was checked against the code and the
|
||||
board before anything was written:
|
||||
|
||||
| claim | measured |
|
||||
|---|---|
|
||||
| "`paOutcome` returned null on 596/725 rows" | **False.** Fire rate is **9,752/9,792 (99.6%)**. The only refusal reason on the whole board is `no_batter_profile: 40`. `pa_outcome_refused` never fires. |
|
||||
| "it needs the batter-vs-pitcher-hand split" *to run* | **False.** `paOutcome` and `hitOnContact` read `k_pct`, `bb_pct`, `barrel_pct`, `hard_hit_pct`, `avg_exit_velo`, `avg_launch_angle` and the pitcher's `k_pct`/`bb_pct`/`hard_hit_pct`. Neither reads `bats` or `throws` at all. A hand split cannot change whether they run. |
|
||||
| "82% FELL BACK to the seasonal rate, i.e. became the counter" | **No fallback path exists.** `baseballChain.chainFn` returns null on any refusal and `chain.applyChainFn` DROPS the atom. Nothing in the module reads a season rate as a substitute. |
|
||||
| "425/425 served-identical" | That figure is from commit `7c8ef8b` (the A1–A7 deploy verification), not from the chain. v1 measured **2/2 served-identical** in the harness and byte-identical served payloads. |
|
||||
|
||||
Chain v1 was also never committed or deployed and migration 038 was never
|
||||
applied, so no chain-shadow rows exist in production — the premise numbers cannot
|
||||
have come from a chain-shadow run.
|
||||
|
||||
**The premise was wrong about the mechanism. It was right about the thing that
|
||||
matters:** the chain was reading a hitter's SEASON rates, which already average
|
||||
his platoon split over whichever hands he happened to face. That is a season read
|
||||
wearing a matchup read's clothes, and un-averaging it is real work. So v2 plumbs
|
||||
the hand split — as a **conditioner on the rate**, not as a fix to the fire rate.
|
||||
|
||||
## 8.2 What was built
|
||||
|
||||
**The split enters at the per-PA hit rate, not at the output probability.**
|
||||
Multiplying `P(hits ≥ 1)` by a platoon factor would scale a number that has
|
||||
already been through the opportunity term — a different and wrong claim. So
|
||||
`baseballChain.chainFn` now writes the chain out explicitly:
|
||||
|
||||
```
|
||||
paOutcome → p_hit_per_pa (season)
|
||||
→ × platoonRead multiplier ← THE MATCHUP CONDITIONER
|
||||
→ Binomial(PA, p_hit) over paDistribution
|
||||
→ P(hits ≥ k)
|
||||
```
|
||||
|
||||
A test asserts that with no split supplied this is **arithmetically identical to
|
||||
`projectSkill`** across three archetypes × three PA values, so the restructuring
|
||||
cannot quietly become a second model.
|
||||
|
||||
**As-of-correct (A4).** The hitter's own split comes from `hitsFactorContext`
|
||||
(whose reads are `lte('as_of_date', asOf)`), the opposing starter's hand through
|
||||
the A6 `matchupKeys` resolve. The probe bounds every read at the row's own
|
||||
`game_date`. Nothing later than the grade can enter.
|
||||
|
||||
**Refuse, never substitute.** `platoonSeverity` already refuses below 60 PA on
|
||||
the smaller side and declares switch hitters unreadable. On a refusal the season
|
||||
rate stands **exactly** untouched (asserted to 6 dp) and the reason is recorded.
|
||||
|
||||
**Refusals are counted by reason** (`context.onRefusal`), and the count is taken
|
||||
off the **stored blocks**, not the legs — a prop with both an over and an under
|
||||
row produces two legs sharing one block, and the first draft reported 48.1% and
|
||||
55.1% for the same fact over two different denominators.
|
||||
|
||||
## 8.3 Measured — real board, 2026-08-07 → 08-12, 9,792 graded hits rows
|
||||
|
||||
```
|
||||
CHAIN FIRE — 9,752/9,792 atoms read (40 refused) across 83/83 games
|
||||
refusal reasons: { no_batter_profile: 40 }
|
||||
|
||||
HAND SPLIT — fired on 288/553 unique props (52.1%)
|
||||
season-rate reasons:
|
||||
insufficient_split_sample 130 the honest 60-PA refusal
|
||||
no_pitcher_hand 71 the only FIXABLE gap
|
||||
switch_hitter_side_value_unknown 60 genuinely unreadable
|
||||
no_splits / missing_split 4
|
||||
|
||||
CHAIN vs COUNTER — n=9,752
|
||||
min p25 median p75 max mean
|
||||
chain_p 0.016 0.359 0.498 0.641 0.984 0.500
|
||||
counter_p 0.050 0.397 0.500 0.603 0.950 0.500
|
||||
|divergence| 0.000 0.044 0.095 0.164 0.509 0.113
|
||||
|
||||
|div| < 0.02 1,142 (11.7%) 0.10–0.20 3,129 (32.1%)
|
||||
0.02–0.05 1,608 (16.5%) >= 0.20 1,538 (15.8%)
|
||||
0.05–0.10 2,335 (23.9%)
|
||||
|
||||
OVER SIDE ONLY (n=4,879): mean +0.045, chain HIGHER on 66.3%
|
||||
|
||||
MATCHUP READ vs SEASON READ
|
||||
split FIRED (n=5,372 rows) median |div| 0.091 · >= 0.10 on 46.0% · within 0.02 on 13.0%
|
||||
split REFUSED (n=4,380 rows) median |div| 0.100 · >= 0.10 on 50.1% · within 0.02 on 10.2%
|
||||
```
|
||||
|
||||
## 8.4 Reading it honestly
|
||||
|
||||
- **The hand split fires on 52.1% of props**, up from 0. Of the 47.9% that do
|
||||
not, **190 of 265 (72%) are principled refusals** — a thin split or a switch
|
||||
hitter. Only `no_pitcher_hand` (71) is a plumbing gap, and it is the A5 shape
|
||||
exactly: the hitter's split is sitting right there and the pitcher hand is
|
||||
missing because that player had no lineup row.
|
||||
- **Feeding the split widened the divergence**: median |div| 0.082 → **0.095**,
|
||||
and the over-side lean +2.9pp → **+4.5pp**. The chain moved further from the
|
||||
counter, which is what a matchup conditioner should do and is *not* evidence
|
||||
it moved in the right direction.
|
||||
- **The rows where the split FIRED disagree with the counter slightly LESS**
|
||||
(median 0.091 vs 0.100) than the rows where it refused. **This comparison is
|
||||
confounded and must not be read as an effect**: a hitter with 60+ PA on both
|
||||
sides is an established regular, and the counter has more game log on him too.
|
||||
It is two different populations, not two treatments.
|
||||
- **Divergence is still not merit.** Nothing here says the chain is closer to
|
||||
what happened. That needs settled outcomes against the stored triple, and it
|
||||
goes through the gate as a new hypothesis against the cumulative denominator.
|
||||
- **The probe remains contaminated** for the batter profile
|
||||
(`statcast_aggregates` keeps one as-of date) and so reports a divergence, never
|
||||
a resolution. The forward cron accrual does not have this problem.
|
||||
|
||||
## 8.5 Still not claimed
|
||||
|
||||
Unchanged from §5: nothing promoted, no head-to-head run, `CALIBRATION_DEPLOYED`
|
||||
still `[]`, A8 untouched, `hitsFactors` untouched, the counter still serves,
|
||||
WNBA untouched.
|
||||
|
||||
## 8.6 Second incidental defect found and fixed
|
||||
|
||||
`hitsFactorContext.build` and `matchupKeys.build` were both gated on
|
||||
`require('../utils/supabase').getSupabaseServiceClient()` called inline, so the
|
||||
entire factor and hand-split path was **unreachable from any test**. "It is
|
||||
wired" could only ever have rested on reading the code — which is precisely how
|
||||
A5 shipped three factors that never fired. Both now read `deps.supabase ||` the
|
||||
real client (additive; `undefined` gives identical behaviour).
|
||||
|
||||
---
|
||||
|
||||
# §9 — chain v3: the opportunity term, and the artifact/signal diagnostic (2026-08-12)
|
||||
|
||||
## 9.1 The premise, again checked first — one half right, one half wrong
|
||||
|
||||
**RIGHT, and a real defect:** the v2 shadow passed **no `lineupSlotFor`**, so
|
||||
every hitter fell to `skillProjection.DEFAULT_PA = 4.1`. "A regular" was asserted
|
||||
about the leadoff man and the nine hole alike. The chain is `rate × opportunity`
|
||||
and the opportunity half was a constant across the entire lineup — a free,
|
||||
known, pre-game fact thrown away. Fixed.
|
||||
|
||||
**WRONG about the conversion, and about the direction:**
|
||||
|
||||
- *"mis-handles the multiple-chances structure"* — it does not.
|
||||
`paDistribution` is a mean-preserving two-point mixture, and
|
||||
`atLeast(Binomial(n, p), 1)` **is** `1 − (1−p)^n` averaged over n. The order's
|
||||
proposed formula is what the code already computes. The defect was the INPUT
|
||||
`E[PA]`, not the conversion.
|
||||
- *"5.7pts BELOW the counter on 92% of rows"* — **measured the other way.** On
|
||||
the over side the chain runs **+4.5pp ABOVE** the counter and is below on
|
||||
**33.6%**. It was never 92%-below, at any point, in any measurement here.
|
||||
|
||||
## 9.2 Phase 1 — the E[PA] fix, A/B on identical rows
|
||||
|
||||
`matchup_keys` now carries `batting_order` (one extra column on a read it
|
||||
already performs — no new query), and the shadow feeds it as the opportunity
|
||||
term. Absent ⇒ `DEFAULT_PA` and the block records `default_regular`, so a
|
||||
league-shaped opportunity term is never mistaken for a posted one.
|
||||
|
||||
```
|
||||
OPPORTUNITY — posted lineup slot on 491/553 props (88.8%)
|
||||
fell to a default regular: 62
|
||||
|
||||
n mean median chain BELOW counter
|
||||
REAL E[PA] (posted slot) 4,879 +0.0450 +0.0499 33.6%
|
||||
CONSTANT 4.1 (the v2 shadow) 4,879 +0.0322 +0.0363 39.6%
|
||||
```
|
||||
|
||||
**The uniform bias did not collapse, because it was never there to collapse.**
|
||||
The fix moved the chain *further above* the counter, not toward it — which is
|
||||
correct behaviour, not a regression: real slots raise E[PA] for the top of the
|
||||
order and lower it for the bottom, and top-of-order hitters are over-represented
|
||||
in the prop board.
|
||||
|
||||
## 9.3 Phase 2 — THE DIAGNOSTIC: artifact or signal?
|
||||
|
||||
Over side, n=3,705 rows carrying an opposing-pitcher profile, quintiles of
|
||||
opposing-pitcher K% (low = soft matchup):
|
||||
|
||||
| bucket | opp K% | n | mean divergence | chain below counter |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 10.8–18.4 | 741 | **+0.0749** | 27.8% |
|
||||
| Q2 | 18.4–20.3 | 741 | +0.0560 | 30.8% |
|
||||
| Q3 | 20.3–23.1 | 741 | +0.0287 | 34.1% |
|
||||
| Q4 | 23.1–26.6 | 741 | +0.0396 | 35.0% |
|
||||
| Q5 | 26.6–40.6 | 741 | **+0.0256** | 38.6% |
|
||||
|
||||
**Q1 − Q5 spread +0.0494 · Pearson r(oppK, divergence) = −0.119**
|
||||
|
||||
### The answer is BOTH, and the order's binary framing does not fit
|
||||
|
||||
The order asked for *uniform (broken)* **or** *difficulty-correlated (signal)*.
|
||||
The data is a **mixture of the two, and both components should be named**:
|
||||
|
||||
- **A difficulty-correlated component, ~+0.049 across the range.** The chain
|
||||
reads soft matchups higher and hard matchups lower than the counter does. The
|
||||
`below %` is **monotone across all five quintiles** (27.8 → 30.8 → 34.1 → 35.0
|
||||
→ 38.6), which is a cleaner signature than the means (Q4 breaks order).
|
||||
- **A uniform positive offset of ~+0.026.** Even in the HARDEST quintile the
|
||||
chain sits +2.6pp above the counter. That floor does not move with the matchup
|
||||
and is therefore not conditioning — it is exactly the shape a residual
|
||||
mechanical bias makes. It is smaller than the conditioning component but it
|
||||
has not been explained, and calling the whole result "signal" would bury it.
|
||||
|
||||
### The caveat that must travel with the correlation
|
||||
|
||||
**This is close to mechanically guaranteed and is NOT evidence of correctness.**
|
||||
The chain reads the opposing pitcher's K rate directly (log5 odds-ratio in
|
||||
`paOutcome`) and the counter reads nothing about the pitcher at all. So a
|
||||
monotone relationship between opposing-pitcher K% and chain-minus-counter is
|
||||
approximately a proof that *the wiring works* — that the pitcher input reaches
|
||||
the number. It says nothing about whether the adjustment is the right size, the
|
||||
right direction on any individual row, or better than ignoring the pitcher.
|
||||
|
||||
**Does the chain see the game, or just miscompute it?** It demonstrably
|
||||
CONDITIONS on the game — the pitcher input reaches the forecast and moves it in
|
||||
the theorised direction. Whether that conditioning is *right* is unanswerable
|
||||
from a divergence and needs settled outcomes against the stored triple. There is
|
||||
also a residual ~2.6pp offset that conditioning does not explain and that should
|
||||
be chased before anyone reads the correlation as a win.
|
||||
|
||||
## 9.4 Open item created by this measurement
|
||||
|
||||
The ~+0.026 floor. Candidates not yet tested: the chain is unclamped where the
|
||||
counter clamps to `[0.10, 0.95]` (chain min 0.016 vs counter min 0.050); the
|
||||
`LEAGUE.babip = 0.291` anchor in `hitOnContact`; the ±35% BABIP bound. Naming it
|
||||
as unexplained is the honest state — it is not yet an artifact and not yet
|
||||
signal.
|
||||
|
||||
## 9.5 Unchanged
|
||||
|
||||
Nothing promoted, no head-to-head, `CALIBRATION_DEPLOYED` still `[]`, A8 and
|
||||
`hitsFactors` untouched, the counter still serves, WNBA untouched. Migration 038
|
||||
is still an unapplied blocking precondition of deploy.
|
||||
@@ -0,0 +1,208 @@
|
||||
# WNBA v1 — the possession / usage feed
|
||||
|
||||
**Status:** DATA INFRA ONLY. No chainFn, no archetype wiring, no shadow, nothing served.
|
||||
**Date:** 2026-08-13
|
||||
**Consumer it exists for:** the chain's basketball `chainFn`
|
||||
(`usage × possessions × efficiency`), which cannot be written without it.
|
||||
|
||||
---
|
||||
|
||||
## 1. Phase 1 — the source, established before any plumbing
|
||||
|
||||
### Candidates and the choice
|
||||
|
||||
| candidate | verdict |
|
||||
|---|---|
|
||||
| Python `nba_api` service (WNBA endpoints) | **Rejected as a dependency.** The service is offline in production; a feed we cannot fetch is not a feed. |
|
||||
| ESPN site API (`site.api.espn.com/.../basketball/wnba`) | **Chosen.** Free, no auth, already the host for WNBA schedules, box scores and live tracking. |
|
||||
|
||||
### What it actually returns — verified live, 2026-08-13
|
||||
|
||||
`summary?event={id}` → `boxscore.players[].statistics[0]`:
|
||||
|
||||
```
|
||||
keys: minutes, points, fieldGoalsMade-fieldGoalsAttempted,
|
||||
threePointFieldGoalsMade-threePointFieldGoalsAttempted,
|
||||
freeThrowsMade-freeThrowsAttempted, rebounds, assists, turnovers,
|
||||
steals, blocks, offensiveRebounds, defensiveRebounds, fouls, plusMinus
|
||||
|
||||
Breanna Stewart: ['30','19','8-21','1-5','2-2','7','5','0','1','0','2','5','2','-2']
|
||||
```
|
||||
|
||||
plus `starter` / `didNotPlay` / `ejected` per athlete, and `boxscore.teams[]` with
|
||||
team totals (`fieldGoalsMade-fieldGoalsAttempted`, `freeThrowsMade-freeThrowsAttempted`,
|
||||
`totalTurnovers`, `offensiveRebounds`). Team minutes sum to 200 in regulation
|
||||
(5 × 40), so overtime enters through the same sum. Play-by-play is also present
|
||||
(378 plays on the sampled game) but is not needed — the box identities suffice.
|
||||
|
||||
### chainFn inputs: delivered, derived, or missing
|
||||
|
||||
| chainFn input | status | how |
|
||||
|---|---|---|
|
||||
| **minutes** | **SERVED** | `minutes` per player |
|
||||
| **usage rate** | **DERIVED (exact)** | `100·((FGA+0.44·FTA+TOV)·(TmMIN/5)) / (MIN·(TmFGA+0.44·TmFTA+TmTOV))` |
|
||||
| **possessions** | **DERIVED (exact)** | `FGA − OREB + TOV + 0.44·FTA` |
|
||||
| **pace** | **DERIVED (exact)** | `possessions × 40 / (TmMIN/5)` |
|
||||
| **efficiency** | **DERIVED (exact)** | TS% `PTS/(2(FGA+0.44·FTA))`, eFG% `(FGM+0.5·FG3M)/FGA` |
|
||||
| **game state** (redistribute hook) | **SERVED** | `starter`, team/opp score → `final_margin` |
|
||||
| shot location / zone | **MISSING** | not in the box score; not a chainFn input today |
|
||||
| on/off, lineup combinations | **MISSING** | would need play-by-play reconstruction |
|
||||
| opponent defensive rating | **MISSING** | derivable later from the same rows (each game is also the opponent's) |
|
||||
|
||||
**"Derived" is not "proxied", and the distinction is load-bearing.** ESPN does not
|
||||
serve a usage rate, a possession count or a pace figure — it serves their
|
||||
components, and these are the standard identities recomputed from counted events.
|
||||
The only estimated term anywhere is the **0.44 free-throw-trip coefficient**,
|
||||
which is the field-standard value.
|
||||
|
||||
### Point-in-time capability
|
||||
|
||||
**Native. No `statcast_history`-style split is needed, and this is a design
|
||||
choice made at line one rather than a retrofit.**
|
||||
|
||||
`statcast_aggregates` had to grow a history twin because it stores a season
|
||||
aggregate upserted in place, destroying every prior version — which is why the
|
||||
first skill backtest was honest only by accident. This stores **per-game rows**.
|
||||
A completed box score never changes, so an as-of profile is
|
||||
`WHERE game_date < asOf`: a filter over immutable facts. Nothing is overwritten,
|
||||
so there is nothing to retain a history *of*.
|
||||
|
||||
**Strictly `<`, never `<=`.** A game on the as-of date may have tipped after
|
||||
grade time; counting it leaks the evening being predicted into the prediction.
|
||||
Same rule as `snapshotSettlementService.isPreGame`.
|
||||
|
||||
---
|
||||
|
||||
## 2. Phase 2 — the feed
|
||||
|
||||
**`wnba_player_game`** (migration 039). Key `(game_id, source_id)`, registered in
|
||||
`src/utils/tableKeys.js`, so `safePaginate` can walk it. Indexed
|
||||
`(sport, player_key, game_date DESC)` — every read is "this player, before this
|
||||
date".
|
||||
|
||||
Scoped to the chainFn's inputs plus the redistribute game-state. **Not a general
|
||||
WNBA stats dump** — `closing_captures` grew to 4.2M rows of something nothing
|
||||
read, and the lesson is to ingest for a named consumer or not at all.
|
||||
|
||||
Every derived rate is stored **alongside its components**, so it is re-derivable
|
||||
and checkable rather than an unfalsifiable number — the `factorFreeze` rule
|
||||
applied to a feed.
|
||||
|
||||
### Coverage — full live season, measured
|
||||
|
||||
```
|
||||
rows 5,085 games 256 players 241
|
||||
teams 15 days 88 errors 0
|
||||
date range 2026-05-01 .. 2026-08-12
|
||||
|
||||
FIELD COVERAGE
|
||||
minutes 5085/5085 starter 5085/5085
|
||||
usage_rate 5072/5085 final_margin 5085/5085
|
||||
team_possessions 5085/5085 ts_pct 4821/5085
|
||||
team_pace 5085/5085 efg_pct 4772/5085
|
||||
```
|
||||
|
||||
The `ts_pct` / `efg_pct` gaps are **correct refusals, not missing data**: a
|
||||
player with zero field-goal and free-throw attempts has no true-shooting
|
||||
percentage, and returning 0 would assert he shot and missed.
|
||||
|
||||
### Two contaminating games found and excluded
|
||||
|
||||
The season pull surfaced two games that are **not league games**:
|
||||
|
||||
- `2026-05-02` — an exhibition against **Nigeria** (`NIGER` vs `IND`)
|
||||
- `2026-07-25` — the **All-Star game** (`SPO` vs `COOP`)
|
||||
|
||||
Their usage and pace context is meaningless for a forward projection (an
|
||||
All-Star game has no defence to speak of), and leaving them in would pollute
|
||||
every profile spanning those dates.
|
||||
|
||||
The filter reads **ESPN's own `/teams` endpoint** rather than a hardcoded
|
||||
fifteen — an expansion franchise is admitted the day the league adds it, and
|
||||
this league has expanded twice in three years. An **empty** teams response is
|
||||
treated as unknown membership and filters nothing, because a failed feed
|
||||
degrading to an empty index looks exactly like an honest absence (the
|
||||
`fielding_oaa` lesson).
|
||||
|
||||
---
|
||||
|
||||
## 3. Phase 3 — verification, on real data through the real functions
|
||||
|
||||
`scripts/wnba-feed-verify.js` drives the production parsers over the live season
|
||||
and then calls `wnbaUsageService.profileAsOf` against those rows. It exercises
|
||||
the shipped code, not a restatement of it.
|
||||
|
||||
### The derivation, checkable by hand
|
||||
|
||||
```
|
||||
Laura Juskaite (TOR) 2026-05-01
|
||||
MIN 28 FGA 10 FTA 2 TOV 3 PTS 6
|
||||
team: MIN 200 FGA 68 FTA 16 TOV 11 OREB 7
|
||||
|
||||
possessions stored 79.04 recomputed 79.040
|
||||
usage_rate stored 23.046 recomputed 23.046
|
||||
```
|
||||
|
||||
### The as-of read
|
||||
|
||||
```
|
||||
player: Natasha Howard — 36 games on record
|
||||
|
||||
asOf 2026-05-12 REFUSED (2 games precede — under the 3-game floor)
|
||||
asOf 2026-06-24 games=18 (expected 18) usage 24.826 min/g 28.11 pace 82.683 TS 0.5998
|
||||
asOf 2026-08-12 games=35 (expected 35) usage 22.057 min/g 28.51 pace 82.275 TS 0.6010
|
||||
asOf 2026-12-31 games=36 (expected 36) usage 22.027 min/g 28.50 pace 82.241 TS 0.6042
|
||||
|
||||
STRICT CUTOFF — a game ON the as-of date is EXCLUDED (correct)
|
||||
THIN HISTORY — asOf before his first game returns null (refused, correct)
|
||||
```
|
||||
|
||||
The profile **moves with the date** (24.8 usage in June, 22.1 by August), which
|
||||
is the property that makes it a point-in-time read rather than a season number
|
||||
wearing a date.
|
||||
|
||||
### Two refusals that are features
|
||||
|
||||
- **Thin history ⇒ `null`, never a league-average player.** Usage feeds the
|
||||
chain's opportunity term directly; a stand-in would assert a usage rate about
|
||||
someone never observed. Same discipline as `platoonSeverity`'s 60-PA floor.
|
||||
- **Usage is MINUTES-WEIGHTED, not a flat mean across games.** The lineup-K-rate
|
||||
lesson: an unweighted aggregate counts a 6-minute cameo like a 34-minute start,
|
||||
and unweighted actively *hurt* that model.
|
||||
|
||||
---
|
||||
|
||||
## 4. Not done, and not claimed
|
||||
|
||||
- **No chainFn.** Basketball's `usage × possessions × efficiency` is the next
|
||||
order. This order stops at the feed.
|
||||
- **No archetype wiring, no shadow, no serving.** Nothing reads this table yet.
|
||||
- **MLB untouched.** A test asserts `chainShadow.js`, `baseballChain.js` and
|
||||
`analyzeViaEngine1.js` contain no reference to the WNBA feed.
|
||||
- `CALIBRATION_DEPLOYED` still `[]`; A8 accrual, the chain's MLB shadow and the
|
||||
served grade are untouched.
|
||||
|
||||
---
|
||||
|
||||
## 5. ⛔ Blocked: the table is NOT created and NOT ingested
|
||||
|
||||
`SUPABASE_DB_PASSWORD` in the local `.env` **fails authentication** against the
|
||||
pooler (`aws-1-us-east-1...` → `password authentication failed for user
|
||||
"postgres"`), and the direct host resolves IPv6-only, which is unreachable from
|
||||
this WSL2 environment. The transposed project ref already documented in
|
||||
CLAUDE.md is the likely cause.
|
||||
|
||||
So:
|
||||
|
||||
- **migration 039 is written and unapplied**
|
||||
- **`ingestRange` is written and has never written a row**
|
||||
- Everything in §2's coverage and §3's verification was measured by running the
|
||||
real parsers and the real as-of read over the live season **in memory**. The
|
||||
parse, the derivations and the as-of logic are verified on real data; the
|
||||
**database round-trip is not**.
|
||||
|
||||
That distinction is stated rather than glossed: a feed that has never been
|
||||
written is not a feed yet.
|
||||
|
||||
**Also still outstanding: migration 038 (`model_snapshots.chain_shadow`)** from
|
||||
the chain orders, for the same reason.
|
||||
@@ -0,0 +1,294 @@
|
||||
# WNBA source survey — possession / play-by-play options
|
||||
|
||||
**Status:** READ-ONLY SURVEY. Nothing built, nothing decided, nothing ingested.
|
||||
**Date:** 2026-08-13
|
||||
**Question:** what is the best source to build the basketball chainFn's possession
|
||||
feed — including the feedback layer — on?
|
||||
|
||||
---
|
||||
|
||||
## 0. Correction to the premise, and the real gap
|
||||
|
||||
The v1 report did **not** conclude "box scores only, no feedback loop". It
|
||||
recorded that ESPN play-by-play *is* present ("378 plays on the sampled game")
|
||||
and that it was not needed for the box identities, and it listed shot location
|
||||
and on/off as MISSING **from the box endpoint**.
|
||||
|
||||
The real gap the order names is correct though: **one source was surveyed.** And
|
||||
that produced a false ceiling — this survey found event-level data, shot
|
||||
coordinates, and pre-parsed per-player possessions across four other sources.
|
||||
|
||||
**The single most important finding is an environment artifact, not a data fact:**
|
||||
|
||||
> `stats.wnba.com` fails over IPv6 and **works over IPv4**. Default curl (and
|
||||
> Node's default resolver order) picks the AAAA record, connects, and hangs. The
|
||||
> same IPv6 pathology that made `db.<ref>.supabase.co` unreachable in WSL2.
|
||||
> Anything in this environment that reports a stats.nba/stats.wnba endpoint as
|
||||
> "unreachable" should be re-tested with `-4` before it is believed.
|
||||
|
||||
---
|
||||
|
||||
## 1. `stats.wnba.com` / `playbyplayv3` — what `nba_api` wraps
|
||||
|
||||
**Reachable: YES, over IPv4 only. Sustained ingest: NO.**
|
||||
|
||||
```
|
||||
curl -4 ... "https://stats.wnba.com/stats/playbyplayv3?GameID=1022600001&StartPeriod=1&EndPeriod=4"
|
||||
→ HTTP 200, 207,769 bytes, 3.78s
|
||||
|
||||
game.actions: 469 events
|
||||
fields: actionId, actionNumber, actionType, clock, description, isFieldGoal,
|
||||
location, period, personId, playerName, playerNameI, pointsTotal,
|
||||
scoreAway, scoreHome, shotDistance, shotResult, shotValue, subType,
|
||||
teamId, teamTricode, xLegacy, yLegacy
|
||||
|
||||
1 PT10M00.00S Jump Ball J. Jones Jump Ball Jones vs. Griner
|
||||
1 PT09M44.00S Missed Shot A. Morrow MISS Morrow 27' 3PT Jump Shot
|
||||
1 PT09M39.00S Rebound M. Johannes Johannes REBOUND (Off:0 Def:1)
|
||||
```
|
||||
|
||||
**Resolution:** full event stream with **shot coordinates** (`xLegacy`/`yLegacy`),
|
||||
`shotDistance`, `shotValue`, `shotResult`, `personId`, running score. Richer raw
|
||||
detail than anything else surveyed.
|
||||
|
||||
**Rate limits — the disqualifier.** One call succeeded. Every subsequent call
|
||||
returned **HTTP 000 (connection accepted, then stalled)**: immediately after,
|
||||
after 5s, and again after a ~3 minute cool-down at a 50s timeout. Two sibling
|
||||
endpoints (`boxscoreadvancedv3`, `boxscoreplayertrackv3`) returned 000 on first
|
||||
attempt. This is the well-known stats.nba.com throttle posture.
|
||||
|
||||
**`nba_api` is not installed locally** (`ModuleNotFoundError`); it is pinned in
|
||||
`src/services/python/requirements.txt:7` — the offline Python service.
|
||||
|
||||
**Point-in-time:** per-game, immutable ⇒ native, same as every source here.
|
||||
**Verdict:** richest raw feed, **hostile to sustained ingest**. Not a dependency.
|
||||
|
||||
---
|
||||
|
||||
## 2. `wehoop` / sportsdataverse — the bulk archive
|
||||
|
||||
**Reachable: YES. Best cost profile of any option.**
|
||||
|
||||
```
|
||||
GET github.com/sportsdataverse/wehoop-wnba-data/raw/main/wnba/pbp/parquet/play_by_play_2026.parquet
|
||||
→ 3,090,207 bytes (the WHOLE 2026 season, one file)
|
||||
|
||||
archive: play_by_play_2021 (2.63 MB) … 2024 (3.39) … 2025 (4.06) … 2026 (3.09 MB)
|
||||
|
||||
columns include: athlete_id_1/2/3, athlete_name_1/2/3, coordinate_x, coordinate_y,
|
||||
coordinate_x_raw, coordinate_y_raw, clock_display_value, clock_minutes,
|
||||
clock_seconds, period_number, home_score, away_score, score_value,
|
||||
shooting_play, team_id, type_id, type_text, wallclock, home_team_spread, …
|
||||
```
|
||||
|
||||
**What it wraps:** ESPN's WNBA feeds, pre-collected and normalised.
|
||||
**Coverage:** 2021 → live 2026, updated through the season.
|
||||
**License:** repo reports `NOASSERTION` (sportsdataverse projects are generally
|
||||
MIT/CC-BY; **the license needs confirming before redistribution** — it does not
|
||||
block internal analytical use, but it is not a clean SPDX tag).
|
||||
**Access:** one HTTP GET per season. No rate limit, GitHub CDN.
|
||||
**Footprint:** ~3 MB/season compressed — the cheapest full-PBP option by an order
|
||||
of magnitude, and materially relevant with the DB at 406/500 MB.
|
||||
|
||||
**Verdict:** the **backfill and cross-check** source. Cannot serve tonight's game
|
||||
(it lags the live feed), so it is not the freshness layer.
|
||||
|
||||
---
|
||||
|
||||
## 3. PBPStats (`api.pbpstats.com`) — possessions already parsed
|
||||
|
||||
**Reachable: YES. Richest *derived* layer. This is the standout.**
|
||||
|
||||
```
|
||||
GET api.pbpstats.com/get-games/wnba?Season=2026&SeasonType=Regular Season
|
||||
→ 250 games, 2026-05-08 .. 2026-08-12, HomePossessions/AwayPossessions on 250/250
|
||||
|
||||
{"GameId":"1022600001","Date":"2026-05-08","HomeTeamAbbreviation":"NYL",
|
||||
"HomePoints":106,"AwayPoints":75,"HomePossessions":87,"AwayPossessions":88}
|
||||
|
||||
GET api.pbpstats.com/get-game-stats?Type=Player&GameId=1022600001&League=wnba
|
||||
→ 47 fields PER PLAYER PER GAME:
|
||||
|
||||
Usage, OffPoss, DefPoss, Minutes, Points, TsPct, EfgPct, ShotQualityAvg,
|
||||
SecondChanceOffPoss, PenaltyOffPoss, PenaltyOffPossPct, PenaltyDefPoss,
|
||||
Arc3FGA, Arc3Frequency, AtRimFG3AFrequency, Avg2ptShotDistance,
|
||||
Avg3ptShotDistance, LongMidRangeAccuracy/FGA/FGM/Frequency,
|
||||
ShortMidRangeFGA/Frequency, Blocked2s, FoulsDrawn, ShootingFouls, …
|
||||
```
|
||||
|
||||
**This delivers the entire chainFn input set pre-computed**, and `OffPoss` is
|
||||
**better than the derived team possessions in v1**: it is the possessions the
|
||||
player was actually on the floor for — the true usage denominator, not a team
|
||||
estimate apportioned by minutes.
|
||||
|
||||
**WOWY (with-or-without-you) — the feedback layer, and it works:**
|
||||
|
||||
```
|
||||
GET api.pbpstats.com/get-wowy-combination-stats/wnba?Season=2026&SeasonType=Regular Season
|
||||
&TeamId=1611661313&PlayerIds=1629568
|
||||
→ HTTP 200
|
||||
{"OffRtg":113.41,"DefRtg":109.77,"NetRtg":3.65,"Minutes":1370.0,
|
||||
"On":"","Off":"Kennedy Burke", …}
|
||||
```
|
||||
|
||||
On/Off splits by player combination — the direct measurement of what changes when
|
||||
a player is off the floor.
|
||||
|
||||
**Rate posture:** 6 rapid sequential calls → **200, 200, 200, 200, 200, 200.** No
|
||||
throttling observed. (Endpoints are undocumented and unversioned; `get-possessions`
|
||||
returned 500 with valid params and `get-lineup-stats` 404 — the surface is
|
||||
uneven, and a 422 helpfully names missing params.)
|
||||
|
||||
**Point-in-time:** per-game rows ⇒ native.
|
||||
**Footprint:** ~5,000 player-game rows/season × 47 fields ≈ single-digit MB.
|
||||
**Verdict:** **the primary source.**
|
||||
|
||||
---
|
||||
|
||||
## 4. Basketball-Reference `/wnba/` — scrapeable, and permitted
|
||||
|
||||
**Reachable: YES. `/wnba/` is NOT disallowed.**
|
||||
|
||||
```
|
||||
robots.txt User-agent: * Crawl-delay: 3
|
||||
Disallow: /basketball/ (…team paths…) ← /wnba/ is absent
|
||||
|
||||
GET /wnba/boxscores/202605080NYL.html → 200 426,971 b
|
||||
GET /wnba/boxscores/pbp/202605080NYL.html → 200 245,516 b
|
||||
GET /wnba/boxscores/shot-chart/202605080NYL.html → 200 174,986 b
|
||||
GET /wnba/boxscores/plus-minus/202605080NYL.html → linked from the boxscore
|
||||
|
||||
3 sequential pbp fetches at the stated 3s crawl-delay → 200, 200, 200
|
||||
```
|
||||
|
||||
**Note:** the earlier 404s in this survey were a wrong URL guess of mine
|
||||
(`…0PHO`), not a BBR limitation. The real ids come off
|
||||
`/wnba/years/2026_games.html`.
|
||||
|
||||
**Coverage:** play-by-play, shot charts AND plus-minus per game.
|
||||
**Cost:** HTML scrape + parse, 3s crawl-delay ⇒ ~13 min for a 250-game season.
|
||||
**ToS:** `Disallow` does not cover `/wnba/`; `Crawl-delay: 3` must be honoured.
|
||||
`GPTBot` is banned outright, so identify honestly and stay slow.
|
||||
**Verdict:** best **independent audit cross-check** (a second opinion on
|
||||
possessions from a different parser). Too slow and too brittle for primary.
|
||||
|
||||
---
|
||||
|
||||
## 5. ESPN WNBA (already wired) — the freshness layer
|
||||
|
||||
**Reachable: YES, already in production use.**
|
||||
|
||||
```
|
||||
summary?event=401857134 → plays: 378
|
||||
fields: awayScore, clock, coordinate, homeScore, id, participants, period,
|
||||
pointsAttempted, scoreValue, scoringPlay, sequenceNumber, shootingPlay,
|
||||
shortDescription, team, text, type, wallclock
|
||||
|
||||
1 9:45 Pullup Jump Shot athlete 4433403 coord {x:16,y:25} 3-0
|
||||
1 9:26 Fade Away Jump Shot athlete 2998928 coord {x:31,y:1} 3-2
|
||||
1 9:17 Driving Floating Jump Shot athlete 4433403 coord {x:30,y:3} 3-2
|
||||
```
|
||||
|
||||
Event-level with coordinates and athlete ids — so **shot location was never
|
||||
actually missing**; it was missing from the *box* endpoint I read in v1.
|
||||
|
||||
**Stability:** undocumented but long-lived, no auth, no observed throttle, and
|
||||
already the host for schedules/box/live-tracking in this codebase.
|
||||
**Verdict:** the **live/tonight** layer.
|
||||
|
||||
---
|
||||
|
||||
## 6. News / context layer — game state, actives, rest
|
||||
|
||||
**No scraper needed. ESPN already serves it.**
|
||||
|
||||
```
|
||||
GET .../basketball/wnba/injuries → HTTP 200, 14 teams, 46 entries
|
||||
Atlanta Dream | Brionna Jones | Out | Leg
|
||||
Chicago Sky | Maddy Westbeld | Out | Coach's Decision
|
||||
|
||||
summary?event=… also carries per-game injuries for both sides:
|
||||
Aliyah Boston Day-To-Day
|
||||
Caitlin Clark Day-To-Day
|
||||
Damiris Dantas Out
|
||||
```
|
||||
|
||||
**The one real limitation: this is a LATEST-ONLY snapshot.** There is no as-of
|
||||
query for injury state, so point-in-time actives require **capturing it daily** —
|
||||
which is exactly what the existing changedetection/cron stack is for. That is a
|
||||
small dated table, not a scraper build.
|
||||
|
||||
Miniflux/SearxNG/RSS would add *narrative* (beat-reporter rest news ahead of the
|
||||
official designation). Useful later; **not required** for the feedback layer,
|
||||
because ESPN injuries + `starter` + `Minutes` already answer "who played, who
|
||||
sat, who started".
|
||||
|
||||
---
|
||||
|
||||
## 7. Scorecard
|
||||
|
||||
| | possession-level for feedback? | as-of? | cost / footprint | reliability & ToS | live freshness |
|
||||
|---|---|---|---|---|---|
|
||||
| **PBPStats** | **YES — per-player `OffPoss`/`DefPoss`/`Usage` + WOWY on/off** | native (per-game) | ~single-digit MB/season | 6/6 rapid 200s; undocumented, uneven surface | good (through 08-12) |
|
||||
| **wehoop** | YES — full event stream | native | **~3 MB/season** (best) | GitHub CDN; license `NOASSERTION` — confirm | lags live |
|
||||
| **ESPN** | YES — 378 events w/ coords | native | small | already in prod, no throttle seen | **best (live)** |
|
||||
| **BBR** | YES — pbp + shot chart + plus-minus | native | scrape, 3s delay ⇒ ~13 min/season | `/wnba/` allowed, crawl-delay 3 | good |
|
||||
| **stats.wnba.com** | YES — richest raw (shot coords, 469 events) | native | small | **HOSTILE — 1 call then HTTP 000, still blocked after 3 min; IPv4-only** | good if you could call it |
|
||||
|
||||
---
|
||||
|
||||
## 8. Ranked recommendation
|
||||
|
||||
**A combination, not one source. Three roles:**
|
||||
|
||||
1. **PBPStats — PRIMARY.** It has already done the possession parsing, and
|
||||
per-player `OffPoss` is the correct usage denominator rather than my v1
|
||||
team-level estimate apportioned by minutes. `Usage`, `TsPct`, `ShotQualityAvg`
|
||||
and the shot-zone splits arrive free. **WOWY gives the feedback layer
|
||||
directly.** Tradeoff: undocumented, unversioned, single maintainer, uneven
|
||||
endpoint surface (a 500 and a 404 in this survey) — so it needs a fallback,
|
||||
and its numbers should be cross-checked once against a second parser.
|
||||
|
||||
2. **ESPN — LIVE + CONTEXT.** Tonight's game before PBPStats has it, plus
|
||||
injuries/actives. Already wired, already trusted in this codebase.
|
||||
|
||||
3. **wehoop — BACKFILL + CROSS-CHECK.** One 3 MB GET replaces 250 API calls for
|
||||
a historical season, and being an independent collection of the same ESPN
|
||||
feed it is a genuine second opinion. Confirm the license before anything
|
||||
leaves the building.
|
||||
|
||||
**Not recommended as a dependency:** `stats.wnba.com`. Richest raw data,
|
||||
unusable throttle. Worth keeping as a *manual* one-off tool now that the IPv4
|
||||
workaround is known.
|
||||
|
||||
**BBR:** hold as an audit path. Its independent possession parse is the best
|
||||
available check on PBPStats, at 3s/request.
|
||||
|
||||
### Can the FEEDBACK LAYER be built now?
|
||||
|
||||
**Yes — with PBPStats, and without deferring.** Two mechanisms are already
|
||||
available:
|
||||
|
||||
- **Usage redistribution:** per-player `Usage` + `OffPoss` per game, joined to
|
||||
who was out that night (ESPN injuries, captured daily). "How does this
|
||||
player's usage move in games where the primary creator sat" is then a direct
|
||||
measurement, not a model.
|
||||
- **Blowout → minutes:** `final_margin` (already in v1's feed) against
|
||||
`Minutes` and `OffPoss` per game.
|
||||
|
||||
**Ingest cost:** 1 call per game (~250/season) + 1 injuries call/day. Storage
|
||||
~single-digit MB — against 94 MB of current headroom.
|
||||
|
||||
**What is genuinely still missing:** *within-game* possession-by-possession
|
||||
lineup state (who was on the floor at each moment). `get-lineup-stats` 404s.
|
||||
Deriving it needs substitution events reconstructed from the raw PBP — available
|
||||
in wehoop/ESPN/BBR, but a real build. Not required for the two mechanisms above.
|
||||
|
||||
---
|
||||
|
||||
## 9. Nothing built
|
||||
|
||||
No ingest, no schema, no chainFn. The v1 feed (`wnba_player_game`, migration 039)
|
||||
remains written and unapplied; this survey is evidence that **its source choice
|
||||
should be revisited before it is applied** — PBPStats supersedes the derived-usage
|
||||
approach with a measured one.
|
||||
@@ -0,0 +1,377 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* espnWnbaAdapter — the WNBA possession/usage tap.
|
||||
*
|
||||
* Spec: specs/wnba-possession-feed.md
|
||||
*
|
||||
* ── WHY THIS SOURCE ──────────────────────────────────────────────────────
|
||||
* The chain's basketball `chainFn` is `usage × possessions × efficiency`, and
|
||||
* none of those three existed for WNBA anywhere in this repo. The candidates
|
||||
* were the Python `nba_api` service (WNBA endpoints) and ESPN's free site API.
|
||||
* The Python service is offline in production, which makes it a source we
|
||||
* cannot depend on — so this is ESPN, the same host already used for WNBA
|
||||
* schedules, box scores and live tracking.
|
||||
*
|
||||
* ── WHAT THE SOURCE ACTUALLY RETURNS (verified live, 2026-08-13) ─────────
|
||||
* `summary?event={id}` → `boxscore.players[].statistics[0]` with
|
||||
*
|
||||
* keys: minutes, points, fieldGoalsMade-fieldGoalsAttempted,
|
||||
* threePointFieldGoalsMade-threePointFieldGoalsAttempted,
|
||||
* freeThrowsMade-freeThrowsAttempted, rebounds, assists, turnovers,
|
||||
* steals, blocks, offensiveRebounds, defensiveRebounds, fouls, plusMinus
|
||||
*
|
||||
* plus `starter` / `didNotPlay` / `ejected` per athlete, and `boxscore.teams[]`
|
||||
* carrying the team totals (FGA, FTA, totalTurnovers, offensiveRebounds).
|
||||
*
|
||||
* ── THE RATES ARE DERIVED, NOT SERVED — AND THAT DISTINCTION IS REAL ─────
|
||||
* ESPN does NOT return a usage rate, a possession count or a pace figure. It
|
||||
* returns their COMPONENTS. Usage and possessions are then exact arithmetic
|
||||
* (the standard box-score identities below), not proxies — every term is a
|
||||
* counted event, and the only approximation is the 0.44 free-throw-trip
|
||||
* coefficient, which is the same constant every public implementation uses.
|
||||
*
|
||||
* That is worth stating precisely because "derived" and "proxied" are different
|
||||
* claims: a proxy stands in for a quantity we cannot see, and these are the
|
||||
* quantity itself, recomputed.
|
||||
*
|
||||
* ── POINT-IN-TIME IS STRUCTURAL HERE, NOT RETROFITTED ────────────────────
|
||||
* `statcast_aggregates` had to grow a `statcast_history` twin because it stores
|
||||
* a SEASON AGGREGATE upserted in place, destroying every prior version. This
|
||||
* stores PER GAME rows instead. A completed box score never changes, so any
|
||||
* as-of profile is `WHERE game_date < asOf` — a filter, not a snapshot table.
|
||||
* There is nothing to retrofit because there is nothing being overwritten.
|
||||
*/
|
||||
|
||||
const SITE = 'https://site.api.espn.com/apis/site/v2/sports/basketball/wnba';
|
||||
|
||||
/** WNBA regulation: 40 minutes, five players on the floor. */
|
||||
const REGULATION_MINUTES = 40;
|
||||
const PLAYERS_ON_FLOOR = 5;
|
||||
/**
|
||||
* The free-throw-trip coefficient. ~44% of free throws end a possession (the
|
||||
* rest are the first of a pair). This is the one estimated constant in the
|
||||
* possession identity and it is the field-standard value.
|
||||
*/
|
||||
const FT_TRIP = 0.44;
|
||||
|
||||
const num = (v) => {
|
||||
if (v == null || v === '') return null;
|
||||
const n = Number(v);
|
||||
return Number.isFinite(n) ? n : null;
|
||||
};
|
||||
|
||||
/** "8-21" → { made: 8, att: 21 }. A malformed pair is absent, never zero. */
|
||||
function madeAtt(v) {
|
||||
const parts = String(v == null ? '' : v).split('-');
|
||||
if (parts.length !== 2) return { made: null, att: null };
|
||||
return { made: num(parts[0]), att: num(parts[1]) };
|
||||
}
|
||||
|
||||
/**
|
||||
* Parse ONE team's player rows out of a summary payload.
|
||||
*
|
||||
* A player who did not play is OMITTED, never recorded with zeros. A zero-minute
|
||||
* row asserts he was available and produced nothing; a DNP is a different fact,
|
||||
* and the usage denominator would divide by his zero minutes anyway.
|
||||
*/
|
||||
function parsePlayers(group) {
|
||||
const st = group && Array.isArray(group.statistics) ? group.statistics[0] : null;
|
||||
if (!st || !Array.isArray(st.keys) || !Array.isArray(st.athletes)) return [];
|
||||
const idx = {};
|
||||
st.keys.forEach((k, i) => { idx[k] = i; });
|
||||
const at = (stats, k) => (idx[k] === undefined ? null : stats[idx[k]]);
|
||||
|
||||
const out = [];
|
||||
for (const a of st.athletes) {
|
||||
if (!a || a.didNotPlay || !Array.isArray(a.stats) || a.stats.length === 0) continue;
|
||||
const s = a.stats;
|
||||
const minutes = num(at(s, 'minutes'));
|
||||
if (minutes === null) continue; // unreadable, not zero
|
||||
const fg = madeAtt(at(s, 'fieldGoalsMade-fieldGoalsAttempted'));
|
||||
const fg3 = madeAtt(at(s, 'threePointFieldGoalsMade-threePointFieldGoalsAttempted'));
|
||||
const ft = madeAtt(at(s, 'freeThrowsMade-freeThrowsAttempted'));
|
||||
out.push({
|
||||
source_id: a.athlete && a.athlete.id != null ? String(a.athlete.id) : null,
|
||||
player_name: (a.athlete && a.athlete.displayName) || null,
|
||||
starter: a.starter === true,
|
||||
minutes,
|
||||
points: num(at(s, 'points')),
|
||||
fgm: fg.made, fga: fg.att,
|
||||
fg3m: fg3.made, fg3a: fg3.att,
|
||||
ftm: ft.made, fta: ft.att,
|
||||
reb: num(at(s, 'rebounds')),
|
||||
oreb: num(at(s, 'offensiveRebounds')),
|
||||
dreb: num(at(s, 'defensiveRebounds')),
|
||||
ast: num(at(s, 'assists')),
|
||||
tov: num(at(s, 'turnovers')),
|
||||
stl: num(at(s, 'steals')),
|
||||
blk: num(at(s, 'blocks')),
|
||||
pf: num(at(s, 'fouls')),
|
||||
plus_minus: num(at(s, 'plusMinus')),
|
||||
});
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
/** Team totals, from `boxscore.teams[]`. Absent stays absent. */
|
||||
function parseTeamTotals(teamEntry) {
|
||||
const m = {};
|
||||
for (const s of (teamEntry && teamEntry.statistics) || []) m[s.name] = s.displayValue;
|
||||
const fg = madeAtt(m['fieldGoalsMade-fieldGoalsAttempted']);
|
||||
const ft = madeAtt(m['freeThrowsMade-freeThrowsAttempted']);
|
||||
return {
|
||||
fga: fg.att,
|
||||
fta: ft.att,
|
||||
// `totalTurnovers` includes team turnovers (a shot-clock violation belongs
|
||||
// to the possession count even though no player committed it); the plain
|
||||
// `turnovers` column does not.
|
||||
tov: num(m.totalTurnovers) ?? num(m.turnovers),
|
||||
oreb: num(m.offensiveRebounds),
|
||||
dreb: num(m.defensiveRebounds),
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* TEAM POSSESSIONS — the standard box-score identity.
|
||||
*
|
||||
* POSS = FGA − OREB + TOV + 0.44 × FTA
|
||||
*
|
||||
* Any missing term ⇒ null. Treating an absent turnover count as zero would
|
||||
* understate possessions and inflate every usage rate computed against it.
|
||||
*/
|
||||
function possessions(t) {
|
||||
if (!t) return null;
|
||||
const { fga, fta, tov, oreb } = t;
|
||||
if ([fga, fta, tov, oreb].some((v) => v == null)) return null;
|
||||
return fga - oreb + tov + FT_TRIP * fta;
|
||||
}
|
||||
|
||||
/**
|
||||
* PACE — possessions per 40 minutes of team play.
|
||||
*
|
||||
* `teamMinutes / PLAYERS_ON_FLOOR` is the number of GAME minutes actually
|
||||
* played, which is how overtime enters without a special case: five players ×
|
||||
* 45 minutes is 225 team-minutes, and the divisor follows.
|
||||
*/
|
||||
function pace(poss, teamMinutes) {
|
||||
if (poss == null || !(teamMinutes > 0)) return null;
|
||||
const gameMinutes = teamMinutes / PLAYERS_ON_FLOOR;
|
||||
if (!(gameMinutes > 0)) return null;
|
||||
return (poss * REGULATION_MINUTES) / gameMinutes;
|
||||
}
|
||||
|
||||
/**
|
||||
* USAGE RATE — the share of his team's possessions a player ends while on the
|
||||
* floor. The standard identity:
|
||||
*
|
||||
* USG% = 100 × ((FGA + 0.44·FTA + TOV) × (TmMIN / 5))
|
||||
* / (MIN × (TmFGA + 0.44·TmFTA + TmTOV))
|
||||
*
|
||||
* Null on any missing term or a zero denominator — a usage rate is a ratio, and
|
||||
* a fabricated one would feed the chain's opportunity term directly.
|
||||
*/
|
||||
function usageRate(p, t, teamMinutes) {
|
||||
if (!p || !t) return null;
|
||||
if ([p.fga, p.fta, p.tov, p.minutes].some((v) => v == null)) return null;
|
||||
if ([t.fga, t.fta, t.tov].some((v) => v == null)) return null;
|
||||
if (!(p.minutes > 0) || !(teamMinutes > 0)) return null;
|
||||
const denom = p.minutes * (t.fga + FT_TRIP * t.fta + t.tov);
|
||||
if (!(denom > 0)) return null;
|
||||
const numer = (p.fga + FT_TRIP * p.fta + p.tov) * (teamMinutes / PLAYERS_ON_FLOOR);
|
||||
return (100 * numer) / denom;
|
||||
}
|
||||
|
||||
/** TRUE SHOOTING — points per scoring possession. Null without attempts. */
|
||||
function trueShooting(p) {
|
||||
if (!p || p.points == null || p.fga == null || p.fta == null) return null;
|
||||
const tsa = 2 * (p.fga + FT_TRIP * p.fta);
|
||||
if (!(tsa > 0)) return null;
|
||||
return p.points / tsa;
|
||||
}
|
||||
|
||||
/** Effective FG% — a three counts for one and a half. */
|
||||
function efgPct(p) {
|
||||
if (!p || p.fgm == null || p.fg3m == null || p.fga == null) return null;
|
||||
if (!(p.fga > 0)) return null;
|
||||
return (p.fgm + 0.5 * p.fg3m) / p.fga;
|
||||
}
|
||||
|
||||
const round = (v, d) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10 ** d) / 10 ** d);
|
||||
|
||||
/**
|
||||
* Turn ONE completed game's summary payload into per-player rows.
|
||||
*
|
||||
* ONLY COMPLETED GAMES. An in-progress box score is a partial fact whose usage
|
||||
* denominator is still moving, and storing it would make a row's meaning depend
|
||||
* on when it was written — the exact property the point-in-time design exists to
|
||||
* avoid.
|
||||
*
|
||||
* @returns {object} { rows, game, skipped }
|
||||
*/
|
||||
function parseSummary(payload, opts = {}) {
|
||||
const comp = payload && payload.header && Array.isArray(payload.header.competitions)
|
||||
? payload.header.competitions[0] : null;
|
||||
const status = comp && comp.status && comp.status.type ? comp.status.type.name : null;
|
||||
const gameId = payload && payload.header ? String(payload.header.id || '') : '';
|
||||
if (!comp || !gameId) return { rows: [], game: null, skipped: 'no_header' };
|
||||
if (status !== 'STATUS_FINAL' && !opts.allowUnfinished) {
|
||||
return { rows: [], game: null, skipped: `not_final:${status || 'unknown'}` };
|
||||
}
|
||||
|
||||
const byId = {};
|
||||
for (const c of comp.competitors || []) {
|
||||
if (c && c.team && c.team.id != null) byId[String(c.team.id)] = c;
|
||||
}
|
||||
const season = payload.header.season ? Number(payload.header.season.year) : null;
|
||||
// The ET calendar date of the game. ESPN dates are UTC and a 23:30Z tip is the
|
||||
// SAME ET evening — using the UTC date would file half the slate a day late,
|
||||
// which is the `isPreGame` lesson in a different costume.
|
||||
const gameDate = etDate(comp.date);
|
||||
|
||||
const groups = (payload.boxscore && payload.boxscore.players) || [];
|
||||
const teamTotals = (payload.boxscore && payload.boxscore.teams) || [];
|
||||
const totalsByTeamId = {};
|
||||
for (const t of teamTotals) {
|
||||
if (t && t.team && t.team.id != null) totalsByTeamId[String(t.team.id)] = parseTeamTotals(t);
|
||||
}
|
||||
|
||||
const rows = [];
|
||||
for (const g of groups) {
|
||||
const teamId = g && g.team && g.team.id != null ? String(g.team.id) : null;
|
||||
if (!teamId) continue;
|
||||
const players = parsePlayers(g);
|
||||
if (players.length === 0) continue;
|
||||
const totals = totalsByTeamId[teamId];
|
||||
const teamMinutes = players.reduce((a, p) => a + (p.minutes || 0), 0);
|
||||
const poss = possessions(totals);
|
||||
|
||||
const me = byId[teamId];
|
||||
const them = Object.values(byId).find((c) => String(c.team.id) !== teamId);
|
||||
const teamAbbr = (me && me.team && me.team.abbreviation) || (g.team && g.team.abbreviation) || null;
|
||||
const oppAbbr = (them && them.team && them.team.abbreviation) || null;
|
||||
const myScore = me ? num(me.score) : null;
|
||||
const oppScore = them ? num(them.score) : null;
|
||||
|
||||
for (const p of players) {
|
||||
rows.push({
|
||||
sport: 'wnba',
|
||||
season,
|
||||
game_id: gameId,
|
||||
game_date: gameDate,
|
||||
source_id: p.source_id,
|
||||
player_name: p.player_name,
|
||||
team: teamAbbr,
|
||||
opponent: oppAbbr,
|
||||
home_away: me ? me.homeAway : null,
|
||||
starter: p.starter,
|
||||
|
||||
minutes: p.minutes,
|
||||
points: p.points,
|
||||
fgm: p.fgm, fga: p.fga, fg3m: p.fg3m, fg3a: p.fg3a, ftm: p.ftm, fta: p.fta,
|
||||
reb: p.reb, oreb: p.oreb, dreb: p.dreb,
|
||||
ast: p.ast, tov: p.tov, stl: p.stl, blk: p.blk, pf: p.pf,
|
||||
plus_minus: p.plus_minus,
|
||||
|
||||
// DERIVED — the three the chainFn is built on.
|
||||
usage_rate: round(usageRate(p, totals, teamMinutes), 3),
|
||||
ts_pct: round(trueShooting(p), 4),
|
||||
efg_pct: round(efgPct(p), 4),
|
||||
|
||||
// TEAM CONTEXT, denormalised so a chainFn read needs ONE table. These
|
||||
// are the possession denominators; recomputing them from a second query
|
||||
// per row is how a per-prop read becomes a per-prop fetch.
|
||||
team_minutes: teamMinutes,
|
||||
team_fga: totals ? totals.fga : null,
|
||||
team_fta: totals ? totals.fta : null,
|
||||
team_tov: totals ? totals.tov : null,
|
||||
team_oreb: totals ? totals.oreb : null,
|
||||
team_possessions: round(poss, 3),
|
||||
team_pace: round(pace(poss, teamMinutes), 3),
|
||||
|
||||
// GAME STATE — what the `redistribute` hook reads. A blowout moves
|
||||
// involvement between archetypes, and in basketball that is live rather
|
||||
// than dormant, so the margin has to be on the row.
|
||||
team_score: myScore,
|
||||
opp_score: oppScore,
|
||||
final_margin: myScore != null && oppScore != null ? myScore - oppScore : null,
|
||||
});
|
||||
}
|
||||
}
|
||||
return { rows, game: { game_id: gameId, game_date: gameDate, season, status }, skipped: null };
|
||||
}
|
||||
|
||||
/** UTC instant → the ET calendar date it belongs to. */
|
||||
function etDate(iso) {
|
||||
if (!iso) return null;
|
||||
const d = new Date(iso);
|
||||
if (Number.isNaN(d.getTime())) return null;
|
||||
return new Intl.DateTimeFormat('en-CA', {
|
||||
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
|
||||
}).format(d);
|
||||
}
|
||||
|
||||
const httpJson = async (url, opts = {}) => {
|
||||
if (typeof opts.fetchImpl === 'function') return opts.fetchImpl(url);
|
||||
const res = await fetch(url, { headers: { accept: 'application/json' } });
|
||||
if (!res.ok) throw new Error(`ESPN ${res.status} for ${url}`);
|
||||
return res.json();
|
||||
};
|
||||
|
||||
/**
|
||||
* LEAGUE MEMBERSHIP, from the source — never a hardcoded table.
|
||||
*
|
||||
* The season pull surfaced two games that are not league games: an exhibition
|
||||
* against Nigeria (2026-05-02, "NIGER") and the All-Star game (2026-07-25,
|
||||
* "SPO" vs "COOP"). Their usage and pace context is meaningless for a forward
|
||||
* projection — an All-Star game has no defence to speak of and an exhibition is
|
||||
* not the league — so leaving them in would quietly pollute every profile that
|
||||
* spans those dates.
|
||||
*
|
||||
* The filter reads ESPN's own `/teams`, so an expansion franchise is admitted
|
||||
* the day the league adds it. A hardcoded fifteen would have to be remembered,
|
||||
* and this league has expanded twice in three years.
|
||||
*/
|
||||
let _teamCache = null;
|
||||
async function getLeagueTeams(opts = {}) {
|
||||
if (_teamCache && !opts.force) return _teamCache;
|
||||
const d = await httpJson(`${SITE}/teams`, opts);
|
||||
const list = ((((d || {}).sports || [])[0] || {}).leagues || [])[0];
|
||||
const abbrs = ((list && list.teams) || [])
|
||||
.map((t) => t && t.team && t.team.abbreviation)
|
||||
.filter(Boolean);
|
||||
// An EMPTY result is a failed fetch wearing an honest-absence costume (the
|
||||
// fielding_oaa lesson). Refuse to cache it, and let the caller decide.
|
||||
if (abbrs.length === 0) return null;
|
||||
_teamCache = new Set(abbrs);
|
||||
return _teamCache;
|
||||
}
|
||||
|
||||
/** Completed game ids for one ET date. */
|
||||
async function getFinalGameIds(dateYYYYMMDD, opts = {}) {
|
||||
const url = `${SITE}/scoreboard?dates=${String(dateYYYYMMDD).replace(/-/g, '')}`;
|
||||
const board = await httpJson(url, opts);
|
||||
return ((board && board.events) || [])
|
||||
.filter((e) => e && e.status && e.status.type && e.status.type.name === 'STATUS_FINAL')
|
||||
.map((e) => String(e.id));
|
||||
}
|
||||
|
||||
async function getGameRows(gameId, opts = {}) {
|
||||
const payload = await httpJson(`${SITE}/summary?event=${encodeURIComponent(gameId)}`, opts);
|
||||
const parsed = parseSummary(payload, opts);
|
||||
if (!parsed.rows.length || opts.skipLeagueFilter) return parsed;
|
||||
const teams = await getLeagueTeams(opts);
|
||||
if (!teams) return parsed; // unknown membership ⇒ no filtering
|
||||
const inLeague = parsed.rows.every((r) => teams.has(r.team) && teams.has(r.opponent));
|
||||
if (inLeague) return parsed;
|
||||
const bad = [...new Set(parsed.rows.flatMap((r) => [r.team, r.opponent]))]
|
||||
.filter((t) => !teams.has(t));
|
||||
return { rows: [], game: parsed.game, skipped: `not_a_league_game:${bad.join(',')}` };
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
parseSummary, parsePlayers, parseTeamTotals,
|
||||
possessions, pace, usageRate, trueShooting, efgPct, madeAtt, etDate,
|
||||
getFinalGameIds, getGameRows, getLeagueTeams,
|
||||
SITE, FT_TRIP, REGULATION_MINUTES, PLAYERS_ON_FLOOR,
|
||||
};
|
||||
@@ -0,0 +1,272 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* baseballChain — BASEBALL'S chainFn. The fixed-opportunity case.
|
||||
*
|
||||
* `chain.js` describes a slot: `chainFn(atom, context) → per-entity
|
||||
* probability`. This fills it for MLB, and it is deliberately the SIMPLE end of
|
||||
* the problem, because the point of the first fill is to prove the slot rather
|
||||
* than to win an argument about basketball.
|
||||
*
|
||||
* ── WHY BASEBALL IS THE CLEAN CASE ───────────────────────────────────────
|
||||
* A chained forecast is always `rate × opportunity`. In basketball the
|
||||
* opportunity term is itself a contested model — possessions are finite and
|
||||
* teammates compete for them, so one player's usage is another's non-usage, and
|
||||
* the opportunity term moves with game script. In baseball it does not. The
|
||||
* batting order is fixed before first pitch, a nine-run lead does not change who
|
||||
* bats next, and plate appearances per game vary over a narrow, stable range set
|
||||
* almost entirely by lineup slot.
|
||||
*
|
||||
* So here opportunity is a STABLE MULTIPLIER, not a model:
|
||||
*
|
||||
* p_hit_per_PA ← skillProjection.paOutcome (the rate — the modelled part)
|
||||
* expected PA ← lineup slot (the opportunity — a lookup)
|
||||
* P(hits ≥ k) ← Binomial(PA, p) mixed over the PA distribution
|
||||
*
|
||||
* ── THE ATOMS THIS ROUTES ────────────────────────────────────────────────
|
||||
* Everything below already existed and was reachable only from `scripts/`:
|
||||
*
|
||||
* skillProjection.fromStatcastRow — the ONE legal units conversion. Statcast
|
||||
* stores PERCENTAGES (0–100); feeding raw
|
||||
* rows in makes `bip = 1 − k − bb` negative
|
||||
* and refuses almost every row.
|
||||
* skillProjection.paOutcome — the PA outcome tree: K and BB via log5
|
||||
* odds-ratio against league, the remainder
|
||||
* balls in play, and hit-on-contact from
|
||||
* the ARCHETYPE-SELECTED skill inputs
|
||||
* (barrel / hard-hit / exit velo / GB-speed)
|
||||
* against the pitcher's contact allowed and
|
||||
* the park.
|
||||
* skillProjection.paDistribution — PA is not fixed at an integer; a hitter
|
||||
* skillProjection.binomialPmf gets 4 or 5 depending on how the lineup
|
||||
* skillProjection.atLeast turns over, so the projection MIXES rather
|
||||
* than pretending PA is known.
|
||||
*
|
||||
* Nothing new is modelled here. This is a router, and saying so is the point:
|
||||
* the reason the "portable core" was not portable is that this routing did not
|
||||
* exist, so the sport-specific work lived in scripts that no pipeline called.
|
||||
*
|
||||
* ── REFUSAL, NOT SUBSTITUTION ────────────────────────────────────────────
|
||||
* No batter profile ⇒ null ⇒ `chain.applyChainFn` DROPS the atom. It never
|
||||
* becomes p=0 (which would zero a whole ticket) and never falls back to a league
|
||||
* hitter (which would assert we had read someone we had not). A missing pitcher
|
||||
* is different and is handled inside `paOutcome`: the batter's own rate stands
|
||||
* untouched rather than being pulled toward league.
|
||||
*
|
||||
* ── SCOPE: HITS ONLY, ON PURPOSE ─────────────────────────────────────────
|
||||
* `projectSkill` also routes total_bases, but TB's head-to-head is on record as
|
||||
* INCONCLUSIVE-under-contamination, and the point-in-time window that would make
|
||||
* it adjudicable only starts once `statcast_history` has accrued. Widening the
|
||||
* first shadow to a stat whose verdict is already muddy would buy noise. Any
|
||||
* other stat returns null and its atom is dropped.
|
||||
*/
|
||||
|
||||
const sk = require('./skillProjection');
|
||||
const pss = require('./platoonSeverity');
|
||||
const { knownNumber, knownRate } = require('../../utils/known');
|
||||
|
||||
/** Only this stat is chained today — see the header. */
|
||||
const CHAINED_STATS = Object.freeze(['hits']);
|
||||
|
||||
/**
|
||||
* EXPECTED PLATE APPEARANCES BY LINEUP SLOT — the opportunity term.
|
||||
*
|
||||
* This is a LOOKUP, not a model, and that is the whole claim being made about
|
||||
* baseball: the leadoff hitter gets roughly three quarters of a plate appearance
|
||||
* more per game than the nine hole, the decline is close to linear at about a
|
||||
* tenth of a PA per slot, and it does not respond to game state. The numbers
|
||||
* below are that documented structure, not a fit — nothing here was tuned on
|
||||
* settled rows, because a tuned opportunity term on 1,741 rows is curve-fitting
|
||||
* dressed as physics.
|
||||
*
|
||||
* An unknown slot falls to `skillProjection.DEFAULT_PA` (4.1, "a regular"),
|
||||
* which is a statement about a typical starter and is applied ONLY when the
|
||||
* lineup has not posted. It is the one substitution in this module and it is
|
||||
* confined to the opportunity term — never to the rate.
|
||||
*/
|
||||
const PA_BY_SLOT = Object.freeze({
|
||||
1: 4.65, 2: 4.55, 3: 4.45, 4: 4.35, 5: 4.25, 6: 4.15, 7: 4.05, 8: 3.95, 9: 3.85,
|
||||
});
|
||||
|
||||
function expectedPaForSlot(slot) {
|
||||
const s = knownNumber(slot);
|
||||
if (s === null) return null; // absent slot ⇒ caller's default
|
||||
const k = Math.round(s);
|
||||
return PA_BY_SLOT[k] ?? null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Normalize a statcast row into the profile `paOutcome` expects.
|
||||
*
|
||||
* `fromStatcastRow` is the single legal chokepoint for the percentage→fraction
|
||||
* conversion; calling it here rather than at each call site is deliberate.
|
||||
*/
|
||||
function profileFrom(row) {
|
||||
if (!row) return null;
|
||||
try { return sk.fromStatcastRow(row) || null; } catch { return null; }
|
||||
}
|
||||
|
||||
/**
|
||||
* P(stat ≥ line) for ONE hitter, chained from the PA rate over expected PA.
|
||||
*
|
||||
* @param {object} atom
|
||||
* statType 'hits'
|
||||
* line the market line (0.5 ⇒ target 1)
|
||||
* batterRow raw `statcast_aggregates` row for the hitter
|
||||
* pitcherRow raw row for the opposing starter (absent ⇒ batter's own rate)
|
||||
* archetype selects WHICH skill inputs drive this hitter
|
||||
* park park multiplier (absent ⇒ 1, i.e. no claim)
|
||||
* lineupSlot batting order position — the opportunity term
|
||||
* expectedPa an explicit override, used ahead of the slot lookup
|
||||
* @param {object} context slate-level context; `allowed` gates which features
|
||||
* may be read (the shadow runs on CANDIDATEs)
|
||||
* @returns {object|null} `{ p, ... }` merged onto the atom, or null ⇒ DROPPED
|
||||
*/
|
||||
/**
|
||||
* THE HAND SPLIT — the matchup conditioner, applied to the RATE.
|
||||
*
|
||||
* `paOutcome` reads a hitter's SEASON rates. Those already contain his platoon
|
||||
* split, averaged over whichever hands he happened to face — which is precisely
|
||||
* the thing a forward matchup read is supposed to un-average. `platoonRead`
|
||||
* gives the severity of THIS hitter's own split (shrunk by the smaller side's
|
||||
* sample, refused outright below 60 PA) in the direction tonight's matchup
|
||||
* actually runs.
|
||||
*
|
||||
* IT ENTERS AT THE PER-PA HIT RATE, NOT AT THE OUTPUT PROBABILITY. Multiplying
|
||||
* P(hits ≥ 1) by a platoon factor would be a different and wrong claim: the
|
||||
* split is a statement about how often a plate appearance becomes a hit, and the
|
||||
* binomial over plate appearances is what turns that into a threshold
|
||||
* probability. Applying it after the chain would scale a number that has already
|
||||
* been through the opportunity term.
|
||||
*
|
||||
* UNREADABLE ⇒ NO ADJUSTMENT, WITH A REASON. The season rate stands untouched —
|
||||
* never nudged toward a league-typical split, which is the claim
|
||||
* `platoonSeverity` refuses to make on a thin sample.
|
||||
*/
|
||||
function handSplitFor(atom) {
|
||||
if (!atom || !atom.platoonSplits) return { multiplier: 1, applied: false, reason: 'no_splits' };
|
||||
if (!atom.throws) return { multiplier: 1, applied: false, reason: 'no_pitcher_hand' };
|
||||
if (!atom.bats) return { multiplier: 1, applied: false, reason: 'no_batter_hand' };
|
||||
const read = pss.platoonRead({ splits: atom.platoonSplits, bats: atom.bats, throws: atom.throws });
|
||||
if (!read) return { multiplier: 1, applied: false, reason: 'unreadable' };
|
||||
if (!read.readable || !Number.isFinite(read.multiplier)) {
|
||||
return { multiplier: 1, applied: false, reason: read.reason || 'unreadable' };
|
||||
}
|
||||
return {
|
||||
multiplier: read.multiplier,
|
||||
applied: true,
|
||||
reason: null,
|
||||
observed_split: read.observed_split,
|
||||
smaller_side_pa: read.smaller_side_pa,
|
||||
facing_opposite_hand: read.facing_opposite_hand,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* P(stat ≥ line) for ONE hitter, chained from the PA rate over expected PA.
|
||||
*
|
||||
* The chain is written out here rather than delegated to `projectSkill` because
|
||||
* the hand split has to enter BETWEEN the rate and the opportunity term, and
|
||||
* `projectSkill` composes those two in one call. With no split supplied this is
|
||||
* arithmetically identical to `projectSkill` — a test asserts that, so the
|
||||
* restructuring cannot quietly become a second model.
|
||||
*
|
||||
* @param {object} atom
|
||||
* statType/line/batterRow/pitcherRow/archetype/park/lineupSlot/expectedPa
|
||||
* bats / throws / platoonSplits — the hand split (A4-dated by the caller)
|
||||
* @param {object} context
|
||||
* allowed which features may be read (the shadow runs on CANDIDATEs)
|
||||
* onRefusal(id, reason) measurement side-channel; refusals are COUNTED, and
|
||||
* a refusal always DROPS the atom — there is no
|
||||
* fallback to a season rate anywhere in this module,
|
||||
* because a silent fallback is indistinguishable from
|
||||
* a read and would make the shadow's agreement with the
|
||||
* counter meaningless.
|
||||
* @returns {object|null} `{ p, ... }` merged onto the atom, or null ⇒ DROPPED
|
||||
*/
|
||||
function chainFn(atom, context = {}) {
|
||||
const refuse = (reason) => {
|
||||
if (typeof context.onRefusal === 'function') {
|
||||
try { context.onRefusal(atom && atom.id, reason); } catch { /* measurement never breaks a read */ }
|
||||
}
|
||||
return null;
|
||||
};
|
||||
if (!atom) return refuse('no_atom');
|
||||
const stat = String(atom.statType || atom.stat_type || '').toLowerCase();
|
||||
if (!CHAINED_STATS.includes(stat)) return refuse('stat_not_chained');
|
||||
|
||||
const line = knownNumber(atom.line);
|
||||
if (line === null) return refuse('no_line');
|
||||
|
||||
const batter = profileFrom(atom.batterRow);
|
||||
if (!batter) return refuse('no_batter_profile'); // absent, never a league hitter
|
||||
|
||||
const pitcher = profileFrom(atom.pitcherRow); // null is fine — paOutcome copes
|
||||
const park = knownRate(atom.park) ?? 1;
|
||||
const allowed = context.allowed || null;
|
||||
const archetype = atom.archetype || null;
|
||||
|
||||
// ── 1. THE RATE ────────────────────────────────────────────────────────
|
||||
const pa = sk.paOutcome({ batter, pitcher, park, archetype, allowed });
|
||||
if (!pa || !Number.isFinite(pa.p_hit_per_pa)) return refuse('pa_outcome_refused');
|
||||
|
||||
// ── 2. THE MATCHUP CONDITIONER ─────────────────────────────────────────
|
||||
const split = handSplitFor(atom);
|
||||
const pHit = Math.min(1, Math.max(0, pa.p_hit_per_pa * split.multiplier));
|
||||
|
||||
// ── 3. THE OPPORTUNITY TERM ────────────────────────────────────────────
|
||||
const expectedPa = knownNumber(atom.expectedPa)
|
||||
?? expectedPaForSlot(atom.lineupSlot)
|
||||
?? null; // null ⇒ paDistribution's DEFAULT_PA
|
||||
const paPmf = sk.paDistribution(expectedPa);
|
||||
|
||||
// ── 4. THE CHAIN ───────────────────────────────────────────────────────
|
||||
const pmf = new Array(sk.PA_CAP + 1).fill(0);
|
||||
for (let n = 0; n < paPmf.length; n += 1) {
|
||||
if (!paPmf[n]) continue;
|
||||
const bp = sk.binomialPmf(n, pHit);
|
||||
for (let x = 0; x < bp.length; x += 1) pmf[x] += paPmf[n] * bp[x];
|
||||
}
|
||||
const target = Math.max(1, Math.ceil(line));
|
||||
const p = sk.atLeast(pmf, target);
|
||||
if (!Number.isFinite(p)) return refuse('chain_produced_no_probability');
|
||||
|
||||
const r3 = (v) => (Number.isFinite(v) ? Math.round(v * 1000) / 1000 : null);
|
||||
return {
|
||||
p,
|
||||
chain_projected_value: r3(pmf.reduce((a, q, i) => a + q * i, 0)),
|
||||
chain_per_pa: {
|
||||
k_rate: r3(pa.k_rate), bb_rate: r3(pa.bb_rate), bip_rate: r3(pa.bip_rate),
|
||||
hit_on_contact: r3(pa.hit_on_contact),
|
||||
p_hit_per_pa_season: r3(pa.p_hit_per_pa),
|
||||
p_hit_per_pa: r3(pHit),
|
||||
},
|
||||
chain_expected_pa: expectedPa,
|
||||
chain_opportunity_source: knownNumber(atom.expectedPa) !== null
|
||||
? 'explicit'
|
||||
: (expectedPaForSlot(atom.lineupSlot) !== null ? 'lineup_slot' : 'default_regular'),
|
||||
chain_pitcher_applied: pa.inputs_used.pitcher_applied,
|
||||
// The hand split, recorded whether or not it fired. A non-application is a
|
||||
// fact about the sample, and hiding it would make the fire rate a fiction.
|
||||
chain_platoon_applied: split.applied,
|
||||
chain_platoon_multiplier: split.applied ? split.multiplier : null,
|
||||
chain_platoon_reason: split.reason,
|
||||
chain_platoon_observed_split: split.observed_split ?? null,
|
||||
chain_platoon_smaller_side_pa: split.smaller_side_pa ?? null,
|
||||
chain_archetype: String(archetype || 'DEFAULT').toUpperCase(),
|
||||
chain_family: 'pa_outcome_tree_binomial',
|
||||
};
|
||||
}
|
||||
|
||||
/** The hook `chain.chainUp`/`chainAcross` take. Dormant in baseball — see chain.js. */
|
||||
function redistribute(legs) {
|
||||
// A nine-run lead does not change who bats next. Returning the legs UNCHANGED
|
||||
// is the honest dormant behaviour; returning null would read as "the hook
|
||||
// failed" rather than "this sport has no redistribution".
|
||||
return legs;
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
chainFn, expectedPaForSlot, redistribute, handSplitFor,
|
||||
PA_BY_SLOT, CHAINED_STATS,
|
||||
};
|
||||
+123
-24
@@ -29,7 +29,14 @@
|
||||
* the bench). DORMANT in baseball — a nine-run lead does
|
||||
* not change who bats next — and LIVE in basketball,
|
||||
* where it is most of the edge. The hook exists here so
|
||||
* basketball is content rather than a rewrite.
|
||||
* basketball is content rather than a rewrite. It runs
|
||||
* on BOTH readings: a redistribution that reached only
|
||||
* the team read would make `selfCheck` flag an
|
||||
* inconsistency the model had itself just created.
|
||||
*
|
||||
* All five are real parameters. `chainFn` was described here and absent from the
|
||||
* code for its whole life, which is why every sport's actual work happened
|
||||
* outside the "portable" core — see `applyChainFn` below.
|
||||
*
|
||||
* ── CALIBRATION IS A HARD PRECONDITION, NOT A WARNING ────────────────────
|
||||
* Errors that are survivable one at a time MULTIPLY when chained. Measured on
|
||||
@@ -55,19 +62,89 @@ function usableAtoms(atoms) {
|
||||
return (atoms || []).filter((a) => a && knownNumber(a.p) !== null);
|
||||
}
|
||||
|
||||
/**
|
||||
* THE chainFn SLOT — atom + context → per-entity probability.
|
||||
*
|
||||
* This is the stage the header always described and the code never had. Without
|
||||
* it, every atom had to arrive with `p` already computed somewhere else, which
|
||||
* is exactly how the sport-specific work ended up living outside the machine
|
||||
* that claims to be portable. Baseball fills it with a fixed-opportunity read
|
||||
* (p per plate appearance, chained over expected PA); basketball will fill it
|
||||
* with a contested-possession read. Neither is a code path here.
|
||||
*
|
||||
* DEFAULT IS IDENTITY-ON-p, so every caller that passes a pre-computed
|
||||
* probability keeps working unchanged.
|
||||
*
|
||||
* A chainFn that returns null — or throws — makes that atom UNREADABLE, which
|
||||
* means DROPPED, never p=0. A zero leg would zero an entire ticket, and "we
|
||||
* could not read him" is not "he cannot do it". The refusals are COUNTED in the
|
||||
* result rather than swallowed, so a chainFn that is quietly failing on the
|
||||
* whole board is visible instead of looking like a thin slate.
|
||||
*/
|
||||
function applyChainFn(atoms, opts = {}) {
|
||||
const fn = typeof opts.chainFn === 'function' ? opts.chainFn : null;
|
||||
if (!fn) return { atoms: atoms || [], applied: false, refused: 0 };
|
||||
const ctx = opts.context || {};
|
||||
let refused = 0;
|
||||
const out = [];
|
||||
for (const a of atoms || []) {
|
||||
if (!a) { refused += 1; continue; }
|
||||
let r;
|
||||
try { r = fn(a, ctx); } catch { r = null; }
|
||||
if (r === null || r === undefined) { refused += 1; continue; }
|
||||
const merged = typeof r === 'object' ? { ...a, ...r } : { ...a, p: r };
|
||||
if (knownNumber(merged.p) === null) { refused += 1; continue; }
|
||||
out.push(merged);
|
||||
}
|
||||
return { atoms: out, applied: true, refused };
|
||||
}
|
||||
|
||||
/** The archetype-redistribution hook, applied identically by BOTH readings. */
|
||||
function applyRedistribute(legs, opts = {}) {
|
||||
if (typeof opts.redistribute !== 'function') return { legs, redistributed: false };
|
||||
const out = opts.redistribute(legs, opts.context || {});
|
||||
if (!Array.isArray(out) || out.length === 0) return { legs, redistributed: false };
|
||||
return { legs: usableAtoms(out), redistributed: true };
|
||||
}
|
||||
|
||||
/**
|
||||
* PREPARE — chainFn, then usability, then redistribution, in that order.
|
||||
*
|
||||
* Exported so a caller can prepare ONCE and hand the SAME legs to both readings.
|
||||
* That matters: `selfCheck` compares the across-read to the up-read, so if the
|
||||
* two were prepared separately and only one saw a redistribution, the check
|
||||
* would flag an inconsistency that the model itself had just created.
|
||||
*/
|
||||
function prepareAtoms(atoms, opts = {}) {
|
||||
const c = applyChainFn(atoms, opts);
|
||||
const usable = usableAtoms(c.atoms);
|
||||
const r = applyRedistribute(usable, opts);
|
||||
return {
|
||||
legs: r.legs,
|
||||
chain_fn_applied: c.applied,
|
||||
chain_fn_refused: c.refused,
|
||||
redistributed: r.redistributed,
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* CHAIN ACROSS — compound probability of every leg landing.
|
||||
*
|
||||
* @param {Array} atoms [{ id, p, calibrated, gameId, entityId, ... }]
|
||||
* @param {object} opts
|
||||
* correlation(a, b) → 0..1 shared-variance estimate between two legs
|
||||
* chainFn(atom, context) → per-entity probability (default: identity on `p`)
|
||||
* correlation(a, b) → -1..1 dependence estimate between two legs
|
||||
* redistribute(legs, context) → reweighted legs (same hook chainUp has)
|
||||
* requireCalibrated (default TRUE) — see the header
|
||||
* @returns {object|null} refusal is explicit and reasoned, never a silent 0
|
||||
*/
|
||||
function chainAcross(atoms, opts = {}) {
|
||||
const requireCalibrated = opts.requireCalibrated !== false;
|
||||
const legs = usableAtoms(atoms);
|
||||
if (legs.length === 0) return { ok: false, reason: 'no_usable_atoms' };
|
||||
const prep = prepareAtoms(atoms, opts);
|
||||
const legs = prep.legs;
|
||||
if (legs.length === 0) {
|
||||
return { ok: false, reason: 'no_usable_atoms', chain_fn_refused: prep.chain_fn_refused };
|
||||
}
|
||||
|
||||
if (requireCalibrated) {
|
||||
const uncal = legs.filter((l) => l.calibrated !== true);
|
||||
@@ -82,10 +159,16 @@ function chainAcross(atoms, opts = {}) {
|
||||
}
|
||||
}
|
||||
|
||||
// Independent product first, then a correlation discount. Same-game legs share
|
||||
// the pitcher, the park and the weather, so treating them as independent
|
||||
// OVERSTATES the ticket — the error runs in the flattering direction, which is
|
||||
// exactly the one to be careful about.
|
||||
// Independent product first, then a correlation adjustment.
|
||||
//
|
||||
// ── THE SIGN IS NOT OPTIONAL ─────────────────────────────────────────────
|
||||
// Correlation used to be clamped to [0, 1], which silently asserted that legs
|
||||
// can only ever land TOGETHER. That is baseball's shape — same-game legs share
|
||||
// the pitcher, the park and the weather — and it is the wrong shape for a
|
||||
// sport where entities compete for the same finite opportunity. Two teammates'
|
||||
// scoring props are negatively correlated through shared possessions: one
|
||||
// player's shot is another player's non-shot. Under the old clamp that case
|
||||
// was not merely mismodelled, it was INEXPRESSIBLE.
|
||||
const corrFn = typeof opts.correlation === 'function' ? opts.correlation : () => 0;
|
||||
let independent = 1;
|
||||
for (const l of legs) independent *= Math.min(1, Math.max(0, Number(l.p)));
|
||||
@@ -95,16 +178,26 @@ function chainAcross(atoms, opts = {}) {
|
||||
for (let i = 0; i < legs.length; i += 1) {
|
||||
for (let j = i + 1; j < legs.length; j += 1) {
|
||||
const c = Number(corrFn(legs[i], legs[j]));
|
||||
if (Number.isFinite(c)) { corrSum += Math.min(1, Math.max(0, c)); pairs += 1; }
|
||||
if (Number.isFinite(c)) { corrSum += Math.min(1, Math.max(-1, c)); pairs += 1; }
|
||||
}
|
||||
}
|
||||
const meanCorr = pairs > 0 ? corrSum / pairs : 0;
|
||||
|
||||
// Positive correlation makes the JOINT more likely than independence implies
|
||||
// (legs tend to land together), so the adjustment moves toward the weakest leg
|
||||
// — bounded, and stated as an approximation rather than a derivation.
|
||||
const weakest = Math.min(...legs.map((l) => Number(l.p)));
|
||||
const compound = independent + meanCorr * (weakest - independent);
|
||||
// The two targets are the FRÉCHET–HOEFFDING bounds on a joint, which is what
|
||||
// makes the direction principled rather than chosen:
|
||||
//
|
||||
// perfectly co-monotone (corr = +1) → joint = min(p_i) — the upper bound
|
||||
// independent (corr = 0) → joint = Π p_i
|
||||
// perfectly counter-monotone (corr = −1) → joint = max(0, Σp − (n−1)) — the lower bound
|
||||
//
|
||||
// So |corr| interpolates from independence toward the bound its SIGN selects.
|
||||
// Positive is unchanged from before (the weakest leg IS the upper bound), so
|
||||
// every previously-correct number stays exactly what it was.
|
||||
const probs = legs.map((l) => Math.min(1, Math.max(0, Number(l.p))));
|
||||
const frechetUpper = Math.min(...probs);
|
||||
const frechetLower = Math.max(0, probs.reduce((a, b) => a + b, 0) - (probs.length - 1));
|
||||
const target = meanCorr >= 0 ? frechetUpper : frechetLower;
|
||||
const compound = independent + Math.abs(meanCorr) * (target - independent);
|
||||
|
||||
return {
|
||||
ok: true,
|
||||
@@ -112,8 +205,12 @@ function chainAcross(atoms, opts = {}) {
|
||||
independent_probability: round4(independent),
|
||||
mean_pairwise_correlation: round4(meanCorr),
|
||||
compound_probability: round4(Math.min(1, Math.max(0, compound))),
|
||||
correlation_direction: meanCorr > 0 ? 'toward_joint' : meanCorr < 0 ? 'apart' : 'independent',
|
||||
cross_game_legs: new Set(legs.map((l) => l.gameId)).size,
|
||||
correlation_caveat: 'pairwise mean, applied as a bounded shift toward the weakest leg — an approximation, not a joint distribution',
|
||||
chain_fn_applied: prep.chain_fn_applied,
|
||||
chain_fn_refused: prep.chain_fn_refused,
|
||||
redistributed: prep.redistributed,
|
||||
correlation_caveat: 'pairwise mean interpolated toward the Fréchet bound its sign selects — an approximation, not a joint distribution',
|
||||
};
|
||||
}
|
||||
|
||||
@@ -124,13 +221,10 @@ function chainAcross(atoms, opts = {}) {
|
||||
* and may return a reweighted set — dormant in baseball, live in basketball.
|
||||
*/
|
||||
function chainUp(atoms, opts = {}) {
|
||||
let legs = usableAtoms(atoms);
|
||||
if (legs.length === 0) return { ok: false, reason: 'no_usable_atoms' };
|
||||
|
||||
let redistributed = false;
|
||||
if (typeof opts.redistribute === 'function') {
|
||||
const out = opts.redistribute(legs, opts.context || {});
|
||||
if (Array.isArray(out) && out.length > 0) { legs = usableAtoms(out); redistributed = true; }
|
||||
const prep = prepareAtoms(atoms, opts);
|
||||
const legs = prep.legs;
|
||||
if (legs.length === 0) {
|
||||
return { ok: false, reason: 'no_usable_atoms', chain_fn_refused: prep.chain_fn_refused };
|
||||
}
|
||||
|
||||
const weightOf = (a) => {
|
||||
@@ -142,7 +236,9 @@ function chainUp(atoms, opts = {}) {
|
||||
ok: true,
|
||||
contributors: legs.length,
|
||||
expected_value: round4(expected),
|
||||
redistributed,
|
||||
chain_fn_applied: prep.chain_fn_applied,
|
||||
chain_fn_refused: prep.chain_fn_refused,
|
||||
redistributed: prep.redistributed,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -223,4 +319,7 @@ function propagate(atom, observation, opts = {}) {
|
||||
|
||||
const round4 = (v) => (Number.isFinite(v) ? Math.round(v * 10000) / 10000 : v);
|
||||
|
||||
module.exports = { chainAcross, chainUp, selfCheck, propagate, usableAtoms };
|
||||
module.exports = {
|
||||
chainAcross, chainUp, selfCheck, propagate,
|
||||
usableAtoms, prepareAtoms, applyChainFn, applyRedistribute,
|
||||
};
|
||||
|
||||
@@ -0,0 +1,359 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* chainShadow — run the chain on the live board and SERVE NONE OF IT.
|
||||
*
|
||||
* This is the A6 shape, one layer up. A6 proved the hits factors could fire,
|
||||
* froze what they WOULD have applied, and changed no served number. This does
|
||||
* the same for the whole forecast: the chain computes a probability for every
|
||||
* MLB hits prop on the slate, it is written to its own column beside
|
||||
* `factor_inputs`, and the counter continues to serve untouched.
|
||||
*
|
||||
* ── WHAT IS RECORDED, AND WHY IT IS THAT SHAPE ───────────────────────────
|
||||
* The unit of evidence is the TRIPLE:
|
||||
*
|
||||
* (chain_p, counter_p, outcome)
|
||||
*
|
||||
* All three on one row, side-aligned. `counter_p` is the served
|
||||
* `estimateProbability` value for that exact side, `outcome` arrives later from
|
||||
* the ordinary settle pass, and `chain_p` is this. Evidence you cannot
|
||||
* adjudicate is not evidence: a chain probability stored without the number it
|
||||
* must beat, or without the result, can only ever be compared to itself — which
|
||||
* is the same defect `factorFreeze` exists to avoid one level down.
|
||||
*
|
||||
* ── UN-SERVABLE, AND SAID SO IN THE DATA ─────────────────────────────────
|
||||
* `chainAcross` is called with `requireCalibrated: false`, which is legitimate
|
||||
* ONLY because nothing downstream reads the result. The chain's probabilities
|
||||
* have never been through a calibration gate, and compounding an uncalibrated
|
||||
* probability is the single most harmful thing this product could ship. So every
|
||||
* row carries `servable: false` and `status: 'UN-SERVABLE'` explicitly, in the
|
||||
* stored payload rather than only in a comment. Flipping that is a separate,
|
||||
* gated event and it requires passing the gate, not editing this file.
|
||||
*
|
||||
* ── THE SELF-CHECK IS VACUOUS TODAY, AND THAT IS REPORTED, NOT HIDDEN ────
|
||||
* `selfCheck` earns its keep by comparing the per-entity reads to an INDEPENDENT
|
||||
* team read. No independent team read exists — the game-script projection was
|
||||
* deliberately not built, because no atom has passed the gate and building it
|
||||
* would be plausibility rather than proof. So the up-read here is assembled from
|
||||
* the SAME atoms as the across-read, and agreement between them is arithmetic,
|
||||
* not evidence. The result is labelled `vacuous: true` so nobody later mistakes
|
||||
* a tautology for a passing consistency test.
|
||||
*/
|
||||
|
||||
const chain = require('./chain');
|
||||
const baseball = require('./baseballChain');
|
||||
const { nameKey } = require('../../utils/playerName');
|
||||
const { knownNumber } = require('../../utils/known');
|
||||
|
||||
const SHADOW_VERSION = 'chain-shadow@1';
|
||||
|
||||
/** Only MLB hits are chained today — see baseballChain's header. */
|
||||
const SHADOW_STAT = 'hits';
|
||||
|
||||
const r4 = (v) => (Number.isFinite(v) ? Math.round(v * 10000) / 10000 : null);
|
||||
|
||||
/** `player_key|stat|line` — direction-free, because the chain reads P(over). */
|
||||
function shadowKey(playerKey, stat, line) {
|
||||
const l = knownNumber(line);
|
||||
return `${playerKey}|${String(stat || '').toLowerCase()}|${l === null ? '' : l}`;
|
||||
}
|
||||
|
||||
/**
|
||||
* The UP reading's chainFn: the same atoms, asked for an expected COUNT rather
|
||||
* than a threshold probability. This is the pluggable slot doing exactly the job
|
||||
* it exists for — one atom set, two questions, no second engine.
|
||||
*/
|
||||
function upChainFn(atom, context) {
|
||||
const across = baseball.chainFn(atom, context);
|
||||
if (!across) return null;
|
||||
const v = knownNumber(across.chain_projected_value);
|
||||
return v === null ? null : { p: v };
|
||||
}
|
||||
|
||||
/** The same question, read off legs the chain has ALREADY computed. */
|
||||
function upFromPrepared(leg) {
|
||||
const v = knownNumber(leg && leg.chain_projected_value);
|
||||
return v === null ? null : { p: v };
|
||||
}
|
||||
|
||||
/**
|
||||
* Build one atom per graded MLB hits prop.
|
||||
*
|
||||
* Every input is READ from what the snapshot already loaded — the statcast rows
|
||||
* the challenger fetched, the archetype the enrichment resolved, the park factor
|
||||
* arch-v1 attached. No new I/O: this runs inside a cron slot that already grades
|
||||
* the whole board, and adding a fetch per prop is how a measurement layer turns
|
||||
* into an outage.
|
||||
*/
|
||||
function buildAtoms(grades, deps = {}) {
|
||||
const statcastByKey = deps.statcastByKey || null;
|
||||
const pitcherRowFor = typeof deps.pitcherRowFor === 'function' ? deps.pitcherRowFor : () => null;
|
||||
const lineupSlotFor = typeof deps.lineupSlotFor === 'function' ? deps.lineupSlotFor : () => null;
|
||||
// THE HAND SPLIT, resolved AS-OF-CORRECT by the caller (A4 discipline: the
|
||||
// split as it was known at grade time, or nothing). This layer never dates
|
||||
// anything itself — it takes what it is handed, so there is exactly one place
|
||||
// the as-of rule lives and it is the place that reads the tables.
|
||||
const handSplitFor = typeof deps.handSplitFor === 'function' ? deps.handSplitFor : () => null;
|
||||
|
||||
const atoms = [];
|
||||
for (const g of grades || []) {
|
||||
const stat = String(g.stat_type || g.stat || '').toLowerCase();
|
||||
if (stat !== SHADOW_STAT) continue;
|
||||
const player = g.player || g.player_name;
|
||||
if (!player) continue;
|
||||
const pk = nameKey(player);
|
||||
// The LOCKED line is the one the grade was made against; falling back to the
|
||||
// current line would score the chain on a number the counter never saw.
|
||||
const line = knownNumber(g.gradedAt && g.gradedAt.line) ?? knownNumber(g.line);
|
||||
if (line === null) continue;
|
||||
|
||||
atoms.push({
|
||||
id: shadowKey(pk, stat, line),
|
||||
entityId: pk,
|
||||
player,
|
||||
player_key: pk,
|
||||
gameId: g.game_id || null,
|
||||
team: g.team || null,
|
||||
statType: stat,
|
||||
line,
|
||||
side: String(g.direction || '').toLowerCase() || null,
|
||||
batterRow: statcastByKey ? statcastByKey.get(pk) || null : null,
|
||||
pitcherRow: pitcherRowFor(g),
|
||||
archetype: g.archetype || null,
|
||||
park: knownNumber(g.env_park_base),
|
||||
lineupSlot: lineupSlotFor(g),
|
||||
// bats/throws/platoonSplits — absent stays absent; `baseballChain` then
|
||||
// leaves the season rate untouched and records WHY, rather than nudging
|
||||
// toward a league-typical split.
|
||||
...(() => {
|
||||
// A failing split resolver degrades to a SEASON read — it must never
|
||||
// take the slate down, and it must never be mistaken for a matchup read
|
||||
// either, which is why the reason is recorded downstream rather than
|
||||
// swallowed here.
|
||||
let hs = null;
|
||||
try { hs = handSplitFor(g); } catch { hs = null; }
|
||||
hs = hs || {};
|
||||
return {
|
||||
bats: hs.bats || null,
|
||||
throws: hs.throws || null,
|
||||
platoonSplits: hs.platoonSplits || null,
|
||||
};
|
||||
})(),
|
||||
// The counter's number for THIS row, carried so the triple is assembled
|
||||
// from what was actually served rather than recomputed later.
|
||||
counter_p_win: knownNumber(g.p_win),
|
||||
counter_side: String(g.direction || '').toLowerCase() || null,
|
||||
// NEVER true here. The chain has not been through a calibration gate.
|
||||
calibrated: false,
|
||||
});
|
||||
}
|
||||
return atoms;
|
||||
}
|
||||
|
||||
/**
|
||||
* Run the shadow over one slate.
|
||||
*
|
||||
* @returns {object} { version, byKey, summary } — `byKey` maps the direction-free
|
||||
* shadow key to the per-prop block; the merge into retention rows side-aligns.
|
||||
*/
|
||||
function runShadow(grades, deps = {}) {
|
||||
const atoms = buildAtoms(grades, deps);
|
||||
|
||||
// REFUSALS ARE COUNTED BY REASON, not just totalled. A bare count cannot tell
|
||||
// "the board has no statcast profiles" from "the PA tree has nothing to read",
|
||||
// and those need opposite fixes. There is NO fallback path: a refused atom is
|
||||
// dropped, never quietly replaced by a season rate — a silent fallback would
|
||||
// make the shadow agree with the counter for a reason that looks like
|
||||
// agreement and is not.
|
||||
const refusalReasons = new Map();
|
||||
const bump = (m, k) => m.set(k, (m.get(k) || 0) + 1);
|
||||
const context = {
|
||||
allowed: deps.allowed || null,
|
||||
onRefusal: (_id, reason) => bump(refusalReasons, reason || 'unknown'),
|
||||
...(deps.context || {}),
|
||||
};
|
||||
|
||||
// ACROSS, per game. A ticket is built from one game's legs far more often than
|
||||
// from the whole board, and the correlation question only means anything
|
||||
// within a game. `requireCalibrated:false` — see the header.
|
||||
const byGame = new Map();
|
||||
for (const a of atoms) {
|
||||
const k = a.gameId || 'unbound';
|
||||
if (!byGame.has(k)) byGame.set(k, []);
|
||||
byGame.get(k).push(a);
|
||||
}
|
||||
|
||||
const perProp = new Map();
|
||||
const gameReads = [];
|
||||
const platoonReasons = new Map();
|
||||
const opportunityReasons = new Map();
|
||||
let readable = 0;
|
||||
let refused = 0;
|
||||
let platoonApplied = 0;
|
||||
let opportunityPosted = 0;
|
||||
|
||||
for (const [gameId, gameAtoms] of byGame) {
|
||||
// PREPARED ONCE, and the two readings run on the PREPARED legs.
|
||||
//
|
||||
// Two reasons, and the first is a correctness bug avoided: chainAcross,
|
||||
// chainUp and prepareAtoms each apply the chainFn, so running the raw atoms
|
||||
// through all three would evaluate every hitter three times and count every
|
||||
// refusal three times — a fire rate off by 3x, in the flattering direction
|
||||
// for the refusal count. The second is that identical legs is precisely what
|
||||
// makes the self-check meaningful at all.
|
||||
const prep = chain.prepareAtoms(gameAtoms, {
|
||||
chainFn: baseball.chainFn, redistribute: baseball.redistribute, context,
|
||||
});
|
||||
// Identity-on-p: the legs already carry the chained probability.
|
||||
const across = chain.chainAcross(prep.legs, {
|
||||
requireCalibrated: false,
|
||||
correlation: typeof deps.correlation === 'function' ? deps.correlation : undefined,
|
||||
});
|
||||
// The UP reading asks the same atoms for an expected COUNT — read off the
|
||||
// value the chain already produced, not recomputed.
|
||||
const up = chain.chainUp(prep.legs, { chainFn: upFromPrepared });
|
||||
const check = chain.selfCheck({
|
||||
perEntity: prep.legs,
|
||||
teamRead: up.ok ? up.expected_value : null,
|
||||
});
|
||||
|
||||
refused += knownNumber(prep.chain_fn_refused) ?? 0;
|
||||
readable += prep.legs.length;
|
||||
|
||||
gameReads.push({
|
||||
game_id: gameId,
|
||||
legs: prep.legs.length,
|
||||
refused: knownNumber(across.chain_fn_refused) ?? 0,
|
||||
compound_probability: across.ok ? across.compound_probability : null,
|
||||
independent_probability: across.ok ? across.independent_probability : null,
|
||||
expected_team_hits: up.ok ? up.expected_value : null,
|
||||
self_check_confidence: check.confidence,
|
||||
});
|
||||
|
||||
const gameBlock = {
|
||||
game_id: gameId,
|
||||
across_compound: across.ok ? across.compound_probability : null,
|
||||
across_legs: across.ok ? across.legs : 0,
|
||||
up_expected_hits: up.ok ? up.expected_value : null,
|
||||
self_check: {
|
||||
confidence: check.confidence,
|
||||
internal_divergence: check.internal_divergence,
|
||||
// Both readings derive from the same atoms today, so agreement is
|
||||
// arithmetic. Labelled rather than quietly reported as a pass.
|
||||
vacuous: true,
|
||||
vacuous_reason: 'no independent team read exists — the up-read is assembled from the same atoms as the across-read',
|
||||
},
|
||||
};
|
||||
|
||||
for (const leg of prep.legs) {
|
||||
const chainP = knownNumber(leg.p);
|
||||
if (chainP === null) continue;
|
||||
perProp.set(leg.id, {
|
||||
v: SHADOW_VERSION,
|
||||
status: 'UN-SERVABLE',
|
||||
servable: false,
|
||||
servable_reason: 'the chain has not passed a calibration gate; requireCalibrated was false for this read',
|
||||
stat: leg.statType,
|
||||
line: leg.line,
|
||||
// The chain's own quantity is P(over the line); the merge side-aligns it.
|
||||
chain_p_over: r4(chainP),
|
||||
chain_projected_value: r4(knownNumber(leg.chain_projected_value)),
|
||||
chain_per_pa: leg.chain_per_pa || null,
|
||||
expected_pa: knownNumber(leg.chain_expected_pa),
|
||||
opportunity_source: leg.chain_opportunity_source || null,
|
||||
pitcher_applied: leg.chain_pitcher_applied === true,
|
||||
// THE HAND SPLIT, recorded either way. Whether the matchup conditioner
|
||||
// fired is the difference between a forward read and a season read, and
|
||||
// a later adjudication that could not tell them apart would be pooling
|
||||
// two different models.
|
||||
platoon_applied: leg.chain_platoon_applied === true,
|
||||
platoon_multiplier: knownNumber(leg.chain_platoon_multiplier),
|
||||
platoon_reason: leg.chain_platoon_reason || null,
|
||||
platoon_observed_split: knownNumber(leg.chain_platoon_observed_split),
|
||||
platoon_smaller_side_pa: knownNumber(leg.chain_platoon_smaller_side_pa),
|
||||
archetype: leg.chain_archetype || null,
|
||||
family: leg.chain_family || null,
|
||||
game: gameBlock,
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
// COUNTED OFF THE STORED BLOCKS, not off the legs.
|
||||
//
|
||||
// A prop with both an over and an under row produces TWO legs with the same
|
||||
// id, and `perProp` collapses them into one block. Counting the hand split per
|
||||
// LEG while the divergence is later sliced per BLOCK gives two rates over two
|
||||
// different denominators that look comparable and are not — the first draft
|
||||
// reported 48.1% and 55.1% for the same fact. One denominator, taken from the
|
||||
// thing that is actually persisted.
|
||||
for (const block of perProp.values()) {
|
||||
if (block.platoon_applied === true) platoonApplied += 1;
|
||||
else bump(platoonReasons, block.platoon_reason || 'unknown');
|
||||
// A POSTED lineup slot is a real opportunity term; `default_regular` is the
|
||||
// league's typical starter asserted about this hitter. Counted apart,
|
||||
// because a chain running on a default opportunity term is doing half the
|
||||
// job it claims and must not report as if it read the lineup.
|
||||
if (block.opportunity_source === 'lineup_slot' || block.opportunity_source === 'explicit') {
|
||||
opportunityPosted += 1;
|
||||
} else bump(opportunityReasons, block.opportunity_source || 'unknown');
|
||||
}
|
||||
|
||||
return {
|
||||
version: SHADOW_VERSION,
|
||||
byKey: perProp,
|
||||
summary: {
|
||||
atoms: atoms.length,
|
||||
readable,
|
||||
refused,
|
||||
// Unique props (a prop's over and under legs share one block). This is the
|
||||
// denominator `platoon_applied` is over.
|
||||
props: perProp.size,
|
||||
// The two rates that matter, and they are DIFFERENT questions:
|
||||
// readable — did the chain produce a probability at all?
|
||||
// platoon — did it read tonight's MATCHUP, or this season's average?
|
||||
// Conflating them is how a season-rate read gets reported as a forward one.
|
||||
platoon_applied: platoonApplied,
|
||||
// The THIRD distinct question: did the chain read a POSTED lineup slot, or
|
||||
// fall to "a regular"? rate x opportunity — reporting only the rate half
|
||||
// as a fire rate is how v2 shipped with a constant 4.1 PA for every hitter.
|
||||
opportunity_posted: opportunityPosted,
|
||||
refusal_reasons: Object.fromEntries(refusalReasons),
|
||||
platoon_reasons: Object.fromEntries(platoonReasons),
|
||||
opportunity_reasons: Object.fromEntries(opportunityReasons),
|
||||
games: byGame.size,
|
||||
games_read: gameReads.filter((g) => g.compound_probability != null).length,
|
||||
per_game: gameReads,
|
||||
},
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* Side-align the chain probability onto ONE settled-or-pending row.
|
||||
*
|
||||
* `p_win` is expressed for the graded side, so an under row must carry
|
||||
* `1 − p_over` or the two halves of the triple would be pointing in opposite
|
||||
* directions and every comparison after it would be inverted. This is the same
|
||||
* trap `snapshotSettlementService` documents for `outcome` vs `actual_value`.
|
||||
*/
|
||||
function alignToSide(block, side, counterP) {
|
||||
if (!block) return null;
|
||||
const s = String(side || '').toLowerCase();
|
||||
const pOver = knownNumber(block.chain_p_over);
|
||||
if (pOver === null || (s !== 'over' && s !== 'under')) return null;
|
||||
const chainP = s === 'under' ? 1 - pOver : pOver;
|
||||
const counter = knownNumber(counterP);
|
||||
return {
|
||||
...block,
|
||||
side: s,
|
||||
chain_p: r4(chainP),
|
||||
counter_p: counter === null ? null : r4(counter),
|
||||
// Recorded, not recomputed downstream: the divergence is the first signal of
|
||||
// whether the chain is finding anything the counter is not.
|
||||
divergence: counter === null ? null : r4(chainP - counter),
|
||||
};
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
runShadow, buildAtoms, alignToSide, shadowKey, upChainFn,
|
||||
SHADOW_VERSION, SHADOW_STAT,
|
||||
};
|
||||
@@ -76,8 +76,12 @@ async function build(deps = {}) {
|
||||
let games = [];
|
||||
try {
|
||||
[lineups, games] = await Promise.all([
|
||||
// `batting_order` rides along on a read that already happens — it is the
|
||||
// OPPORTUNITY term the chain needs (PA per game is set almost entirely by
|
||||
// lineup slot), and fetching it separately would be a second query for a
|
||||
// column already in the row. Nothing on the served path reads it.
|
||||
paginate(() => sb.from('lineup_context')
|
||||
.select('as_of_date, game_date, sport, game_pk, team, side, player_key')
|
||||
.select('as_of_date, game_date, sport, game_pk, team, side, player_key, batting_order')
|
||||
.eq('sport', 'mlb').eq('game_date', gameDate).lte('as_of_date', asOf),
|
||||
{ key: uniqueKeyFor('lineup_context'), pageSize: 1000, label: 'matchupKeys:lineup_context' }),
|
||||
getSchedule(gameDate),
|
||||
@@ -122,15 +126,27 @@ async function build(deps = {}) {
|
||||
*/
|
||||
function resolve(prop) {
|
||||
const key = nameKey(prop && (prop.player || prop.player_name));
|
||||
const empty = { opponent: null, opposing_pitcher: null, source: null, refused: 'no_lineup_row' };
|
||||
const empty = { opponent: null, opposing_pitcher: null, batting_order: null, source: null, refused: 'no_lineup_row' };
|
||||
if (!key) return { ...empty, refused: 'no_player' };
|
||||
const lu = teamByPlayer.get(key);
|
||||
if (!lu || !lu.team) return empty;
|
||||
// The posted batting slot, carried whether or not the matchup resolves — a
|
||||
// hitter with no probable pitcher declared still has a lineup position, and
|
||||
// that is the opportunity term. Absent stays absent (never slot 4 by
|
||||
// default, which would assert he bats cleanup).
|
||||
const slot = lu.batting_order == null ? null : Number(lu.batting_order);
|
||||
const battingOrder = Number.isFinite(slot) ? slot : null;
|
||||
const m = matchup.get(lu.team) || matchup.get(String(lu.team).split(' ').pop());
|
||||
if (!m) return { opponent: null, opposing_pitcher: null, source: null, refused: 'team_not_in_schedule' };
|
||||
if (!m) {
|
||||
return {
|
||||
opponent: null, opposing_pitcher: null, batting_order: battingOrder,
|
||||
team: lu.team, source: null, refused: 'team_not_in_schedule',
|
||||
};
|
||||
}
|
||||
return {
|
||||
opponent: m.opponent || null,
|
||||
opposing_pitcher: m.opposing_pitcher || null,
|
||||
batting_order: battingOrder,
|
||||
team: lu.team,
|
||||
source: 'lineup_context+schedule',
|
||||
refused: m.opposing_pitcher ? null : 'no_probable_pitcher',
|
||||
|
||||
@@ -156,6 +156,13 @@ function rowsFromSides(base, sides, ctx = {}) {
|
||||
// the multiplier is recomputable from them and is deliberately not stored.
|
||||
// Null on every non-hits row and on any row with no factor context.
|
||||
factor_inputs: s.factor_inputs || null,
|
||||
|
||||
// The SHADOW CHAIN read. Declared here (always present, usually null) so
|
||||
// every row in a batch carries the same keys — PostgREST builds a bulk
|
||||
// insert from the FIRST row's shape, so a column that appears only on some
|
||||
// rows is silently dropped for the whole batch. Filled by
|
||||
// `mergeChainShadow` after enrichment; never read by anything served.
|
||||
chain_shadow: null,
|
||||
});
|
||||
}
|
||||
return rows;
|
||||
@@ -246,6 +253,43 @@ function mergeEnrichment(rows, enrichedGrades) {
|
||||
});
|
||||
}
|
||||
|
||||
/**
|
||||
* Attach the SHADOW CHAIN read to collected rows — the (chain_p, counter_p,
|
||||
* outcome) triple's first two thirds.
|
||||
*
|
||||
* Like `mergeEnrichment` this fills ONE field and touches nothing else. It runs
|
||||
* later than grade time for the same structural reason the archetype does: the
|
||||
* chain needs the statcast rows and park factors the enrichment pass loaded, and
|
||||
* pulling that forward into the grader would put per-prop I/O on the serving
|
||||
* path.
|
||||
*
|
||||
* THE SIDE ALIGNMENT IS THE LOAD-BEARING PART. The chain computes P(over the
|
||||
* line); `p_win` on the row is expressed for the graded SIDE. Storing the raw
|
||||
* over-probability against an under row's `p_win` would invert every comparison
|
||||
* made from it afterwards, silently — so `alignToSide` is given the row's own
|
||||
* side and the row's own counter probability, and both halves of the triple end
|
||||
* up pointing the same way. `outcome` is written by the ordinary settle pass and
|
||||
* is already side-aligned, which completes it.
|
||||
*
|
||||
* A row the chain could not read is left NULL. It is not a zero probability and
|
||||
* not an average hitter — the chain refusing to read someone is a fact worth
|
||||
* keeping, and a fabricated third of a triple would poison the adjudication this
|
||||
* column exists to enable.
|
||||
*/
|
||||
function mergeChainShadow(rows, shadow) {
|
||||
if (!Array.isArray(rows) || !rows.length) return rows || [];
|
||||
const byKey = shadow && shadow.byKey;
|
||||
if (!byKey || typeof byKey.get !== 'function') return rows;
|
||||
const cs = require('./model/chainShadow');
|
||||
|
||||
return rows.map((r) => {
|
||||
const block = byKey.get(cs.shadowKey(r.player_key, r.stat, r.line));
|
||||
if (!block) return r;
|
||||
const aligned = cs.alignToSide(block, r.side, r.p_win);
|
||||
return aligned ? { ...r, chain_shadow: aligned } : r;
|
||||
});
|
||||
}
|
||||
|
||||
function newSnapshotId() {
|
||||
return crypto.randomUUID();
|
||||
}
|
||||
@@ -257,6 +301,7 @@ module.exports = {
|
||||
rowsFromSides,
|
||||
createCollector,
|
||||
mergeEnrichment,
|
||||
mergeChainShadow,
|
||||
persist,
|
||||
newSnapshotId,
|
||||
__internals: { numOrNull, intOrNull, boolOrNull },
|
||||
|
||||
@@ -326,6 +326,18 @@ const CALIBRATION_BASIS = Object.freeze({});
|
||||
async function runSnapshot(sport, opts = {}) {
|
||||
const sp = String(sport || '').toLowerCase();
|
||||
const deps = {
|
||||
// ── THE SEAMS THIS OBJECT WAS AN ALLOWLIST FOR ──────────────────────────
|
||||
// Fourteen call sites below read `deps.challenger`, `deps.loadStatcast`,
|
||||
// `deps.environmentContext`, `deps.lineupContext`, `deps.hitsFactorContext`,
|
||||
// `deps.matchupKeys`, `deps.gameBinder` and friends, each documented as
|
||||
// injectable. None of them were in this literal, so every one of them was
|
||||
// permanently `undefined` and always fell through to the real module — the
|
||||
// injection seam existed in the comment and not in the code.
|
||||
//
|
||||
// Spreading FIRST rather than adding fourteen entries: every explicit key
|
||||
// below still wins (they are declared after and already read `opts.X`), so
|
||||
// no existing behaviour moves, while an unlisted dep now actually arrives.
|
||||
...opts,
|
||||
getOdds: opts.getOdds || require('./oddsService').getOdds,
|
||||
gradeAndCacheSlate: opts.gradeAndCacheSlate || require('./gradeSlateService').gradeAndCacheSlate,
|
||||
resolveStats: opts.resolveStats || require('./playerIntelService').resolvePlayerStats,
|
||||
@@ -474,7 +486,11 @@ async function runSnapshot(sport, opts = {}) {
|
||||
if (sp === 'mlb') {
|
||||
try {
|
||||
const ctxSvc = deps.hitsFactorContext || require('./model/hitsFactorContext');
|
||||
const sbc = require('../utils/supabase').getSupabaseServiceClient();
|
||||
// Injectable so the seam is REACHABLE. Without this the whole factor +
|
||||
// hand-split path is gated behind a real Supabase client, so no test can
|
||||
// reach it and "it is wired" would rest on reading the code — which is
|
||||
// precisely how A5 shipped three factors that never fired.
|
||||
const sbc = deps.supabase || require('../utils/supabase').getSupabaseServiceClient();
|
||||
factorContext = sbc ? await ctxSvc.build(sbc) : null;
|
||||
if (factorContext) console.log(`[factors] ${sp} hits context loaded — ${JSON.stringify(factorContext.__stats)}`);
|
||||
else console.log(`[factors] ${sp} — no factor context; grading unadjusted`);
|
||||
@@ -493,7 +509,7 @@ async function runSnapshot(sport, opts = {}) {
|
||||
if (sp === 'mlb') {
|
||||
try {
|
||||
const mk = deps.matchupKeys || require('./model/matchupKeys');
|
||||
const sbm = require('../utils/supabase').getSupabaseServiceClient();
|
||||
const sbm = deps.supabase || require('../utils/supabase').getSupabaseServiceClient();
|
||||
const mlbAdapter = deps.mlbAdapter || require('./adapters/mlbStatsAdapter');
|
||||
const resolver = sbm ? await mk.build({
|
||||
sb: sbm,
|
||||
@@ -531,12 +547,20 @@ async function runSnapshot(sport, opts = {}) {
|
||||
// empty-slate early return: a slate that graded nothing but refused
|
||||
// everything is exactly the case worth recording.
|
||||
let retentionRows = 0;
|
||||
const persistRetention = async (enrichedGrades) => {
|
||||
const persistRetention = async (enrichedGrades, chainShadow = null) => {
|
||||
if (!retention || !collector || !collector.rows.length) return;
|
||||
try {
|
||||
const rows = retention.mergeEnrichment
|
||||
let rows = retention.mergeEnrichment
|
||||
? retention.mergeEnrichment(collector.rows, enrichedGrades || [])
|
||||
: collector.rows;
|
||||
// CHAIN v1 SHADOW — attaches the (chain_p, counter_p) pair; `outcome`
|
||||
// completes the triple at settle. Own guard: the shadow is a measurement
|
||||
// layer and must never cost the retention write, which is the record.
|
||||
if (chainShadow && retention.mergeChainShadow) {
|
||||
try { rows = retention.mergeChainShadow(rows, chainShadow); } catch (e) {
|
||||
console.warn('[chain-shadow] merge skipped (retention continues):', e.message);
|
||||
}
|
||||
}
|
||||
const r = await retention.persist(rows);
|
||||
retentionRows = r.written || 0;
|
||||
console.log(`[snapshot] retention ${sp}: ${r.written}/${r.attempted} rows${r.skipped ? ' (skipped — no supabase env)' : ''}${r.error ? ` ERROR: ${r.error}` : ''}`);
|
||||
@@ -708,10 +732,17 @@ async function runSnapshot(sport, opts = {}) {
|
||||
// `enriched` (champion p_win) is READ, never written: the serving projection
|
||||
// is untouched, and the challenger rides alongside it to the ledger.
|
||||
let withChallenger = enriched;
|
||||
// CHAIN v1 SHADOW reuses what the challenger pass already fetched — the
|
||||
// statcast profiles and the resolved opposing starter. Captured out here so
|
||||
// the shadow costs ZERO additional I/O; a measurement layer that adds a fetch
|
||||
// per prop inside a cron slot is how measurement becomes an outage.
|
||||
let statcastRowsByKey = null;
|
||||
let oppPitcherByTeam = null;
|
||||
try {
|
||||
const challenger = deps.challenger || require('./challengerProjection');
|
||||
const axes = deps.archetypeAxes || require('./archetypeAxes');
|
||||
const rowsByKey = await (deps.loadStatcast || loadStatcastRows)(sp);
|
||||
statcastRowsByKey = rowsByKey;
|
||||
if (rowsByKey && rowsByKey.size) {
|
||||
const classifyFor = (playerName) => {
|
||||
const row = rowsByKey.get(nameKey(playerName || ''));
|
||||
@@ -728,6 +759,7 @@ async function runSnapshot(sport, opts = {}) {
|
||||
origin: process.env.BACKEND_SELF_ORIGIN || 'http://localhost:3000',
|
||||
});
|
||||
ctxInternals = ctx._internals || null;
|
||||
oppPitcherByTeam = (ctxInternals && ctxInternals.oppPitcherByTeam) || null;
|
||||
// Enrich each grade with the hitter hand the platoon estimate needs
|
||||
// (statcast_aggregates.bats, already loaded above).
|
||||
const handOf = (name) => {
|
||||
@@ -868,6 +900,119 @@ async function runSnapshot(sport, opts = {}) {
|
||||
}
|
||||
}
|
||||
|
||||
// ── CHAIN v1 — THE SHADOW READ ──────────────────────────────────────────
|
||||
// The chain computes a probability the way the model was always described as
|
||||
// working: base-event atoms chained over opportunity, rather than a count of
|
||||
// how often the player has cleared this number before. It runs on the live
|
||||
// board and REACHES NOTHING. `enriched` is read, never written; the served
|
||||
// `p_win` is byte-identical whether or not this block executes, which is the
|
||||
// same fence A6 used for the factor shadow and the same reason it is safe to
|
||||
// run an uncalibrated forecast at all.
|
||||
//
|
||||
// `requireCalibrated:false` is legitimate here and ONLY here: no consumer
|
||||
// exists. Every stored block says so in its own payload (`servable: false`),
|
||||
// because a caveat that lives only in a comment is not attached to the data
|
||||
// once the data is queried by something else.
|
||||
//
|
||||
// It writes ONE column on the retention row — the chain's probability next to
|
||||
// the counter's, so the settle pass completes an adjudicable triple.
|
||||
let chainShadow = null;
|
||||
if (sp === 'mlb') {
|
||||
try {
|
||||
const cs = deps.chainShadow || require('./model/chainShadow');
|
||||
const reg = deps.featureRegistry || require('./model/featureRegistry');
|
||||
// CANDIDATE features, not PROVEN. The registry's live gate returns PROVEN
|
||||
// only and would refuse every row — correctly, for anything served. A
|
||||
// challenger must be allowed to READ what it is being measured on, which
|
||||
// is exactly what `candidateFeatures` is for.
|
||||
const allowed = reg.candidateFeatures('mlb');
|
||||
const pitcherRowFor = (g) => {
|
||||
if (!oppPitcherByTeam || !statcastRowsByKey) return null;
|
||||
try {
|
||||
const { abbrOf } = require('./environmentContext');
|
||||
const pid = oppPitcherByTeam.get(abbrOf(g.team));
|
||||
if (pid == null) return null;
|
||||
for (const row of statcastRowsByKey.values()) {
|
||||
if (Number(row.source_id) === Number(pid) && row.role === 'pitcher') return row;
|
||||
}
|
||||
return null;
|
||||
} catch { return null; }
|
||||
};
|
||||
// ── CHAIN v2 — THE HAND SPLIT, AS-OF-CORRECT ────────────────────────
|
||||
// `paOutcome` reads a hitter's SEASON rates, which already average his
|
||||
// platoon split over whichever hands he happened to face. Un-averaging it
|
||||
// is what makes the read a MATCHUP rather than a season.
|
||||
//
|
||||
// Both halves are ALREADY LOADED, so this costs nothing new: the batter's
|
||||
// hand and his vs-LHP/vs-RHP splits come from `factorContext`
|
||||
// (`hitsFactorContext.build`, whose reads are A4-dated), and the opposing
|
||||
// starter's hand comes through the A6 `matchupKeys` resolve. Without those
|
||||
// keys the resolver's `throws` is null — the A5 finding — so the split
|
||||
// would silently never fire, which is the whole failure mode this order
|
||||
// exists inside. Absent ⇒ NO adjustment and a recorded reason; never a
|
||||
// league-typical split asserted about a hitter we cannot read.
|
||||
// Gated on `factorContext` ALONE, not on both. Without the keys the
|
||||
// hitter's own split is still readable and only the pitcher hand is
|
||||
// missing — so the recorded reason is `no_pitcher_hand`, which names the
|
||||
// input that is actually absent, rather than `no_splits`, which would
|
||||
// point at the half that was there all along. That mis-naming is how A5's
|
||||
// real cause stayed hidden for months.
|
||||
const handSplitFor = factorContext
|
||||
? (g) => {
|
||||
try {
|
||||
const keys = matchupKeys ? (matchupKeys(g) || null) : null;
|
||||
const ctx = factorContext(g, sp, keys);
|
||||
if (!ctx) return null;
|
||||
return { bats: ctx.bats || null, throws: ctx.throws || null, platoonSplits: ctx.platoonSplits || null };
|
||||
} catch { return null; }
|
||||
}
|
||||
: null;
|
||||
if (!handSplitFor) {
|
||||
console.log(`[chain-shadow] ${sp} — no factor context; the chain reads SEASON rates`);
|
||||
} else if (!matchupKeys) {
|
||||
console.log(`[chain-shadow] ${sp} — factor context loaded but NO join keys; the hand split will refuse with no_pitcher_hand`);
|
||||
}
|
||||
|
||||
// ── CHAIN v3 — THE OPPORTUNITY TERM, REALLY FED ─────────────────────
|
||||
// v2 shipped the shadow WITHOUT a `lineupSlotFor`, so every hitter fell to
|
||||
// `skillProjection.DEFAULT_PA` (4.1) — "a regular" asserted about the
|
||||
// leadoff man and the nine hole alike. The chain is rate x opportunity;
|
||||
// holding opportunity constant across the lineup throws away the half of
|
||||
// the model that is a free, known, pre-game fact.
|
||||
//
|
||||
// The slot comes off the A6 lineup read (`matchupKeys`), which already
|
||||
// fetches the row — one extra column, no extra query, as-of dated. Absent
|
||||
// ⇒ DEFAULT_PA and the block records `default_regular`, so a season-shaped
|
||||
// opportunity term is never mistaken for a posted one.
|
||||
const lineupSlotFor = matchupKeys
|
||||
? (g) => {
|
||||
try {
|
||||
const k = matchupKeys(g);
|
||||
return k && k.batting_order != null ? k.batting_order : null;
|
||||
} catch { return null; }
|
||||
}
|
||||
: null;
|
||||
|
||||
chainShadow = cs.runShadow(withChallenger, {
|
||||
statcastByKey: statcastRowsByKey,
|
||||
pitcherRowFor,
|
||||
handSplitFor,
|
||||
lineupSlotFor,
|
||||
allowed,
|
||||
});
|
||||
const s = chainShadow.summary;
|
||||
console.log(`[chain-shadow] ${sp} — ${s.readable}/${s.atoms} atoms read (${s.refused} refused) across ${s.games_read}/${s.games} games; `
|
||||
+ `hand split fired on ${s.platoon_applied}, posted lineup slot on ${s.opportunity_posted}/${s.props}; `
|
||||
+ 'UN-SERVABLE, served p_win untouched');
|
||||
if (s.refused) console.log(`[chain-shadow] ${sp} refusals: ${JSON.stringify(s.refusal_reasons)}`);
|
||||
if (s.readable > s.platoon_applied) console.log(`[chain-shadow] ${sp} season-rate rows: ${JSON.stringify(s.platoon_reasons)}`);
|
||||
} catch (e) {
|
||||
// A measurement layer must never break the pipeline it measures inside.
|
||||
console.warn(`[chain-shadow] ${sp} skipped:`, e.message);
|
||||
chainShadow = null;
|
||||
}
|
||||
}
|
||||
|
||||
// LINEUP + BASERUNNER CONTEXT — the input RBI and runs have always needed.
|
||||
// Best-effort and dated: a context failure must never break a snapshot, and a
|
||||
// lineup is a PRE-GAME fact that changes by the hour, so what we knew at grade
|
||||
@@ -885,7 +1030,7 @@ async function runSnapshot(sport, opts = {}) {
|
||||
}
|
||||
}
|
||||
|
||||
await persistRetention(enriched);
|
||||
await persistRetention(enriched, chainShadow);
|
||||
|
||||
// Line deltas vs the previous snapshot's locked lines.
|
||||
const prev = await deps.cacheGet(`snapshot:${sp}:latest`);
|
||||
@@ -946,6 +1091,24 @@ async function runSnapshot(sport, opts = {}) {
|
||||
// writes nothing. Retention is best-effort by design, which makes a broken
|
||||
// write invisible without this.
|
||||
retentionRows,
|
||||
// Chain v1 shadow — counts only, so a run that read nothing is visible
|
||||
// instead of looking like a slate with no hits props. Never a probability:
|
||||
// this is the return value of a snapshot, which surfaces are allowed to read.
|
||||
chainShadow: chainShadow ? {
|
||||
atoms: chainShadow.summary.atoms,
|
||||
readable: chainShadow.summary.readable,
|
||||
refused: chainShadow.summary.refused,
|
||||
props: chainShadow.summary.props,
|
||||
// Reported SEPARATELY from `readable` on purpose: "the chain produced a
|
||||
// number" and "the chain read tonight's matchup" are different claims, and
|
||||
// collapsing them is how a season read gets reported as a forward one.
|
||||
platoon_applied: chainShadow.summary.platoon_applied,
|
||||
// rate x opportunity — BOTH halves reported. v2 shipped with the
|
||||
// opportunity half a constant, and nothing in the summary could show it.
|
||||
opportunity_posted: chainShadow.summary.opportunity_posted,
|
||||
games_read: chainShadow.summary.games_read,
|
||||
servable: false,
|
||||
} : null,
|
||||
topGrades: enriched.filter((g) => isTopGrade(g.grade)).slice(0, 5).map((g) => ({
|
||||
player: g.player || g.player_name, stat: g.stat_type || g.stat, grade: g.grade, archetype: g.archetype,
|
||||
})),
|
||||
|
||||
@@ -0,0 +1,215 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* wnbaUsageService — ingest the WNBA possession/usage feed, and read it AS-OF.
|
||||
*
|
||||
* Spec: specs/wnba-possession-feed.md
|
||||
*
|
||||
* Two responsibilities, deliberately in one place because they must agree about
|
||||
* what "as of" means:
|
||||
*
|
||||
* ingestRange() pull completed games and persist per-player rows
|
||||
* profileAsOf() the point-in-time read the chainFn will consume
|
||||
*
|
||||
* ── THE AS-OF RULE, STATED ONCE ──────────────────────────────────────────
|
||||
* A profile as of date D is built from games with `game_date < D`, STRICTLY.
|
||||
* Not `<=`. A game on D may have tipped after grade time, so including it would
|
||||
* feed tonight's result into tonight's forecast — the same leak
|
||||
* `snapshotSettlementService.isPreGame` exists to stop, and the reason its drain
|
||||
* had to filter by pre-game rather than by calendar date.
|
||||
*
|
||||
* This is the whole point-in-time story. There is no history table because
|
||||
* nothing is overwritten: a completed box score is immutable, so the filter IS
|
||||
* the retention.
|
||||
*
|
||||
* ── WHAT IT REFUSES ──────────────────────────────────────────────────────
|
||||
* No games before the cutoff ⇒ `null`, not a league-average player. A replacement
|
||||
* profile would silently assert a usage rate about someone we have never seen,
|
||||
* and usage feeds the chain's opportunity term directly — the same reason
|
||||
* `platoonSeverity` refuses a thin split rather than shrinking it to league.
|
||||
*/
|
||||
|
||||
const adapter = require('./adapters/espnWnbaAdapter');
|
||||
const { paginate } = require('../utils/safePaginate');
|
||||
const { uniqueKeyFor } = require('../utils/tableKeys');
|
||||
const { nameKey } = require('../utils/playerName');
|
||||
|
||||
const TABLE = 'wnba_player_game';
|
||||
/** Minimum games before a profile is worth calling a profile. */
|
||||
const MIN_GAMES = 3;
|
||||
|
||||
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
|
||||
const round = (v, d = 3) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10 ** d) / 10 ** d);
|
||||
|
||||
/** ET calendar date, the only calendar this pipeline uses. */
|
||||
function dateET(d = new Date()) {
|
||||
return new Intl.DateTimeFormat('en-CA', {
|
||||
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
|
||||
}).format(d instanceof Date ? d : new Date(d));
|
||||
}
|
||||
|
||||
function* eachDate(from, to) {
|
||||
const start = new Date(`${from}T12:00:00Z`);
|
||||
const end = new Date(`${to}T12:00:00Z`);
|
||||
for (let d = start; d <= end; d = new Date(d.getTime() + 86_400_000)) {
|
||||
yield d.toISOString().slice(0, 10);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Ingest every completed game in an ET date range.
|
||||
*
|
||||
* Idempotent: the primary key is (game_id, source_id) and this upserts, so a
|
||||
* re-run over the same dates rewrites identical rows rather than duplicating.
|
||||
* That matters because the natural operation here is "re-run yesterday", and a
|
||||
* feed that doubles on a retry is a feed nobody dares re-run.
|
||||
*/
|
||||
async function ingestRange(sb, { from, to, fetchImpl, onProgress } = {}) {
|
||||
const out = {
|
||||
dates: 0, games: 0, rows: 0, written: 0, skipped: [], errors: [],
|
||||
};
|
||||
if (!sb) return { ...out, errors: ['no supabase client'] };
|
||||
if (!from || !to) return { ...out, errors: ['from and to required'] };
|
||||
|
||||
for (const date of eachDate(from, to)) {
|
||||
out.dates += 1;
|
||||
let ids = [];
|
||||
try {
|
||||
ids = await adapter.getFinalGameIds(date, { fetchImpl });
|
||||
} catch (e) {
|
||||
// A failed DATE is surfaced, never swallowed as "no games" — an empty feed
|
||||
// and a broken feed look identical downstream, and that costume has cost
|
||||
// this codebase two outages (the fielding_oaa 404, the settle 500-row URL).
|
||||
out.errors.push(`${date}: ${e.message}`);
|
||||
continue;
|
||||
}
|
||||
for (const gameId of ids) {
|
||||
let parsed;
|
||||
try {
|
||||
parsed = await adapter.getGameRows(gameId, { fetchImpl });
|
||||
} catch (e) {
|
||||
out.errors.push(`${date}/${gameId}: ${e.message}`);
|
||||
continue;
|
||||
}
|
||||
if (parsed.skipped) { out.skipped.push({ game_id: gameId, reason: parsed.skipped }); continue; }
|
||||
if (!parsed.rows.length) { out.skipped.push({ game_id: gameId, reason: 'no_player_rows' }); continue; }
|
||||
out.games += 1;
|
||||
|
||||
const rows = parsed.rows
|
||||
.filter((r) => r.source_id && r.player_name)
|
||||
.map((r) => ({ ...r, player_key: nameKey(r.player_name) }));
|
||||
out.rows += rows.length;
|
||||
|
||||
const { error } = await sb.from(TABLE).upsert(rows, { onConflict: 'game_id,source_id' });
|
||||
if (error) out.errors.push(`${gameId}: ${error.message}`);
|
||||
else out.written += rows.length;
|
||||
if (typeof onProgress === 'function') onProgress({ date, gameId, rows: rows.length });
|
||||
}
|
||||
}
|
||||
return out;
|
||||
}
|
||||
|
||||
/**
|
||||
* Every stored game for one player STRICTLY BEFORE `asOf`.
|
||||
*
|
||||
* Read through `safePaginate`: an unordered `.range()` returns the right COUNT
|
||||
* and the wrong ROWS, measured at up to 33.6% duplication on this database, and
|
||||
* duplicated games would double-weight a player's usage.
|
||||
*/
|
||||
async function gamesBefore(sb, { playerKey, asOf, season = null } = {}) {
|
||||
if (!sb || !playerKey || !asOf) return [];
|
||||
return paginate(
|
||||
() => {
|
||||
let q = sb.from(TABLE).select('*').eq('sport', 'wnba')
|
||||
.eq('player_key', playerKey).lt('game_date', asOf);
|
||||
if (season != null) q = q.eq('season', season);
|
||||
return q;
|
||||
},
|
||||
{ key: uniqueKeyFor(TABLE), label: 'wnbaUsage:gamesBefore' },
|
||||
);
|
||||
}
|
||||
|
||||
/**
|
||||
* THE POINT-IN-TIME PROFILE the chainFn will consume.
|
||||
*
|
||||
* @returns {object|null} null when fewer than `minGames` games precede `asOf` —
|
||||
* never a league-average stand-in.
|
||||
*/
|
||||
async function profileAsOf(sb, { playerKey, asOf, season = null, minGames = MIN_GAMES, recent = 5 } = {}) {
|
||||
const games = await gamesBefore(sb, { playerKey, asOf, season });
|
||||
if (games.length < minGames) {
|
||||
return null;
|
||||
}
|
||||
// Most recent first — the recency window is the head of this list.
|
||||
games.sort((a, b) => String(b.game_date).localeCompare(String(a.game_date)));
|
||||
|
||||
const played = games.filter((g) => Number(g.minutes) > 0);
|
||||
const usable = played.filter((g) => g.usage_rate != null);
|
||||
const recentGames = played.slice(0, Math.max(1, recent));
|
||||
|
||||
// MINUTES-WEIGHTED usage, not a flat mean. A 6-minute cameo and a 34-minute
|
||||
// start are one row each; weighting by playing time is the same lesson the
|
||||
// lineup K-rate learned — an unweighted team aggregate counts a 12-PA callup
|
||||
// like an everyday starter, and it HURT the model until it was PA-weighted.
|
||||
const wUsage = (() => {
|
||||
const w = usable.reduce((a, g) => a + Number(g.minutes), 0);
|
||||
if (!(w > 0)) return null;
|
||||
return usable.reduce((a, g) => a + Number(g.usage_rate) * Number(g.minutes), 0) / w;
|
||||
})();
|
||||
|
||||
return {
|
||||
player_key: playerKey,
|
||||
as_of: asOf,
|
||||
// The cutoff is reported so a caller can see the read was bounded, and by what.
|
||||
games: games.length,
|
||||
games_played: played.length,
|
||||
last_game_date: games[0] ? games[0].game_date : null,
|
||||
|
||||
// ── the chainFn's three inputs ──
|
||||
usage_rate: round(wUsage, 3),
|
||||
usage_rate_recent: round(mean(recentGames.filter((g) => g.usage_rate != null).map((g) => Number(g.usage_rate))), 3),
|
||||
minutes_per_game: round(mean(played.map((g) => Number(g.minutes))), 2),
|
||||
minutes_recent: round(mean(recentGames.map((g) => Number(g.minutes))), 2),
|
||||
team_pace: round(mean(played.filter((g) => g.team_pace != null).map((g) => Number(g.team_pace))), 3),
|
||||
team_possessions: round(mean(played.filter((g) => g.team_possessions != null).map((g) => Number(g.team_possessions))), 3),
|
||||
ts_pct: round(mean(played.filter((g) => g.ts_pct != null).map((g) => Number(g.ts_pct))), 4),
|
||||
efg_pct: round(mean(played.filter((g) => g.efg_pct != null).map((g) => Number(g.efg_pct))), 4),
|
||||
|
||||
// ── game-state, for the redistribute hook ──
|
||||
starter_rate: round(played.length ? played.filter((g) => g.starter).length / played.length : null, 3),
|
||||
avg_final_margin: round(mean(played.filter((g) => g.final_margin != null).map((g) => Number(g.final_margin))), 2),
|
||||
|
||||
team: games[0] ? games[0].team : null,
|
||||
source: 'espn_wnba_boxscore',
|
||||
};
|
||||
}
|
||||
|
||||
/** Coverage of what is stored — rows, players, games, date range. */
|
||||
async function coverage(sb, { season = null } = {}) {
|
||||
if (!sb) return null;
|
||||
const rows = await paginate(
|
||||
() => {
|
||||
let q = sb.from(TABLE).select('game_id,source_id,player_key,game_date,season,usage_rate,team_pace,minutes')
|
||||
.eq('sport', 'wnba');
|
||||
if (season != null) q = q.eq('season', season);
|
||||
return q;
|
||||
},
|
||||
{ key: uniqueKeyFor(TABLE), label: 'wnbaUsage:coverage' },
|
||||
);
|
||||
if (!rows.length) return { rows: 0 };
|
||||
const dates = rows.map((r) => r.game_date).filter(Boolean).sort();
|
||||
return {
|
||||
rows: rows.length,
|
||||
players: new Set(rows.map((r) => r.player_key)).size,
|
||||
games: new Set(rows.map((r) => r.game_id)).size,
|
||||
from: dates[0],
|
||||
to: dates[dates.length - 1],
|
||||
with_usage: rows.filter((r) => r.usage_rate != null).length,
|
||||
with_pace: rows.filter((r) => r.team_pace != null).length,
|
||||
};
|
||||
}
|
||||
|
||||
module.exports = {
|
||||
ingestRange, profileAsOf, gamesBefore, coverage, dateET, eachDate,
|
||||
TABLE, MIN_GAMES,
|
||||
};
|
||||
@@ -38,6 +38,12 @@ const UNIQUE_KEY = Object.freeze({
|
||||
park_dimensions: Object.freeze(['as_of_date', 'sport', 'venue_id']),
|
||||
hitter_opportunity: Object.freeze(['as_of_date', 'sport', 'season', 'player_key']),
|
||||
lineup_context: Object.freeze(['as_of_date', 'sport', 'game_pk', 'player_key']),
|
||||
|
||||
// PER-GAME grain, not a dated snapshot. A completed box score is immutable, so
|
||||
// point-in-time here is `game_date < asOf` over facts that never change —
|
||||
// there is no aggregate being overwritten and therefore no history table to
|
||||
// pair it with. (migration 039)
|
||||
wnba_player_game: Object.freeze(['game_id', 'source_id']),
|
||||
});
|
||||
|
||||
/**
|
||||
|
||||
@@ -0,0 +1,40 @@
|
||||
-- Migration 038: model_snapshots.chain_shadow — the SHADOW CHAIN read.
|
||||
--
|
||||
-- The chain (src/services/model/chain.js) computes a probability by chaining
|
||||
-- base-event atoms: p per plate appearance from the PA outcome tree, chained
|
||||
-- over expected plate appearances. It is SERVED BY NOTHING. The counter
|
||||
-- (probabilityEstimator) remains the only thing a user ever sees.
|
||||
--
|
||||
-- WHY A COLUMN AT ALL. Evidence you cannot adjudicate is not evidence. A chain
|
||||
-- probability stored on its own could only ever be compared to itself. This
|
||||
-- column stores the chain's number NEXT TO the number it must beat, on a row
|
||||
-- that already carries the result:
|
||||
--
|
||||
-- chain_shadow.chain_p the chain's probability, SIDE-ALIGNED
|
||||
-- chain_shadow.counter_p the served counter's p_win for that same side
|
||||
-- model_snapshots.outcome written by the ordinary settle pass
|
||||
--
|
||||
-- Those three make the triple. Without all three on one row, the chain could be
|
||||
-- described but never judged.
|
||||
--
|
||||
-- SIDE ALIGNMENT IS LOAD-BEARING. The chain computes P(over the line); p_win is
|
||||
-- expressed for the graded side. An under row stores 1 - p_over. Storing the raw
|
||||
-- over-probability against an under row would invert every later comparison,
|
||||
-- silently.
|
||||
--
|
||||
-- IT IS LABELLED UN-SERVABLE IN THE DATA, not only in a comment: every block
|
||||
-- carries status 'UN-SERVABLE' and servable=false, because it was produced with
|
||||
-- requireCalibrated=false and the chain has never passed a calibration gate.
|
||||
--
|
||||
-- A SEPARATE COLUMN, not extra keys inside `features` — champion-ablation.js
|
||||
-- iterates every key of `features` for its residual scan, so widening it would
|
||||
-- silently enlarge that multiple-comparisons denominator. Same reasoning as 034.
|
||||
--
|
||||
-- NULLABLE and populated only on MLB `hits` rows the chain could read, so the
|
||||
-- storage cost is bounded. Adding a nullable column is metadata-only in
|
||||
-- Postgres — no table rewrite — which matters while this database sits near its
|
||||
-- free-tier size cap.
|
||||
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS chain_shadow jsonb;
|
||||
|
||||
COMMENT ON COLUMN model_snapshots.chain_shadow IS
|
||||
'Chain v1 SHADOW: side-aligned chain_p + the served counter_p, forming the (chain_p, counter_p, outcome) triple with model_snapshots.outcome. UN-SERVABLE (requireCalibrated=false); read by no serving path.';
|
||||
@@ -0,0 +1,106 @@
|
||||
-- Migration 039: wnba_player_game — the WNBA possession / usage / minutes feed.
|
||||
--
|
||||
-- The chain's basketball chainFn is `usage x possessions x efficiency`. None of
|
||||
-- those three existed for WNBA anywhere in this codebase; this is the feed that
|
||||
-- makes the slot fillable. It is scoped to exactly those inputs plus the
|
||||
-- game-state the `redistribute` hook reads — NOT a general WNBA stats dump.
|
||||
-- (closing_captures grew to 4.2M rows of something nothing read; the lesson is
|
||||
-- to ingest for a named consumer or not at all.)
|
||||
--
|
||||
-- ── POINT-IN-TIME IS STRUCTURAL, NOT A LATER RETROFIT ──────────────────────
|
||||
-- `statcast_aggregates` stores a SEASON AGGREGATE upserted in place, so every
|
||||
-- prior version is destroyed and a point-in-time question is unanswerable from
|
||||
-- it — which is why `statcast_history` had to be built afterwards, and why the
|
||||
-- skill backtest was honest only by accident.
|
||||
--
|
||||
-- This stores PER GAME rows. A completed box score never changes, so an as-of
|
||||
-- profile is `WHERE game_date < asOf` — a filter over immutable facts, not a
|
||||
-- dated snapshot of a mutable aggregate. There is nothing to overwrite, so there
|
||||
-- is nothing to retain a history OF.
|
||||
--
|
||||
-- STRICTLY `<`, never `<=`: a game played ON the as-of date may have tipped
|
||||
-- after grade time, and counting it would leak the evening being predicted into
|
||||
-- the prediction. Same rule as isPreGame in snapshotSettlementService.
|
||||
--
|
||||
-- ── THE KEY ────────────────────────────────────────────────────────────────
|
||||
-- (game_id, source_id) is unique: one row per player per game. `game_date` and
|
||||
-- `sport` lead the index because every read is date-bounded and single-sport,
|
||||
-- and safePaginate orders on the full tuple. Registered in src/utils/tableKeys.js
|
||||
-- — a stale entry there surfaces as a runtime duplicate-tuple throw.
|
||||
--
|
||||
-- ── DERIVED vs SERVED ──────────────────────────────────────────────────────
|
||||
-- ESPN returns the COMPONENTS (minutes, FGA, FTA, TOV, OREB and team totals),
|
||||
-- not the rates. `usage_rate`, `team_possessions`, `team_pace`, `ts_pct` and
|
||||
-- `efg_pct` are computed here from the standard box-score identities and stored
|
||||
-- so a chainFn read is one query. Every component is stored ALONGSIDE its
|
||||
-- derived value, so each rate is re-derivable and checkable rather than an
|
||||
-- unfalsifiable number (the factorFreeze rule: store inputs, not just outputs).
|
||||
CREATE TABLE IF NOT EXISTS wnba_player_game (
|
||||
sport text NOT NULL DEFAULT 'wnba',
|
||||
season integer,
|
||||
game_id text NOT NULL,
|
||||
game_date date NOT NULL,
|
||||
source_id text NOT NULL,
|
||||
player_key text NOT NULL,
|
||||
player_name text,
|
||||
team text,
|
||||
opponent text,
|
||||
home_away text,
|
||||
starter boolean,
|
||||
|
||||
-- raw box components (the usage/possession inputs)
|
||||
minutes numeric,
|
||||
points integer,
|
||||
fgm integer,
|
||||
fga integer,
|
||||
fg3m integer,
|
||||
fg3a integer,
|
||||
ftm integer,
|
||||
fta integer,
|
||||
reb integer,
|
||||
oreb integer,
|
||||
dreb integer,
|
||||
ast integer,
|
||||
tov integer,
|
||||
stl integer,
|
||||
blk integer,
|
||||
pf integer,
|
||||
plus_minus integer,
|
||||
|
||||
-- derived, per the identities in espnWnbaAdapter
|
||||
usage_rate numeric,
|
||||
ts_pct numeric,
|
||||
efg_pct numeric,
|
||||
|
||||
-- team possession context, denormalised so a chainFn read is ONE query
|
||||
team_minutes numeric,
|
||||
team_fga integer,
|
||||
team_fta integer,
|
||||
team_tov integer,
|
||||
team_oreb integer,
|
||||
team_possessions numeric,
|
||||
team_pace numeric,
|
||||
|
||||
-- game state — what the archetype `redistribute` hook reads (live in
|
||||
-- basketball: a blowout fades the star and feeds the bench)
|
||||
team_score integer,
|
||||
opp_score integer,
|
||||
final_margin integer,
|
||||
|
||||
ingested_at timestamptz NOT NULL DEFAULT now(),
|
||||
|
||||
PRIMARY KEY (game_id, source_id)
|
||||
);
|
||||
|
||||
-- Every read is "this player, before this date" or "this slate's date range".
|
||||
CREATE INDEX IF NOT EXISTS wnba_player_game_asof_idx
|
||||
ON wnba_player_game (sport, player_key, game_date DESC);
|
||||
CREATE INDEX IF NOT EXISTS wnba_player_game_date_idx
|
||||
ON wnba_player_game (sport, game_date);
|
||||
|
||||
COMMENT ON TABLE wnba_player_game IS
|
||||
'WNBA per-player-per-game possession/usage/minutes feed for the chain basketball chainFn. PER-GAME grain makes point-in-time native: an as-of profile is WHERE game_date < asOf (strictly), over immutable completed box scores. Scoped to chainFn inputs + redistribute game-state.';
|
||||
COMMENT ON COLUMN wnba_player_game.usage_rate IS
|
||||
'DERIVED (not served by ESPN): 100*((FGA+0.44*FTA+TOV)*(TmMIN/5))/(MIN*(TmFGA+0.44*TmFTA+TmTOV)). Components stored alongside so it is re-derivable.';
|
||||
COMMENT ON COLUMN wnba_player_game.team_possessions IS
|
||||
'DERIVED: FGA - OREB + TOV + 0.44*FTA. The 0.44 free-throw-trip coefficient is the only estimated term.';
|
||||
@@ -0,0 +1,394 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* BASEBALL'S chainFn + the shadow that runs it on the live board.
|
||||
*
|
||||
* Two things are being protected here. First, that the chain actually COMPUTES
|
||||
* a probability from base events rather than counting how often the player has
|
||||
* cleared this number — that is the whole difference between the chain and the
|
||||
* incumbent counter. Second, that none of it can reach a served grade.
|
||||
*/
|
||||
|
||||
const bc = require('../../src/services/model/baseballChain');
|
||||
const cs = require('../../src/services/model/chainShadow');
|
||||
const chain = require('../../src/services/model/chain');
|
||||
const sk = require('../../src/services/model/skillProjection');
|
||||
|
||||
/**
|
||||
* A real-shaped `statcast_aggregates` row. UNITS MATTER: the table stores
|
||||
* PERCENTAGES (0-100), and feeding fractions in is the mistake that made
|
||||
* `bip = 1 - k - bb` come out wrong and refused 568 of 576 rows on the first
|
||||
* Stage A run.
|
||||
*/
|
||||
const batterRow = (over = {}) => ({
|
||||
player_key: 'test batter',
|
||||
role: 'batter',
|
||||
k_pct: 22.2,
|
||||
bb_pct: 8.5,
|
||||
barrel_pct: 9.5,
|
||||
hard_hit_pct: 44.0,
|
||||
avg_exit_velo: 91.2,
|
||||
avg_launch_angle: 14.0,
|
||||
bats: 'R',
|
||||
sample_pa: 400,
|
||||
...over,
|
||||
});
|
||||
|
||||
const pitcherRow = (over = {}) => ({
|
||||
player_key: 'test pitcher',
|
||||
role: 'pitcher',
|
||||
source_id: 4242,
|
||||
k_pct: 29.6,
|
||||
bb_pct: 6.1,
|
||||
hard_hit_pct: 33.0,
|
||||
sample_ip: 120,
|
||||
...over,
|
||||
});
|
||||
|
||||
describe('BASEBALL chainFn — rate x opportunity, and opportunity is a LOOKUP', () => {
|
||||
it('produces a probability by chaining the PA rate over expected PA', () => {
|
||||
const out = bc.chainFn({ statType: 'hits', line: 0.5, batterRow: batterRow(), archetype: 'TORCH' });
|
||||
expect(out).not.toBeNull();
|
||||
expect(out.p).toBeGreaterThan(0);
|
||||
expect(out.p).toBeLessThan(1);
|
||||
// It is a CHAINED read, not a frequency table: the PA tree is on the row.
|
||||
expect(out.chain_per_pa.p_hit_per_pa).toBeGreaterThan(0);
|
||||
expect(out.chain_family).toBe('pa_outcome_tree_binomial');
|
||||
});
|
||||
|
||||
it('the probability is genuinely the binomial over the PA distribution', () => {
|
||||
// Recomputed independently from the same atoms — if these diverge, the
|
||||
// chainFn is doing something other than what it claims.
|
||||
const atom = { statType: 'hits', line: 0.5, batterRow: batterRow(), archetype: 'TORCH', expectedPa: 4 };
|
||||
const out = bc.chainFn(atom);
|
||||
const pa = sk.paOutcome({ batter: sk.fromStatcastRow(batterRow()), pitcher: null, park: 1, archetype: 'TORCH' });
|
||||
const pmf = sk.binomialPmf(4, pa.p_hit_per_pa);
|
||||
expect(out.p).toBeCloseTo(sk.atLeast(pmf, 1), 2);
|
||||
});
|
||||
|
||||
it('LINEUP SLOT is the opportunity term — leadoff outranks the nine hole', () => {
|
||||
const at = (slot) => bc.chainFn({
|
||||
statType: 'hits', line: 0.5, batterRow: batterRow(), archetype: 'TORCH', lineupSlot: slot,
|
||||
});
|
||||
const lead = at(1);
|
||||
const nine = at(9);
|
||||
expect(lead.chain_expected_pa).toBeGreaterThan(nine.chain_expected_pa);
|
||||
expect(lead.p).toBeGreaterThan(nine.p);
|
||||
expect(lead.chain_opportunity_source).toBe('lineup_slot');
|
||||
});
|
||||
|
||||
it('an unposted lineup falls to a REGULAR, and says which it used', () => {
|
||||
const out = bc.chainFn({ statType: 'hits', line: 0.5, batterRow: batterRow(), archetype: 'TORCH' });
|
||||
expect(out.chain_opportunity_source).toBe('default_regular');
|
||||
const explicit = bc.chainFn({ statType: 'hits', line: 0.5, batterRow: batterRow(), expectedPa: 5 });
|
||||
expect(explicit.chain_opportunity_source).toBe('explicit');
|
||||
});
|
||||
|
||||
it('the slot table is monotone — the decline is the documented structure', () => {
|
||||
const slots = Object.keys(bc.PA_BY_SLOT).map(Number).sort((a, b) => a - b);
|
||||
for (let i = 1; i < slots.length; i += 1) {
|
||||
expect(bc.PA_BY_SLOT[slots[i]]).toBeLessThan(bc.PA_BY_SLOT[slots[i - 1]]);
|
||||
}
|
||||
expect(bc.expectedPaForSlot(null)).toBeNull();
|
||||
expect(bc.expectedPaForSlot(12)).toBeNull();
|
||||
});
|
||||
|
||||
it('ARCHETYPE SELECTS THE FEATURES — the same hitter reads differently', () => {
|
||||
// A BOMBER's hits ride on barrels; a GHOST's ride on beating out grounders.
|
||||
// If these agreed, the archetype would be a label rather than a selector.
|
||||
const row = batterRow({ barrel_pct: 16.0, avg_launch_angle: 20.0 });
|
||||
const bomber = bc.chainFn({ statType: 'hits', line: 0.5, batterRow: row, archetype: 'BOMBER' });
|
||||
const ghost = bc.chainFn({ statType: 'hits', line: 0.5, batterRow: row, archetype: 'GHOST' });
|
||||
expect(Math.abs(bomber.p - ghost.p)).toBeGreaterThan(0.01);
|
||||
});
|
||||
|
||||
it('the OPPOSING PITCHER moves the read, and its absence leaves it untouched', () => {
|
||||
const solo = bc.chainFn({ statType: 'hits', line: 0.5, batterRow: batterRow(), archetype: 'TORCH' });
|
||||
const vsAce = bc.chainFn({
|
||||
statType: 'hits', line: 0.5, batterRow: batterRow(), archetype: 'TORCH', pitcherRow: pitcherRow(),
|
||||
});
|
||||
expect(solo.chain_pitcher_applied).toBe(false);
|
||||
expect(vsAce.chain_pitcher_applied).toBe(true);
|
||||
// A high-K, contact-suppressing arm lowers it.
|
||||
expect(vsAce.p).toBeLessThan(solo.p);
|
||||
});
|
||||
|
||||
it('NO BATTER PROFILE IS A REFUSAL, never a league hitter', () => {
|
||||
expect(bc.chainFn({ statType: 'hits', line: 0.5, batterRow: null })).toBeNull();
|
||||
expect(bc.chainFn({ statType: 'hits', line: 0.5, batterRow: {} })).toBeNull();
|
||||
});
|
||||
|
||||
it('refuses every stat but hits, so a muddy verdict is not widened into', () => {
|
||||
expect(bc.chainFn({ statType: 'total_bases', line: 1.5, batterRow: batterRow() })).toBeNull();
|
||||
expect(bc.chainFn({ statType: 'rbi', line: 0.5, batterRow: batterRow() })).toBeNull();
|
||||
expect(bc.chainFn({ statType: 'hits', line: null, batterRow: batterRow() })).toBeNull();
|
||||
});
|
||||
|
||||
it('plugs into the core as a chainFn with no special-casing in chain.js', () => {
|
||||
const atoms = [
|
||||
{ id: 'a', statType: 'hits', line: 0.5, batterRow: batterRow(), gameId: 'G1', calibrated: false },
|
||||
{ id: 'b', statType: 'hits', line: 0.5, batterRow: null, gameId: 'G1', calibrated: false },
|
||||
];
|
||||
const out = chain.chainAcross(atoms, { chainFn: bc.chainFn, requireCalibrated: false });
|
||||
expect(out.ok).toBe(true);
|
||||
expect(out.legs).toBe(1); // the unreadable atom was dropped, not zeroed
|
||||
expect(out.chain_fn_refused).toBe(1);
|
||||
});
|
||||
|
||||
it('redistribution is DORMANT in baseball — a nine-run lead changes no batting order', () => {
|
||||
const legs = [{ id: 'a', p: 0.5 }];
|
||||
expect(bc.redistribute(legs)).toBe(legs);
|
||||
});
|
||||
});
|
||||
|
||||
describe('THE HAND SPLIT — a matchup read, not a season read', () => {
|
||||
/** A real split: .284 vs LHP, .221 vs RHP, both sides well over the floor. */
|
||||
const SPLIT = {
|
||||
vl: { pa: 220, atBats: 200, hits: 57 }, // .285
|
||||
vr: { pa: 480, atBats: 430, hits: 95 }, // .221
|
||||
};
|
||||
const base = (over = {}) => ({
|
||||
statType: 'hits', line: 0.5, batterRow: batterRow(), archetype: 'TORCH', expectedPa: 4, ...over,
|
||||
});
|
||||
|
||||
it('WITHOUT a split the chain is arithmetically identical to projectSkill', () => {
|
||||
// The chain is written out in baseballChain so the split can enter between
|
||||
// the rate and the opportunity term. This is what stops that restructuring
|
||||
// from quietly becoming a second, different model.
|
||||
for (const archetype of ['TORCH', 'BOMBER', 'GHOST', null]) {
|
||||
for (const pa of [3, 4, 5]) {
|
||||
const mine = bc.chainFn(base({ archetype, expectedPa: pa }));
|
||||
const theirs = sk.projectSkill({
|
||||
batter: sk.fromStatcastRow(batterRow()), pitcher: null, park: 1,
|
||||
archetype, statType: 'hits', line: 0.5, expectedPa: pa,
|
||||
});
|
||||
expect(mine.p).toBeCloseTo(theirs.p_over_line, 3);
|
||||
expect(mine.chain_projected_value).toBeCloseTo(theirs.projected_value, 3);
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
it('FIRES when bats + throws + a thick split are all present', () => {
|
||||
const out = bc.chainFn(base({ bats: 'R', throws: 'L', platoonSplits: SPLIT }));
|
||||
expect(out.chain_platoon_applied).toBe(true);
|
||||
expect(out.chain_platoon_reason).toBeNull();
|
||||
expect(out.chain_platoon_multiplier).toBeGreaterThan(1); // RHB vs LHP = the edge
|
||||
expect(out.chain_platoon_observed_split).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
it('it enters at the RATE, not at the output probability', () => {
|
||||
// p_hit_per_pa moves; the season rate is preserved beside it so the
|
||||
// adjustment stays re-derivable.
|
||||
const out = bc.chainFn(base({ bats: 'R', throws: 'L', platoonSplits: SPLIT }));
|
||||
const season = out.chain_per_pa.p_hit_per_pa_season;
|
||||
const applied = out.chain_per_pa.p_hit_per_pa;
|
||||
expect(applied).not.toBeCloseTo(season, 6);
|
||||
expect(applied).toBeCloseTo(season * out.chain_platoon_multiplier, 3);
|
||||
});
|
||||
|
||||
it('and it runs in the DIRECTION tonight\'s matchup runs', () => {
|
||||
const edge = bc.chainFn(base({ bats: 'R', throws: 'L', platoonSplits: SPLIT }));
|
||||
const wrongSide = bc.chainFn(base({ bats: 'R', throws: 'R', platoonSplits: SPLIT }));
|
||||
expect(edge.p).toBeGreaterThan(wrongSide.p);
|
||||
});
|
||||
|
||||
it('REFUSES on a thin split — with the reason, not a league-typical guess', () => {
|
||||
const thin = { vl: { pa: 40, atBats: 36, hits: 11 }, vr: { pa: 480, atBats: 430, hits: 95 } };
|
||||
const out = bc.chainFn(base({ bats: 'R', throws: 'L', platoonSplits: thin }));
|
||||
expect(out.chain_platoon_applied).toBe(false);
|
||||
expect(out.chain_platoon_reason).toBe('insufficient_split_sample');
|
||||
// The season rate stands EXACTLY untouched — a refusal is not a nudge.
|
||||
expect(out.chain_per_pa.p_hit_per_pa).toBeCloseTo(out.chain_per_pa.p_hit_per_pa_season, 6);
|
||||
expect(out.p).toBeCloseTo(bc.chainFn(base()).p, 6); // a refusal is EXACT, not approximate
|
||||
});
|
||||
|
||||
it('a switch hitter is UNREADABLE, never credited with an automatic edge', () => {
|
||||
const out = bc.chainFn(base({ bats: 'S', throws: 'L', platoonSplits: SPLIT }));
|
||||
expect(out.chain_platoon_applied).toBe(false);
|
||||
expect(out.chain_platoon_reason).toBe('switch_hitter_side_value_unknown');
|
||||
});
|
||||
|
||||
it('names WHICH input is missing, so the fix is diagnosable', () => {
|
||||
const r = (over) => bc.chainFn(base(over)).chain_platoon_reason;
|
||||
expect(r({})).toBe('no_splits');
|
||||
expect(r({ platoonSplits: SPLIT, bats: 'R' })).toBe('no_pitcher_hand');
|
||||
expect(r({ platoonSplits: SPLIT, throws: 'L' })).toBe('no_batter_hand');
|
||||
});
|
||||
|
||||
it('THERE IS NO FALLBACK PATH — a refused read is dropped, never a season rate', () => {
|
||||
// The failure this guards is subtle: if an unreadable atom silently became a
|
||||
// season-rate read, the shadow would agree with the counter for a reason
|
||||
// that LOOKS like agreement, and the whole comparison would be worthless.
|
||||
expect(bc.chainFn(base({ batterRow: null }))).toBeNull();
|
||||
expect(bc.chainFn(base({ statType: 'rbi' }))).toBeNull();
|
||||
});
|
||||
|
||||
it('refusals are REPORTED to the caller with a reason', () => {
|
||||
const seen = [];
|
||||
const ctx = { onRefusal: (id, reason) => seen.push([id, reason]) };
|
||||
bc.chainFn({ ...base({ batterRow: null }), id: 'a' }, ctx);
|
||||
bc.chainFn({ ...base({ statType: 'rbi' }), id: 'b' }, ctx);
|
||||
bc.chainFn({ ...base({ line: null }), id: 'c' }, ctx);
|
||||
expect(seen).toEqual([['a', 'no_batter_profile'], ['b', 'stat_not_chained'], ['c', 'no_line']]);
|
||||
});
|
||||
|
||||
it('a throwing onRefusal never breaks the read', () => {
|
||||
expect(() => bc.chainFn({ ...base({ batterRow: null }) }, {
|
||||
onRefusal: () => { throw new Error('boom'); },
|
||||
})).not.toThrow();
|
||||
});
|
||||
});
|
||||
|
||||
describe('SHADOW — fire rate and matchup rate are DIFFERENT questions', () => {
|
||||
const SPLIT = { vl: { pa: 220, atBats: 200, hits: 57 }, vr: { pa: 480, atBats: 430, hits: 95 } };
|
||||
const grade = (over = {}) => ({
|
||||
player: 'Test Batter', stat_type: 'hits', direction: 'over', p_win: 0.62,
|
||||
game_id: 'G1', team: 'New York Yankees', archetype: 'TORCH',
|
||||
gradedAt: { line: 0.5 }, ...over,
|
||||
});
|
||||
const deps = (over = {}) => ({
|
||||
statcastByKey: new Map([['test batter', batterRow()]]),
|
||||
pitcherRowFor: () => pitcherRow({ throws: 'L' }),
|
||||
...over,
|
||||
});
|
||||
|
||||
it('counts the hand split separately from the fire rate', () => {
|
||||
const withSplit = cs.runShadow([grade()], deps({
|
||||
handSplitFor: () => ({ bats: 'R', throws: 'L', platoonSplits: SPLIT }),
|
||||
}));
|
||||
expect(withSplit.summary.readable).toBe(1);
|
||||
expect(withSplit.summary.platoon_applied).toBe(1);
|
||||
|
||||
const without = cs.runShadow([grade()], deps());
|
||||
expect(without.summary.readable).toBe(1); // it still FIRES...
|
||||
expect(without.summary.platoon_applied).toBe(0); // ...but on the SEASON rate
|
||||
expect(without.summary.platoon_reasons.no_splits).toBe(1);
|
||||
});
|
||||
|
||||
it('a refused atom is counted ONCE, not once per reading', () => {
|
||||
// chainAcross, chainUp and prepareAtoms each apply the chainFn. Running raw
|
||||
// atoms through all three would report a 3x refusal count.
|
||||
const out = cs.runShadow([grade({ player: 'Nobody' })], deps());
|
||||
expect(out.summary.refused).toBe(1);
|
||||
expect(out.summary.refusal_reasons).toEqual({ no_batter_profile: 1 });
|
||||
});
|
||||
|
||||
it('the stored block records the split state either way', () => {
|
||||
const on = [...cs.runShadow([grade()], deps({
|
||||
handSplitFor: () => ({ bats: 'R', throws: 'L', platoonSplits: SPLIT }),
|
||||
})).byKey.values()][0];
|
||||
expect(on.platoon_applied).toBe(true);
|
||||
expect(on.platoon_multiplier).toBeGreaterThan(1);
|
||||
|
||||
const off = [...cs.runShadow([grade()], deps()).byKey.values()][0];
|
||||
expect(off.platoon_applied).toBe(false);
|
||||
expect(off.platoon_reason).toBe('no_splits');
|
||||
expect(off.platoon_multiplier).toBeNull();
|
||||
});
|
||||
|
||||
it('a throwing handSplitFor degrades to a season read, never a broken slate', () => {
|
||||
const out = cs.runShadow([grade()], deps({
|
||||
handSplitFor: () => { throw new Error('boom'); },
|
||||
}));
|
||||
expect(out.summary.readable).toBe(1);
|
||||
expect(out.summary.platoon_applied).toBe(0);
|
||||
});
|
||||
});
|
||||
|
||||
describe('CHAIN SHADOW — runs on the board, reaches nothing', () => {
|
||||
const grade = (over = {}) => ({
|
||||
player: 'Test Batter',
|
||||
stat_type: 'hits',
|
||||
direction: 'over',
|
||||
p_win: 0.62,
|
||||
grade: 'B',
|
||||
game_id: 'mlb:2026-08-12:NYY@BOS',
|
||||
team: 'New York Yankees',
|
||||
archetype: 'TORCH',
|
||||
gradedAt: { line: 0.5, odds: -130, timestamp: '2026-08-12T18:00:00Z' },
|
||||
...over,
|
||||
});
|
||||
|
||||
const deps = () => ({
|
||||
statcastByKey: new Map([['test batter', batterRow()]]),
|
||||
pitcherRowFor: () => pitcherRow(),
|
||||
});
|
||||
|
||||
it('produces the (chain_p, counter_p) pair the triple needs', () => {
|
||||
const out = cs.runShadow([grade()], deps());
|
||||
const block = [...out.byKey.values()][0];
|
||||
expect(block).toBeTruthy();
|
||||
const aligned = cs.alignToSide(block, 'over', 0.62);
|
||||
expect(aligned.chain_p).toBeGreaterThan(0);
|
||||
expect(aligned.counter_p).toBeCloseTo(0.62, 6);
|
||||
expect(aligned.divergence).toBeCloseTo(aligned.chain_p - 0.62, 4);
|
||||
});
|
||||
|
||||
it('EVERY block declares itself UN-SERVABLE in the data, not just in a comment', () => {
|
||||
const out = cs.runShadow([grade()], deps());
|
||||
for (const block of out.byKey.values()) {
|
||||
expect(block.servable).toBe(false);
|
||||
expect(block.status).toBe('UN-SERVABLE');
|
||||
expect(block.servable_reason).toMatch(/calibration gate/);
|
||||
}
|
||||
});
|
||||
|
||||
it('SIDE ALIGNMENT: an under row carries 1 - p_over, or every later read inverts', () => {
|
||||
const out = cs.runShadow([grade({ direction: 'under', p_win: 0.38 })], deps());
|
||||
const block = [...out.byKey.values()][0];
|
||||
const over = cs.alignToSide(block, 'over', 0.62);
|
||||
const under = cs.alignToSide(block, 'under', 0.38);
|
||||
expect(over.chain_p + under.chain_p).toBeCloseTo(1, 4);
|
||||
expect(under.side).toBe('under');
|
||||
});
|
||||
|
||||
it('scores against the LOCKED line, not the current one', () => {
|
||||
// Scoring the chain on a line the counter never saw would compare two
|
||||
// forecasts of two different questions.
|
||||
const g = grade({ line: 1.5, gradedAt: { line: 0.5 } });
|
||||
const atoms = cs.buildAtoms([g], deps());
|
||||
expect(atoms[0].line).toBe(0.5);
|
||||
});
|
||||
|
||||
it('ignores every stat but hits, and any row with no locked line', () => {
|
||||
const rows = [
|
||||
grade({ stat_type: 'total_bases' }),
|
||||
grade({ stat_type: 'rbi' }),
|
||||
grade({ gradedAt: null, line: null }),
|
||||
];
|
||||
expect(cs.buildAtoms(rows, deps())).toHaveLength(0);
|
||||
});
|
||||
|
||||
it('a player with no statcast profile yields NO BLOCK — absent, not zero', () => {
|
||||
const out = cs.runShadow([grade({ player: 'Unknown Guy' })], deps());
|
||||
expect(out.byKey.size).toBe(0);
|
||||
expect(out.summary.refused).toBe(1);
|
||||
expect(out.summary.atoms).toBe(1); // it was attempted, and that is visible
|
||||
});
|
||||
|
||||
it('an absent block never fabricates a triple', () => {
|
||||
expect(cs.alignToSide(null, 'over', 0.5)).toBeNull();
|
||||
expect(cs.alignToSide({ chain_p_over: null }, 'over', 0.5)).toBeNull();
|
||||
expect(cs.alignToSide({ chain_p_over: 0.4 }, '', 0.5)).toBeNull();
|
||||
// An absent counter is absent — not zero, which would read as a 40pt edge.
|
||||
expect(cs.alignToSide({ chain_p_over: 0.4 }, 'over', null).counter_p).toBeNull();
|
||||
expect(cs.alignToSide({ chain_p_over: 0.4 }, 'over', null).divergence).toBeNull();
|
||||
});
|
||||
|
||||
it('runs BOTH readings and reports the self-check as VACUOUS today', () => {
|
||||
// Both readings derive from the same atoms because no independent team read
|
||||
// exists. Agreement is arithmetic. Labelling it is the honest move.
|
||||
const out = cs.runShadow([grade(), grade({ player: 'Test Batter', stat_type: 'hits' })], deps());
|
||||
const block = [...out.byKey.values()][0];
|
||||
expect(block.game.up_expected_hits).toBeGreaterThan(0);
|
||||
expect(block.game.across_compound).toBeGreaterThan(0);
|
||||
expect(block.game.self_check.vacuous).toBe(true);
|
||||
expect(block.game.self_check.vacuous_reason).toMatch(/same atoms/);
|
||||
});
|
||||
|
||||
it('an empty slate is an empty result, never a thrown snapshot', () => {
|
||||
const out = cs.runShadow([], deps());
|
||||
expect(out.byKey.size).toBe(0);
|
||||
expect(out.summary.atoms).toBe(0);
|
||||
});
|
||||
});
|
||||
@@ -61,6 +61,167 @@ describe('CHAIN ACROSS — the calibration gate is structural', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('THE chainFn SLOT — the stage the header described and the code lacked', () => {
|
||||
it('defaults to identity-on-p, so every pre-existing caller is unchanged', () => {
|
||||
const out = chain.chainAcross([atom('a', 0.8), atom('b', 0.5)]);
|
||||
expect(out.independent_probability).toBeCloseTo(0.4, 6);
|
||||
expect(out.chain_fn_applied).toBe(false);
|
||||
});
|
||||
|
||||
it('computes the per-entity probability from the atom + context when supplied', () => {
|
||||
// The sport-specific work is an INPUT, not a code path: rate x opportunity.
|
||||
const atoms = [
|
||||
{ id: 'a', rate: 0.25, pa: 4, calibrated: true, gameId: 'G1' },
|
||||
{ id: 'b', rate: 0.30, pa: 4, calibrated: true, gameId: 'G2' },
|
||||
];
|
||||
const chainFn = (at) => 1 - (1 - at.rate) ** at.pa;
|
||||
const out = chain.chainAcross(atoms, { chainFn });
|
||||
expect(out.chain_fn_applied).toBe(true);
|
||||
const pa = 1 - 0.75 ** 4;
|
||||
const pb = 1 - 0.70 ** 4;
|
||||
expect(out.independent_probability).toBeCloseTo(pa * pb, 4);
|
||||
});
|
||||
|
||||
it('reads context, so slate-level modifiers reach the atom', () => {
|
||||
const chainFn = (at, ctx) => at.rate * (ctx.parkBoost || 1);
|
||||
const legs = [{ id: 'a', rate: 0.4, calibrated: true }];
|
||||
const plain = chain.chainUp(legs, { chainFn });
|
||||
const boosted = chain.chainUp(legs, { chainFn, context: { parkBoost: 1.5 } });
|
||||
expect(plain.expected_value).toBeCloseTo(0.4, 6);
|
||||
expect(boosted.expected_value).toBeCloseTo(0.6, 6);
|
||||
});
|
||||
|
||||
it('an atom the chainFn cannot read is DROPPED and COUNTED, never p=0', () => {
|
||||
// A zero leg would zero an entire ticket, and "we could not read him" is not
|
||||
// "he cannot do it". The count is what stops a silently-failing chainFn from
|
||||
// looking like a thin slate.
|
||||
const chainFn = (at) => (at.id === 'b' ? null : 0.8);
|
||||
const out = chain.chainAcross([atom('a', 0), atom('b', 0)], { chainFn });
|
||||
expect(out.legs).toBe(1);
|
||||
expect(out.chain_fn_refused).toBe(1);
|
||||
expect(out.compound_probability).toBeCloseTo(0.8, 6);
|
||||
});
|
||||
|
||||
it('a THROWING chainFn refuses that atom rather than breaking the read', () => {
|
||||
const chainFn = (at) => { if (at.id === 'b') throw new Error('no profile'); return 0.5; };
|
||||
const out = chain.chainAcross([atom('a', 0), atom('b', 0)], { chainFn });
|
||||
expect(out.ok).toBe(true);
|
||||
expect(out.chain_fn_refused).toBe(1);
|
||||
});
|
||||
|
||||
it('a chainFn returning an object merges its metadata onto the atom', () => {
|
||||
const out = chain.chainUp([{ id: 'a', calibrated: true }], {
|
||||
chainFn: () => ({ p: 0.3, weight: 2 }),
|
||||
});
|
||||
expect(out.expected_value).toBeCloseTo(0.6, 6);
|
||||
});
|
||||
|
||||
it('refusing EVERY atom is an explicit refusal carrying the count', () => {
|
||||
const out = chain.chainAcross([atom('a', 0.5), atom('b', 0.5)], { chainFn: () => null });
|
||||
expect(out.ok).toBe(false);
|
||||
expect(out.reason).toBe('no_usable_atoms');
|
||||
expect(out.chain_fn_refused).toBe(2);
|
||||
});
|
||||
});
|
||||
|
||||
describe('CORRELATION IS SIGNED — basketball is not baseball with a hook', () => {
|
||||
const pair = (p) => [atom('a', p, { gameId: 'G1' }), atom('b', p, { gameId: 'G1' })];
|
||||
|
||||
it('zero correlation is exactly the independent product', () => {
|
||||
const out = chain.chainAcross(pair(0.7), { correlation: () => 0 });
|
||||
expect(out.compound_probability).toBeCloseTo(0.49, 4);
|
||||
expect(out.correlation_direction).toBe('independent');
|
||||
});
|
||||
|
||||
it('POSITIVE moves the joint toward the weakest leg — unchanged from before', () => {
|
||||
const out = chain.chainAcross(pair(0.7), { correlation: () => 0.6 });
|
||||
expect(out.compound_probability).toBeGreaterThan(0.49);
|
||||
// Fréchet upper bound: two co-monotone 0.7s land together at most 0.7.
|
||||
expect(out.compound_probability).toBeLessThanOrEqual(0.7);
|
||||
expect(out.correlation_direction).toBe('toward_joint');
|
||||
});
|
||||
|
||||
it('NEGATIVE moves the joint APART — the case the old clamp made inexpressible', () => {
|
||||
// Teammates competing for finite possessions: one player's shot is another
|
||||
// player's non-shot, so they land together LESS often than independence says.
|
||||
const out = chain.chainAcross(pair(0.7), { correlation: () => -0.6 });
|
||||
expect(out.compound_probability).toBeLessThan(0.49);
|
||||
expect(out.correlation_direction).toBe('apart');
|
||||
});
|
||||
|
||||
it('the extremes are the FRÉCHET BOUNDS, which is why the direction is principled', () => {
|
||||
const hi = chain.chainAcross(pair(0.7), { correlation: () => 1 });
|
||||
const lo = chain.chainAcross(pair(0.7), { correlation: () => -1 });
|
||||
expect(hi.compound_probability).toBeCloseTo(0.7, 4); // min(p_i)
|
||||
expect(lo.compound_probability).toBeCloseTo(0.4, 4); // max(0, Sum p - (n-1))
|
||||
});
|
||||
|
||||
it('the lower bound never goes below zero', () => {
|
||||
const out = chain.chainAcross(
|
||||
[atom('a', 0.2, { gameId: 'G1' }), atom('b', 0.3, { gameId: 'G1' })],
|
||||
{ correlation: () => -1 },
|
||||
);
|
||||
expect(out.compound_probability).toBe(0);
|
||||
});
|
||||
|
||||
it('a correlation outside [-1,1] is clamped, not trusted', () => {
|
||||
const wild = chain.chainAcross(pair(0.7), { correlation: () => -50 });
|
||||
const unit = chain.chainAcross(pair(0.7), { correlation: () => -1 });
|
||||
expect(wild.compound_probability).toBeCloseTo(unit.compound_probability, 6);
|
||||
});
|
||||
});
|
||||
|
||||
describe('REDISTRIBUTE REACHES BOTH READINGS — or selfCheck lies about itself', () => {
|
||||
const blowout = (as, ctx) => (ctx.blowout
|
||||
? as.map((a) => (a.id === 'star' ? { ...a, p: a.p * 0.6 } : { ...a, p: a.p * 2 }))
|
||||
: as);
|
||||
|
||||
it('chainAcross now takes the hook chainUp always had', () => {
|
||||
const legs = [atom('star', 0.6, { gameId: 'G1' }), atom('bench', 0.1, { gameId: 'G1' })];
|
||||
const out = chain.chainAcross(legs, { redistribute: blowout, context: { blowout: true } });
|
||||
expect(out.redistributed).toBe(true);
|
||||
expect(out.independent_probability).toBeCloseTo(0.36 * 0.2, 6);
|
||||
});
|
||||
|
||||
it('and stays dormant unless a sport supplies it', () => {
|
||||
const legs = [atom('star', 0.6, { gameId: 'G1' }), atom('bench', 0.1, { gameId: 'G1' })];
|
||||
expect(chain.chainAcross(legs).redistributed).toBe(false);
|
||||
});
|
||||
|
||||
it('THE POINT: a redistribution applied to both readings does not false-flag', () => {
|
||||
// Before, redistribute reached only the team read. The across-read would
|
||||
// then be built from pre-redistribution atoms, and selfCheck would report an
|
||||
// INTERNAL_INCONSISTENCY the model had itself just manufactured.
|
||||
const legs = [atom('star', 0.6, { gameId: 'G1' }), atom('bench', 0.1, { gameId: 'G1' })];
|
||||
const opts = { redistribute: blowout, context: { blowout: true } };
|
||||
|
||||
const prep = chain.prepareAtoms(legs, opts);
|
||||
const up = chain.chainUp(legs, opts);
|
||||
const check = chain.selfCheck({ perEntity: prep.legs, teamRead: up.expected_value });
|
||||
|
||||
expect(prep.redistributed).toBe(true);
|
||||
expect(check.confidence).toBe('NORMAL');
|
||||
expect(check.flags).toEqual([]);
|
||||
});
|
||||
|
||||
it('prepareAtoms is the shared preparation — same legs, both readings', () => {
|
||||
const legs = [atom('star', 0.6), atom('bench', 0.1)];
|
||||
const opts = { redistribute: blowout, context: { blowout: true } };
|
||||
const prep = chain.prepareAtoms(legs, opts);
|
||||
expect(prep.legs.map((l) => l.p)).toEqual([0.36, 0.2]);
|
||||
// chainAcross and chainUp must see EXACTLY these.
|
||||
expect(chain.chainUp(legs, opts).expected_value).toBeCloseTo(0.56, 6);
|
||||
expect(chain.chainAcross(legs, opts).independent_probability).toBeCloseTo(0.072, 6);
|
||||
});
|
||||
|
||||
it('a redistribute that returns nothing usable leaves the legs alone', () => {
|
||||
const legs = [atom('a', 0.5)];
|
||||
const out = chain.chainUp(legs, { redistribute: () => [] });
|
||||
expect(out.redistributed).toBe(false);
|
||||
expect(out.expected_value).toBeCloseTo(0.5, 6);
|
||||
});
|
||||
});
|
||||
|
||||
describe('CHAIN UP — same atoms, team reading', () => {
|
||||
it('sums atoms into an expected value', () => {
|
||||
const out = chain.chainUp([atom('a', 0.3), atom('b', 0.4), atom('c', 0.5)]);
|
||||
|
||||
@@ -0,0 +1,467 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* CHAIN v1 — does the shadow accrue on the SCHEDULED path, and does it leak?
|
||||
*
|
||||
* Two questions, and the second one is the gate.
|
||||
*
|
||||
* A5's whole finding was a factor that was built, correct, and never invoked.
|
||||
* A6's answer was to drive the real path rather than read the code. Same here:
|
||||
* this runs the REAL `snapshotService.runSnapshot` with injected deps and checks
|
||||
* that a `chain_shadow` block lands on a retention row — because evidence that
|
||||
* is generated and never persisted is that same failure one layer along.
|
||||
*
|
||||
* Then it proves the thing that actually matters: the served `p_win` is
|
||||
* BYTE-IDENTICAL with the shadow on and off. A shadow that moves a served number
|
||||
* is not a shadow, and no amount of labelling fixes it.
|
||||
*/
|
||||
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const svc = require('../../src/services/snapshotService');
|
||||
const retention = require('../../src/services/retentionService');
|
||||
const cs = require('../../src/services/model/chainShadow');
|
||||
const { nameKey } = require('../../src/utils/playerName');
|
||||
|
||||
const ROOT = path.join(__dirname, '..', '..');
|
||||
|
||||
/** A chainable Supabase stub — the context builders are gated on a client. */
|
||||
function fakeSbTop() {
|
||||
const q = {};
|
||||
for (const m of ['select', 'eq', 'lte', 'gte', 'in', 'order', 'not', 'is']) q[m] = () => q;
|
||||
q.range = async () => ({ data: [], error: null });
|
||||
return { from: () => q };
|
||||
}
|
||||
|
||||
function memCache() {
|
||||
const store = {};
|
||||
return {
|
||||
store,
|
||||
cacheGet: async (k) => (k in store ? store[k] : null),
|
||||
cacheSet: async (k, v) => { store[k] = v; },
|
||||
};
|
||||
}
|
||||
|
||||
const PROPS = [
|
||||
{ player: 'Mookie Betts', stat_type: 'hits', line: 0.5, over_odds: -130, under_odds: 100, book: 'dk' },
|
||||
{ player: 'Aaron Judge', stat_type: 'hits', line: 0.5, over_odds: -140, under_odds: 110, book: 'dk' },
|
||||
];
|
||||
|
||||
/** Served grades, with the counter's p_win — the number that must not move. */
|
||||
const GRADES = [
|
||||
{
|
||||
player: 'Mookie Betts', stat_type: 'hits', line: 0.5, direction: 'over',
|
||||
grade: 'B', confidence: 61, p_win: 0.61, game_id: 'mlb:2026-08-12:LAD@SFG',
|
||||
},
|
||||
{
|
||||
player: 'Aaron Judge', stat_type: 'hits', line: 0.5, direction: 'under',
|
||||
grade: 'C', confidence: 42, p_win: 0.42, game_id: 'mlb:2026-08-12:NYY@BOS',
|
||||
},
|
||||
];
|
||||
|
||||
const statcastRow = (key, over = {}) => ({
|
||||
player_key: key, role: 'batter',
|
||||
k_pct: 20.0, bb_pct: 9.0, barrel_pct: 10.0, hard_hit_pct: 43.0,
|
||||
avg_exit_velo: 90.5, avg_launch_angle: 13.0, bats: 'R', sample_pa: 420,
|
||||
...over,
|
||||
});
|
||||
|
||||
const STATCAST = new Map([
|
||||
[nameKey('Mookie Betts'), statcastRow(nameKey('Mookie Betts'))],
|
||||
[nameKey('Aaron Judge'), statcastRow(nameKey('Aaron Judge'), { barrel_pct: 16.0 })],
|
||||
]);
|
||||
|
||||
/** Retention stub that captures rows instead of writing them. */
|
||||
function captureRetention(sink) {
|
||||
return {
|
||||
...retention,
|
||||
persist: async (rows) => { sink.push(...rows); return { attempted: rows.length, written: rows.length }; },
|
||||
};
|
||||
}
|
||||
|
||||
/**
|
||||
* The challenger pass fetches park/weather and probable pitchers over the
|
||||
* network. Stubbed to no-ops so this suite is hermetic — the shadow reuses the
|
||||
* SAME statcast handle either way, which is the coupling under test.
|
||||
*/
|
||||
const inertChallengerDeps = () => ({
|
||||
notify: async () => {},
|
||||
gameBinder: { attachGameTimes: async () => ({ bound: 0, alreadyHad: 2, unresolved: 0, ambiguous: 0 }) },
|
||||
environmentContext: {
|
||||
buildContext: async () => ({
|
||||
_internals: { oppPitcherByTeam: new Map() },
|
||||
contextFor: () => null,
|
||||
stats: {},
|
||||
}),
|
||||
},
|
||||
challenger: { attachChallenger: async (g) => g },
|
||||
archetypeAxes: { classifyPlayer: () => null },
|
||||
contactChallenger: { buildRefs: () => ({}), attachContactChallenger: async (g) => g },
|
||||
projectionChallenger: { attachProjection: async (g) => g },
|
||||
loadArsenals: async () => new Map(),
|
||||
mlbAdapter: { getPlayerGameLog: async () => [] },
|
||||
lineupContext: { refreshContext: async () => ({ lineups: 0, opportunity: 0 }) },
|
||||
ledger: { recordPipelineGrades: async () => ({ written: 0 }), captureClosing: async () => {} },
|
||||
buildEspnIndex: async () => ({}),
|
||||
captureBookPrices: () => ({}),
|
||||
lockLineCapture: { buildLockRows: () => [], persist: async () => ({ retired: true, attempted: 0 }) },
|
||||
});
|
||||
|
||||
function deps(cache, sink, over = {}) {
|
||||
return {
|
||||
...inertChallengerDeps(),
|
||||
getOdds: async () => ({ sport: 'mlb', props: PROPS, provider: 'propline' }),
|
||||
gradeAndCacheSlate: async (_sport, props, opts) => {
|
||||
// The REAL collector path: rowsFromSides is what writes model_snapshots.
|
||||
for (const base of props) {
|
||||
const sides = GRADES.filter((g) => g.player === base.player)
|
||||
.map((g) => ({ ...g, stat_type: g.stat_type, line: g.line }));
|
||||
if (opts.onGraded) opts.onGraded({ ...base, game_id: sides[0] && sides[0].game_id }, sides);
|
||||
}
|
||||
await opts.cacheSet('grades:x', { grades: GRADES, updated_at: opts.now(), source: 'test' });
|
||||
},
|
||||
resolveStats: async () => ({ found: true, classifierInput: { hr: 30, avg: 0.29, ops: 0.93, k_rate: 20 } }),
|
||||
classify: require('../../src/services/archetypeService').classify,
|
||||
retention: captureRetention(sink),
|
||||
loadStatcast: async () => STATCAST,
|
||||
cacheGet: cache.cacheGet,
|
||||
cacheSet: cache.cacheSet,
|
||||
now: () => '2026-08-12T18:00:00Z',
|
||||
nowMs: () => 1000,
|
||||
...over,
|
||||
};
|
||||
}
|
||||
|
||||
describe('the chain shadow accrues on the REAL scheduled path', () => {
|
||||
it('a graded hits row lands with chain_shadow on it', async () => {
|
||||
const sink = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), sink));
|
||||
const row = sink.find((r) => r.chain_shadow);
|
||||
expect(row).toBeDefined();
|
||||
expect(row.chain_shadow.chain_p).toBeGreaterThan(0);
|
||||
expect(row.chain_shadow.chain_p).toBeLessThan(1);
|
||||
});
|
||||
|
||||
it('THE TRIPLE: chain_p and counter_p on the same row, both side-aligned', async () => {
|
||||
const sink = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), sink));
|
||||
for (const r of sink.filter((x) => x.chain_shadow)) {
|
||||
expect(r.chain_shadow.side).toBe(r.side);
|
||||
expect(r.chain_shadow.counter_p).toBe(r.p_win);
|
||||
expect(r.chain_shadow.divergence)
|
||||
.toBeCloseTo(r.chain_shadow.chain_p - r.chain_shadow.counter_p, 4);
|
||||
}
|
||||
// `outcome` is the third leg and is written by the ordinary settle pass —
|
||||
// it is null here because these games have not been played.
|
||||
expect(sink.every((r) => r.outcome === undefined || r.outcome === null)).toBe(true);
|
||||
});
|
||||
|
||||
it('an UNDER row carries 1 - p_over, so no later comparison inverts', async () => {
|
||||
const sink = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), sink));
|
||||
const pair = sink.filter((r) => r.player_key === nameKey('Aaron Judge') && r.chain_shadow);
|
||||
const over = pair.find((r) => r.side === 'over');
|
||||
const under = pair.find((r) => r.side === 'under');
|
||||
if (over && under) expect(over.chain_shadow.chain_p + under.chain_shadow.chain_p).toBeCloseTo(1, 4);
|
||||
});
|
||||
|
||||
it('every stored block declares itself UN-SERVABLE', async () => {
|
||||
const sink = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), sink));
|
||||
for (const r of sink.filter((x) => x.chain_shadow)) {
|
||||
expect(r.chain_shadow.servable).toBe(false);
|
||||
expect(r.chain_shadow.status).toBe('UN-SERVABLE');
|
||||
}
|
||||
});
|
||||
|
||||
it('the run reports its own counts, so a shadow reading NOTHING is visible', async () => {
|
||||
const r = await svc.runSnapshot('mlb', deps(memCache(), []));
|
||||
expect(r.chainShadow).toBeTruthy();
|
||||
expect(r.chainShadow.atoms).toBeGreaterThan(0);
|
||||
expect(r.chainShadow.servable).toBe(false);
|
||||
});
|
||||
|
||||
it('no statcast profiles => no shadow, and the snapshot still succeeds', async () => {
|
||||
const sink = [];
|
||||
const r = await svc.runSnapshot('mlb', deps(memCache(), sink, { loadStatcast: async () => new Map() }));
|
||||
expect(r.status).toBe('ok');
|
||||
expect(sink.every((x) => !x.chain_shadow)).toBe(true);
|
||||
});
|
||||
|
||||
it('a THROWING shadow never breaks the snapshot', async () => {
|
||||
const sink = [];
|
||||
const r = await svc.runSnapshot('mlb', deps(memCache(), sink, {
|
||||
chainShadow: { runShadow: () => { throw new Error('boom'); } },
|
||||
}));
|
||||
expect(r.status).toBe('ok');
|
||||
expect(r.gradeCount).toBe(2);
|
||||
expect(r.chainShadow).toBeNull();
|
||||
});
|
||||
});
|
||||
|
||||
describe('CHAIN v2 — the hand split reaches the shadow through the REAL path', () => {
|
||||
const SPLIT = { vl: { pa: 220, atBats: 200, hits: 57 }, vr: { pa: 480, atBats: 430, hits: 95 } };
|
||||
|
||||
/**
|
||||
* The A4 context resolver and the A6 key resolver, in miniature. The shadow
|
||||
* composes them: keys give the opposing starter's hand, the context gives the
|
||||
* hitter's own split. WITHOUT the keys the resolver's `throws` is null — the
|
||||
* A5 finding — and the split silently never fires, which is the exact failure
|
||||
* this order exists inside.
|
||||
*/
|
||||
const factorContext = (prop, sport, keys) => ({
|
||||
spray: null, positionOaa: null, bats: 'R',
|
||||
throws: keys && keys.opposing_pitcher ? 'L' : null,
|
||||
pitcherHardHit: null, platoonSplits: SPLIT, as_of: '2026-08-12',
|
||||
});
|
||||
const KEYS = { opponent: 'Boston Red Sox', opposing_pitcher: 'Chris Sale', source: 'lineup_context+schedule', refused: null };
|
||||
|
||||
/** A chainable Supabase stub — the context builders are gated on a client. */
|
||||
const fakeSb = () => {
|
||||
const q = {};
|
||||
for (const m of ['select', 'eq', 'lte', 'gte', 'in', 'order', 'not', 'is']) q[m] = () => q;
|
||||
q.range = async () => ({ data: [], error: null });
|
||||
q.then = undefined;
|
||||
return { from: () => q };
|
||||
};
|
||||
|
||||
it('FIRES when both the context and the join keys are present', async () => {
|
||||
const sink = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), sink, {
|
||||
supabase: fakeSb(),
|
||||
hitsFactorContext: { build: async () => factorContext },
|
||||
matchupKeys: { build: async () => Object.assign(() => KEYS, { stats: {} }) },
|
||||
}));
|
||||
const row = sink.find((r) => r.chain_shadow);
|
||||
expect(row).toBeDefined();
|
||||
expect(row.chain_shadow.platoon_applied).toBe(true);
|
||||
expect(row.chain_shadow.platoon_multiplier).toBeGreaterThan(0);
|
||||
expect(row.chain_shadow.platoon_reason).toBeNull();
|
||||
});
|
||||
|
||||
it('WITHOUT the join keys it falls to a SEASON read — and says which', async () => {
|
||||
// This is the A5 shape reproduced deliberately: the context loads, the
|
||||
// hitter's own split is right there, and the pitcher hand is missing, so
|
||||
// the matchup conditioner cannot run. It must be visible, not silent.
|
||||
const sink = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), sink, {
|
||||
supabase: fakeSb(),
|
||||
hitsFactorContext: { build: async () => factorContext },
|
||||
matchupKeys: { build: async () => null },
|
||||
}));
|
||||
const row = sink.find((r) => r.chain_shadow);
|
||||
expect(row.chain_shadow.platoon_applied).toBe(false);
|
||||
expect(row.chain_shadow.platoon_reason).toBe('no_pitcher_hand');
|
||||
});
|
||||
|
||||
it('the run reports the matchup rate SEPARATELY from the fire rate', async () => {
|
||||
const r = await svc.runSnapshot('mlb', deps(memCache(), [], {
|
||||
supabase: fakeSb(),
|
||||
hitsFactorContext: { build: async () => factorContext },
|
||||
matchupKeys: { build: async () => Object.assign(() => KEYS, { stats: {} }) },
|
||||
}));
|
||||
// Two different questions: did the chain produce a probability at all, and
|
||||
// did it read tonight's matchup or this season's average?
|
||||
expect(r.chainShadow.readable).toBeGreaterThan(0);
|
||||
expect(r.chainShadow.platoon_applied).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
it('THE HAND SPLIT MOVES THE NUMBER — otherwise plumbing it changed nothing', async () => {
|
||||
const withSplit = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), withSplit, {
|
||||
supabase: fakeSb(),
|
||||
hitsFactorContext: { build: async () => factorContext },
|
||||
matchupKeys: { build: async () => Object.assign(() => KEYS, { stats: {} }) },
|
||||
}));
|
||||
const seasonOnly = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), seasonOnly, {
|
||||
supabase: fakeSb(),
|
||||
hitsFactorContext: { build: async () => null },
|
||||
matchupKeys: { build: async () => null },
|
||||
}));
|
||||
const a = withSplit.find((r) => r.chain_shadow).chain_shadow.chain_p;
|
||||
const b = seasonOnly.find((r) => r.chain_shadow).chain_shadow.chain_p;
|
||||
expect(a).not.toBeCloseTo(b, 3);
|
||||
});
|
||||
});
|
||||
|
||||
describe('CHAIN v3 — the opportunity term is REALLY fed, not a constant 4.1', () => {
|
||||
const KEYS_WITH_SLOT = (slot) => ({
|
||||
opponent: 'Boston Red Sox', opposing_pitcher: 'Chris Sale',
|
||||
batting_order: slot, source: 'lineup_context+schedule', refused: null,
|
||||
});
|
||||
const ctxFor = (prop, sport, keys) => ({
|
||||
spray: null, positionOaa: null, bats: 'R',
|
||||
throws: keys && keys.opposing_pitcher ? 'L' : null,
|
||||
pitcherHardHit: null, platoonSplits: null, as_of: '2026-08-12',
|
||||
});
|
||||
const withSlot = (slot, sink) => svc.runSnapshot('mlb', deps(memCache(), sink, {
|
||||
supabase: fakeSbTop(),
|
||||
hitsFactorContext: { build: async () => ctxFor },
|
||||
matchupKeys: { build: async () => Object.assign(() => KEYS_WITH_SLOT(slot), { stats: {} }) },
|
||||
}));
|
||||
|
||||
it('a POSTED lineup slot reaches the chain and sets E[PA]', async () => {
|
||||
const sink = [];
|
||||
await withSlot(1, sink);
|
||||
const row = sink.find((r) => r.chain_shadow);
|
||||
expect(row.chain_shadow.opportunity_source).toBe('lineup_slot');
|
||||
expect(row.chain_shadow.expected_pa).toBe(4.65); // leadoff, not 4.1
|
||||
});
|
||||
|
||||
it('LEADOFF and the NINE HOLE do not get the same probability', async () => {
|
||||
// This is the v2 defect made a test: without the slot every hitter ran on
|
||||
// DEFAULT_PA (4.1), so the opportunity half of `rate x opportunity` was a
|
||||
// constant across the entire lineup.
|
||||
const lead = []; const nine = [];
|
||||
await withSlot(1, lead);
|
||||
await withSlot(9, nine);
|
||||
const a = lead.find((r) => r.chain_shadow).chain_shadow;
|
||||
const b = nine.find((r) => r.chain_shadow).chain_shadow;
|
||||
expect(a.expected_pa).toBeGreaterThan(b.expected_pa);
|
||||
expect(a.chain_p).not.toBeCloseTo(b.chain_p, 3);
|
||||
});
|
||||
|
||||
it('an UNPOSTED lineup falls to a regular AND SAYS SO', async () => {
|
||||
const sink = [];
|
||||
await svc.runSnapshot('mlb', deps(memCache(), sink, {
|
||||
supabase: fakeSbTop(),
|
||||
hitsFactorContext: { build: async () => ctxFor },
|
||||
matchupKeys: { build: async () => Object.assign(() => KEYS_WITH_SLOT(null), { stats: {} }) },
|
||||
}));
|
||||
const row = sink.find((r) => r.chain_shadow);
|
||||
expect(row.chain_shadow.opportunity_source).toBe('default_regular');
|
||||
// 4.1 is "a regular", which is a claim about the league, not about him. It
|
||||
// must never be reported as if the lineup had been read.
|
||||
expect(row.chain_shadow.expected_pa).toBeNull();
|
||||
});
|
||||
|
||||
it('the run reports posted-slot coverage separately from the fire rate', async () => {
|
||||
const r = await svc.runSnapshot('mlb', deps(memCache(), [], {
|
||||
supabase: fakeSbTop(),
|
||||
hitsFactorContext: { build: async () => ctxFor },
|
||||
matchupKeys: { build: async () => Object.assign(() => KEYS_WITH_SLOT(3), { stats: {} }) },
|
||||
}));
|
||||
expect(r.chainShadow.opportunity_posted).toBeGreaterThan(0);
|
||||
expect(r.chainShadow.props).toBeGreaterThan(0);
|
||||
});
|
||||
|
||||
it('matchupKeys carries batting_order on EVERY return shape, including refusals', () => {
|
||||
// A field present only on the happy path is a field consumers cannot rely
|
||||
// on; the refusal shapes must carry it as an explicit null.
|
||||
const src = fs.readFileSync(path.join(ROOT, 'src/services/model/matchupKeys.js'), 'utf8');
|
||||
const returns = src.slice(src.indexOf('function resolve(prop)'));
|
||||
expect(returns).toMatch(/const empty = \{[^}]*batting_order: null/);
|
||||
expect((returns.match(/batting_order/g) || []).length).toBeGreaterThanOrEqual(4);
|
||||
});
|
||||
});
|
||||
|
||||
describe('SERVED-GRADE-UNCHANGED — the gate this whole order sits behind', () => {
|
||||
/** Everything a user can see, pulled from what the snapshot actually wrote. */
|
||||
const servedShape = (store) => JSON.stringify({
|
||||
grades: (store['grades:mlb'] || {}).grades,
|
||||
snapshot: (store['snapshot:mlb:latest'] || {}).grades,
|
||||
});
|
||||
|
||||
it('the served payload is BYTE-IDENTICAL with the shadow on and off', async () => {
|
||||
const on = memCache();
|
||||
await svc.runSnapshot('mlb', deps(on, []));
|
||||
|
||||
const off = memCache();
|
||||
// `chainShadow: null` cannot be injected (the ?? default would re-require
|
||||
// it), so the OFF arm is a shadow service that reads and returns nothing.
|
||||
await svc.runSnapshot('mlb', deps(off, [], {
|
||||
chainShadow: { runShadow: () => ({ version: 'off', byKey: new Map(), summary: { atoms: 0, readable: 0, refused: 0, games: 0, games_read: 0, per_game: [] } }) },
|
||||
}));
|
||||
|
||||
expect(servedShape(on)).toBe(servedShape(off));
|
||||
});
|
||||
|
||||
it('p_win on every served grade is exactly the counter value', async () => {
|
||||
const cache = memCache();
|
||||
await svc.runSnapshot('mlb', deps(cache, []));
|
||||
const byPlayer = Object.fromEntries(GRADES.map((g) => [g.player, g.p_win]));
|
||||
for (const g of cache.store['grades:mlb'].grades) {
|
||||
expect(g.p_win).toBe(byPlayer[g.player]);
|
||||
}
|
||||
});
|
||||
|
||||
it('NO served grade carries a chain field at all', async () => {
|
||||
const cache = memCache();
|
||||
await svc.runSnapshot('mlb', deps(cache, []));
|
||||
for (const g of cache.store['grades:mlb'].grades) {
|
||||
for (const k of Object.keys(g)) {
|
||||
expect(k).not.toMatch(/^chain/);
|
||||
}
|
||||
}
|
||||
});
|
||||
});
|
||||
|
||||
describe('0 LEAKS — measured against the source, not asserted in prose', () => {
|
||||
const read = (p) => fs.readFileSync(path.join(ROOT, p), 'utf8');
|
||||
const walk = (dir, out = []) => {
|
||||
for (const e of fs.readdirSync(path.join(ROOT, dir), { withFileTypes: true })) {
|
||||
const rel = `${dir}/${e.name}`;
|
||||
if (e.isDirectory()) { if (e.name !== 'node_modules' && e.name !== '.next') walk(rel, out); }
|
||||
else if (/\.(js|ts|tsx)$/.test(e.name)) out.push(rel);
|
||||
}
|
||||
return out;
|
||||
};
|
||||
|
||||
it('nothing under src/routes reads the chain shadow', () => {
|
||||
for (const f of walk('src/routes')) {
|
||||
const src = read(f);
|
||||
expect(src).not.toMatch(/chain_shadow|chainShadow|chain_p\b/);
|
||||
}
|
||||
});
|
||||
|
||||
it('nothing in the web app reads the chain shadow', () => {
|
||||
for (const f of walk('web/src')) {
|
||||
const src = read(f);
|
||||
expect(src).not.toMatch(/chain_shadow|chainShadow|chain_p\b/);
|
||||
}
|
||||
});
|
||||
|
||||
it('the ONLY writers are the shadow service, retention and the snapshot', () => {
|
||||
const writers = walk('src').filter((f) => /chain_shadow|chainShadow/.test(read(f)));
|
||||
expect(writers.sort()).toEqual([
|
||||
'src/services/model/chainShadow.js',
|
||||
'src/services/retentionService.js',
|
||||
'src/services/snapshotService.js',
|
||||
]);
|
||||
});
|
||||
|
||||
it('the shadow runs with requireCalibrated FALSE and says so where it is stored', () => {
|
||||
expect(read('src/services/model/chainShadow.js')).toMatch(/requireCalibrated:\s*false/);
|
||||
expect(cs.SHADOW_VERSION).toBe('chain-shadow@1');
|
||||
});
|
||||
|
||||
it('CALIBRATION_DEPLOYED is still empty — nothing is served calibrated', () => {
|
||||
expect(svc.CALIBRATION_DEPLOYED).toEqual([]);
|
||||
expect(Object.isFrozen(svc.CALIBRATION_DEPLOYED)).toBe(true);
|
||||
});
|
||||
|
||||
it('the counter still serves: probabilityEstimator is what analyzeViaEngine1 calls', () => {
|
||||
const engine = read('src/services/intelligence/analyzeViaEngine1.js');
|
||||
expect(engine).toMatch(/require\('\.\/probabilityEstimator'\)/);
|
||||
// And the engine knows nothing about the chain — the shadow is downstream.
|
||||
expect(engine).not.toMatch(/chainShadow|baseballChain/);
|
||||
});
|
||||
|
||||
it('A8 and the factor freeze are untouched — factor_inputs still lands', async () => {
|
||||
const rows = require('../../src/services/retentionService').rowsFromSides(
|
||||
{ player: 'X', stat_type: 'hits', line: 0.5 },
|
||||
[{ player: 'X', stat_type: 'hits', line: 0.5, direction: 'over', grade: 'B', factor_inputs: { v: 'hits-factors@1' } }],
|
||||
{ snapshotId: 's', capturedAt: 't', sport: 'mlb', gameDate: '2026-08-12' },
|
||||
);
|
||||
expect(rows[0].factor_inputs).toEqual({ v: 'hits-factors@1' });
|
||||
expect(rows[0].chain_shadow).toBeNull(); // declared, so batches keep one shape
|
||||
});
|
||||
|
||||
it('WNBA is untouched — the shadow is MLB-gated', async () => {
|
||||
const src = read('src/services/snapshotService.js');
|
||||
const i = src.indexOf('CHAIN v1 — THE SHADOW READ');
|
||||
const block = src.slice(i, src.indexOf('LINEUP + BASERUNNER CONTEXT', i));
|
||||
expect(block).toMatch(/if \(sp === 'mlb'\)/);
|
||||
});
|
||||
});
|
||||
@@ -0,0 +1,319 @@
|
||||
'use strict';
|
||||
|
||||
/**
|
||||
* The WNBA possession/usage feed — the source, the derivations, the as-of rule.
|
||||
*
|
||||
* Two failures this suite exists to prevent. First, a derived rate that is not
|
||||
* actually the identity it claims to be: usage and possessions feed the chain's
|
||||
* opportunity term directly, so a wrong constant or a transposed term would be
|
||||
* invisible downstream and wrong everywhere. Second — and this is the one that
|
||||
* has bitten this codebase repeatedly — a read that quietly includes the game it
|
||||
* is trying to predict.
|
||||
*/
|
||||
|
||||
const adapter = require('../../src/services/adapters/espnWnbaAdapter');
|
||||
const usage = require('../../src/services/wnbaUsageService');
|
||||
|
||||
/** A real-shaped ESPN summary payload, trimmed to what the parser reads. */
|
||||
const PLAYER_KEYS = [
|
||||
'minutes', 'points', 'fieldGoalsMade-fieldGoalsAttempted',
|
||||
'threePointFieldGoalsMade-threePointFieldGoalsAttempted',
|
||||
'freeThrowsMade-freeThrowsAttempted', 'rebounds', 'assists', 'turnovers',
|
||||
'steals', 'blocks', 'offensiveRebounds', 'defensiveRebounds', 'fouls', 'plusMinus',
|
||||
];
|
||||
|
||||
const athlete = (id, name, stats, over = {}) => ({
|
||||
athlete: { id, displayName: name }, starter: true, stats, ...over,
|
||||
});
|
||||
|
||||
const teamStats = (fg, ft, tov, oreb) => [
|
||||
{ name: 'fieldGoalsMade-fieldGoalsAttempted', displayValue: fg },
|
||||
{ name: 'freeThrowsMade-freeThrowsAttempted', displayValue: ft },
|
||||
{ name: 'totalTurnovers', displayValue: String(tov) },
|
||||
{ name: 'offensiveRebounds', displayValue: String(oreb) },
|
||||
];
|
||||
|
||||
function summary(over = {}) {
|
||||
return {
|
||||
header: {
|
||||
id: '401857134',
|
||||
season: { year: 2026 },
|
||||
competitions: [{
|
||||
date: '2026-08-11T23:30Z',
|
||||
status: { type: { name: 'STATUS_FINAL' } },
|
||||
competitors: [
|
||||
{ homeAway: 'home', score: '106', team: { id: '5', abbreviation: 'IND' } },
|
||||
{ homeAway: 'away', score: '92', team: { id: '9', abbreviation: 'NY' } },
|
||||
],
|
||||
}],
|
||||
...over.header,
|
||||
},
|
||||
boxscore: {
|
||||
players: [
|
||||
{
|
||||
team: { id: '9', abbreviation: 'NY' },
|
||||
statistics: [{
|
||||
keys: PLAYER_KEYS,
|
||||
athletes: [
|
||||
// MIN PTS FG 3PT FT REB AST TO STL BLK OREB DREB PF +/-
|
||||
athlete('1', 'Breanna Stewart', ['30', '19', '8-21', '1-5', '2-2', '7', '5', '0', '1', '0', '2', '5', '2', '-2']),
|
||||
athlete('2', 'Sabrina Ionescu', ['34', '22', '7-18', '4-9', '4-4', '4', '6', '3', '2', '0', '0', '4', '1', '-5']),
|
||||
athlete('3', 'Deep Bench', null, { didNotPlay: true, stats: [] }),
|
||||
],
|
||||
}],
|
||||
},
|
||||
],
|
||||
teams: [
|
||||
{ team: { id: '9', abbreviation: 'NY' }, statistics: teamStats('35-76', '11-13', 13, 13) },
|
||||
{ team: { id: '5', abbreviation: 'IND' }, statistics: teamStats('37-69', '16-19', 11, 4) },
|
||||
],
|
||||
},
|
||||
...over,
|
||||
};
|
||||
}
|
||||
|
||||
describe('THE SOURCE — ESPN returns components, and the parser reads them', () => {
|
||||
it('parses a real-shaped box score into per-player rows', () => {
|
||||
const { rows, game } = adapter.parseSummary(summary());
|
||||
expect(game.game_id).toBe('401857134');
|
||||
expect(rows).toHaveLength(2);
|
||||
const s = rows.find((r) => r.player_name === 'Breanna Stewart');
|
||||
expect(s.minutes).toBe(30);
|
||||
expect(s.fga).toBe(21);
|
||||
expect(s.fta).toBe(2);
|
||||
expect(s.tov).toBe(0);
|
||||
expect(s.team).toBe('NY');
|
||||
expect(s.opponent).toBe('IND');
|
||||
expect(s.home_away).toBe('away');
|
||||
});
|
||||
|
||||
it('a DNP is OMITTED, not recorded as a zero line', () => {
|
||||
// A zero-minute row asserts he was available and produced nothing. That is a
|
||||
// different fact from "he did not play", and the usage denominator would
|
||||
// divide by his zero minutes anyway.
|
||||
const { rows } = adapter.parseSummary(summary());
|
||||
expect(rows.find((r) => r.player_name === 'Deep Bench')).toBeUndefined();
|
||||
});
|
||||
|
||||
it('ET dating — a 23:30Z tip is the SAME ET evening, not the next day', () => {
|
||||
// Using the UTC date would file half of every slate a day late.
|
||||
const { rows } = adapter.parseSummary(summary());
|
||||
expect(rows[0].game_date).toBe('2026-08-11');
|
||||
});
|
||||
|
||||
it('an UNFINISHED game is refused — a partial box score is a moving denominator', () => {
|
||||
const live = summary();
|
||||
live.header.competitions[0].status.type.name = 'STATUS_IN_PROGRESS';
|
||||
const out = adapter.parseSummary(live);
|
||||
expect(out.rows).toHaveLength(0);
|
||||
expect(out.skipped).toMatch(/not_final/);
|
||||
});
|
||||
|
||||
it('a malformed made-attempted pair is ABSENT, never zero', () => {
|
||||
expect(adapter.madeAtt('8-21')).toEqual({ made: 8, att: 21 });
|
||||
expect(adapter.madeAtt('--')).toEqual({ made: null, att: null });
|
||||
expect(adapter.madeAtt(null)).toEqual({ made: null, att: null });
|
||||
});
|
||||
});
|
||||
|
||||
describe('THE DERIVATIONS — exact identities, not proxies', () => {
|
||||
const t = { fga: 76, fta: 13, tov: 13, oreb: 13 };
|
||||
|
||||
it('possessions = FGA - OREB + TOV + 0.44*FTA', () => {
|
||||
expect(adapter.possessions(t)).toBeCloseTo(76 - 13 + 13 + 0.44 * 13, 6);
|
||||
});
|
||||
|
||||
it('a missing term makes possessions ABSENT — never a partial sum', () => {
|
||||
// Treating an absent turnover count as zero understates possessions and
|
||||
// inflates every usage rate computed against it.
|
||||
expect(adapter.possessions({ ...t, tov: null })).toBeNull();
|
||||
expect(adapter.possessions(null)).toBeNull();
|
||||
});
|
||||
|
||||
it('usage matches the standard identity, recomputed independently', () => {
|
||||
const p = { fga: 21, fta: 2, tov: 0, minutes: 30 };
|
||||
const teamMinutes = 200;
|
||||
const expected = 100 * ((21 + 0.44 * 2 + 0) * (teamMinutes / 5))
|
||||
/ (30 * (76 + 0.44 * 13 + 13));
|
||||
expect(adapter.usageRate(p, t, teamMinutes)).toBeCloseTo(expected, 9);
|
||||
});
|
||||
|
||||
it('usage is ABSENT on a zero-minute or missing-term row, never Infinity or 0', () => {
|
||||
expect(adapter.usageRate({ fga: 5, fta: 0, tov: 1, minutes: 0 }, t, 200)).toBeNull();
|
||||
expect(adapter.usageRate({ fga: null, fta: 0, tov: 1, minutes: 20 }, t, 200)).toBeNull();
|
||||
expect(adapter.usageRate({ fga: 5, fta: 0, tov: 1, minutes: 20 }, t, 0)).toBeNull();
|
||||
});
|
||||
|
||||
it('pace handles OVERTIME through team minutes, with no special case', () => {
|
||||
// 5 players x 40 min = 200 team-minutes in regulation; an OT game has more,
|
||||
// and the divisor follows rather than being branched on.
|
||||
const reg = adapter.pace(100, 200);
|
||||
const ot = adapter.pace(100, 225);
|
||||
expect(reg).toBeCloseTo(100, 6); // 100 poss per 40 min
|
||||
expect(ot).toBeLessThan(reg); // same possessions over more time
|
||||
});
|
||||
|
||||
it('true shooting and eFG are absent without attempts', () => {
|
||||
expect(adapter.trueShooting({ points: 19, fga: 21, fta: 2 })).toBeCloseTo(19 / (2 * (21 + 0.88)), 9);
|
||||
expect(adapter.trueShooting({ points: 0, fga: 0, fta: 0 })).toBeNull();
|
||||
expect(adapter.efgPct({ fgm: 8, fg3m: 1, fga: 21 })).toBeCloseTo(8.5 / 21, 9);
|
||||
expect(adapter.efgPct({ fgm: 0, fg3m: 0, fga: 0 })).toBeNull();
|
||||
});
|
||||
|
||||
it('END TO END: the stored rate equals the identity recomputed from the stored components', () => {
|
||||
// This is the factorFreeze rule applied to a feed: every component is stored
|
||||
// beside its derived value, so the rate is checkable rather than asserted.
|
||||
const { rows } = adapter.parseSummary(summary());
|
||||
for (const r of rows) {
|
||||
const expected = 100 * ((r.fga + 0.44 * r.fta + r.tov) * (r.team_minutes / 5))
|
||||
/ (r.minutes * (r.team_fga + 0.44 * r.team_fta + r.team_tov));
|
||||
expect(r.usage_rate).toBeCloseTo(expected, 3);
|
||||
expect(r.team_possessions)
|
||||
.toBeCloseTo(r.team_fga - r.team_oreb + r.team_tov + 0.44 * r.team_fta, 3);
|
||||
}
|
||||
});
|
||||
|
||||
it('carries the game state the redistribute hook reads', () => {
|
||||
const { rows } = adapter.parseSummary(summary());
|
||||
const ny = rows[0];
|
||||
expect(ny.team_score).toBe(92);
|
||||
expect(ny.opp_score).toBe(106);
|
||||
expect(ny.final_margin).toBe(-14); // a blowout the hook can see
|
||||
expect(ny.starter).toBe(true);
|
||||
});
|
||||
});
|
||||
|
||||
describe('LEAGUE MEMBERSHIP — exhibitions and All-Star are not league games', () => {
|
||||
const teamsPayload = {
|
||||
sports: [{ leagues: [{ teams: [{ team: { abbreviation: 'NY' } }, { team: { abbreviation: 'IND' } }] }] }],
|
||||
};
|
||||
const fetchImpl = (url) => Promise.resolve(url.includes('/teams') ? teamsPayload : summary());
|
||||
|
||||
it('a league game passes through', async () => {
|
||||
const out = await adapter.getGameRows('401857134', { fetchImpl, force: true });
|
||||
expect(out.rows.length).toBe(2);
|
||||
});
|
||||
|
||||
it('a game against a NON-LEAGUE side is refused, and names who', async () => {
|
||||
// The season pull surfaced exactly two: an exhibition vs Nigeria and the
|
||||
// All-Star game. Their usage/pace context is meaningless for a forward
|
||||
// projection, so counting them would pollute every profile spanning them.
|
||||
const exhibition = summary();
|
||||
exhibition.header.competitions[0].competitors[0].team.abbreviation = 'NIGER';
|
||||
exhibition.boxscore.teams[1].team.abbreviation = 'NIGER';
|
||||
const f = (url) => Promise.resolve(url.includes('/teams') ? teamsPayload : exhibition);
|
||||
const out = await adapter.getGameRows('401867794', { fetchImpl: f, force: true });
|
||||
expect(out.rows).toHaveLength(0);
|
||||
expect(out.skipped).toMatch(/^not_a_league_game:/);
|
||||
expect(out.skipped).toContain('NIGER');
|
||||
});
|
||||
|
||||
it('an EMPTY teams response does not silently filter the whole league away', async () => {
|
||||
// A failed feed that degrades to an empty index looks exactly like an honest
|
||||
// absence — the fielding_oaa lesson. Unknown membership means NO filtering,
|
||||
// never "nothing is in the league".
|
||||
const empty = { sports: [{ leagues: [{ teams: [] }] }] };
|
||||
const f = (url) => Promise.resolve(url.includes('/teams') ? empty : summary());
|
||||
const out = await adapter.getGameRows('401857134', { fetchImpl: f, force: true });
|
||||
expect(out.rows.length).toBe(2);
|
||||
});
|
||||
});
|
||||
|
||||
describe('THE AS-OF RULE — point-in-time is structural here', () => {
|
||||
const games = [
|
||||
{ game_date: '2026-08-01', minutes: 30, usage_rate: 24, team_pace: 96, team_possessions: 80, ts_pct: 0.55, efg_pct: 0.5, starter: true, final_margin: 5, team: 'NY', game_id: 'g1', source_id: '1' },
|
||||
{ game_date: '2026-08-04', minutes: 34, usage_rate: 28, team_pace: 98, team_possessions: 82, ts_pct: 0.58, efg_pct: 0.52, starter: true, final_margin: -3, team: 'NY', game_id: 'g2', source_id: '1' },
|
||||
{ game_date: '2026-08-07', minutes: 20, usage_rate: 30, team_pace: 94, team_possessions: 78, ts_pct: 0.50, efg_pct: 0.48, starter: false, final_margin: 12, team: 'NY', game_id: 'g3', source_id: '1' },
|
||||
{ game_date: '2026-08-11', minutes: 32, usage_rate: 26, team_pace: 99, team_possessions: 83, ts_pct: 0.60, efg_pct: 0.55, starter: true, final_margin: -14, team: 'NY', game_id: 'g4', source_id: '1' },
|
||||
].map((g) => ({ sport: 'wnba', player_key: 'test player', season: 2026, ...g }));
|
||||
|
||||
/** A stub client driving the REAL paginate + read path. */
|
||||
const client = (rows) => ({
|
||||
from() {
|
||||
const st = { rows: rows.slice(), order: [] };
|
||||
const q = {
|
||||
select: () => q,
|
||||
eq: (c, v) => { st.rows = st.rows.filter((r) => r[c] === v); return q; },
|
||||
lt: (c, v) => { st.rows = st.rows.filter((r) => String(r[c]) < String(v)); return q; },
|
||||
order: (c, o) => { st.order.push([c, !o || o.ascending !== false]); return q; },
|
||||
range: async (a, b) => {
|
||||
const sorted = st.rows.slice().sort((x, y) => {
|
||||
for (const [c, asc] of st.order) {
|
||||
const cmp = String(x[c]).localeCompare(String(y[c]));
|
||||
if (cmp) return asc ? cmp : -cmp;
|
||||
}
|
||||
return 0;
|
||||
});
|
||||
return { data: sorted.slice(a, b + 1), error: null };
|
||||
},
|
||||
};
|
||||
return q;
|
||||
},
|
||||
});
|
||||
|
||||
it('reads STRICTLY before the cutoff — a game ON the date is excluded', async () => {
|
||||
// The leak this prevents: a 2026-08-11 game may tip after grade time, so
|
||||
// counting it feeds the evening being predicted into the prediction. Same
|
||||
// rule as snapshotSettlementService.isPreGame.
|
||||
const p = await usage.profileAsOf(client(games), { playerKey: 'test player', asOf: '2026-08-11' });
|
||||
expect(p.games).toBe(3);
|
||||
expect(p.last_game_date).toBe('2026-08-07');
|
||||
});
|
||||
|
||||
it('a later cutoff sees more games — the profile MOVES with the date', async () => {
|
||||
const early = await usage.profileAsOf(client(games), { playerKey: 'test player', asOf: '2026-08-08' });
|
||||
const late = await usage.profileAsOf(client(games), { playerKey: 'test player', asOf: '2026-08-12' });
|
||||
expect(early.games).toBe(3);
|
||||
expect(late.games).toBe(4);
|
||||
// The whole point of a point-in-time read: the same player reads differently
|
||||
// on different dates, because the history behind him is different.
|
||||
expect(early.usage_rate).not.toBe(late.usage_rate);
|
||||
expect(early.last_game_date).toBe('2026-08-07');
|
||||
expect(late.last_game_date).toBe('2026-08-11');
|
||||
});
|
||||
|
||||
it('REFUSES on thin history rather than returning a league-average player', async () => {
|
||||
// Usage feeds the chain's opportunity term directly; a stand-in profile
|
||||
// would assert a usage rate about someone never observed.
|
||||
expect(await usage.profileAsOf(client(games), { playerKey: 'test player', asOf: '2026-08-04' })).toBeNull();
|
||||
expect(await usage.profileAsOf(client(games), { playerKey: 'nobody', asOf: '2026-12-01' })).toBeNull();
|
||||
});
|
||||
|
||||
it('usage is MINUTES-WEIGHTED, not a flat mean across games', async () => {
|
||||
// The lineup-K-rate lesson: an unweighted aggregate counts a 20-minute
|
||||
// cameo like a 34-minute start, and unweighted HURT that model.
|
||||
const p = await usage.profileAsOf(client(games), { playerKey: 'test player', asOf: '2026-12-01' });
|
||||
const flat = (24 + 28 + 30 + 26) / 4;
|
||||
const weighted = (24 * 30 + 28 * 34 + 30 * 20 + 26 * 32) / (30 + 34 + 20 + 32);
|
||||
expect(p.usage_rate).toBeCloseTo(weighted, 3);
|
||||
expect(p.usage_rate).not.toBeCloseTo(flat, 3);
|
||||
});
|
||||
|
||||
it('delivers every chainFn input, and the redistribute game-state', async () => {
|
||||
const p = await usage.profileAsOf(client(games), { playerKey: 'test player', asOf: '2026-12-01' });
|
||||
for (const f of ['usage_rate', 'minutes_per_game', 'team_pace', 'team_possessions', 'ts_pct', 'efg_pct']) {
|
||||
expect(p[f]).not.toBeNull();
|
||||
}
|
||||
expect(p.starter_rate).toBeCloseTo(0.75, 3);
|
||||
expect(p.avg_final_margin).toBeCloseTo(0, 3);
|
||||
expect(p.as_of).toBe('2026-12-01'); // the cutoff is reported, not implied
|
||||
});
|
||||
});
|
||||
|
||||
describe('the feed is registered and scoped', () => {
|
||||
it('its unique key is in tableKeys, so safePaginate can walk it', () => {
|
||||
const { uniqueKeyFor } = require('../../src/utils/tableKeys');
|
||||
expect(uniqueKeyFor('wnba_player_game')).toEqual(['game_id', 'source_id']);
|
||||
});
|
||||
|
||||
it('MLB and the chain shadow are untouched by this order', () => {
|
||||
const fs = require('fs');
|
||||
const path = require('path');
|
||||
const root = path.join(__dirname, '..', '..');
|
||||
// The WNBA feed must not have reached into the MLB engine or the served path.
|
||||
for (const f of ['src/services/model/chainShadow.js', 'src/services/model/baseballChain.js',
|
||||
'src/services/intelligence/analyzeViaEngine1.js']) {
|
||||
expect(fs.readFileSync(path.join(root, f), 'utf8')).not.toMatch(/wnbaUsage|espnWnbaAdapter|wnba_player_game/);
|
||||
}
|
||||
});
|
||||
});
|
||||
Reference in New Issue
Block a user