checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion recon (Tony resurrection, STEP 0). This work existed only on the laptop disk. - chain shadow accrual + probe script (038_chain_shadow.sql) - WNBA possession feed: ESPN adapter, usage service, verify script (039_wnba_player_game.sql) - baseball chain - retention/snapshot service updates, tableKeys, matchupKeys - specs: chain-v1, wnba-possession-feed, wnba-source-survey - unit tests for the above Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
This commit is contained in:
@@ -0,0 +1,503 @@
|
||||
# chain-v1 / v2 — the portable engine, made whole, fed, and shadowed on MLB
|
||||
|
||||
> **v2 (2026-08-12) is appended at §8.** It corrects the order's premise (the
|
||||
> chain was already firing on 99.6%, `paOutcome` reads no handedness, and there
|
||||
> is no fallback path anywhere) and plumbs the hand split as a matchup
|
||||
> conditioner on the RATE. Read §8 for the current numbers; §4 below is the v1
|
||||
> measurement and is superseded.
|
||||
|
||||
|
||||
|
||||
**Status:** SHADOW. Served by nothing. Nothing promoted.
|
||||
**Date:** 2026-08-12
|
||||
**Predecessors:** chaining-v1 (Session 90 — the design, which never had a spec
|
||||
file), A5 (factor freeze), A6 (the join keys, shadow), A7 (shadow accrual).
|
||||
|
||||
---
|
||||
|
||||
## 0. The finding this order started from
|
||||
|
||||
The extraction pass found `chain.js` was a **shell of its own header**:
|
||||
|
||||
| the header promised | the code had |
|
||||
|---|---|
|
||||
| `atoms` | a positional array ✅ |
|
||||
| `context` | passed to `redistribute` only, partially |
|
||||
| **`chainFn`** — atom → per-entity probability | **ABSENT. No parameter, no call site, no export.** |
|
||||
| `aggregator` — ACROSS / UP | two functions ✅ |
|
||||
| `redistribute` | **on `chainUp` only** |
|
||||
|
||||
And it was inert for a second, undocumented reason: `chainAcross` requires
|
||||
`calibrated: true`, and the only writer of that flag is a loop over
|
||||
`CALIBRATION_DEPLOYED`, which is `Object.freeze([])`. **No grade in the system
|
||||
carries the flag**, so `chainAcross` had zero callers *and* would have refused
|
||||
any caller it had.
|
||||
|
||||
The missing `chainFn` is why the "portable core" was not portable: with no slot
|
||||
for the atom→probability stage, every sport's real work had to live somewhere
|
||||
else. For MLB it lived in `scripts/`, reachable from no pipeline.
|
||||
|
||||
---
|
||||
|
||||
## 1. The three core fixes (`src/services/model/chain.js`)
|
||||
|
||||
### 1.1 `chainFn(atom, context) → per-entity probability`
|
||||
|
||||
- `applyChainFn` runs before usability filtering. **Default is identity-on-`p`**,
|
||||
so every pre-existing caller is byte-identical.
|
||||
- Returning a **number** sets `p`; returning an **object** merges (so a chainFn
|
||||
can attach its own trace); returning **null or throwing** makes the atom
|
||||
UNREADABLE ⇒ **DROPPED, counted, never `p=0`**. A zero leg would zero an entire
|
||||
ticket, and "we could not read him" is not "he cannot do it".
|
||||
- The refusals are surfaced as `chain_fn_refused` rather than swallowed, so a
|
||||
chainFn quietly failing across the board is visible instead of looking like a
|
||||
thin slate.
|
||||
|
||||
### 1.2 Correlation is SIGNED, and the direction follows the sign
|
||||
|
||||
Was clamped `[0, 1]` with the joint always shifted toward the weakest leg. That
|
||||
is baseball's shape — same-game legs share the pitcher, the park and the weather
|
||||
— and it made basketball's case **inexpressible**: teammates compete for finite
|
||||
possessions, so one player's shot is another's non-shot and their props are
|
||||
NEGATIVELY correlated.
|
||||
|
||||
Correlation is now clamped `[-1, 1]` and `|corr|` interpolates from independence
|
||||
toward the **Fréchet–Hoeffding bound its sign selects**:
|
||||
|
||||
```
|
||||
corr = +1 → joint = min(p_i) (upper bound, co-monotone)
|
||||
corr = 0 → joint = Π p_i (independent)
|
||||
corr = −1 → joint = max(0, Σp − (n−1)) (lower bound, counter-monotone)
|
||||
```
|
||||
|
||||
This is why the direction is principled rather than chosen. The positive branch
|
||||
is arithmetically unchanged — the weakest leg **is** the upper bound — so every
|
||||
previously-correct number stays exactly what it was.
|
||||
|
||||
### 1.3 `redistribute` reaches BOTH readings
|
||||
|
||||
Previously on `chainUp` only. A redistribution that reached the team read and not
|
||||
the across read would leave `selfCheck` comparing post-redistribution to
|
||||
pre-redistribution atoms and flagging an INTERNAL_INCONSISTENCY **the model had
|
||||
itself just manufactured**.
|
||||
|
||||
`prepareAtoms(atoms, opts)` is exported so a caller prepares ONCE — chainFn, then
|
||||
usability, then redistribution — and hands the identical legs to both readings.
|
||||
|
||||
---
|
||||
|
||||
## 2. Baseball's chainFn (`src/services/model/baseballChain.js`)
|
||||
|
||||
A chained forecast is always `rate × opportunity`. Baseball is the clean case
|
||||
because **opportunity is fixed**: the batting order is set before first pitch, a
|
||||
nine-run lead does not change who bats next, and PA/game varies over a narrow
|
||||
range set almost entirely by lineup slot.
|
||||
|
||||
```
|
||||
p_hit_per_PA ← skillProjection.paOutcome the RATE (modelled)
|
||||
expected PA ← lineup slot the OPPORTUNITY (a lookup)
|
||||
P(hits ≥ k) ← Binomial(PA, p) mixed over paDistribution
|
||||
```
|
||||
|
||||
**Atoms routed** — all pre-existing, all previously reachable only from
|
||||
`scripts/`:
|
||||
|
||||
| atom | source | role |
|
||||
|---|---|---|
|
||||
| `fromStatcastRow` | skillProjection:161 | the ONE legal units conversion (statcast stores PERCENTAGES 0–100) |
|
||||
| `paOutcome` | skillProjection:268 | K / BB via log5 odds-ratio vs league; remainder = balls in play |
|
||||
| `hitOnContact` | skillProjection:213 | archetype-selected barrel / hard-hit / exit-velo / GB-speed × pitcher contact allowed × park |
|
||||
| `paDistribution`, `binomialPmf`, `atLeast` | skillProjection:327/308/339 | the chain over opportunity |
|
||||
|
||||
`PA_BY_SLOT` is a **lookup, not a fit** — 4.65 (leadoff) down to 3.85 (nine hole),
|
||||
the documented ~0.1-PA-per-slot decline. Nothing was tuned on settled rows; a
|
||||
tuned opportunity term on 1,741 rows is curve-fitting dressed as physics.
|
||||
|
||||
**Refusals:** no batter profile ⇒ null ⇒ dropped (never a league hitter). A
|
||||
missing pitcher is different and handled inside `paOutcome` — the batter's own
|
||||
rate stands rather than being pulled toward average.
|
||||
|
||||
**Scope: hits only.** `projectSkill` also routes `total_bases`, but TB's
|
||||
head-to-head is on record as INCONCLUSIVE-under-contamination. Widening the first
|
||||
shadow to a stat whose verdict is already muddy buys noise.
|
||||
|
||||
**`redistribute` is dormant** and returns the legs unchanged — the honest dormant
|
||||
behaviour. Returning null would read as "the hook failed".
|
||||
|
||||
---
|
||||
|
||||
## 3. The shadow (`src/services/model/chainShadow.js`)
|
||||
|
||||
Runs at the post-enriched / pre-persist point in `snapshotService` — the A6
|
||||
position, where the slate exists as a set. Reuses the statcast rows and resolved
|
||||
opposing starter the challenger pass already fetched: **zero new I/O**.
|
||||
|
||||
### The unit of evidence is the TRIPLE
|
||||
|
||||
```
|
||||
(chain_p, counter_p, outcome)
|
||||
```
|
||||
|
||||
all three on one row, all side-aligned. `counter_p` is the served
|
||||
`estimateProbability` value for that exact side; `outcome` arrives from the
|
||||
ordinary settle pass; `chain_p` is the chain's. **Evidence you cannot adjudicate
|
||||
is not evidence** — a chain probability stored without the number it must beat,
|
||||
or without the result, can only ever be compared to itself.
|
||||
|
||||
**Side alignment is load-bearing.** The chain computes P(over the line); `p_win`
|
||||
is expressed for the graded SIDE. An under row stores `1 − p_over`. Storing the
|
||||
raw over-probability against an under row would invert every later comparison,
|
||||
silently.
|
||||
|
||||
### UN-SERVABLE, said in the data
|
||||
|
||||
`requireCalibrated: false` is legitimate **only** because nothing downstream
|
||||
reads the result. Every stored block carries `status: 'UN-SERVABLE'` and
|
||||
`servable: false` in its own payload, not merely in a comment — a caveat that
|
||||
lives only in a comment is not attached to the data once something else queries
|
||||
it. Flipping that requires passing the gate, not editing a file.
|
||||
|
||||
### The self-check is VACUOUS today, and says so
|
||||
|
||||
`selfCheck` earns its keep by comparing per-entity reads to an **independent**
|
||||
team read. None exists — the game-script projection was deliberately not built
|
||||
(no atom has passed the gate). So the up-read is assembled from the SAME atoms as
|
||||
the across-read and agreement between them is arithmetic. Every block carries
|
||||
`self_check.vacuous: true` with its reason, so nobody later mistakes a tautology
|
||||
for a passing consistency test.
|
||||
|
||||
### Storage
|
||||
|
||||
`model_snapshots.chain_shadow jsonb` (migration 038). A separate column, not
|
||||
extra keys inside `features` — `champion-ablation.js` iterates every `features`
|
||||
key for its residual scan, so widening it would silently enlarge that
|
||||
multiple-comparisons denominator. Same reasoning as 034.
|
||||
|
||||
---
|
||||
|
||||
## 4. First measurement (`scripts/chain-shadow-probe.js`)
|
||||
|
||||
Real board, 2026-08-07 → 2026-08-12, MLB hits, **9,376 graded rows**:
|
||||
|
||||
```
|
||||
CHAIN FIRE — 9,340/9,376 atoms read (36 refused) across 82/82 games [99.6%]
|
||||
|
||||
min p25 median p75 max mean
|
||||
chain_p 0.013 0.382 0.500 0.618 0.987 0.500
|
||||
counter_p 0.050 0.396 0.500 0.604 0.950 0.500
|
||||
divergence (signed) −0.556 −0.082 0.000 0.082 0.556 −0.000
|
||||
divergence (abs) 0.000 0.038 0.082 0.142 0.556 0.100
|
||||
|
||||
|divergence| < 0.02 1,203 (12.9%) agrees with the counter
|
||||
0.02 – 0.05 1,760 (18.8%)
|
||||
0.05 – 0.10 2,577 (27.6%)
|
||||
0.10 – 0.20 2,835 (30.4%)
|
||||
≥ 0.20 965 (10.3%) a different read entirely
|
||||
|
||||
OVER SIDE ONLY (n=4,673)
|
||||
divergence (signed) −0.526 −0.045 0.031 0.104 0.556 0.029
|
||||
chain HIGHER on 2,837 (60.7%)
|
||||
```
|
||||
|
||||
**Reading it honestly:**
|
||||
|
||||
- The chain **fires**, on 99.6% of the board. It is not the A5 case (built,
|
||||
correct, never invoked).
|
||||
- It is **not a relabelled counter**: 40.7% of rows differ by ≥0.10, and 10.3%
|
||||
by ≥0.20. It is also not noise — 12.9% agree inside 0.02.
|
||||
- The **symmetry of the two-sided signed distribution is arithmetic**, not a
|
||||
finding: both sides of every prop are in the sample, so each pair contributes
|
||||
`+d` and `−d`. Reading `mean −0.000` as "unbiased" would be reading the
|
||||
sampling scheme. The over-side slice is the one that can lean, and it does:
|
||||
**+2.9pp mean, higher on 60.7%**.
|
||||
- **Divergence is not merit.** A challenger that disagrees is interesting, not
|
||||
right. Which of the two is closer to what happened is the settle pass's
|
||||
question, and it is exactly why the triple is stored.
|
||||
- **The probe is contaminated by construction** and reports only a divergence
|
||||
(a property of two forecasts) rather than a resolution (a property of a
|
||||
forecast against an outcome): `statcast_aggregates` is upserted in place and
|
||||
keeps one as-of date, so profiles read for a row graded three days ago are
|
||||
today's. The forward accrual on the cron does not have this problem.
|
||||
- The probe resolves **no opposing pitcher** (offline), so it measures the
|
||||
batter-side read. The live shadow does resolve it.
|
||||
|
||||
---
|
||||
|
||||
## 5. What is NOT claimed
|
||||
|
||||
- **Nothing is promoted.** The proven set remains empty.
|
||||
- **The chain is not calibrated** and has passed no gate.
|
||||
- **No head-to-head has been run.** That needs settled outcomes against the
|
||||
stored triples and must go through `factorGate` / the cumulative Bonferroni
|
||||
denominator, as a NEW hypothesis.
|
||||
- **`CALIBRATION_DEPLOYED` stays `[]`.** Nothing is served calibrated.
|
||||
- **WNBA is untouched.** Its contested-possession chainFn and its possession feed
|
||||
are later orders. The core changes (signed correlation, redistribute on both
|
||||
readings) were built now because they are engine honesty, not because MLB needs
|
||||
them — MLB exercises neither.
|
||||
|
||||
---
|
||||
|
||||
## 6. The pre-registered next step
|
||||
|
||||
Once settled outcomes accrue against `chain_shadow`:
|
||||
|
||||
1. Score `chain_p` vs `counter_p` on the SAME rows (paired bootstrap — comparing
|
||||
independent SEs overstates uncertainty and has previously read a reliable
|
||||
−0.022 as noise).
|
||||
2. Per stat, never pooled (pooled resolution is inflated by base-rate structure).
|
||||
3. Through the cumulative test ledger; the CI widens to `1 − 0.05/tests`.
|
||||
4. **Pre-registered fallback, stated before the answer is known:** if the chain
|
||||
moves ~87% of the board and does not improve Brier, it is THEATER by the
|
||||
`factorGate` definition and the correct action is to leave it off — not to
|
||||
re-tune the opportunity term until it passes.
|
||||
|
||||
---
|
||||
|
||||
## 7. Incidental defect found and fixed
|
||||
|
||||
`runSnapshot`'s `deps` object was an **allowlist of 17 keys**, but fourteen call
|
||||
sites read `deps.challenger`, `deps.loadStatcast`, `deps.environmentContext`,
|
||||
`deps.lineupContext`, `deps.hitsFactorContext`, `deps.matchupKeys`,
|
||||
`deps.gameBinder`, `deps.archetypeAxes`, `deps.contactChallenger`,
|
||||
`deps.projectionChallenger`, `deps.loadArsenals`, `deps.mlbAdapter` — each
|
||||
documented as injectable, each **permanently `undefined`**, each always falling
|
||||
through to the real module. The seam existed in the comment and not in the code.
|
||||
|
||||
Fixed by spreading `...opts` FIRST in the literal: every explicit key is declared
|
||||
after and already reads `opts.X`, so no existing behaviour moves, while an
|
||||
unlisted dep now actually arrives. Found because the chain shadow's own test
|
||||
could not inject a statcast map — the test would have passed while measuring
|
||||
nothing, which is the failure this whole line of work exists to stop repeating.
|
||||
|
||||
---
|
||||
|
||||
# §8 — chain v2: FEED THE ENGINE (2026-08-12)
|
||||
|
||||
## 8.1 The order's premise, checked before building
|
||||
|
||||
The v2 order opened from four numbers. Each was checked against the code and the
|
||||
board before anything was written:
|
||||
|
||||
| claim | measured |
|
||||
|---|---|
|
||||
| "`paOutcome` returned null on 596/725 rows" | **False.** Fire rate is **9,752/9,792 (99.6%)**. The only refusal reason on the whole board is `no_batter_profile: 40`. `pa_outcome_refused` never fires. |
|
||||
| "it needs the batter-vs-pitcher-hand split" *to run* | **False.** `paOutcome` and `hitOnContact` read `k_pct`, `bb_pct`, `barrel_pct`, `hard_hit_pct`, `avg_exit_velo`, `avg_launch_angle` and the pitcher's `k_pct`/`bb_pct`/`hard_hit_pct`. Neither reads `bats` or `throws` at all. A hand split cannot change whether they run. |
|
||||
| "82% FELL BACK to the seasonal rate, i.e. became the counter" | **No fallback path exists.** `baseballChain.chainFn` returns null on any refusal and `chain.applyChainFn` DROPS the atom. Nothing in the module reads a season rate as a substitute. |
|
||||
| "425/425 served-identical" | That figure is from commit `7c8ef8b` (the A1–A7 deploy verification), not from the chain. v1 measured **2/2 served-identical** in the harness and byte-identical served payloads. |
|
||||
|
||||
Chain v1 was also never committed or deployed and migration 038 was never
|
||||
applied, so no chain-shadow rows exist in production — the premise numbers cannot
|
||||
have come from a chain-shadow run.
|
||||
|
||||
**The premise was wrong about the mechanism. It was right about the thing that
|
||||
matters:** the chain was reading a hitter's SEASON rates, which already average
|
||||
his platoon split over whichever hands he happened to face. That is a season read
|
||||
wearing a matchup read's clothes, and un-averaging it is real work. So v2 plumbs
|
||||
the hand split — as a **conditioner on the rate**, not as a fix to the fire rate.
|
||||
|
||||
## 8.2 What was built
|
||||
|
||||
**The split enters at the per-PA hit rate, not at the output probability.**
|
||||
Multiplying `P(hits ≥ 1)` by a platoon factor would scale a number that has
|
||||
already been through the opportunity term — a different and wrong claim. So
|
||||
`baseballChain.chainFn` now writes the chain out explicitly:
|
||||
|
||||
```
|
||||
paOutcome → p_hit_per_pa (season)
|
||||
→ × platoonRead multiplier ← THE MATCHUP CONDITIONER
|
||||
→ Binomial(PA, p_hit) over paDistribution
|
||||
→ P(hits ≥ k)
|
||||
```
|
||||
|
||||
A test asserts that with no split supplied this is **arithmetically identical to
|
||||
`projectSkill`** across three archetypes × three PA values, so the restructuring
|
||||
cannot quietly become a second model.
|
||||
|
||||
**As-of-correct (A4).** The hitter's own split comes from `hitsFactorContext`
|
||||
(whose reads are `lte('as_of_date', asOf)`), the opposing starter's hand through
|
||||
the A6 `matchupKeys` resolve. The probe bounds every read at the row's own
|
||||
`game_date`. Nothing later than the grade can enter.
|
||||
|
||||
**Refuse, never substitute.** `platoonSeverity` already refuses below 60 PA on
|
||||
the smaller side and declares switch hitters unreadable. On a refusal the season
|
||||
rate stands **exactly** untouched (asserted to 6 dp) and the reason is recorded.
|
||||
|
||||
**Refusals are counted by reason** (`context.onRefusal`), and the count is taken
|
||||
off the **stored blocks**, not the legs — a prop with both an over and an under
|
||||
row produces two legs sharing one block, and the first draft reported 48.1% and
|
||||
55.1% for the same fact over two different denominators.
|
||||
|
||||
## 8.3 Measured — real board, 2026-08-07 → 08-12, 9,792 graded hits rows
|
||||
|
||||
```
|
||||
CHAIN FIRE — 9,752/9,792 atoms read (40 refused) across 83/83 games
|
||||
refusal reasons: { no_batter_profile: 40 }
|
||||
|
||||
HAND SPLIT — fired on 288/553 unique props (52.1%)
|
||||
season-rate reasons:
|
||||
insufficient_split_sample 130 the honest 60-PA refusal
|
||||
no_pitcher_hand 71 the only FIXABLE gap
|
||||
switch_hitter_side_value_unknown 60 genuinely unreadable
|
||||
no_splits / missing_split 4
|
||||
|
||||
CHAIN vs COUNTER — n=9,752
|
||||
min p25 median p75 max mean
|
||||
chain_p 0.016 0.359 0.498 0.641 0.984 0.500
|
||||
counter_p 0.050 0.397 0.500 0.603 0.950 0.500
|
||||
|divergence| 0.000 0.044 0.095 0.164 0.509 0.113
|
||||
|
||||
|div| < 0.02 1,142 (11.7%) 0.10–0.20 3,129 (32.1%)
|
||||
0.02–0.05 1,608 (16.5%) >= 0.20 1,538 (15.8%)
|
||||
0.05–0.10 2,335 (23.9%)
|
||||
|
||||
OVER SIDE ONLY (n=4,879): mean +0.045, chain HIGHER on 66.3%
|
||||
|
||||
MATCHUP READ vs SEASON READ
|
||||
split FIRED (n=5,372 rows) median |div| 0.091 · >= 0.10 on 46.0% · within 0.02 on 13.0%
|
||||
split REFUSED (n=4,380 rows) median |div| 0.100 · >= 0.10 on 50.1% · within 0.02 on 10.2%
|
||||
```
|
||||
|
||||
## 8.4 Reading it honestly
|
||||
|
||||
- **The hand split fires on 52.1% of props**, up from 0. Of the 47.9% that do
|
||||
not, **190 of 265 (72%) are principled refusals** — a thin split or a switch
|
||||
hitter. Only `no_pitcher_hand` (71) is a plumbing gap, and it is the A5 shape
|
||||
exactly: the hitter's split is sitting right there and the pitcher hand is
|
||||
missing because that player had no lineup row.
|
||||
- **Feeding the split widened the divergence**: median |div| 0.082 → **0.095**,
|
||||
and the over-side lean +2.9pp → **+4.5pp**. The chain moved further from the
|
||||
counter, which is what a matchup conditioner should do and is *not* evidence
|
||||
it moved in the right direction.
|
||||
- **The rows where the split FIRED disagree with the counter slightly LESS**
|
||||
(median 0.091 vs 0.100) than the rows where it refused. **This comparison is
|
||||
confounded and must not be read as an effect**: a hitter with 60+ PA on both
|
||||
sides is an established regular, and the counter has more game log on him too.
|
||||
It is two different populations, not two treatments.
|
||||
- **Divergence is still not merit.** Nothing here says the chain is closer to
|
||||
what happened. That needs settled outcomes against the stored triple, and it
|
||||
goes through the gate as a new hypothesis against the cumulative denominator.
|
||||
- **The probe remains contaminated** for the batter profile
|
||||
(`statcast_aggregates` keeps one as-of date) and so reports a divergence, never
|
||||
a resolution. The forward cron accrual does not have this problem.
|
||||
|
||||
## 8.5 Still not claimed
|
||||
|
||||
Unchanged from §5: nothing promoted, no head-to-head run, `CALIBRATION_DEPLOYED`
|
||||
still `[]`, A8 untouched, `hitsFactors` untouched, the counter still serves,
|
||||
WNBA untouched.
|
||||
|
||||
## 8.6 Second incidental defect found and fixed
|
||||
|
||||
`hitsFactorContext.build` and `matchupKeys.build` were both gated on
|
||||
`require('../utils/supabase').getSupabaseServiceClient()` called inline, so the
|
||||
entire factor and hand-split path was **unreachable from any test**. "It is
|
||||
wired" could only ever have rested on reading the code — which is precisely how
|
||||
A5 shipped three factors that never fired. Both now read `deps.supabase ||` the
|
||||
real client (additive; `undefined` gives identical behaviour).
|
||||
|
||||
---
|
||||
|
||||
# §9 — chain v3: the opportunity term, and the artifact/signal diagnostic (2026-08-12)
|
||||
|
||||
## 9.1 The premise, again checked first — one half right, one half wrong
|
||||
|
||||
**RIGHT, and a real defect:** the v2 shadow passed **no `lineupSlotFor`**, so
|
||||
every hitter fell to `skillProjection.DEFAULT_PA = 4.1`. "A regular" was asserted
|
||||
about the leadoff man and the nine hole alike. The chain is `rate × opportunity`
|
||||
and the opportunity half was a constant across the entire lineup — a free,
|
||||
known, pre-game fact thrown away. Fixed.
|
||||
|
||||
**WRONG about the conversion, and about the direction:**
|
||||
|
||||
- *"mis-handles the multiple-chances structure"* — it does not.
|
||||
`paDistribution` is a mean-preserving two-point mixture, and
|
||||
`atLeast(Binomial(n, p), 1)` **is** `1 − (1−p)^n` averaged over n. The order's
|
||||
proposed formula is what the code already computes. The defect was the INPUT
|
||||
`E[PA]`, not the conversion.
|
||||
- *"5.7pts BELOW the counter on 92% of rows"* — **measured the other way.** On
|
||||
the over side the chain runs **+4.5pp ABOVE** the counter and is below on
|
||||
**33.6%**. It was never 92%-below, at any point, in any measurement here.
|
||||
|
||||
## 9.2 Phase 1 — the E[PA] fix, A/B on identical rows
|
||||
|
||||
`matchup_keys` now carries `batting_order` (one extra column on a read it
|
||||
already performs — no new query), and the shadow feeds it as the opportunity
|
||||
term. Absent ⇒ `DEFAULT_PA` and the block records `default_regular`, so a
|
||||
league-shaped opportunity term is never mistaken for a posted one.
|
||||
|
||||
```
|
||||
OPPORTUNITY — posted lineup slot on 491/553 props (88.8%)
|
||||
fell to a default regular: 62
|
||||
|
||||
n mean median chain BELOW counter
|
||||
REAL E[PA] (posted slot) 4,879 +0.0450 +0.0499 33.6%
|
||||
CONSTANT 4.1 (the v2 shadow) 4,879 +0.0322 +0.0363 39.6%
|
||||
```
|
||||
|
||||
**The uniform bias did not collapse, because it was never there to collapse.**
|
||||
The fix moved the chain *further above* the counter, not toward it — which is
|
||||
correct behaviour, not a regression: real slots raise E[PA] for the top of the
|
||||
order and lower it for the bottom, and top-of-order hitters are over-represented
|
||||
in the prop board.
|
||||
|
||||
## 9.3 Phase 2 — THE DIAGNOSTIC: artifact or signal?
|
||||
|
||||
Over side, n=3,705 rows carrying an opposing-pitcher profile, quintiles of
|
||||
opposing-pitcher K% (low = soft matchup):
|
||||
|
||||
| bucket | opp K% | n | mean divergence | chain below counter |
|
||||
|---|---|---|---|---|
|
||||
| Q1 | 10.8–18.4 | 741 | **+0.0749** | 27.8% |
|
||||
| Q2 | 18.4–20.3 | 741 | +0.0560 | 30.8% |
|
||||
| Q3 | 20.3–23.1 | 741 | +0.0287 | 34.1% |
|
||||
| Q4 | 23.1–26.6 | 741 | +0.0396 | 35.0% |
|
||||
| Q5 | 26.6–40.6 | 741 | **+0.0256** | 38.6% |
|
||||
|
||||
**Q1 − Q5 spread +0.0494 · Pearson r(oppK, divergence) = −0.119**
|
||||
|
||||
### The answer is BOTH, and the order's binary framing does not fit
|
||||
|
||||
The order asked for *uniform (broken)* **or** *difficulty-correlated (signal)*.
|
||||
The data is a **mixture of the two, and both components should be named**:
|
||||
|
||||
- **A difficulty-correlated component, ~+0.049 across the range.** The chain
|
||||
reads soft matchups higher and hard matchups lower than the counter does. The
|
||||
`below %` is **monotone across all five quintiles** (27.8 → 30.8 → 34.1 → 35.0
|
||||
→ 38.6), which is a cleaner signature than the means (Q4 breaks order).
|
||||
- **A uniform positive offset of ~+0.026.** Even in the HARDEST quintile the
|
||||
chain sits +2.6pp above the counter. That floor does not move with the matchup
|
||||
and is therefore not conditioning — it is exactly the shape a residual
|
||||
mechanical bias makes. It is smaller than the conditioning component but it
|
||||
has not been explained, and calling the whole result "signal" would bury it.
|
||||
|
||||
### The caveat that must travel with the correlation
|
||||
|
||||
**This is close to mechanically guaranteed and is NOT evidence of correctness.**
|
||||
The chain reads the opposing pitcher's K rate directly (log5 odds-ratio in
|
||||
`paOutcome`) and the counter reads nothing about the pitcher at all. So a
|
||||
monotone relationship between opposing-pitcher K% and chain-minus-counter is
|
||||
approximately a proof that *the wiring works* — that the pitcher input reaches
|
||||
the number. It says nothing about whether the adjustment is the right size, the
|
||||
right direction on any individual row, or better than ignoring the pitcher.
|
||||
|
||||
**Does the chain see the game, or just miscompute it?** It demonstrably
|
||||
CONDITIONS on the game — the pitcher input reaches the forecast and moves it in
|
||||
the theorised direction. Whether that conditioning is *right* is unanswerable
|
||||
from a divergence and needs settled outcomes against the stored triple. There is
|
||||
also a residual ~2.6pp offset that conditioning does not explain and that should
|
||||
be chased before anyone reads the correlation as a win.
|
||||
|
||||
## 9.4 Open item created by this measurement
|
||||
|
||||
The ~+0.026 floor. Candidates not yet tested: the chain is unclamped where the
|
||||
counter clamps to `[0.10, 0.95]` (chain min 0.016 vs counter min 0.050); the
|
||||
`LEAGUE.babip = 0.291` anchor in `hitOnContact`; the ±35% BABIP bound. Naming it
|
||||
as unexplained is the honest state — it is not yet an artifact and not yet
|
||||
signal.
|
||||
|
||||
## 9.5 Unchanged
|
||||
|
||||
Nothing promoted, no head-to-head, `CALIBRATION_DEPLOYED` still `[]`, A8 and
|
||||
`hitsFactors` untouched, the counter still serves, WNBA untouched. Migration 038
|
||||
is still an unapplied blocking precondition of deploy.
|
||||
@@ -0,0 +1,208 @@
|
||||
# WNBA v1 — the possession / usage feed
|
||||
|
||||
**Status:** DATA INFRA ONLY. No chainFn, no archetype wiring, no shadow, nothing served.
|
||||
**Date:** 2026-08-13
|
||||
**Consumer it exists for:** the chain's basketball `chainFn`
|
||||
(`usage × possessions × efficiency`), which cannot be written without it.
|
||||
|
||||
---
|
||||
|
||||
## 1. Phase 1 — the source, established before any plumbing
|
||||
|
||||
### Candidates and the choice
|
||||
|
||||
| candidate | verdict |
|
||||
|---|---|
|
||||
| Python `nba_api` service (WNBA endpoints) | **Rejected as a dependency.** The service is offline in production; a feed we cannot fetch is not a feed. |
|
||||
| ESPN site API (`site.api.espn.com/.../basketball/wnba`) | **Chosen.** Free, no auth, already the host for WNBA schedules, box scores and live tracking. |
|
||||
|
||||
### What it actually returns — verified live, 2026-08-13
|
||||
|
||||
`summary?event={id}` → `boxscore.players[].statistics[0]`:
|
||||
|
||||
```
|
||||
keys: minutes, points, fieldGoalsMade-fieldGoalsAttempted,
|
||||
threePointFieldGoalsMade-threePointFieldGoalsAttempted,
|
||||
freeThrowsMade-freeThrowsAttempted, rebounds, assists, turnovers,
|
||||
steals, blocks, offensiveRebounds, defensiveRebounds, fouls, plusMinus
|
||||
|
||||
Breanna Stewart: ['30','19','8-21','1-5','2-2','7','5','0','1','0','2','5','2','-2']
|
||||
```
|
||||
|
||||
plus `starter` / `didNotPlay` / `ejected` per athlete, and `boxscore.teams[]` with
|
||||
team totals (`fieldGoalsMade-fieldGoalsAttempted`, `freeThrowsMade-freeThrowsAttempted`,
|
||||
`totalTurnovers`, `offensiveRebounds`). Team minutes sum to 200 in regulation
|
||||
(5 × 40), so overtime enters through the same sum. Play-by-play is also present
|
||||
(378 plays on the sampled game) but is not needed — the box identities suffice.
|
||||
|
||||
### chainFn inputs: delivered, derived, or missing
|
||||
|
||||
| chainFn input | status | how |
|
||||
|---|---|---|
|
||||
| **minutes** | **SERVED** | `minutes` per player |
|
||||
| **usage rate** | **DERIVED (exact)** | `100·((FGA+0.44·FTA+TOV)·(TmMIN/5)) / (MIN·(TmFGA+0.44·TmFTA+TmTOV))` |
|
||||
| **possessions** | **DERIVED (exact)** | `FGA − OREB + TOV + 0.44·FTA` |
|
||||
| **pace** | **DERIVED (exact)** | `possessions × 40 / (TmMIN/5)` |
|
||||
| **efficiency** | **DERIVED (exact)** | TS% `PTS/(2(FGA+0.44·FTA))`, eFG% `(FGM+0.5·FG3M)/FGA` |
|
||||
| **game state** (redistribute hook) | **SERVED** | `starter`, team/opp score → `final_margin` |
|
||||
| shot location / zone | **MISSING** | not in the box score; not a chainFn input today |
|
||||
| on/off, lineup combinations | **MISSING** | would need play-by-play reconstruction |
|
||||
| opponent defensive rating | **MISSING** | derivable later from the same rows (each game is also the opponent's) |
|
||||
|
||||
**"Derived" is not "proxied", and the distinction is load-bearing.** ESPN does not
|
||||
serve a usage rate, a possession count or a pace figure — it serves their
|
||||
components, and these are the standard identities recomputed from counted events.
|
||||
The only estimated term anywhere is the **0.44 free-throw-trip coefficient**,
|
||||
which is the field-standard value.
|
||||
|
||||
### Point-in-time capability
|
||||
|
||||
**Native. No `statcast_history`-style split is needed, and this is a design
|
||||
choice made at line one rather than a retrofit.**
|
||||
|
||||
`statcast_aggregates` had to grow a history twin because it stores a season
|
||||
aggregate upserted in place, destroying every prior version — which is why the
|
||||
first skill backtest was honest only by accident. This stores **per-game rows**.
|
||||
A completed box score never changes, so an as-of profile is
|
||||
`WHERE game_date < asOf`: a filter over immutable facts. Nothing is overwritten,
|
||||
so there is nothing to retain a history *of*.
|
||||
|
||||
**Strictly `<`, never `<=`.** A game on the as-of date may have tipped after
|
||||
grade time; counting it leaks the evening being predicted into the prediction.
|
||||
Same rule as `snapshotSettlementService.isPreGame`.
|
||||
|
||||
---
|
||||
|
||||
## 2. Phase 2 — the feed
|
||||
|
||||
**`wnba_player_game`** (migration 039). Key `(game_id, source_id)`, registered in
|
||||
`src/utils/tableKeys.js`, so `safePaginate` can walk it. Indexed
|
||||
`(sport, player_key, game_date DESC)` — every read is "this player, before this
|
||||
date".
|
||||
|
||||
Scoped to the chainFn's inputs plus the redistribute game-state. **Not a general
|
||||
WNBA stats dump** — `closing_captures` grew to 4.2M rows of something nothing
|
||||
read, and the lesson is to ingest for a named consumer or not at all.
|
||||
|
||||
Every derived rate is stored **alongside its components**, so it is re-derivable
|
||||
and checkable rather than an unfalsifiable number — the `factorFreeze` rule
|
||||
applied to a feed.
|
||||
|
||||
### Coverage — full live season, measured
|
||||
|
||||
```
|
||||
rows 5,085 games 256 players 241
|
||||
teams 15 days 88 errors 0
|
||||
date range 2026-05-01 .. 2026-08-12
|
||||
|
||||
FIELD COVERAGE
|
||||
minutes 5085/5085 starter 5085/5085
|
||||
usage_rate 5072/5085 final_margin 5085/5085
|
||||
team_possessions 5085/5085 ts_pct 4821/5085
|
||||
team_pace 5085/5085 efg_pct 4772/5085
|
||||
```
|
||||
|
||||
The `ts_pct` / `efg_pct` gaps are **correct refusals, not missing data**: a
|
||||
player with zero field-goal and free-throw attempts has no true-shooting
|
||||
percentage, and returning 0 would assert he shot and missed.
|
||||
|
||||
### Two contaminating games found and excluded
|
||||
|
||||
The season pull surfaced two games that are **not league games**:
|
||||
|
||||
- `2026-05-02` — an exhibition against **Nigeria** (`NIGER` vs `IND`)
|
||||
- `2026-07-25` — the **All-Star game** (`SPO` vs `COOP`)
|
||||
|
||||
Their usage and pace context is meaningless for a forward projection (an
|
||||
All-Star game has no defence to speak of), and leaving them in would pollute
|
||||
every profile spanning those dates.
|
||||
|
||||
The filter reads **ESPN's own `/teams` endpoint** rather than a hardcoded
|
||||
fifteen — an expansion franchise is admitted the day the league adds it, and
|
||||
this league has expanded twice in three years. An **empty** teams response is
|
||||
treated as unknown membership and filters nothing, because a failed feed
|
||||
degrading to an empty index looks exactly like an honest absence (the
|
||||
`fielding_oaa` lesson).
|
||||
|
||||
---
|
||||
|
||||
## 3. Phase 3 — verification, on real data through the real functions
|
||||
|
||||
`scripts/wnba-feed-verify.js` drives the production parsers over the live season
|
||||
and then calls `wnbaUsageService.profileAsOf` against those rows. It exercises
|
||||
the shipped code, not a restatement of it.
|
||||
|
||||
### The derivation, checkable by hand
|
||||
|
||||
```
|
||||
Laura Juskaite (TOR) 2026-05-01
|
||||
MIN 28 FGA 10 FTA 2 TOV 3 PTS 6
|
||||
team: MIN 200 FGA 68 FTA 16 TOV 11 OREB 7
|
||||
|
||||
possessions stored 79.04 recomputed 79.040
|
||||
usage_rate stored 23.046 recomputed 23.046
|
||||
```
|
||||
|
||||
### The as-of read
|
||||
|
||||
```
|
||||
player: Natasha Howard — 36 games on record
|
||||
|
||||
asOf 2026-05-12 REFUSED (2 games precede — under the 3-game floor)
|
||||
asOf 2026-06-24 games=18 (expected 18) usage 24.826 min/g 28.11 pace 82.683 TS 0.5998
|
||||
asOf 2026-08-12 games=35 (expected 35) usage 22.057 min/g 28.51 pace 82.275 TS 0.6010
|
||||
asOf 2026-12-31 games=36 (expected 36) usage 22.027 min/g 28.50 pace 82.241 TS 0.6042
|
||||
|
||||
STRICT CUTOFF — a game ON the as-of date is EXCLUDED (correct)
|
||||
THIN HISTORY — asOf before his first game returns null (refused, correct)
|
||||
```
|
||||
|
||||
The profile **moves with the date** (24.8 usage in June, 22.1 by August), which
|
||||
is the property that makes it a point-in-time read rather than a season number
|
||||
wearing a date.
|
||||
|
||||
### Two refusals that are features
|
||||
|
||||
- **Thin history ⇒ `null`, never a league-average player.** Usage feeds the
|
||||
chain's opportunity term directly; a stand-in would assert a usage rate about
|
||||
someone never observed. Same discipline as `platoonSeverity`'s 60-PA floor.
|
||||
- **Usage is MINUTES-WEIGHTED, not a flat mean across games.** The lineup-K-rate
|
||||
lesson: an unweighted aggregate counts a 6-minute cameo like a 34-minute start,
|
||||
and unweighted actively *hurt* that model.
|
||||
|
||||
---
|
||||
|
||||
## 4. Not done, and not claimed
|
||||
|
||||
- **No chainFn.** Basketball's `usage × possessions × efficiency` is the next
|
||||
order. This order stops at the feed.
|
||||
- **No archetype wiring, no shadow, no serving.** Nothing reads this table yet.
|
||||
- **MLB untouched.** A test asserts `chainShadow.js`, `baseballChain.js` and
|
||||
`analyzeViaEngine1.js` contain no reference to the WNBA feed.
|
||||
- `CALIBRATION_DEPLOYED` still `[]`; A8 accrual, the chain's MLB shadow and the
|
||||
served grade are untouched.
|
||||
|
||||
---
|
||||
|
||||
## 5. ⛔ Blocked: the table is NOT created and NOT ingested
|
||||
|
||||
`SUPABASE_DB_PASSWORD` in the local `.env` **fails authentication** against the
|
||||
pooler (`aws-1-us-east-1...` → `password authentication failed for user
|
||||
"postgres"`), and the direct host resolves IPv6-only, which is unreachable from
|
||||
this WSL2 environment. The transposed project ref already documented in
|
||||
CLAUDE.md is the likely cause.
|
||||
|
||||
So:
|
||||
|
||||
- **migration 039 is written and unapplied**
|
||||
- **`ingestRange` is written and has never written a row**
|
||||
- Everything in §2's coverage and §3's verification was measured by running the
|
||||
real parsers and the real as-of read over the live season **in memory**. The
|
||||
parse, the derivations and the as-of logic are verified on real data; the
|
||||
**database round-trip is not**.
|
||||
|
||||
That distinction is stated rather than glossed: a feed that has never been
|
||||
written is not a feed yet.
|
||||
|
||||
**Also still outstanding: migration 038 (`model_snapshots.chain_shadow`)** from
|
||||
the chain orders, for the same reason.
|
||||
@@ -0,0 +1,294 @@
|
||||
# WNBA source survey — possession / play-by-play options
|
||||
|
||||
**Status:** READ-ONLY SURVEY. Nothing built, nothing decided, nothing ingested.
|
||||
**Date:** 2026-08-13
|
||||
**Question:** what is the best source to build the basketball chainFn's possession
|
||||
feed — including the feedback layer — on?
|
||||
|
||||
---
|
||||
|
||||
## 0. Correction to the premise, and the real gap
|
||||
|
||||
The v1 report did **not** conclude "box scores only, no feedback loop". It
|
||||
recorded that ESPN play-by-play *is* present ("378 plays on the sampled game")
|
||||
and that it was not needed for the box identities, and it listed shot location
|
||||
and on/off as MISSING **from the box endpoint**.
|
||||
|
||||
The real gap the order names is correct though: **one source was surveyed.** And
|
||||
that produced a false ceiling — this survey found event-level data, shot
|
||||
coordinates, and pre-parsed per-player possessions across four other sources.
|
||||
|
||||
**The single most important finding is an environment artifact, not a data fact:**
|
||||
|
||||
> `stats.wnba.com` fails over IPv6 and **works over IPv4**. Default curl (and
|
||||
> Node's default resolver order) picks the AAAA record, connects, and hangs. The
|
||||
> same IPv6 pathology that made `db.<ref>.supabase.co` unreachable in WSL2.
|
||||
> Anything in this environment that reports a stats.nba/stats.wnba endpoint as
|
||||
> "unreachable" should be re-tested with `-4` before it is believed.
|
||||
|
||||
---
|
||||
|
||||
## 1. `stats.wnba.com` / `playbyplayv3` — what `nba_api` wraps
|
||||
|
||||
**Reachable: YES, over IPv4 only. Sustained ingest: NO.**
|
||||
|
||||
```
|
||||
curl -4 ... "https://stats.wnba.com/stats/playbyplayv3?GameID=1022600001&StartPeriod=1&EndPeriod=4"
|
||||
→ HTTP 200, 207,769 bytes, 3.78s
|
||||
|
||||
game.actions: 469 events
|
||||
fields: actionId, actionNumber, actionType, clock, description, isFieldGoal,
|
||||
location, period, personId, playerName, playerNameI, pointsTotal,
|
||||
scoreAway, scoreHome, shotDistance, shotResult, shotValue, subType,
|
||||
teamId, teamTricode, xLegacy, yLegacy
|
||||
|
||||
1 PT10M00.00S Jump Ball J. Jones Jump Ball Jones vs. Griner
|
||||
1 PT09M44.00S Missed Shot A. Morrow MISS Morrow 27' 3PT Jump Shot
|
||||
1 PT09M39.00S Rebound M. Johannes Johannes REBOUND (Off:0 Def:1)
|
||||
```
|
||||
|
||||
**Resolution:** full event stream with **shot coordinates** (`xLegacy`/`yLegacy`),
|
||||
`shotDistance`, `shotValue`, `shotResult`, `personId`, running score. Richer raw
|
||||
detail than anything else surveyed.
|
||||
|
||||
**Rate limits — the disqualifier.** One call succeeded. Every subsequent call
|
||||
returned **HTTP 000 (connection accepted, then stalled)**: immediately after,
|
||||
after 5s, and again after a ~3 minute cool-down at a 50s timeout. Two sibling
|
||||
endpoints (`boxscoreadvancedv3`, `boxscoreplayertrackv3`) returned 000 on first
|
||||
attempt. This is the well-known stats.nba.com throttle posture.
|
||||
|
||||
**`nba_api` is not installed locally** (`ModuleNotFoundError`); it is pinned in
|
||||
`src/services/python/requirements.txt:7` — the offline Python service.
|
||||
|
||||
**Point-in-time:** per-game, immutable ⇒ native, same as every source here.
|
||||
**Verdict:** richest raw feed, **hostile to sustained ingest**. Not a dependency.
|
||||
|
||||
---
|
||||
|
||||
## 2. `wehoop` / sportsdataverse — the bulk archive
|
||||
|
||||
**Reachable: YES. Best cost profile of any option.**
|
||||
|
||||
```
|
||||
GET github.com/sportsdataverse/wehoop-wnba-data/raw/main/wnba/pbp/parquet/play_by_play_2026.parquet
|
||||
→ 3,090,207 bytes (the WHOLE 2026 season, one file)
|
||||
|
||||
archive: play_by_play_2021 (2.63 MB) … 2024 (3.39) … 2025 (4.06) … 2026 (3.09 MB)
|
||||
|
||||
columns include: athlete_id_1/2/3, athlete_name_1/2/3, coordinate_x, coordinate_y,
|
||||
coordinate_x_raw, coordinate_y_raw, clock_display_value, clock_minutes,
|
||||
clock_seconds, period_number, home_score, away_score, score_value,
|
||||
shooting_play, team_id, type_id, type_text, wallclock, home_team_spread, …
|
||||
```
|
||||
|
||||
**What it wraps:** ESPN's WNBA feeds, pre-collected and normalised.
|
||||
**Coverage:** 2021 → live 2026, updated through the season.
|
||||
**License:** repo reports `NOASSERTION` (sportsdataverse projects are generally
|
||||
MIT/CC-BY; **the license needs confirming before redistribution** — it does not
|
||||
block internal analytical use, but it is not a clean SPDX tag).
|
||||
**Access:** one HTTP GET per season. No rate limit, GitHub CDN.
|
||||
**Footprint:** ~3 MB/season compressed — the cheapest full-PBP option by an order
|
||||
of magnitude, and materially relevant with the DB at 406/500 MB.
|
||||
|
||||
**Verdict:** the **backfill and cross-check** source. Cannot serve tonight's game
|
||||
(it lags the live feed), so it is not the freshness layer.
|
||||
|
||||
---
|
||||
|
||||
## 3. PBPStats (`api.pbpstats.com`) — possessions already parsed
|
||||
|
||||
**Reachable: YES. Richest *derived* layer. This is the standout.**
|
||||
|
||||
```
|
||||
GET api.pbpstats.com/get-games/wnba?Season=2026&SeasonType=Regular Season
|
||||
→ 250 games, 2026-05-08 .. 2026-08-12, HomePossessions/AwayPossessions on 250/250
|
||||
|
||||
{"GameId":"1022600001","Date":"2026-05-08","HomeTeamAbbreviation":"NYL",
|
||||
"HomePoints":106,"AwayPoints":75,"HomePossessions":87,"AwayPossessions":88}
|
||||
|
||||
GET api.pbpstats.com/get-game-stats?Type=Player&GameId=1022600001&League=wnba
|
||||
→ 47 fields PER PLAYER PER GAME:
|
||||
|
||||
Usage, OffPoss, DefPoss, Minutes, Points, TsPct, EfgPct, ShotQualityAvg,
|
||||
SecondChanceOffPoss, PenaltyOffPoss, PenaltyOffPossPct, PenaltyDefPoss,
|
||||
Arc3FGA, Arc3Frequency, AtRimFG3AFrequency, Avg2ptShotDistance,
|
||||
Avg3ptShotDistance, LongMidRangeAccuracy/FGA/FGM/Frequency,
|
||||
ShortMidRangeFGA/Frequency, Blocked2s, FoulsDrawn, ShootingFouls, …
|
||||
```
|
||||
|
||||
**This delivers the entire chainFn input set pre-computed**, and `OffPoss` is
|
||||
**better than the derived team possessions in v1**: it is the possessions the
|
||||
player was actually on the floor for — the true usage denominator, not a team
|
||||
estimate apportioned by minutes.
|
||||
|
||||
**WOWY (with-or-without-you) — the feedback layer, and it works:**
|
||||
|
||||
```
|
||||
GET api.pbpstats.com/get-wowy-combination-stats/wnba?Season=2026&SeasonType=Regular Season
|
||||
&TeamId=1611661313&PlayerIds=1629568
|
||||
→ HTTP 200
|
||||
{"OffRtg":113.41,"DefRtg":109.77,"NetRtg":3.65,"Minutes":1370.0,
|
||||
"On":"","Off":"Kennedy Burke", …}
|
||||
```
|
||||
|
||||
On/Off splits by player combination — the direct measurement of what changes when
|
||||
a player is off the floor.
|
||||
|
||||
**Rate posture:** 6 rapid sequential calls → **200, 200, 200, 200, 200, 200.** No
|
||||
throttling observed. (Endpoints are undocumented and unversioned; `get-possessions`
|
||||
returned 500 with valid params and `get-lineup-stats` 404 — the surface is
|
||||
uneven, and a 422 helpfully names missing params.)
|
||||
|
||||
**Point-in-time:** per-game rows ⇒ native.
|
||||
**Footprint:** ~5,000 player-game rows/season × 47 fields ≈ single-digit MB.
|
||||
**Verdict:** **the primary source.**
|
||||
|
||||
---
|
||||
|
||||
## 4. Basketball-Reference `/wnba/` — scrapeable, and permitted
|
||||
|
||||
**Reachable: YES. `/wnba/` is NOT disallowed.**
|
||||
|
||||
```
|
||||
robots.txt User-agent: * Crawl-delay: 3
|
||||
Disallow: /basketball/ (…team paths…) ← /wnba/ is absent
|
||||
|
||||
GET /wnba/boxscores/202605080NYL.html → 200 426,971 b
|
||||
GET /wnba/boxscores/pbp/202605080NYL.html → 200 245,516 b
|
||||
GET /wnba/boxscores/shot-chart/202605080NYL.html → 200 174,986 b
|
||||
GET /wnba/boxscores/plus-minus/202605080NYL.html → linked from the boxscore
|
||||
|
||||
3 sequential pbp fetches at the stated 3s crawl-delay → 200, 200, 200
|
||||
```
|
||||
|
||||
**Note:** the earlier 404s in this survey were a wrong URL guess of mine
|
||||
(`…0PHO`), not a BBR limitation. The real ids come off
|
||||
`/wnba/years/2026_games.html`.
|
||||
|
||||
**Coverage:** play-by-play, shot charts AND plus-minus per game.
|
||||
**Cost:** HTML scrape + parse, 3s crawl-delay ⇒ ~13 min for a 250-game season.
|
||||
**ToS:** `Disallow` does not cover `/wnba/`; `Crawl-delay: 3` must be honoured.
|
||||
`GPTBot` is banned outright, so identify honestly and stay slow.
|
||||
**Verdict:** best **independent audit cross-check** (a second opinion on
|
||||
possessions from a different parser). Too slow and too brittle for primary.
|
||||
|
||||
---
|
||||
|
||||
## 5. ESPN WNBA (already wired) — the freshness layer
|
||||
|
||||
**Reachable: YES, already in production use.**
|
||||
|
||||
```
|
||||
summary?event=401857134 → plays: 378
|
||||
fields: awayScore, clock, coordinate, homeScore, id, participants, period,
|
||||
pointsAttempted, scoreValue, scoringPlay, sequenceNumber, shootingPlay,
|
||||
shortDescription, team, text, type, wallclock
|
||||
|
||||
1 9:45 Pullup Jump Shot athlete 4433403 coord {x:16,y:25} 3-0
|
||||
1 9:26 Fade Away Jump Shot athlete 2998928 coord {x:31,y:1} 3-2
|
||||
1 9:17 Driving Floating Jump Shot athlete 4433403 coord {x:30,y:3} 3-2
|
||||
```
|
||||
|
||||
Event-level with coordinates and athlete ids — so **shot location was never
|
||||
actually missing**; it was missing from the *box* endpoint I read in v1.
|
||||
|
||||
**Stability:** undocumented but long-lived, no auth, no observed throttle, and
|
||||
already the host for schedules/box/live-tracking in this codebase.
|
||||
**Verdict:** the **live/tonight** layer.
|
||||
|
||||
---
|
||||
|
||||
## 6. News / context layer — game state, actives, rest
|
||||
|
||||
**No scraper needed. ESPN already serves it.**
|
||||
|
||||
```
|
||||
GET .../basketball/wnba/injuries → HTTP 200, 14 teams, 46 entries
|
||||
Atlanta Dream | Brionna Jones | Out | Leg
|
||||
Chicago Sky | Maddy Westbeld | Out | Coach's Decision
|
||||
|
||||
summary?event=… also carries per-game injuries for both sides:
|
||||
Aliyah Boston Day-To-Day
|
||||
Caitlin Clark Day-To-Day
|
||||
Damiris Dantas Out
|
||||
```
|
||||
|
||||
**The one real limitation: this is a LATEST-ONLY snapshot.** There is no as-of
|
||||
query for injury state, so point-in-time actives require **capturing it daily** —
|
||||
which is exactly what the existing changedetection/cron stack is for. That is a
|
||||
small dated table, not a scraper build.
|
||||
|
||||
Miniflux/SearxNG/RSS would add *narrative* (beat-reporter rest news ahead of the
|
||||
official designation). Useful later; **not required** for the feedback layer,
|
||||
because ESPN injuries + `starter` + `Minutes` already answer "who played, who
|
||||
sat, who started".
|
||||
|
||||
---
|
||||
|
||||
## 7. Scorecard
|
||||
|
||||
| | possession-level for feedback? | as-of? | cost / footprint | reliability & ToS | live freshness |
|
||||
|---|---|---|---|---|---|
|
||||
| **PBPStats** | **YES — per-player `OffPoss`/`DefPoss`/`Usage` + WOWY on/off** | native (per-game) | ~single-digit MB/season | 6/6 rapid 200s; undocumented, uneven surface | good (through 08-12) |
|
||||
| **wehoop** | YES — full event stream | native | **~3 MB/season** (best) | GitHub CDN; license `NOASSERTION` — confirm | lags live |
|
||||
| **ESPN** | YES — 378 events w/ coords | native | small | already in prod, no throttle seen | **best (live)** |
|
||||
| **BBR** | YES — pbp + shot chart + plus-minus | native | scrape, 3s delay ⇒ ~13 min/season | `/wnba/` allowed, crawl-delay 3 | good |
|
||||
| **stats.wnba.com** | YES — richest raw (shot coords, 469 events) | native | small | **HOSTILE — 1 call then HTTP 000, still blocked after 3 min; IPv4-only** | good if you could call it |
|
||||
|
||||
---
|
||||
|
||||
## 8. Ranked recommendation
|
||||
|
||||
**A combination, not one source. Three roles:**
|
||||
|
||||
1. **PBPStats — PRIMARY.** It has already done the possession parsing, and
|
||||
per-player `OffPoss` is the correct usage denominator rather than my v1
|
||||
team-level estimate apportioned by minutes. `Usage`, `TsPct`, `ShotQualityAvg`
|
||||
and the shot-zone splits arrive free. **WOWY gives the feedback layer
|
||||
directly.** Tradeoff: undocumented, unversioned, single maintainer, uneven
|
||||
endpoint surface (a 500 and a 404 in this survey) — so it needs a fallback,
|
||||
and its numbers should be cross-checked once against a second parser.
|
||||
|
||||
2. **ESPN — LIVE + CONTEXT.** Tonight's game before PBPStats has it, plus
|
||||
injuries/actives. Already wired, already trusted in this codebase.
|
||||
|
||||
3. **wehoop — BACKFILL + CROSS-CHECK.** One 3 MB GET replaces 250 API calls for
|
||||
a historical season, and being an independent collection of the same ESPN
|
||||
feed it is a genuine second opinion. Confirm the license before anything
|
||||
leaves the building.
|
||||
|
||||
**Not recommended as a dependency:** `stats.wnba.com`. Richest raw data,
|
||||
unusable throttle. Worth keeping as a *manual* one-off tool now that the IPv4
|
||||
workaround is known.
|
||||
|
||||
**BBR:** hold as an audit path. Its independent possession parse is the best
|
||||
available check on PBPStats, at 3s/request.
|
||||
|
||||
### Can the FEEDBACK LAYER be built now?
|
||||
|
||||
**Yes — with PBPStats, and without deferring.** Two mechanisms are already
|
||||
available:
|
||||
|
||||
- **Usage redistribution:** per-player `Usage` + `OffPoss` per game, joined to
|
||||
who was out that night (ESPN injuries, captured daily). "How does this
|
||||
player's usage move in games where the primary creator sat" is then a direct
|
||||
measurement, not a model.
|
||||
- **Blowout → minutes:** `final_margin` (already in v1's feed) against
|
||||
`Minutes` and `OffPoss` per game.
|
||||
|
||||
**Ingest cost:** 1 call per game (~250/season) + 1 injuries call/day. Storage
|
||||
~single-digit MB — against 94 MB of current headroom.
|
||||
|
||||
**What is genuinely still missing:** *within-game* possession-by-possession
|
||||
lineup state (who was on the floor at each moment). `get-lineup-stats` 404s.
|
||||
Deriving it needs substitution events reconstructed from the raw PBP — available
|
||||
in wehoop/ESPN/BBR, but a real build. Not required for the two mechanisms above.
|
||||
|
||||
---
|
||||
|
||||
## 9. Nothing built
|
||||
|
||||
No ingest, no schema, no chainFn. The v1 feed (`wnba_player_game`, migration 039)
|
||||
remains written and unapplied; this survey is evidence that **its source choice
|
||||
should be revisited before it is applied** — PBPStats supersedes the derived-usage
|
||||
approach with a measured one.
|
||||
Reference in New Issue
Block a user