Files
vyndr/specs/wnba-source-survey.md
builtbykev 6c34af3414 checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion
recon (Tony resurrection, STEP 0). This work existed only on the
laptop disk.

- chain shadow accrual + probe script (038_chain_shadow.sql)
- WNBA possession feed: ESPN adapter, usage service, verify script
  (039_wnba_player_game.sql)
- baseball chain
- retention/snapshot service updates, tableKeys, matchupKeys
- specs: chain-v1, wnba-possession-feed, wnba-source-survey
- unit tests for the above

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
2026-08-14 16:53:37 -04:00

13 KiB
Raw Permalink Blame History

WNBA source survey — possession / play-by-play options

Status: READ-ONLY SURVEY. Nothing built, nothing decided, nothing ingested. Date: 2026-08-13 Question: what is the best source to build the basketball chainFn's possession feed — including the feedback layer — on?


0. Correction to the premise, and the real gap

The v1 report did not conclude "box scores only, no feedback loop". It recorded that ESPN play-by-play is present ("378 plays on the sampled game") and that it was not needed for the box identities, and it listed shot location and on/off as MISSING from the box endpoint.

The real gap the order names is correct though: one source was surveyed. And that produced a false ceiling — this survey found event-level data, shot coordinates, and pre-parsed per-player possessions across four other sources.

The single most important finding is an environment artifact, not a data fact:

stats.wnba.com fails over IPv6 and works over IPv4. Default curl (and Node's default resolver order) picks the AAAA record, connects, and hangs. The same IPv6 pathology that made db.<ref>.supabase.co unreachable in WSL2. Anything in this environment that reports a stats.nba/stats.wnba endpoint as "unreachable" should be re-tested with -4 before it is believed.


1. stats.wnba.com / playbyplayv3 — what nba_api wraps

Reachable: YES, over IPv4 only. Sustained ingest: NO.

curl -4 ... "https://stats.wnba.com/stats/playbyplayv3?GameID=1022600001&StartPeriod=1&EndPeriod=4"
→ HTTP 200, 207,769 bytes, 3.78s

game.actions: 469 events
fields: actionId, actionNumber, actionType, clock, description, isFieldGoal,
        location, period, personId, playerName, playerNameI, pointsTotal,
        scoreAway, scoreHome, shotDistance, shotResult, shotValue, subType,
        teamId, teamTricode, xLegacy, yLegacy

  1  PT10M00.00S  Jump Ball      J. Jones     Jump Ball Jones vs. Griner
  1  PT09M44.00S  Missed Shot    A. Morrow    MISS Morrow 27' 3PT Jump Shot
  1  PT09M39.00S  Rebound        M. Johannes  Johannes REBOUND (Off:0 Def:1)

Resolution: full event stream with shot coordinates (xLegacy/yLegacy), shotDistance, shotValue, shotResult, personId, running score. Richer raw detail than anything else surveyed.

Rate limits — the disqualifier. One call succeeded. Every subsequent call returned HTTP 000 (connection accepted, then stalled): immediately after, after 5s, and again after a ~3 minute cool-down at a 50s timeout. Two sibling endpoints (boxscoreadvancedv3, boxscoreplayertrackv3) returned 000 on first attempt. This is the well-known stats.nba.com throttle posture.

nba_api is not installed locally (ModuleNotFoundError); it is pinned in src/services/python/requirements.txt:7 — the offline Python service.

Point-in-time: per-game, immutable ⇒ native, same as every source here. Verdict: richest raw feed, hostile to sustained ingest. Not a dependency.


2. wehoop / sportsdataverse — the bulk archive

Reachable: YES. Best cost profile of any option.

GET github.com/sportsdataverse/wehoop-wnba-data/raw/main/wnba/pbp/parquet/play_by_play_2026.parquet
→ 3,090,207 bytes (the WHOLE 2026 season, one file)

archive: play_by_play_2021 (2.63 MB) … 2024 (3.39) … 2025 (4.06) … 2026 (3.09 MB)

columns include: athlete_id_1/2/3, athlete_name_1/2/3, coordinate_x, coordinate_y,
  coordinate_x_raw, coordinate_y_raw, clock_display_value, clock_minutes,
  clock_seconds, period_number, home_score, away_score, score_value,
  shooting_play, team_id, type_id, type_text, wallclock, home_team_spread, …

What it wraps: ESPN's WNBA feeds, pre-collected and normalised. Coverage: 2021 → live 2026, updated through the season. License: repo reports NOASSERTION (sportsdataverse projects are generally MIT/CC-BY; the license needs confirming before redistribution — it does not block internal analytical use, but it is not a clean SPDX tag). Access: one HTTP GET per season. No rate limit, GitHub CDN. Footprint: ~3 MB/season compressed — the cheapest full-PBP option by an order of magnitude, and materially relevant with the DB at 406/500 MB.

Verdict: the backfill and cross-check source. Cannot serve tonight's game (it lags the live feed), so it is not the freshness layer.


3. PBPStats (api.pbpstats.com) — possessions already parsed

Reachable: YES. Richest derived layer. This is the standout.

GET api.pbpstats.com/get-games/wnba?Season=2026&SeasonType=Regular Season
→ 250 games, 2026-05-08 .. 2026-08-12, HomePossessions/AwayPossessions on 250/250

  {"GameId":"1022600001","Date":"2026-05-08","HomeTeamAbbreviation":"NYL",
   "HomePoints":106,"AwayPoints":75,"HomePossessions":87,"AwayPossessions":88}

GET api.pbpstats.com/get-game-stats?Type=Player&GameId=1022600001&League=wnba
→ 47 fields PER PLAYER PER GAME:

  Usage, OffPoss, DefPoss, Minutes, Points, TsPct, EfgPct, ShotQualityAvg,
  SecondChanceOffPoss, PenaltyOffPoss, PenaltyOffPossPct, PenaltyDefPoss,
  Arc3FGA, Arc3Frequency, AtRimFG3AFrequency, Avg2ptShotDistance,
  Avg3ptShotDistance, LongMidRangeAccuracy/FGA/FGM/Frequency,
  ShortMidRangeFGA/Frequency, Blocked2s, FoulsDrawn, ShootingFouls, …

This delivers the entire chainFn input set pre-computed, and OffPoss is better than the derived team possessions in v1: it is the possessions the player was actually on the floor for — the true usage denominator, not a team estimate apportioned by minutes.

WOWY (with-or-without-you) — the feedback layer, and it works:

GET api.pbpstats.com/get-wowy-combination-stats/wnba?Season=2026&SeasonType=Regular Season
    &TeamId=1611661313&PlayerIds=1629568
→ HTTP 200
  {"OffRtg":113.41,"DefRtg":109.77,"NetRtg":3.65,"Minutes":1370.0,
   "On":"","Off":"Kennedy Burke", …}

On/Off splits by player combination — the direct measurement of what changes when a player is off the floor.

Rate posture: 6 rapid sequential calls → 200, 200, 200, 200, 200, 200. No throttling observed. (Endpoints are undocumented and unversioned; get-possessions returned 500 with valid params and get-lineup-stats 404 — the surface is uneven, and a 422 helpfully names missing params.)

Point-in-time: per-game rows ⇒ native. Footprint: ~5,000 player-game rows/season × 47 fields ≈ single-digit MB. Verdict: the primary source.


4. Basketball-Reference /wnba/ — scrapeable, and permitted

Reachable: YES. /wnba/ is NOT disallowed.

robots.txt  User-agent: *   Crawl-delay: 3
            Disallow: /basketball/   (…team paths…)   ← /wnba/ is absent

GET /wnba/boxscores/202605080NYL.html            → 200  426,971 b
GET /wnba/boxscores/pbp/202605080NYL.html        → 200  245,516 b
GET /wnba/boxscores/shot-chart/202605080NYL.html → 200  174,986 b
GET /wnba/boxscores/plus-minus/202605080NYL.html → linked from the boxscore

3 sequential pbp fetches at the stated 3s crawl-delay → 200, 200, 200

Note: the earlier 404s in this survey were a wrong URL guess of mine (…0PHO), not a BBR limitation. The real ids come off /wnba/years/2026_games.html.

Coverage: play-by-play, shot charts AND plus-minus per game. Cost: HTML scrape + parse, 3s crawl-delay ⇒ ~13 min for a 250-game season. ToS: Disallow does not cover /wnba/; Crawl-delay: 3 must be honoured. GPTBot is banned outright, so identify honestly and stay slow. Verdict: best independent audit cross-check (a second opinion on possessions from a different parser). Too slow and too brittle for primary.


5. ESPN WNBA (already wired) — the freshness layer

Reachable: YES, already in production use.

summary?event=401857134 → plays: 378
fields: awayScore, clock, coordinate, homeScore, id, participants, period,
        pointsAttempted, scoreValue, scoringPlay, sequenceNumber, shootingPlay,
        shortDescription, team, text, type, wallclock

  1  9:45  Pullup Jump Shot          athlete 4433403  coord {x:16,y:25}  3-0
  1  9:26  Fade Away Jump Shot       athlete 2998928  coord {x:31,y:1}   3-2
  1  9:17  Driving Floating Jump Shot athlete 4433403 coord {x:30,y:3}   3-2

Event-level with coordinates and athlete ids — so shot location was never actually missing; it was missing from the box endpoint I read in v1.

Stability: undocumented but long-lived, no auth, no observed throttle, and already the host for schedules/box/live-tracking in this codebase. Verdict: the live/tonight layer.


6. News / context layer — game state, actives, rest

No scraper needed. ESPN already serves it.

GET .../basketball/wnba/injuries → HTTP 200, 14 teams, 46 entries
   Atlanta Dream | Brionna Jones   | Out | Leg
   Chicago Sky   | Maddy Westbeld  | Out | Coach's Decision

summary?event=… also carries per-game injuries for both sides:
   Aliyah Boston  Day-To-Day
   Caitlin Clark  Day-To-Day
   Damiris Dantas Out

The one real limitation: this is a LATEST-ONLY snapshot. There is no as-of query for injury state, so point-in-time actives require capturing it daily — which is exactly what the existing changedetection/cron stack is for. That is a small dated table, not a scraper build.

Miniflux/SearxNG/RSS would add narrative (beat-reporter rest news ahead of the official designation). Useful later; not required for the feedback layer, because ESPN injuries + starter + Minutes already answer "who played, who sat, who started".


7. Scorecard

possession-level for feedback? as-of? cost / footprint reliability & ToS live freshness
PBPStats YES — per-player OffPoss/DefPoss/Usage + WOWY on/off native (per-game) ~single-digit MB/season 6/6 rapid 200s; undocumented, uneven surface good (through 08-12)
wehoop YES — full event stream native ~3 MB/season (best) GitHub CDN; license NOASSERTION — confirm lags live
ESPN YES — 378 events w/ coords native small already in prod, no throttle seen best (live)
BBR YES — pbp + shot chart + plus-minus native scrape, 3s delay ⇒ ~13 min/season /wnba/ allowed, crawl-delay 3 good
stats.wnba.com YES — richest raw (shot coords, 469 events) native small HOSTILE — 1 call then HTTP 000, still blocked after 3 min; IPv4-only good if you could call it

8. Ranked recommendation

A combination, not one source. Three roles:

  1. PBPStats — PRIMARY. It has already done the possession parsing, and per-player OffPoss is the correct usage denominator rather than my v1 team-level estimate apportioned by minutes. Usage, TsPct, ShotQualityAvg and the shot-zone splits arrive free. WOWY gives the feedback layer directly. Tradeoff: undocumented, unversioned, single maintainer, uneven endpoint surface (a 500 and a 404 in this survey) — so it needs a fallback, and its numbers should be cross-checked once against a second parser.

  2. ESPN — LIVE + CONTEXT. Tonight's game before PBPStats has it, plus injuries/actives. Already wired, already trusted in this codebase.

  3. wehoop — BACKFILL + CROSS-CHECK. One 3 MB GET replaces 250 API calls for a historical season, and being an independent collection of the same ESPN feed it is a genuine second opinion. Confirm the license before anything leaves the building.

Not recommended as a dependency: stats.wnba.com. Richest raw data, unusable throttle. Worth keeping as a manual one-off tool now that the IPv4 workaround is known.

BBR: hold as an audit path. Its independent possession parse is the best available check on PBPStats, at 3s/request.

Can the FEEDBACK LAYER be built now?

Yes — with PBPStats, and without deferring. Two mechanisms are already available:

  • Usage redistribution: per-player Usage + OffPoss per game, joined to who was out that night (ESPN injuries, captured daily). "How does this player's usage move in games where the primary creator sat" is then a direct measurement, not a model.
  • Blowout → minutes: final_margin (already in v1's feed) against Minutes and OffPoss per game.

Ingest cost: 1 call per game (~250/season) + 1 injuries call/day. Storage ~single-digit MB — against 94 MB of current headroom.

What is genuinely still missing: within-game possession-by-possession lineup state (who was on the floor at each moment). get-lineup-stats 404s. Deriving it needs substitution events reconstructed from the raw PBP — available in wehoop/ESPN/BBR, but a real build. Not required for the two mechanisms above.


9. Nothing built

No ingest, no schema, no chainFn. The v1 feed (wnba_player_game, migration 039) remains written and unapplied; this survey is evidence that its source choice should be revisited before it is applied — PBPStats supersedes the derived-usage approach with a measured one.