diff --git a/BUILD-STATE.md b/BUILD-STATE.md index 0d3e49a..49111b2 100755 --- a/BUILD-STATE.md +++ b/BUILD-STATE.md @@ -1,7 +1,182 @@ # VYNDR — Build State ## Last Updated -2026-08-03 +2026-08-09 + +## Fix A0 (2026-08-09) — Withdraw unknowable verdicts + read-integrity harness ✅ +4,569 tests / 363 suites green (repo config). **Measure-only: no production +reader edited, no eligibility opened, no delete, no settle, no schema.** +- **SHIPPED** `specs/read-integrity-harness.md` (spec-first), + `src/utils/readIntegrity.js` (pure + injectable core), + `scripts/read-integrity.js` (CLI + declarative 25-reader registry), + `src/services/model/withdrawnVerdicts.js` (append-only withdrawal record), + 30 new unit tests across two suites. +- **AC5 verified on prod:** the harness reproduces the probe — known-corrupt + `prove-hit-factors:136` at 16.5%, known-clean `challenger-scoreboard:112` at 0%. +- **Baseline: 11 FAIL / 13 PASS / 1 KEY_NOT_UNIQUE of 25.** Worst + `champion-ablation:173` 33.6%, `proven-status:76` 32.6%, + `calibrationService.fromLedger:110` (the one LIVE reader) 24.8%. +- **WITHDRAWN:** `defense_by_direction` + `pitcher_contact_profile` on hits → + `WITHDRAWN_PENDING_REAUDIT`. Nothing is PROVEN on hits from the factor gate now. +- **OPEN, urgent, NOT in this order's scope:** (1) the live + `calibrationService.fromLedger` reader is 24.8% corrupt in production; (2) + `hitsFactors.js` still serves both withdrawn factors AND `platoon_severity`, + which never passed the gate at all. Both are A1/A2 decisions. + +## Fix A1 (2026-08-09) — safePaginate + the first corrected reader ✅ +4,592 tests / 365 suites green, web build exit 0. **Served grade FROZEN; +`CALIBRATION_DEPLOYED` still `[]`; nothing re-certified, re-deployed or +un-withdrawn.** +- **SHIPPED** `src/utils/safePaginate.js` (stable ORDER BY on a unique key + + runtime uniqueness guard + error propagation, built on the tested + `readIntegrity.walk`), 13 unit tests; `calibrationService.loadSettledRows` + extracted + converted, 11 unit tests. +- **HARNESS: FAIL 24.8% → PASS (0 dup / 0 missing of 2,490)**, measured through + the REAL function via the new `readerRows` arm (`arm_a: real_reader_function`). + The unfixed twin `lowParamService.fromLedger:78` still measures 24.8% — the + control that proves the PASS is not a registry edit. +- **CORRECTION to A0:** this reader was never live. `CALIBRATION_DEPLOYED = []` + (snapshotService.js:301) means the loop never runs; calibrationService is only + the SHADOW; `chainAcross` has no callers. **A1 changed no served number.** +- **MAP DELTA (reported, NOT deployed):** certified-band error −0.046 → +0.010, + ceiling 0.833 → 0.810, span unchanged at 0.50–0.70, `fitted_through` shifted a + day because duplicates moved the time split. + +## Fix A2 (2026-08-09) — safePaginate rolled across every FAIL reader ✅ +4,628 tests / 366 suites, web build exit 0. **Served grade FROZEN; +`CALIBRATION_DEPLOYED` still `[]`; both hits verdicts still withdrawn; no +eligibility opened, no delete/settle.** +- **11 FAIL → 0 FAIL.** Full 26-reader baseline: **25 PASS + 1 KEY_NOT_UNIQUE** + (the `batter_spray` composite-key data fact, unchanged and correctly named). + Worst corruption 33.6% → **0%**. +- Every fixed reader verified through the harness's `real_reader_function` arm, + i.e. the code that actually runs — not a restated query, never a `fixed:true`. +- **PRIMARY calibrator first:** `lowParamService.fromLedger:78` 24.8% → PASS. +- **7 scripts converted** via `pageSafe` + named `READS` exports + a + `require.main === module` guard (needed because the harness imports them and + several write to `mc_test_ledger`; verified unchanged at 168 rows). +- **RESISTED THE HELPER:** the context tables have composite PKs with no single + unique column, so `paginate()` cannot order them. They measure 0% and are left + on `page()`, documented in-file. Composite-key ordering is the open item. +- **OPEN:** an unresolved intermittent in `snapshotService.test.js` (2/30 on + branch, 0/29 at HEAD — not statistically distinguishable, assertion never + captured). Recorded, not dismissed. + +## Fix A2b (2026-08-09) — every read clean by construction ✅ +4,699 tests / 366 suites, web build exit 0. **Grade FROZEN; `CALIBRATION_DEPLOYED` +still `[]`; verdicts still withdrawn; no eligibility, no settlement, no +version-stamp, no delete. No context data altered.** +- **34 readers, 34 PASS, 0 FAIL, 0 KEY_NOT_UNIQUE** — up from 26 readers / 1 + KEY_NOT_UNIQUE. 30 of 34 measured through the REAL reader function. +- **safePaginate takes composite keys**; `src/utils/tableKeys.js` holds the real + constraints from `pg_index`. All 7 context tables converted; `batter_spray` + resolved (wrong key, not bad data). +- **The 3 A2 legacy reads converted**; `KNOWN_LEGACY_READS` is now empty. +- **`calibrate-hits.js` found at 24.7%** during the sweep and converted. +- **The unordered `page()` helper is DELETED from all 15 scripts.** +- **Write-scripts assessed, data untouched:** omission-only failure mode; read is + 0% today; historical exposure bounded at ~6.6 players / ~0.16 games. + +## Fix A3 (2026-08-09) — eligibility can honestly count ✅ +4,723 tests / 367 suites, web build exit 0. **Grade FROZEN; `CALIBRATION_DEPLOYED` +still `[]`; `hitsFactors`, `hitsFactorContext`, `withdrawnVerdicts`, +`reAuditEligibility`, `snapshotService` all EMPTY DIFF. No verdict re-run, no +calibration re-fit, no historical re-stamp.** +- **ONE version source:** `src/config/modelVersion.js` = `engine1@2026-08-07-fullwindow`. + Both `ledgerService` and `retentionService` import it; the twin defaults that + silently blocked eligibility are gone. +- **Snapshot settlement on the cron:** `snapshotSettlementService` wired into + `snapshotScheduler` beside the ledger settle. Outcomes only, no context + reconstruction. Settled **15,484 → 28,894**; repaired-champion **0 → 7,556**; + 0 rows settled with a null actual_value. 24 new unit tests. +- **Eligibility (counted, not run):** 2 eligible dates against 10/10/14/14. + All four measurements remain BLOCKED — now for the honest reason. + +## Fix A4 (2026-08-09) — as-of-correct context ✅ +4,735 tests / 368 suites, web build exit 0, harness still 34/34 PASS. +**Grade output UNCHANGED (proven 400/400); inputs NOT frozen (that is A5); no +re-audit run, no verdict re-run, no calibration re-fit; `CALIBRATION_DEPLOYED` +still `[]`.** +- `hitsFactorContext.build(sb, {asOf})` — dated reads bounded `.lte('as_of_date')`, + latest-within-bound, **refusal on absence** (never nearest/latest). +- Dated statcast comes from `statcast_history` (aggregates keeps one date); + verified identical to aggregates at the head, 1,414 rows / 0 diffs. +- **Coverage on the two re-audit dates: spray/platoon/handedness 99% and 98.9%**, + 2 players short per date. **Defence 0% from the row** — `opponent` is NULL on + every settled repaired-champion row; it needs the game-log join. +- 12 new unit tests incl. the contamination lock (never a row dated after asOf). + +## Fix A5 (2026-08-10) — input freeze, and a finding that reframes it ✅ +4,751 tests / 369 suites, web build exit 0. **Grade UNCHANGED (496/496); no +re-audit, no verdict re-run, no calibration; `CALIBRATION_DEPLOYED` still `[]`.** +- **FINDING: the three hits factors NEVER FIRE.** 596/596 real props skipped, + multiplier 1 every time, because `prop.opponent` / `prop.opposing_pitcher` are + never set by anything. They are wired and inert. Plumbing them is a MODEL + CHANGE and needs its own order. +- **`model_snapshots.factor_inputs` jsonb** (migration 034, APPLIED) freezes the + raw inputs at grade time — never the multiplier, which stays recomputable via + `factorFreeze.recompute()`. Proven: frozen == read (496/496), recompute == live + (496/496). +- **Forward-only.** Existing 7,556 settled rows keep NULL `factor_inputs` and + NULL `opponent`; defence re-audit on them still needs the game-log join. + +## Fix A6 (2026-08-10) — join keys plumbed into a SHADOW resolve ✅ +4,762 tests / 370 suites, web build exit 0. **Served grade FROZEN; no re-audit, +no verdict reinstated, no calibration; `CALIBRATION_DEPLOYED` still `[]`.** +- `matchupKeys` resolves opponent + opposing pitcher from `lineup_context` + (as-of) + schedule probables. Refuses rather than guessing. +- **Key resolution 248/308 (80.5%)** on 2026-08-09; 60 refused (no lineup row). +- **Shadow fire: pitcher_contact 248, defense_by_direction 185, + platoon_severity 147** — up from 0/0/0. Multiplier median 1.017, range + 0.803–1.250; 245/248 would move the served number. +- Served p_win identical on every evaluated row; shadow recompute-check 248/248. + +## Fix A7 (2026-08-11) — shadow accrual proven + A8 pre-registered ✅ +4,772 tests / 371 suites, web build exit 0. **Nothing turned on. Served grade +frozen; `CALIBRATION_DEPLOYED` still `[]`; no verdict reinstated; gate NOT run.** +- **Accrual is automatic on the scheduled pass** — proven end-to-end through the + real `gradeAndCacheSlate → onGraded → rowsFromSides` chain (10 new tests), + plus assertions that `runSnapshot` builds the key index pre-grade and degrades + safely. +- **Complete (would_fire, outcome) pairs today: 0.** `factor_inputs` is + 0/122,276 — A5/A6 are not deployed. Accrual begins at deploy. +- **`specs/a8-shadow-factor-gate.md`** pre-registers the two-part gate (movement + AND Brier, cluster-resampled, cumulative-Bonferroni, new-test α) with named + fallbacks. Not run. +- **Binding constraint is CLUSTERS: ~10.5 settled games/slate ⇒ `MIN_CLUSTERS` + 40 in ~4 slates**, while `MIN_N` 500 is met in 1–2. A8 runnable ≈5 slates + after deploy. + +### Next +- **A8 — run the pre-registered gate** once ≈5 slates have accrued. +- **A9 — turn the factors on LIVE**, only if A8 returns PROVES, and with its own + before/after on the served board. +- **(superseded) A7 — turn the factors on LIVE.** This is the biggest served-grade change in + VYNDR's history: ~80% of hits props would move, median +1.7%, tails ±20-25%. + Gate it on (a) accrued shadow evidence, (b) a re-proof of the withdrawn + verdicts on rows where the factors actually fire. +- **(superseded) A6 — plumb `opponent` + `opposing_pitcher` into the graded prop.** This is a + MODEL CHANGE (grades will move): measure the delta against the frozen + `available` fields first, then decide. +- **(superseded) A5 — freeze the factor inputs onto the graded row** so a re-audit does not + depend on context tables at all (and so defence stops needing a game-log join). +- **(done in A4) as-of-correct context** (`hitsFactorContext.build` has no as-of cutoff; + it also orders context walks on a non-unique prefix). **The re-audit RUN is + gated on this, not on clean reads.** +- **A2c (was A2b)** — composite-key ordering in `safePaginate` so the context tables can be + made safe by construction rather than clean by luck; then convert the 3 known + legacy reads and register `backfill-context` / `reconstruct-game-environment`. +- **A2c** — chase the `snapshotService.test.js` intermittent to an assertion. +- **(superseded)** roll `safePaginate` across the remaining FAIL readers, starting with + `lowParamService.fromLedger:78` (the PRIMARY calibrator, 24.8%). Each must move + FAIL → PASS on a measured re-run through `readerRows`, never on the presence of + an `.order()` clause. Remaining: `champion-ablation:173` 33.6%, + `proven-status:76` 32.6%, `prove-hit-factors:131/136`, `prove-tb-factors:149/154`, + `cluster-prove:337`, `tb-solo-and-interactions:260`, `build-grade-bands:49/54`. +- **A3** — decide what to do about `hitsFactors.js` serving two withdrawn factors + plus `platoon_severity`, which never passed the gate. +- **Re-certification of calibration is a separate gated order** on the accrual + clock — not unlocked by A1. ## Session 94 (2026-08-04) — Causally-correct platoon + park inputs ✅ 4,307 tests / 344 suites green, build exit 0. Counter + frozen clusters diff --git a/CLAUDE.md b/CLAUDE.md index e35372d..cce4673 100755 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1749,6 +1749,342 @@ phased plan in the Session-57 conversation / BUILD-STATE Next section). - **Proven factors for hits remain: `pitcher_contact_profile`, `defense_by_direction`.** Crude `defense` and `platoon` both still NOT_PROVEN. +## Read integrity + withdrawn verdicts (Fix A0 — non-obvious) +- **THE SHARED `page()` HELPER RETURNS THE RIGHT COUNT AND THE WRONG ROWS.** + `.range(from, from+PAGE-1)` with no `ORDER BY` was measured on prod returning + 410–617 duplicates out of 2,490 on the hits gate's ledger read, with an equal + number of real rows NEVER returned. `rows.length` matches the server count + exactly, which is why it hid for months. `settle-model-snapshots.js:92-98` + found and fixed this in ONE script and it was never carried across. +- **EXPOSURE CANNOT BE REASONED FROM SOURCE — only measured, per reader, per + run.** `challenger-scoreboard` walks 12 pages of `ledger_entries` clean; + `calibrationService.fromLedger` walks 3 pages of the SAME table at 24.8% + corrupt. The difference is the query PLAN (filters × table stats × concurrent + writes), which nothing in the code predicts. Same query re-measured minutes + apart moved 16.5% → 24.8%. **Never infer a reader is safe; run the harness.** +- **`src/utils/readIntegrity.js` + `scripts/read-integrity.js` are the + instrument.** Spec `specs/read-integrity-harness.md`. A reader PASSES only on + MEASURED set identity vs an ordered control — **the presence of an `.order()` + clause is never evidence** (a grep-for-the-clause guard would verify its own + intention, the meta-scar). The control validates itself: ordered-distinct ≠ + server `count(*)` ⇒ `CONTROL_INVALID`, never PASS. A non-unique key ⇒ + `KEY_NOT_UNIQUE`, so duplicate DATA is never blamed on the walk. + `walk()` PROPAGATES a fetch error — swallowing it as end-of-data is the + `calibrationService.fromLedger:117` defect and the 2026-08-01 settlement shape. +- **Baseline 2026-08-09: 11 of 25 registered readers FAIL, worst 33.6%** + (`champion-ablation:173`), `proven-status:76` 32.6%. Duplicates are WORSE than + a random subsample: they inflate `factorGate.movement.n` against `MIN_N=500` + AND corrupt the leave-one-out per-player baseline that is the gate's null. +- **BOTH hits PROVES are WITHDRAWN_PENDING_REAUDIT** + (`src/services/model/withdrawnVerdicts.js`, append-only — originals never + edited). `defense_by_direction` proved at n=528; at ~20% duplication distinct-n + ≈422 < MIN_N 500, so it was likely never eligible (arithmetic on the recorded + n, NOT a re-measurement). `pitcher_contact_profile` n=741 was plausibly + eligible but its CI was computed on a corrupt sample against a corrupt + baseline. Reinstatement requires re-running the gate through readers the + harness reports PASS for ON THE DAY. +- **WITHDRAWAL ≠ DISARMAMENT.** `hitsFactors.js` HARDCODES its three factors and + consults no registry, so both withdrawn factors still move served `p_win`. + Worse: its header calls all three "the three PROVEN hits factors" but + **`platoon_severity` never passed the gate** (n=452 < 500) — a factor that was + never proven is live in the served forecast. +- **There is NO verdict table.** `featureRegistry.statVerdicts` is an in-memory + Map that resets per process; `mc_test_ledger` holds hypothesis COUNTS for the + Bonferroni denominator, not verdicts. Standing verdicts live in `specs/` + + CLAUDE.md, which is why the withdrawal is a committed file. + +## safePaginate + the calibration fence (Fix A1 — non-obvious) +- **`src/utils/safePaginate.js` is the ONE way to walk a paginated PostgREST + read.** `paginate(makeQuery, {key:'id', label})` — takes a query FACTORY (a + builder is single-use), applies a stable ORDER BY on a UNIQUE key on every + page, VERIFIES uniqueness at runtime (a repeated key throws — the guard that + runs every time, vs the harness which runs on demand), and THROWS on a query + error. It reuses `readIntegrity.walk`; never hand-roll a second walk. + **An `.order()` on a NON-UNIQUE column is not a fix** — ties still scramble. +- **`null` and `throw` mean different things and must stay different.** + `fromLedger` returns `null` for "not enough settled history" and PROPAGATES a + throw for "the read failed". Collapsing them (`if (error || !data) break`) is + what let a 1,000-row fragment look like a complete 2,490-row fit. +- **A fixed reader is measured through the REAL FUNCTION.** The harness spec + takes `readerRows` (a call into production code), not a restatement of the + corrected query and never a `fixed:true` flag — those verify a restatement and + a comment respectively. `arm_a` on the result says which was used. +- **CORRECTION to the A0 report: `calibrationService.fromLedger` was NOT live.** + `snapshotService.js:301` is `CALIBRATION_DEPLOYED = Object.freeze([])` with no + env override, so `for (const stat of CALIBRATION_DEPLOYED)` never executes. + It is fenced THREE deep: empty deploy list; `calibrationService` is only the + SHADOW (the primary is `lowParamService`); and `chain.chainAcross` has ZERO + callers anywhere. `p_win_calibrated` appears in no route and no web file. + **Nothing serves a calibrated number.** `calibrationDeployGate.test.js` locks it. +- **The corrupt fit was wrong in a way row-count could not show.** Corrupt vs + corrected on the same 2,490 rows: certified-band observed error **−0.046 → + +0.010** (4.6x closer), ceiling 0.833 → 0.810, map output moving up to ±0.043. + The certified span stayed 0.50–0.70 either way. And the duplicates shifted the + TIME split: `fitted_through` 2026-08-06 → 2026-08-05, so the corrupt read + fitted a different window than intended. +- **`lowParamService.fromLedger:78` carries the byte-identical defect and is the + PRIMARY calibrator** (calibrationService is its shadow). Left broken on purpose + in A1 as the harness's control — it still measures 24.8%, which is how we know + the PASS is real and not a registry edit. First target for A2. + +## A2 rollout — 0 FAIL readers (non-obvious) +- **`pageSafe()` is the per-script wrapper**, present in the 7 gate scripts: + `pageSafe(sb, table, select, apply, key='id')` → `safePaginate.paginate`. The + legacy `page()` SURVIVES in the same files **on purpose** — the context tables + (`statcast_aggregates`, `batter_spray`, `team_defense`, `platoon_splits`, + `park_dimensions`, `hitter_opportunity`, `lineup_context`) have **COMPOSITE + primary keys with no single unique column**, which `paginate()` cannot express. + They measure 0% today. **Never use `page()` for `ledger_entries` or + `model_snapshots`** — `tests/unit/readerPagination.test.js` fails if you do. +- **`ledger_entries` and `model_snapshots` are the only two tables with a + single-column PK (`id`)** — which is why all 11 corrupt readers were fixable + and the context tables were not. Composite-key ordering is the open item. +- **Every fixed read is a named `READS` export**, and `main()` is wrapped in + `if (require.main === module)`. Both are load-bearing: the harness `require()`s + these scripts to call the REAL function, and several of them WRITE to + `mc_test_ledger` (the Bonferroni denominator) via `tl.recordAndCount`. Without + the guard, measuring a reader would inflate the multiple-comparisons + correction. Verified: `mc_test_ledger` stayed 168 rows / last write 2026-08-06 + across the full harness run. +- **3 reads deliberately left on the legacy walk** (`proven-status:66`, `:85`, + `champion-ablation:176`) — they measured 0% and the A2 order scoped to the FAIL + set. They are listed in `KNOWN_LEGACY_READS` in the test so the gap is tracked, + not invisible. They are clean **by plan, not by construction**. +- **`backfill-context.js` and `reconstruct-game-environment.js` still walk + `ledger_entries` unordered and are NOT in the harness registry** — unmeasured, + not proven clean. +- **SUITE FLAKE, UNRESOLVED:** `snapshotService.test.js` "Wave 2A — threads the + resolved athlete id" failed 2 times in 30 full-suite runs on this branch and + 0 in 29 at HEAD. That difference is NOT statistically significant (Fisher + ≈ p 0.5) and the assertion could not be captured in 24 instrumented runs. The + plausible mechanism is that +89 tests reshuffles jest worker sharding and + exposes a latent cross-suite interference. Isolated: 5/5 clean. **Do not treat + it as resolved.** + +## A2b — clean by construction (non-obvious) +- **`safePaginate` takes a COMPOSITE key**: `key: ['as_of_date','sport','season','player_key']`. + Every column is ordered, in order, and the uniqueness guard checks the FULL + tuple. Ordering on a PREFIX is not a fix — the trailing columns are exactly + where the ties live. String keys still work; the 12 A2 readers were unaffected. +- **`src/utils/tableKeys.js` is the single source of truth for row identity**, + pulled from `pg_index WHERE indisunique` on 2026-08-09 — never assumed. It + THROWS for an unknown table rather than defaulting to `id`, because a silent + `id` default would order by a column that may not exist and would look fixed + while doing nothing. `pageSafe(sb, table, ...)` defaults its key from here. + **If a migration changes a constraint, change it here** — a stale entry + surfaces as a runtime duplicate-tuple throw, which is the intended loud failure. +- **Only `ledger_entries`, `model_snapshots` and `game_context` have a + single-column key.** Every other context table is dated-composite + (`as_of_date, sport, season, `) by design — one as-of snapshot per + entity per day. That is why A2 could not reach them. +- **`batter_spray`'s KEY_NOT_UNIQUE was a wrong key, not bad data** — 6 "duplicate" + `player_key|as_of_date` pairs resolve cleanly on the real + `as_of_date,sport,season,source_id`. The harness naming it rather than + swallowing it is what made that diagnosable. +- **THE UNORDERED `page()` HELPER IS DELETED FROM ALL 15 SCRIPTS**, not parked + next to `pageSafe`. It had become dead code, and a dead broken helper is an + invitation. `readerPagination.test.js` fails if one reappears. +- **A converted read MUST select its key columns.** 10 reads selected a subset + that omitted `id` or the composite columns; the uniqueness guard would have + thrown at runtime. A checker script verified every converted select covers its + key — do that after any conversion. +- **`calibrate-hits.js:40` was an unregistered 24.7%-corrupt ledger walker** + found during the sweep — it fits the calibration reliability curve, so a fifth + of the history was double-weighted and another fifth absent from every bin. + Converted. +- **WRITE-SCRIPT VERDICT (data untouched):** `backfill-context` and + `reconstruct-game-environment` build DISTINCT SETS from their read, so + corruption causes OMISSION, never a wrong value — the values come from + statsapi/Open-Meteo. Their shared read measures **0% today** (5,107 rows), and + a random-drop sensitivity at 24.8% bounds historical exposure at **~6.6 of 419 + players** (18 have a single row) and **~0.16 of 153 games** (median 37 rows). + Small, bounded, and NOT a reason to rewrite context data. +- **STILL OPEN — a clean read is not an as-of-correct read.** + `hitsFactorContext.build` has no as-of cutoff (it takes `latestBy(as_of_date)`) + AND orders its context walks by a NON-UNIQUE prefix (`player_key`/`team`). Same + prefix problem in `factor-wiring-audit.js:62-65`. Reading today's context for a + past row is contamination even when the read is perfect. Next order. + +## A3 — version alignment + snapshot settlement (non-obvious) +- **`src/config/modelVersion.js` is the ONLY place a model version is declared.** + `ledgerService` and `retentionService` both read `process.env.MODEL_VERSION` + but carried DIFFERENT hardcoded defaults (`engine1@2026-07-20` vs + `engine1@2026-08-07-fullwindow`). Prod sets no env var, so the two tables took + different defaults and **no ledger row ever carried the marker + `reAuditEligibility` requires** — the accrual clock read "blocked" when the + real state was "mis-stamped". Never re-declare a version default beside its + consumer; it will drift from the other consumer and both sides look locally + correct. +- **`REPAIRED_CHAMPION_VERSION` is deliberately NOT `MODEL_VERSION`.** The + eligibility bar is a fixed fact about one repair; tying it to "whatever we + stamp today" would make every future version silently re-qualify itself. +- **History is NOT re-stamped.** Rows keep the marker they were written with — + rewriting them destroys the only record of which forecast produced them. Old + rows stay honestly old. +- **`snapshotSettlementService` settles `model_snapshots` on the cron** that + already settles the ledger, best-effort (the retention table is a measurement + asset, not the public record — a failure must not take the ledger settle or the + grade with it). **Outcomes only: it never reconstructs context** (a test greps + for every context table name). +- **THE DRAIN ORDER IS A CORRECTNESS PROPERTY, and oldest-first is wrong.** + Measured: 308 rows across the eight oldest dates are structurally unsettleable + (post-hoc-logged or orphaned), so an oldest-first window re-processed dead + dates forever and `dates_remaining` never moved. It is now (1) NEWEST-first, + (2) date selection filtered by `isPreGame` — pure, so unsettleable rows are + known without fetching — and (3) window-advancing when a window produces + nothing. All three were needed to make it terminate. +- **A row logged after first pitch is not a prediction.** A 01:00-UTC cycle is + 21:00 ET the *previous* evening — same game date, three hours into the slate. + A UTC date compare would keep those rows. **20,064 rows are permanently + unsettleable** for this reason and are counted, not hidden. +- **`outcome` is SIDE-ALIGNED, `actual_value` is raw.** `p_win` is expressed for + the graded side, so a raw `realized > line` indicator would invert the target + on every under row. +- **RESULT:** settled 15,484 → **28,894**; repaired-champion **0 → 7,556**; + `settled_but_null_actual = 0` on every row. +- **ELIGIBILITY NOW COUNTS BUT IS STILL SHORT: 2 eligible dates vs 10/10/14/14 + needed.** And the re-audit must not RUN on clean reads alone — it is gated on + the **as-of-correct context fix** (`hitsFactorContext.build` has no as-of + cutoff), which is a separate order. + +## A4 — as-of-correct context (non-obvious) +- **`hitsFactorContext.build(sb, { asOf })`.** No `asOf` = LIVE = the exact code + that ran before, so the served grade cannot move. Only an explicitly dated call + takes the audit path. **Proven, not asserted: 400 real graded rows resolved + both ways gave identical context, identical multiplier and identical + p_adjusted — 400/400.** +- **REFUSAL OVER RECONSTRUCTION.** No row at-or-before `asOf` for an entity ⇒ its + factor input is null ⇒ the factor does not apply and the base rate is + untouched. NEVER the nearest or latest row: a substituted row is a plausible + wrong value wearing a date, and nothing downstream can see it. +- **`statcast_aggregates` CANNOT answer an as-of question** — it is upserted in + place and keeps one date. The dated path reads **`statcast_history`**. Verified + equivalent at the head before switching: on 2026-08-09 the two agree on all + 1,414 rows with **0 differences**, so changing source does not itself move a + number. History starts **2026-08-03**; an as-of before that refuses. +- **Dated-table coverage starts 2026-08-04** (`batter_spray`, `team_defense`, + `platoon_splits`). An as-of earlier than that refuses everything — correct, and + it means a re-audit cannot reach back past early August no matter how many + dates accrue. +- **`model_snapshots.opponent` is NULL on 100% of settled repaired-champion rows** + (0 of 3,064 and 0 of 4,492). `team` is present on ~98%, `archetype` on ~63%. + So `defense_by_direction` — one of the two withdrawn PROVES — **cannot be + re-audited from the snapshot row alone**; the opponent must come from the + player's own game log, as `prove-hit-factors` already does. Spray/platoon/ + handedness reach 99%; defence reaches 0% without that join. +- **Archetype is CARRIED, never invented** — no context table holds it; it lives + on `model_snapshots.archetype`, so an audit caller supplies it and the resolver + passes it through for per-archetype conditioning. +- `factor-wiring-audit.js` gets the same `asOf` (env `FWA_AS_OF`) and its context + walks moved off a NON-UNIQUE prefix order (`player_key`/`team`) onto the real + composite key. + +## A5 — THE HITS FACTORS NEVER FIRE (the finding) + input freeze +- **MEASURED 2026-08-10, 596 real graded hits props: ALL THREE hits factors are + skipped on 596/596 and the multiplier is exactly 1 every time.** The context + loads fine (618 spray players, 31 defence teams, 789 pitcher profiles) and + resolves batter-side inputs — spray 590/596, bats 594, platoon 594 — but + **`positionOaa` 0, `throws` 0, `pitcherHardHit` 0.** +- **CAUSE: the two join keys are never set.** `prop.opponent` and + `prop.opposing_pitcher` are read by `hitsFactorContext` and written by NOTHING + — `oddsNormalizer` never sets them, and `snapshotService` attaches team/opponent + AFTER grading (`:500`). `grep opposing_pitcher src/` returns exactly one hit: + the line that reads it. +- **CONSEQUENCE, and it corrects several earlier notes:** no hits factor has ever + moved a served `p_win`. So (a) A0/A2's "hitsFactors serves two withdrawn + factors plus platoon_severity" is WRONG — they are wired but inert; (b) S94's + "PROVEN FACTORS, LOADED BEFORE THE GRADE — this ordering IS the fix" was + necessary but NOT sufficient; the ordering was fixed and the join keys were + never plumbed; (c) `defense_by_direction` did not "do the thing without + recording it" — it never did the thing. +- **Plumbing the keys WOULD change served grades**, so it is not a bug-fix, it is + a model change and needs its own order with a before/after. +- **`factorFreeze.freeze(ctx, prop)` records INPUTS, never the multiplier** + (`model_snapshots.factor_inputs` jsonb, migration 034). A stored multiplier can + only be compared to itself; stored inputs re-run through `hitsFactors` via + `factorFreeze.recompute()` and get CHECKED. Captured in the same pass that + reads the context (`analyzeViaEngine1`), so recorded cannot drift from used. +- **A separate column, not extra keys in `features`** — `champion-ablation.js` + iterates every `features` key for its residual scan, so widening it would + silently enlarge that multiple-comparisons denominator. +- **`available` (team/game_id/game_date) is recorded but NEVER fed to a factor.** + It is the evidence a later order needs to measure what plumbing the join keys + would change, before changing it. +- Proven on 496 real rows: frozen == what the factor read (496/496), recompute + from frozen == live multiplier (496/496), `p_win` unchanged (496/496). +- **FORWARD-ONLY.** Rows graded before deploy keep `factor_inputs` NULL and + `opponent` NULL; `defense_by_direction` re-audit on them still needs the + game-log opponent join. No backfill — that would reconstruct. + +## A6 — the join keys, SHADOW (non-obvious) +- **`matchupKeys.build({sb,getSchedule,gameDate,asOf})`** resolves the two keys + the hits factors have always needed: `player -> team` from **`lineup_context`** + (as-of dated) and `team -> {opponent, opposing_pitcher}` from + `getScheduleWithPitchers`. **At grade time the game has not happened**, so + `prove-hit-factors`' game-log source is useless here — the probable pitcher is + the only honest pre-game source. +- **NEVER guess the opponent from the prop.** A prop carries `home_team` and + `away_team`, so inferring which side a hitter bats for would be right about + half the time and wrong invisibly — and a wrong opponent feeds the defence + factor a real team's fielders against the wrong hitter. No lineup row ⇒ + `refused: 'no_lineup_row'`, no fire. +- **`hitsFactorContext`'s resolver takes an OPTIONAL third arg `keys`.** Omitted + on the live path ⇒ byte-identical to before it existed. Supplied only by the + shadow resolve. That is what keeps the served grade frozen while the factors + compute. +- **MEASURED 2026-08-09, 308 hits props:** keys resolved on **248 (80.5%)**, 60 + refused for no lineup row. Live fire **0/0/0** (unchanged). **Shadow fire: + `pitcher_contact_profile` 248, `defense_by_direction` 185, + `platoon_severity` 147** — from 0/596 in A5. Would-be multiplier: min 0.803, + p25 0.966, **median 1.017**, p75 1.064, max 1.250; **245 of 248 would move the + number**. Served p_win identical on every evaluated row; shadow recompute-check + 248/248. +- **`factor_inputs.would_fire` stores a multiplier — deliberately breaking A5's + inputs-only rule, and only because `shadow_inputs` is stored beside it**, so it + stays re-derivable. A number you cannot re-derive is what that rule forbids. +- **The live frozen inputs still record the refusal** (`position_oaa: null`, + `throws: null`) while `would_fire.shadow_inputs` records what it WOULD have + seen. Both, side by side, on the same row. +- A6 is EVIDENCE, not a verdict. The withdrawn factors are still withdrawn; A7 + (turning them on live) is a model change gated on this evidence plus a re-proof + on rows where they actually fire. +- **Test-window gotcha:** the A5 test sliced a fixed 700 bytes from its marker + and A6's insertion pushed the assertion out of the window — a test failing + because a neighbour grew. Window to a syntactic landmark, never a byte count. + +## A7 — shadow accrual + the pre-registered A8 gate (non-obvious) +- **`specs/a8-shadow-factor-gate.md` is PRE-REGISTERED and NOT RUN.** Written + before any accrual exists, so the decision rule cannot be picked after seeing + the answer. It names the fallback out loud: a factor that moves ~80% of the + board and does not improve Brier is THEATER, and the correct action is to leave + it off — not to re-tune the multiplier until it passes. +- **A8 is a NEW hypothesis, not a re-test**, and counts fresh against the + cumulative Bonferroni denominator. The S87 "re-asking on more data is not a new + shot on goal" rule does NOT apply: the old question was "does this improve a + reconstructed audit"; the new one is "does the multiplier this factor actually + produces, on rows where it actually fires, improve the served forecast". +- **THE BINDING CONSTRAINT IS CLUSTERS, NOT ROWS.** For all three factors the + treatment entity (`player|opponent`, `starter_id`, `player_key`) outnumbers the + games, so `factorGate`'s coarser-of rule resolves to `game_id`. Measured: + ~925 settled hits rows but only **~10.5 distinct settled GAMES per slate**, so + `MIN_N` 500 is met in 1–2 slates and `MIN_CLUSTERS` 40 takes **~4**. Reading + the row count alone would have said "ready tomorrow" and been wrong. +- **A `would_fire` with no settled outcome is NOT evidence.** Complete pairs + today: **0** — `factor_inputs` is 0/122,276 because A5/A6 are not deployed. + Accrual starts at deploy, not at merge; add a one-day settle lag. **A8 is + runnable ≈5 slates after deploy.** +- **`tests/unit/shadowAccrual.test.js` drives the REAL + `gradeAndCacheSlate → onGraded → rowsFromSides` chain** rather than reading + code, because A5's whole finding was a factor that was built, correct, and + never invoked — evidence that is generated but never persisted is that same + failure one layer along. It also asserts the SCHEDULED path (`runSnapshot` + builds the key index before the grade call, `runAllSnapshots` routes through + it, a key-resolve failure degrades to no-shadow). +- **Settlement lags the slate**: 2026-08-09/10/11 had 1,438–1,804 graded hits + rows and **0 settled** at the time of writing, because snapshot settlement runs + on the cron and had last drained through 08-08. Accrual counts must be read off + SETTLED rows, never graded ones. + ## Active Skills - vyndr-voice (all user-facing output) - prop-analysis (grading methodology) diff --git a/scripts/backfill-context.js b/scripts/backfill-context.js index 9f76e31..f9a4849 100644 --- a/scripts/backfill-context.js +++ b/scripts/backfill-context.js @@ -29,30 +29,45 @@ const { createClient } = require('@supabase/supabase-js'); const ctx = require('../src/services/lineupContextService'); const mlb = require('../src/services/adapters/mlbStatsAdapter'); const { knownNumber } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; const SEASON = Number(process.env.BF_SEASON || 2026); const PAGE = 1000; -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2/A2b) ──────────────────────────────────────────── +// This script used to walk pages with `.range()` and NO ORDER BY. Measured on +// production, that returned the correct row COUNT and the wrong ROWS: up to +// 33.6% of a read came back twice while an equal share never came back at all, +// so `rows.length` looked perfect while a fifth of the sample was missing. +// +// `pageSafe` routes every read through `src/utils/safePaginate`: a stable ORDER +// BY on the table's real UNIQUE key — single OR composite, looked up from +// `src/utils/tableKeys` rather than assumed — a runtime tuple-uniqueness check, +// and a THROWN error instead of a silent stop. The old unordered helper is gone +// rather than left beside it, because a dead broken helper is an invitation. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + ledger: (sb) => pageSafe(sb, 'ledger_entries', 'id, player_key, player_name, stat, outcome, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases']) + .in('outcome', ['hit', 'miss'])), + platoonHeld: (sb) => pageSafe(sb, 'platoon_splits', 'as_of_date, sport, season, player_key', + (q) => q.eq('sport', 'mlb')), +}; + async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); // Every hitter who appears on a CLEAN settled row — the true denominator. - const led = await page(sb, 'ledger_entries', 'player_key, player_name, stat, outcome, quarantine_reason', + const led = await pageSafe(sb, 'ledger_entries', 'id, player_key, player_name, stat, outcome, quarantine_reason', (q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases']) .in('outcome', ['hit', 'miss'])); const need = new Map(); @@ -61,7 +76,7 @@ async function main() { if (!need.has(r.player_key)) need.set(r.player_key, r.player_name); } - const have = new Set((await page(sb, 'platoon_splits', 'player_key', (q) => q.eq('sport', 'mlb'))) + const have = new Set((await pageSafe(sb, 'platoon_splits', 'as_of_date, sport, season, player_key', (q) => q.eq('sport', 'mlb'))) .map((r) => r.player_key)); const missing = [...need.entries()].filter(([k]) => !have.has(k)); @@ -100,4 +115,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS: (typeof READS !== 'undefined' ? READS : undefined), READS_LEDGER: (typeof READS_LEDGER !== 'undefined' ? READS_LEDGER : undefined) }; diff --git a/scripts/build-grade-bands.js b/scripts/build-grade-bands.js index 2b95a74..5f2a6dd 100644 --- a/scripts/build-grade-bands.js +++ b/scripts/build-grade-bands.js @@ -18,6 +18,8 @@ const gb = require('../src/services/model/gradeBands'); const tl = require('../src/services/model/testLedger'); const cal = require('../src/services/model/calibration'); const { knownNumber } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -31,29 +33,43 @@ const PAGE = 1000; */ const PROVEN_BY_ARCHETYPE = Object.freeze({}); -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── FIX A2 (2026-08-09) — THE SAFE WALK ─────────────────────────────────── +// `page()` above walks with no ORDER BY. Measured on production, that returned +// the correct row COUNT and the wrong ROWS: up to 33.6% of a read came back +// twice while an equal share never came back at all. `pageSafe` routes the same +// call through `src/utils/safePaginate`, which orders on a UNIQUE key on every +// page, verifies uniqueness at runtime, and THROWS on a query error instead of +// treating it as end-of-data. +// +// `page()` SURVIVES only for the context tables (statcast_aggregates, +// batter_spray, team_defense, platoon_splits, park_dimensions, game_context...). +// Those have COMPOSITE primary keys with no single unique column, so +// safePaginate cannot express them. They measure 0% corruption today; making +// them safe needs a composite-key ordering the helper does not yet have. Do not +// use `page()` for ledger_entries or model_snapshots. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + snaps: (sb, stat = STAT) => pageSafe(sb, 'model_snapshots', 'id, player_key, game_date, archetype, stat', + (q) => q.eq('sport', 'mlb').eq('stat', stat).not('archetype', 'is', null)), + ledger: (sb, stat = STAT) => pageSafe(sb, 'ledger_entries', + 'id, player_key, game_date, outcome, p_win, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', stat) + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)), +}; + async function main() { const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat', - (q) => q.eq('sport', 'mlb').eq('stat', STAT).not('archetype', 'is', null)); + const snaps = await READS.snaps(sb); const archOf = new Map(); for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype); - const led = await page(sb, 'ledger_entries', 'player_key, game_date, outcome, p_win, quarantine_reason', - (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', STAT) - .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); + const led = await READS.ledger(sb); const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')); const byArch = new Map(); @@ -97,4 +113,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/calibrate-hits.js b/scripts/calibrate-hits.js index cb13cac..aeeb2c6 100644 --- a/scripts/calibrate-hits.js +++ b/scripts/calibrate-hits.js @@ -29,6 +29,8 @@ require('dotenv').config(); const { createClient } = require('@supabase/supabase-js'); const cal = require('../src/services/model/calibration'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -37,19 +39,24 @@ const PAGE = 1000; const r3 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 1000) / 1000); -async function page(sb, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from('ledger_entries') - .select('p_win, outcome, game_date, quarantine_reason')).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2b) ─────────────────────────────────────────────── +// This walked pages with no ORDER BY and measured 24.7% corrupt on production — +// 616 of 2,490 rows returned twice, an equal share never returned — while +// `rows.length` matched the server count exactly. It fits the calibration +// reliability curve, so a fifth of the history was double-weighted and another +// fifth absent from every bin. +async function pageSafe(sb, apply) { + return paginate(() => apply(sb.from('ledger_entries') + .select('id, p_win, outcome, game_date, quarantine_reason')), + { key: uniqueKeyFor('ledger_entries'), pageSize: PAGE, label: 'calibrate-hits' }); } +/** THE MEASURED READ — main() and the harness call the same function. */ +const READS = { + ledger: (sb) => pageSafe(sb, (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits') + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)), +}; + /** Reliability rendered per bin with n — the only honest way to read this. */ function curve(rows, label) { return cal.reliability(rows, 10) @@ -67,8 +74,7 @@ async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const raw = await page(sb, (q) => q.eq('sport', 'mlb').is('user_id', null) - .eq('stat', 'hits').in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); + const raw = await READS.ledger(sb); const all = raw .filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')) .map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, d: String(r.game_date) })); @@ -133,4 +139,10 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; + diff --git a/scripts/champion-ablation.js b/scripts/champion-ablation.js index b349148..0c1af91 100644 --- a/scripts/champion-ablation.js +++ b/scripts/champion-ablation.js @@ -51,6 +51,8 @@ require('dotenv').config(); const { createClient } = require('@supabase/supabase-js'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -143,20 +145,39 @@ function bootstrapCorr(rows, key, seed = 20260804, iters = 3000) { }; } -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - let q = sb.from(table).select(select); - q = apply(q).range(from, from + PAGE - 1); - const { data, error } = await q; - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── FIX A2 (2026-08-09) — THE SAFE WALK ─────────────────────────────────── +// `page()` above walks with no ORDER BY. Measured on production, that returned +// the correct row COUNT and the wrong ROWS: up to 33.6% of a read came back +// twice while an equal share never came back at all. `pageSafe` routes the same +// call through `src/utils/safePaginate`, which orders on a UNIQUE key on every +// page, verifies uniqueness at runtime, and THROWS on a query error instead of +// treating it as end-of-data. +// +// `page()` SURVIVES only for the context tables (statcast_aggregates, +// batter_spray, team_defense, platoon_splits, park_dimensions, game_context...). +// Those have COMPOSITE primary keys with no single unique column, so +// safePaginate cannot express them. They measure 0% corruption today; making +// them safe needs a composite-key ordering the helper does not yet have. Do not +// use `page()` for ledger_entries or model_snapshots. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS_LEDGER = { + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, player_key, stat, line, side, game_date, outcome, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).in('outcome', ['hit', 'miss'])), +}; + +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + snaps: (sb) => pageSafe(sb, 'model_snapshots', + 'id, player_key, stat, line, side, game_date, p_win, features, quarantine_reason, captured_at, archetype', + (q) => q.eq('sport', 'mlb').not('p_win', 'is', null).not('features', 'is', null)), +}; + const propKey = (r) => `${r.player_key}|${r.stat}|${Number(r.line)}|${String(r.side).toLowerCase()}|${r.game_date}`; /** @@ -170,10 +191,8 @@ const propKey = (r) => `${r.player_key}|${r.stat}|${Number(r.line)}|${String(r.s * this run is read-only. */ async function fetchAll(sb) { - const snaps = await page(sb, 'model_snapshots', - 'player_key, stat, line, side, game_date, p_win, features, quarantine_reason, captured_at, archetype', - (q) => q.eq('sport', 'mlb').not('p_win', 'is', null).not('features', 'is', null)); - const led = await page(sb, 'ledger_entries', 'player_key, stat, line, side, game_date, outcome, quarantine_reason', + const snaps = await READS.snaps(sb); + const led = await pageSafe(sb, 'ledger_entries', 'id, player_key, stat, line, side, game_date, outcome, quarantine_reason', (q) => q.eq('sport', 'mlb').is('user_id', null).in('outcome', ['hit', 'miss'])); const outcomeBy = new Map(); @@ -371,4 +390,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS, READS_LEDGER }; diff --git a/scripts/cluster-prove.js b/scripts/cluster-prove.js index f5321bf..a7fc835 100644 --- a/scripts/cluster-prove.js +++ b/scripts/cluster-prove.js @@ -44,6 +44,8 @@ const sk = require('../src/services/model/skillProjection'); const reg = require('../src/services/model/featureRegistry'); const mlb = require('../src/services/adapters/mlbStatsAdapter'); const { knownRate, knownNumber } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -156,18 +158,35 @@ function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260804) { }; } -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── FIX A2 (2026-08-09) — THE SAFE WALK ─────────────────────────────────── +// `page()` above walks with no ORDER BY. Measured on production, that returned +// the correct row COUNT and the wrong ROWS: up to 33.6% of a read came back +// twice while an equal share never came back at all. `pageSafe` routes the same +// call through `src/utils/safePaginate`, which orders on a UNIQUE key on every +// page, verifies uniqueness at runtime, and THROWS on a query error instead of +// treating it as end-of-data. +// +// `page()` SURVIVES only for the context tables (statcast_aggregates, +// batter_spray, team_defense, platoon_splits, park_dimensions, game_context...). +// Those have COMPOSITE primary keys with no single unique column, so +// safePaginate cannot express them. They measure 0% corruption today; making +// them safe needs a composite-key ordering the helper does not yet have. Do not +// use `page()` for ledger_entries or model_snapshots. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + snaps: (sb) => pageSafe(sb, 'model_snapshots', 'id, player_key, game_date, archetype', + (q) => q.eq('sport', 'mlb').eq('stat', STAT).not('archetype', 'is', null)), + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', STAT) + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)), +}; + async function opposingStarters(dates) { const m = new Map(); for (const d of dates) { @@ -305,7 +324,7 @@ async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10); const batters = new Map(); const pitchersById = new Map(); for (const r of statcast) { @@ -329,15 +348,11 @@ async function main() { // model_snapshots carries the actual classification per prop, so the // conditioning variable is now a genuine BOMBER indicator, which is // categorical and therefore not a transform of barrel at all. - const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype', - (q) => q.eq('sport', 'mlb').eq('stat', STAT).not('archetype', 'is', null)); + const snaps = await READS.snaps(sb); const archetypeBy = new Map(); for (const r of snaps) if (r.player_key && r.game_date) archetypeBy.set(`${r.player_key}|${r.game_date}`, r.archetype); - const led = await page(sb, 'ledger_entries', - 'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', - (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', STAT) - .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); + const led = await READS.ledger(sb); // POINT-IN-TIME IS NO LONGER AVAILABLE FROM THIS TABLE. // // `statcast_aggregates` is upserted in place and keeps one as-of date. The @@ -369,7 +384,7 @@ async function main() { console.error(`[debug] matches in sample=${sampleKeys.filter((k) => batters.has(k)).length}/5`); } // TEAM DEFENCE — the newly-ingested Statcast OAA, per team, dated. - const defRows = await page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb')); + const defRows = await pageSafe(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb')); const defByTeam = new Map(); for (const d of defRows) { const prev = defByTeam.get(d.team); @@ -550,4 +565,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/factor-wiring-audit.js b/scripts/factor-wiring-audit.js index 12a7879..7fa8018 100644 --- a/scripts/factor-wiring-audit.js +++ b/scripts/factor-wiring-audit.js @@ -11,6 +11,8 @@ require('dotenv').config(); const fs = require('fs'); const path = require('path'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const { createClient } = require('@supabase/supabase-js'); const hf = require('../src/services/model/hitsFactors'); const lp = require('../src/services/model/lowParamCalibrator'); @@ -22,15 +24,25 @@ const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json'); const SEQ = path.join(process.cwd(), '.seq-cache', 'sequences.json'); const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null); -async function page(sb, t, s, f, orderBy = 'id') { - const o = []; - for (let i = 0; ; i += 1000) { - const { data, error } = await f(sb.from(t).select(s)).order(orderBy, { ascending: true }).range(i, i + 999); - if (error) throw new Error(`${t}: ${error.message}`); - if (!data || !data.length) break; o.push(...data); if (data.length < 1000) break; - } - return o; +// FIX A4 — the context walks here ordered by a NON-UNIQUE prefix +// ('player_key' / 'team'), which leaves the trailing key columns tied and lets +// pages overlap. They now order by the table's real unique tuple. +async function page(sb, t, s, f, key) { + return paginate(() => f(sb.from(t).select(s)), + { key: key || uniqueKeyFor(t), pageSize: 1000, label: `factor-wiring-audit:${t}` }); } + +/** + * AS-OF. `FWA_AS_OF=YYYY-MM-DD` bounds every dated context table to rows at or + * before that date, so an audit of a past slate cannot join a profile built from + * games played after it. Unset = today = the live read, unchanged. + * + * REFUSAL OVER RECONSTRUCTION: an entity with no row at or before the cutoff is + * simply absent from the index, so its factor does not apply. Never the nearest + * row — a substituted row is a plausible wrong value wearing a date. + */ +const AS_OF = process.env.FWA_AS_OF || null; +const dated = (q) => (AS_OF ? q.lte('as_of_date', AS_OF) : q); const isPreGame = (c, g) => { const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000); const d = et.toISOString().slice(0, 10); @@ -59,15 +71,33 @@ function decompose(rows, bins = 10) { // Factor inputs. const [spray, defense, platoon, statcast] = await Promise.all([ - page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb'), 'player_key'), - page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb'), 'team'), - page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb'), 'player_key'), - page(sb, 'statcast_aggregates', 'player_key, role, bats, throws, hard_hit_pct', (q) => q.eq('sport', 'mlb'), 'player_key'), + page(sb, 'batter_spray', '*', (q) => dated(q.eq('sport', 'mlb'))), + page(sb, 'team_defense', '*', (q) => dated(q.eq('sport', 'mlb'))), + page(sb, 'platoon_splits', '*', (q) => dated(q.eq('sport', 'mlb'))), + // `statcast_aggregates` keeps ONE as-of date, so it cannot answer an as-of + // question; the dated path reads the retained history instead. + AS_OF + ? page(sb, 'statcast_history', 'player_key, source_id, role, bats, throws, hard_hit_pct, as_of_date', + (q) => q.eq('sport', 'mlb').lte('as_of_date', AS_OF)) + : page(sb, 'statcast_aggregates', 'player_key, source_id, role, bats, throws, hard_hit_pct', + (q) => q.eq('sport', 'mlb')), ]); const latest = (rows, k) => { const m = new Map(); for (const r of rows) { const key = r[k]; if (!key) continue; const p = m.get(key); if (!p || String(r.as_of_date) > String(p.as_of_date)) m.set(key, r); } return m; }; const sprayBy = latest(spray, 'player_key'); const defBy = latest(defense, 'team'); const platBy = latest(platoon, 'player_key'); const batBy = new Map(); const pitBy = new Map(); - for (const r of statcast) { if (!r.player_key) continue; (r.role === 'pitcher' ? pitBy : batBy).set(r.player_key, r); } + // On the dated path several as-of rows per player are in scope; collapse to + // the latest WITHIN the bound, keyed on (source_id, role) so a two-way player + // keeps both profiles. No-op on the live path (one row per player). + const statcastLatest = AS_OF ? (() => { + const m = new Map(); + for (const r of statcast) { + const k = `${r.source_id}|${r.role}`; + const prev = m.get(k); + if (!prev || String(r.as_of_date) > String(prev.as_of_date)) m.set(k, r); + } + return [...m.values()]; + })() : statcast; + for (const r of statcastLatest) { if (!r.player_key) continue; (r.role === 'pitcher' ? pitBy : batBy).set(r.player_key, r); } const frac = (v) => { const n = knownNumber(v); return n === null ? null : (n > 1 ? n / 100 : n); }; // Opponent + starter per (player,date) from the sequence cache. diff --git a/scripts/pitcher-prove-k.js b/scripts/pitcher-prove-k.js index 75d444e..782d661 100644 --- a/scripts/pitcher-prove-k.js +++ b/scripts/pitcher-prove-k.js @@ -30,6 +30,8 @@ const pe = require('../src/services/model/pitcherEngine'); const sk = require('../src/services/model/skillProjection'); const mlb = require('../src/services/adapters/mlbStatsAdapter'); const { knownRate, knownNumber } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -86,17 +88,30 @@ function bootstrapDiff(rows, kA, kB, iters = 4000, seed = 20260805) { const ci = [q(0.025), q(0.975)]; return { point: r4(cv.pearson(rows.map((r) => r[kA]), rows.map((r) => r.won)).r - cv.pearson(rows.map((r) => r[kB]), rows.map((r) => r.won)).r), ci95: ci, ci_excludes_zero: ci[0] > 0 || ci[1] < 0 }; } -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2/A2b) ──────────────────────────────────────────── +// This script used to walk pages with `.range()` and NO ORDER BY. Measured on +// production, that returned the correct row COUNT and the wrong ROWS: up to +// 33.6% of a read came back twice while an equal share never came back at all, +// so `rows.length` looked perfect while a fifth of the sample was missing. +// +// `pageSafe` routes every read through `src/utils/safePaginate`: a stable ORDER +// BY on the table's real UNIQUE key — single OR composite, looked up from +// `src/utils/tableKeys` rather than assumed — a runtime tuple-uniqueness check, +// and a THROWN error instead of a silent stop. The old unordered helper is gone +// rather than left beside it, because a dead broken helper is an invitation. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, game_id', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'strikeouts') + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)), +}; + const SOLO = ['pitcher_whiff_pct', 'pitcher_k_pct', 'pitcher_chase_pct', 'pitcher_gb_pct', 'pitcher_arm_angle', 'opposing_lineup_k_rate']; @@ -125,7 +140,7 @@ async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10); const pitchByKey = new Map(); const batterByKey = new Map(); for (const r of statcast) { @@ -134,8 +149,8 @@ async function main() { if (r.role === 'batter' && r.player_key) batterByKey.set(r.player_key, prof); } - const led = await page(sb, 'ledger_entries', - 'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', + const led = await pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'strikeouts') .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')); @@ -357,4 +372,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/prove-hit-factors.js b/scripts/prove-hit-factors.js index fa25cea..974a0c6 100644 --- a/scripts/prove-hit-factors.js +++ b/scripts/prove-hit-factors.js @@ -27,24 +27,53 @@ const { knownNumber, knownRate } = require('../src/utils/known'); const { nameKey } = require('../src/utils/playerName'); const sd = require('../src/services/model/sprayDefense'); const pss = require('../src/services/model/platoonSeverity'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; const PAGE = 1000; const ARCHS = (process.env.HF_ARCHETYPES || 'BOMBER,GHOST,ALL').split(','); -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── FIX A2 (2026-08-09) — THE SAFE WALK ─────────────────────────────────── +// `page()` above walks with no ORDER BY. Measured on production, that returned +// the correct row COUNT and the wrong ROWS: up to 33.6% of a read came back +// twice while an equal share never came back at all. `pageSafe` routes the same +// call through `src/utils/safePaginate`, which orders on a UNIQUE key on every +// page, verifies uniqueness at runtime, and THROWS on a query error instead of +// treating it as end-of-data. +// +// `page()` SURVIVES only for the context tables (statcast_aggregates, +// batter_spray, team_defense, platoon_splits, park_dimensions, game_context...). +// Those have COMPOSITE primary keys with no single unique column, so +// safePaginate cannot express them. They measure 0% corruption today; making +// them safe needs a composite-key ordering the helper does not yet have. Do not +// use `page()` for ledger_entries or model_snapshots. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** + * THE MEASURED READS, named once. + * + * `main()` calls these and so does the read-integrity harness, so a PASS there + * is a statement about the code that actually runs. A registry that restated + * the query would verify the restatement instead. + */ +const READS = { + statcast: (sb) => pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')), + spray: (sb) => pageSafe(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb')), + platoon: (sb) => pageSafe(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb')), + defense: (sb) => pageSafe(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb')), + snaps: (sb) => pageSafe(sb, 'model_snapshots', 'id, player_key, game_date, archetype, stat', + (q) => q.eq('sport', 'mlb').eq('stat', 'hits').not('archetype', 'is', null)), + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits') + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)), +}; + /** * THE FACTORS. Each returns a MULTIPLIER on the base rate, or null when the * input is absent — an absent factor must leave the baseline untouched rather @@ -101,14 +130,14 @@ async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); const batters = new Map(); const pitchersById = new Map(); for (const r of statcast) { const prof = sk.fromStatcastRow(r); if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), prof); if (r.role === 'batter' && r.player_key) batters.set(r.player_key, prof); } - const sprayRows = await page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb')); + const sprayRows = await pageSafe(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb')); const sprayByKey = new Map(); for (const r of sprayRows) { if (!r.player_key) continue; @@ -116,7 +145,7 @@ async function main() { if (!prev || String(r.as_of_date) > String(prev.as_of_date)) sprayByKey.set(r.player_key, r); } - const platRows = await page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb')); + const platRows = await pageSafe(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb')); const platByKey = new Map(); for (const r of platRows) { if (!r.player_key) continue; @@ -124,19 +153,15 @@ async function main() { if (!prev || String(r.as_of_date) > String(prev.as_of_date)) platByKey.set(r.player_key, r); } - const defRows = await page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb')); + const defRows = await pageSafe(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb')); const defByTeam = new Map(); for (const d of defRows) defByTeam.set(d.team, d); - const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat', - (q) => q.eq('sport', 'mlb').eq('stat', 'hits').not('archetype', 'is', null)); + const snaps = await READS.snaps(sb); const archOf = new Map(); for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype); - const led = await page(sb, 'ledger_entries', - 'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base', - (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits') - .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); + const led = await READS.ledger(sb); const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')); // Opponent faced, from each hitter's own game log. @@ -308,4 +333,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/prove-park-weather.js b/scripts/prove-park-weather.js index 713a424..c46a2ed 100644 --- a/scripts/prove-park-weather.js +++ b/scripts/prove-park-weather.js @@ -20,39 +20,53 @@ const pw = require('../src/services/model/parkWeather'); const fg = require('../src/services/model/factorGate'); const tl = require('../src/services/model/testLedger'); const { knownNumber } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; const PAGE = 1000; -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2/A2b) ──────────────────────────────────────────── +// This script used to walk pages with `.range()` and NO ORDER BY. Measured on +// production, that returned the correct row COUNT and the wrong ROWS: up to +// 33.6% of a read came back twice while an equal share never came back at all, +// so `rows.length` looked perfect while a fifth of the sample was missing. +// +// `pageSafe` routes every read through `src/utils/safePaginate`: a stable ORDER +// BY on the table's real UNIQUE key — single OR composite, looked up from +// `src/utils/tableKeys` rather than assumed — a runtime tuple-uniqueness check, +// and a THROWN error instead of a silent stop. The old unordered helper is gone +// rather than left beside it, because a dead broken helper is an invitation. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases') + .in('outcome', ['hit', 'miss'])), +}; + /** Hit-type shares → expected total bases, so a reshape has a consequence. */ const tbFromShares = (s) => s.single + 2 * s.double + 3 * s.triple + 4 * s.home_run; async function main() { const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const parks = await page(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb')); + const parks = await pageSafe(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb')); const byVenue = new Map(); for (const p of parks) if (!byVenue.has(p.venue_id)) byVenue.set(p.venue_id, p); const league = pw.leagueGeometry([...byVenue.values()]); - const ctx = await page(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q); + const ctx = await pageSafe(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q); const ctxBy = new Map(ctx.map((c) => [c.game_id, c])); - const led = await page(sb, 'ledger_entries', - 'game_id, game_date, player_key, stat, line, side, outcome, quarantine_reason, p_win, proj_hits_p_over', + const led = await pageSafe(sb, 'ledger_entries', + 'id, game_id, game_date, player_key, stat, line, side, outcome, quarantine_reason, p_win, proj_hits_p_over', (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases').in('outcome', ['hit', 'miss'])); const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')); @@ -118,4 +132,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/prove-runs-rbi.js b/scripts/prove-runs-rbi.js index af8e248..ed16540 100644 --- a/scripts/prove-runs-rbi.js +++ b/scripts/prove-runs-rbi.js @@ -37,6 +37,8 @@ const tl = require('../src/services/model/testLedger'); const sk = require('../src/services/model/skillProjection'); const { knownNumber, knownRate } = require('../src/utils/known'); const { nameKey } = require('../src/utils/playerName'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -52,18 +54,27 @@ const PA_EVENT = new Set([...HIT, 'field_out', 'strikeout', 'grounded_into_doubl const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null); -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2/A2b) ──────────────────────────────────────────── +// This script used to walk pages with `.range()` and NO ORDER BY. Measured on +// production, that returned the correct row COUNT and the wrong ROWS: up to +// 33.6% of a read came back twice while an equal share never came back at all, +// so `rows.length` looked perfect while a fifth of the sample was missing. +// +// `pageSafe` routes every read through `src/utils/safePaginate`: a stable ORDER +// BY on the table's real UNIQUE key — single OR composite, looked up from +// `src/utils/tableKeys` rather than assumed — a runtime tuple-uniqueness check, +// and a THROWN error instead of a silent stop. The old unordered helper is gone +// rather than left beside it, because a dead broken helper is an invitation. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + opportunity: (sb) => pageSafe(sb, 'hitter_opportunity', '*', (q) => q.eq('sport', 'mlb')), +}; + /** * Reconstruct, point-in-time, from play-by-play: * order[date|nameKey] the hitter's batting slot that game @@ -170,7 +181,7 @@ const FACTORS = { async function main() { const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb').eq('role', 'batter')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb').eq('role', 'batter')); const batByKey = new Map(); const barrelByKey = new Map(); for (const r of statcast) { @@ -180,7 +191,7 @@ async function main() { if (prof.barrel_pct != null) barrelByKey.set(r.player_key, prof.barrel_pct); } - const oppRows = await page(sb, 'hitter_opportunity', '*', (q) => q.eq('sport', 'mlb')); + const oppRows = await pageSafe(sb, 'hitter_opportunity', '*', (q) => q.eq('sport', 'mlb')); const oppByKey = new Map(); for (const r of oppRows) { const prev = oppByKey.get(r.player_key); @@ -192,12 +203,12 @@ async function main() { const out = { generated_note: 'null is the ARCHETYPE base rate, leave-one-out' }; for (const stat of ['rbi', 'runs']) { - const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype', + const snaps = await pageSafe(sb, 'model_snapshots', 'id, player_key, game_date, archetype', (q) => q.eq('sport', 'mlb').eq('stat', stat).not('archetype', 'is', null)); const archOf = new Map(); for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype); - const led = await page(sb, 'ledger_entries', + const led = await pageSafe(sb, 'ledger_entries', 'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason', (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', stat).in('outcome', ['hit', 'miss'])); const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book') @@ -298,4 +309,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS: (typeof READS !== 'undefined' ? READS : undefined), READS_LEDGER: (typeof READS_LEDGER !== 'undefined' ? READS_LEDGER : undefined) }; diff --git a/scripts/prove-tb-factors.js b/scripts/prove-tb-factors.js index 3120e45..26aef20 100644 --- a/scripts/prove-tb-factors.js +++ b/scripts/prove-tb-factors.js @@ -38,6 +38,8 @@ const { nameKey } = require('../src/utils/playerName'); const sd = require('../src/services/model/sprayDefense'); const pss = require('../src/services/model/platoonSeverity'); const pw = require('../src/services/model/parkWeather'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); /** League-typical hit-type shares; the atom reshapes these and TB follows. */ const BASE_SHARES = { single: 0.655, double: 0.195, triple: 0.017, home_run: 0.133 }; @@ -48,18 +50,37 @@ const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SER const PAGE = 1000; const ARCHS = (process.env.TB_ARCHETYPES || 'ALL,BOMBER,GHOST,BRUSH,DRIVER').split(','); -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── FIX A2 (2026-08-09) — THE SAFE WALK ─────────────────────────────────── +// `page()` above walks with no ORDER BY. Measured on production, that returned +// the correct row COUNT and the wrong ROWS: up to 33.6% of a read came back +// twice while an equal share never came back at all. `pageSafe` routes the same +// call through `src/utils/safePaginate`, which orders on a UNIQUE key on every +// page, verifies uniqueness at runtime, and THROWS on a query error instead of +// treating it as end-of-data. +// +// `page()` SURVIVES only for the context tables (statcast_aggregates, +// batter_spray, team_defense, platoon_splits, park_dimensions, game_context...). +// Those have COMPOSITE primary keys with no single unique column, so +// safePaginate cannot express them. They measure 0% corruption today; making +// them safe needs a composite-key ordering the helper does not yet have. Do not +// use `page()` for ledger_entries or model_snapshots. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + parks: (sb) => pageSafe(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb')), + gameCtx: (sb) => pageSafe(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q), + snaps: (sb) => pageSafe(sb, 'model_snapshots', 'id, player_key, game_date, archetype, stat', + (q) => q.eq('sport', 'mlb').eq('stat', 'total_bases').not('archetype', 'is', null)), + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases') + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)), +}; + /** * THE FACTORS. Each returns a MULTIPLIER on the base rate, or null when the * input is absent — an absent factor must leave the baseline untouched rather @@ -112,14 +133,14 @@ async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); const batters = new Map(); const pitchersById = new Map(); for (const r of statcast) { const prof = sk.fromStatcastRow(r); if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), prof); if (r.role === 'batter' && r.player_key) batters.set(r.player_key, prof); } - const sprayRows = await page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb')); + const sprayRows = await pageSafe(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb')); const sprayByKey = new Map(); for (const r of sprayRows) { if (!r.player_key) continue; @@ -127,7 +148,7 @@ async function main() { if (!prev || String(r.as_of_date) > String(prev.as_of_date)) sprayByKey.set(r.player_key, r); } - const platRows = await page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb')); + const platRows = await pageSafe(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb')); const platByKey = new Map(); for (const r of platRows) { if (!r.player_key) continue; @@ -135,26 +156,22 @@ async function main() { if (!prev || String(r.as_of_date) > String(prev.as_of_date)) platByKey.set(r.player_key, r); } - const parkRows = await page(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb')); + const parkRows = await pageSafe(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb')); const parkByVenue = new Map(); for (const p of parkRows) if (!parkByVenue.has(p.venue_id)) parkByVenue.set(p.venue_id, p); const parkLeague = pw.leagueGeometry([...parkByVenue.values()]); - const ctxRows = await page(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q); + const ctxRows = await pageSafe(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q); const ctxBy = new Map(ctxRows.map((c) => [c.game_id, c])); - const defRows = await page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb')); + const defRows = await pageSafe(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb')); const defByTeam = new Map(); for (const d of defRows) defByTeam.set(d.team, d); - const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat', - (q) => q.eq('sport', 'mlb').eq('stat', 'total_bases').not('archetype', 'is', null)); + const snaps = await READS.snaps(sb); const archOf = new Map(); for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype); - const led = await page(sb, 'ledger_entries', - 'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base', - (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases') - .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); + const led = await READS.ledger(sb); const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')); // Opponent faced, from each hitter's own game log. @@ -342,4 +359,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/proven-status.js b/scripts/proven-status.js index e6830bd..64a2be8 100644 --- a/scripts/proven-status.js +++ b/scripts/proven-status.js @@ -28,6 +28,8 @@ require('dotenv').config(); const { createClient } = require('@supabase/supabase-js'); const cv = require('../src/services/model/correlateValidator'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -47,23 +49,45 @@ const RECORDED = [ verdict: 'INCONCLUSIVE', spec: 'specs/lineup-k-rate-rung1.md' }, ]; -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── FIX A2 (2026-08-09) — THE SAFE WALK ─────────────────────────────────── +// `page()` above walks with no ORDER BY. Measured on production, that returned +// the correct row COUNT and the wrong ROWS: up to 33.6% of a read came back +// twice while an equal share never came back at all. `pageSafe` routes the same +// call through `src/utils/safePaginate`, which orders on a UNIQUE key on every +// page, verifies uniqueness at runtime, and THROWS on a query error instead of +// treating it as end-of-data. +// +// `page()` SURVIVES only for the context tables (statcast_aggregates, +// batter_spray, team_defense, platoon_splits, park_dimensions, game_context...). +// Those have COMPOSITE primary keys with no single unique column, so +// safePaginate cannot express them. They measure 0% corruption today; making +// them safe needs a composite-key ordering the helper does not yet have. Do not +// use `page()` for ledger_entries or model_snapshots. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS_LEDGER = { + ledgerAll: (sb) => pageSafe(sb, 'ledger_entries', 'id, stat, outcome, quarantine_reason, p_win', + (q) => q.eq('sport', 'mlb').is('user_id', null)), + ledgerSettled: (sb) => pageSafe(sb, 'ledger_entries', + 'id, stat, outcome, quarantine_reason, player_key, line, side, game_date', + (q) => q.eq('sport', 'mlb').is('user_id', null).in('outcome', ['hit', 'miss'])), +}; + +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + snaps: (sb) => pageSafe(sb, 'model_snapshots', 'id, stat, archetype, player_key, line, side, game_date', + (q) => q.eq('sport', 'mlb').not('archetype', 'is', null)), +}; + async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const led = await page(sb, 'ledger_entries', 'stat, outcome, quarantine_reason, p_win', + const led = await pageSafe(sb, 'ledger_entries', 'id, stat, outcome, quarantine_reason, p_win', (q) => q.eq('sport', 'mlb').is('user_id', null)); const settled = {}; for (const r of led) { @@ -73,8 +97,7 @@ async function main() { settled[r.stat] = (settled[r.stat] || 0) + 1; } - const snaps = await page(sb, 'model_snapshots', 'stat, archetype, player_key, line, side, game_date', - (q) => q.eq('sport', 'mlb').not('archetype', 'is', null)); + const snaps = await READS.snaps(sb); const archOf = new Map(); for (const s of snaps) archOf.set(`${s.player_key}|${s.stat}|${s.line}|${String(s.side).toLowerCase()}|${s.game_date}`, s.archetype); @@ -82,7 +105,7 @@ async function main() { // SNAPSHOT CYCLE, so a naive join fans out and inflates the count — it read // BOMBER x hits as 641 when the true figure is 287, which is the difference // between "gate-ready" and "not close". Dedupe on the ledger row's identity. - const led2 = await page(sb, 'ledger_entries', 'id, stat, outcome, quarantine_reason, player_key, line, side, game_date', + const led2 = await pageSafe(sb, 'ledger_entries', 'id, stat, outcome, quarantine_reason, player_key, line, side, game_date', (q) => q.eq('sport', 'mlb').is('user_id', null).in('outcome', ['hit', 'miss'])); const byArch = {}; const seen = new Set(); @@ -123,4 +146,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS, READS_LEDGER }; diff --git a/scripts/read-integrity.js b/scripts/read-integrity.js new file mode 100644 index 0000000..7729ac1 --- /dev/null +++ b/scripts/read-integrity.js @@ -0,0 +1,297 @@ +'use strict'; + +/** + * read-integrity — measure whether each learning reader returns the rows it thinks. + * + * Spec: specs/read-integrity-harness.md + * + * READ-ONLY. It issues SELECTs and a head-count. It writes nothing, anywhere — + * a test asserts this file contains no insert/update/upsert/delete/rpc. + * + * USAGE + * node scripts/read-integrity.js # every registered reader + * node scripts/read-integrity.js --id prove-hit-factors:136 + * node scripts/read-integrity.js --acceptance # AC5 only: known-corrupt + known-clean + * + * WHAT IT IS FOR: the 2026-08-09 probe proved exposure cannot be reasoned from + * source — `challenger-scoreboard` walks 12 pages of a live table and is clean, + * `calibrationService.fromLedger` walks 3 pages of the same table and is a + * quarter corrupt. So every reader fix must be verified by MEASUREMENT, on the + * day, against an ordered control. Never by the presence of an ORDER BY clause. + */ + +const path = require('path'); +require('dotenv').config({ path: path.join(__dirname, '..', '.env') }); +const { createClient } = require('@supabase/supabase-js'); +const ri = require('../src/utils/readIntegrity'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); + +// The local .env carries a transposed project ref (CLAUDE.md, S76). Allow an +// explicit override so a local run reaches the real project. +const SB_URL = process.env.READ_INTEGRITY_SUPABASE_URL || process.env.SUPABASE_URL; +const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; + +const MLB = ['eq', 'sport', 'mlb']; +const PUBLIC = ['is', 'user_id', null]; +const SETTLED = ['in', 'outcome', ['hit', 'miss']]; +const HAS_PWIN = ['not', 'p_win', 'is', null]; +const HAS_ARCH = ['not', 'archetype', 'is', null]; + +const LEDGER_KEY = ['id']; +// model_snapshots rows are one per prop PER CYCLE, so identity needs the cycle. +const SNAP_KEY = ['id']; + +/** + * THE REGISTRY — every reader that walks pages without a stable order, as it + * stands today. Declarative so it is auditable against the cited source line. + */ +const READERS = [ + // ── the hits gate ────────────────────────────────────────────────────── + { id: 'prove-hit-factors:136', source: 'scripts/prove-hit-factors.js:136', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/prove-hit-factors').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'hits'], SETTLED, HAS_PWIN] }, + { id: 'prove-hit-factors:131', source: 'scripts/prove-hit-factors.js:131', table: 'model_snapshots', + select: 'id', key: SNAP_KEY, + readerRows: (sb) => require('../scripts/prove-hit-factors').READS.snaps(sb), + filters: [MLB, ['eq', 'stat', 'hits'], HAS_ARCH] }, + + // ── the other stat gates ─────────────────────────────────────────────── + { id: 'prove-tb-factors:154', source: 'scripts/prove-tb-factors.js:154', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/prove-tb-factors').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'total_bases'], SETTLED, HAS_PWIN] }, + { id: 'prove-tb-factors:149', source: 'scripts/prove-tb-factors.js:149', table: 'model_snapshots', + select: 'id', key: SNAP_KEY, + readerRows: (sb) => require('../scripts/prove-tb-factors').READS.snaps(sb), + filters: [MLB, ['eq', 'stat', 'total_bases'], HAS_ARCH] }, + { id: 'prove-runs-rbi:200(rbi)', source: 'scripts/prove-runs-rbi.js:200', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + filters: [MLB, PUBLIC, ['eq', 'stat', 'rbi'], SETTLED] }, + { id: 'prove-runs-rbi:200(runs)', source: 'scripts/prove-runs-rbi.js:200', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + filters: [MLB, PUBLIC, ['eq', 'stat', 'runs'], SETTLED] }, + { id: 'prove-park-weather:54', source: 'scripts/prove-park-weather.js:54', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/prove-park-weather').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'total_bases'], SETTLED] }, + { id: 'pitcher-prove-k:137', source: 'scripts/pitcher-prove-k.js:137', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/pitcher-prove-k').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'strikeouts'], SETTLED, HAS_PWIN] }, + { id: 'cluster-prove:337', source: 'scripts/cluster-prove.js:337', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/cluster-prove').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'total_bases'], SETTLED, HAS_PWIN] }, + { id: 'tb-solo-and-interactions:260', source: 'scripts/tb-solo-and-interactions.js:260', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/tb-solo-and-interactions').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'total_bases'], SETTLED, HAS_PWIN] }, + { id: 'stagea-gate-run:142', source: 'scripts/stagea-gate-run.js:142', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/stagea-gate-run').READS.ledger(sb), + filters: [MLB, PUBLIC] }, + { id: 'skill-v1-stagea:185', source: 'scripts/skill-v1-stagea.js:185', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/skill-v1-stagea').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'hits'], SETTLED] }, + + // ── the status/ablation/board readers ────────────────────────────────── + { id: 'proven-status:76', source: 'scripts/proven-status.js:76', table: 'model_snapshots', + select: 'id', key: SNAP_KEY, + readerRows: (sb) => require('../scripts/proven-status').READS.snaps(sb), filters: [MLB, HAS_ARCH] }, + { id: 'champion-ablation:173', source: 'scripts/champion-ablation.js:173', table: 'model_snapshots', + select: 'id', key: SNAP_KEY, + readerRows: (sb) => require('../scripts/champion-ablation').READS.snaps(sb), + filters: [MLB, HAS_PWIN, ['not', 'features', 'is', null]] }, + { id: 'build-grade-bands:54', source: 'scripts/build-grade-bands.js:54', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, + readerRows: (sb) => require('../scripts/build-grade-bands').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'hits'], SETTLED, HAS_PWIN] }, + { id: 'build-grade-bands:49', source: 'scripts/build-grade-bands.js:49', table: 'model_snapshots', + select: 'id', key: SNAP_KEY, + readerRows: (sb) => require('../scripts/build-grade-bands').READS.snaps(sb), filters: [MLB, ['eq', 'stat', 'hits'], HAS_ARCH] }, + { id: 'challenger-scoreboard:112', source: 'scripts/challenger-scoreboard.js:112', table: 'ledger_entries', + select: 'id', key: LEDGER_KEY, filters: [MLB, PUBLIC, SETTLED, HAS_PWIN] }, + + // ── calibrationService: FIXED in A1, and measured through the REAL function ── + // `readerRows` calls the production `loadSettledRows`, so a PASS here means the + // shipped code returned a set identical to the control — not that its query was + // re-declared correctly in this registry. + { id: 'calibrationService.fromLedger:110', source: 'src/services/model/calibrationService.js:110', + table: 'ledger_entries', select: 'id', key: LEDGER_KEY, live: true, + readerRows: (sb) => require('../src/services/model/calibrationService') + .loadSettledRows(sb, { sport: 'mlb', stat: 'hits' }), + filters: [MLB, PUBLIC, ['eq', 'stat', 'hits'], SETTLED, HAS_PWIN, + ['lt', 'game_date', todayEt()]] }, + + // The PRIMARY calibrator (calibrationService is only its shadow). FIXED in A2. + { id: 'lowParamService.fromLedger:78', source: 'src/services/model/lowParamService.js:78', + table: 'ledger_entries', select: 'id', key: LEDGER_KEY, live: true, + readerRows: (sb) => require('../src/services/model/lowParamService') + .loadSettledRows(sb, { sport: 'mlb', stat: 'hits' }), + filters: [MLB, PUBLIC, ['eq', 'stat', 'hits'], SETTLED, HAS_PWIN, + ['lt', 'game_date', todayEt()]] }, + + // ── context tables — A2b: composite keys from the live schema ────────── + { id: 'context:statcast_aggregates', source: 'scripts/prove-hit-factors.js:104', table: 'statcast_aggregates', + select: '*', key: uniqueKeyFor('statcast_aggregates'), + readerRows: (sb) => require('../scripts/prove-hit-factors').READS.statcast(sb), filters: [MLB] }, + { id: 'context:batter_spray', source: 'scripts/prove-hit-factors.js:111', table: 'batter_spray', + select: '*', key: uniqueKeyFor('batter_spray'), + readerRows: (sb) => require('../scripts/prove-hit-factors').READS.spray(sb), filters: [MLB] }, + { id: 'context:platoon_splits', source: 'scripts/prove-hit-factors.js:119', table: 'platoon_splits', + select: '*', key: uniqueKeyFor('platoon_splits'), + readerRows: (sb) => require('../scripts/prove-hit-factors').READS.platoon(sb), filters: [MLB] }, + { id: 'context:team_defense', source: 'scripts/prove-hit-factors.js:127', table: 'team_defense', + select: '*', key: uniqueKeyFor('team_defense'), + readerRows: (sb) => require('../scripts/prove-hit-factors').READS.defense(sb), filters: [MLB] }, + { id: 'context:park_dimensions', source: 'scripts/prove-tb-factors.js:138', table: 'park_dimensions', + select: '*', key: uniqueKeyFor('park_dimensions'), + readerRows: (sb) => require('../scripts/prove-tb-factors').READS.parks(sb), filters: [MLB] }, + { id: 'context:game_context', source: 'scripts/prove-tb-factors.js:142', table: 'game_context', + select: 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', + key: uniqueKeyFor('game_context'), + readerRows: (sb) => require('../scripts/prove-tb-factors').READS.gameCtx(sb), filters: [] }, + { id: 'context:hitter_opportunity', source: 'scripts/prove-runs-rbi.js:183', table: 'hitter_opportunity', + select: '*', key: uniqueKeyFor('hitter_opportunity'), + readerRows: (sb) => require('../scripts/prove-runs-rbi').READS.opportunity(sb), filters: [MLB] }, + // lineup_context has no paginated READER today (lineupContextService writes it); + // measured as-written so its exposure is on the record rather than assumed. + { id: 'context:lineup_context', source: '(no paginated reader — write-only today)', table: 'lineup_context', + select: 'as_of_date, sport, game_pk, player_key', key: uniqueKeyFor('lineup_context'), filters: [MLB] }, + + // ── A2b: the three legacy reads, now converted ──────────────────────── + { id: 'proven-status:66', source: 'scripts/proven-status.js:66', table: 'ledger_entries', + select: 'id', key: uniqueKeyFor('ledger_entries'), + readerRows: (sb) => require('../scripts/proven-status').READS_LEDGER.ledgerAll(sb), + filters: [MLB, PUBLIC] }, + { id: 'proven-status:85', source: 'scripts/proven-status.js:85', table: 'ledger_entries', + select: 'id', key: uniqueKeyFor('ledger_entries'), + readerRows: (sb) => require('../scripts/proven-status').READS_LEDGER.ledgerSettled(sb), + filters: [MLB, PUBLIC, SETTLED] }, + { id: 'champion-ablation:176', source: 'scripts/champion-ablation.js:176', table: 'ledger_entries', + select: 'id', key: uniqueKeyFor('ledger_entries'), + readerRows: (sb) => require('../scripts/champion-ablation').READS_LEDGER.ledger(sb), + filters: [MLB, PUBLIC, SETTLED] }, + + // ── A2b: found during the sweep — the last unordered ledger walker ──── + // Measured 24.7% before conversion; converted in the same pass because it is + // the identical defect and the fix is the identical one-liner. + { id: 'calibrate-hits:40', source: 'scripts/calibrate-hits.js:40', table: 'ledger_entries', + select: 'id', key: uniqueKeyFor('ledger_entries'), + readerRows: (sb) => require('../scripts/calibrate-hits').READS.ledger(sb), + filters: [MLB, PUBLIC, ['eq', 'stat', 'hits'], SETTLED, HAS_PWIN] }, + + // ── A2b: the two WRITE-scripts, previously unmeasured ───────────────── + { id: 'backfill-context:55', source: 'scripts/backfill-context.js:55', table: 'ledger_entries', + select: 'id', key: uniqueKeyFor('ledger_entries'), writes: 'platoon_splits', + readerRows: (sb) => require('../scripts/backfill-context').READS.ledger(sb), + filters: [MLB, PUBLIC, ['in', 'stat', ['hits', 'total_bases']], SETTLED] }, + { id: 'backfill-context:64', source: 'scripts/backfill-context.js:64', table: 'platoon_splits', + select: 'as_of_date, sport, season, player_key', key: uniqueKeyFor('platoon_splits'), + writes: 'platoon_splits', + readerRows: (sb) => require('../scripts/backfill-context').READS.platoonHeld(sb), filters: [MLB] }, + { id: 'reconstruct-game-environment:61', source: 'scripts/reconstruct-game-environment.js:61', + table: 'ledger_entries', select: 'id', key: uniqueKeyFor('ledger_entries'), writes: 'game_context', + readerRows: (sb) => require('../scripts/reconstruct-game-environment').READS.ledger(sb), + filters: [MLB, PUBLIC, ['in', 'stat', ['hits', 'total_bases']], SETTLED] }, +]; + +function todayEt() { + return new Intl.DateTimeFormat('en-CA', { + timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit', + }).format(new Date()); +} + +/** Build the three queries a spec needs. Read-only by construction. */ +function bindings(sb, spec) { + const base = () => ri.applyFilters(sb.from(spec.table).select(spec.select), spec.filters); + // A2b — the CONTROL orders by the FULL unique tuple. Ordering it on a prefix + // would leave the control itself tie-scrambled, and a control that cannot be + // trusted must never issue a PASS. + const orderCols = spec.key || ['id']; + return { + exactCount: async () => { + const { count, error } = await ri.applyFilters( + sb.from(spec.table).select(orderCols[0], { count: 'exact', head: true }), spec.filters); + if (error) throw new Error(`count: ${error.message}`); + return count; + }, + // ARM (a): exactly as the reader stands — no order. + fetchUnordered: async (from, to) => { + const { data, error } = await base().range(from, to); + if (error) throw new Error(`unordered: ${error.message}`); + return data; + }, + // ARM (b): identical + a stable order. The CONTROL, which must itself + // reconcile to the server count before any PASS is issued. + fetchOrdered: async (from, to) => { + let q = base(); + for (const c of orderCols) q = q.order(c, { ascending: true }); + const { data, error } = await q.range(from, to); + if (error) throw new Error(`ordered: ${error.message}`); + return data; + }, + // Once a reader is fixed, arm (a) becomes the REAL function's output. + ...(spec.readerRows ? { readerRows: () => spec.readerRows(sb) } : {}), + }; +} + +async function main() { + if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); + const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); + + const argv = process.argv.slice(2); + const only = argv.includes('--id') ? argv[argv.indexOf('--id') + 1] : null; + const acceptance = argv.includes('--acceptance'); + + let specs = READERS; + if (only) specs = READERS.filter((s) => s.id === only); + if (acceptance) { + specs = READERS.filter((s) => ['prove-hit-factors:136', 'challenger-scoreboard:112'].includes(s.id)); + } + + const results = []; + for (const spec of specs) { + const r = await ri.measure(spec, bindings(sb, spec)); + results.push(r); + process.stderr.write(` ${r.verdict.padEnd(15)} ${r.id} — ${r.corruption_pct ?? '?'}%\n`); + } + + const summary = { + run_at: new Date().toISOString(), + readers_measured: results.length, + pass: results.filter((r) => r.verdict === ri.VERDICT.PASS).length, + fail: results.filter((r) => r.verdict === ri.VERDICT.FAIL).length, + other: results.filter((r) => ![ri.VERDICT.PASS, ri.VERDICT.FAIL].includes(r.verdict)).length, + worst_corruption_pct: results.reduce((m, r) => Math.max(m, r.corruption_pct || 0), 0), + }; + + if (acceptance) { + // AC5 — the harness must reproduce the probe or it is not trustworthy. + const corrupt = results.find((r) => r.id === 'prove-hit-factors:136'); + const clean = results.find((r) => r.id === 'challenger-scoreboard:112'); + summary.acceptance = { + known_corrupt_reader: corrupt && corrupt.id, + known_corrupt_pct: corrupt && corrupt.corruption_pct, + known_corrupt_in_16_to_25_band: !!corrupt && corrupt.corruption_pct >= 16 && corrupt.corruption_pct <= 25, + known_clean_reader: clean && clean.id, + known_clean_pct: clean && clean.corruption_pct, + known_clean_is_zero: !!clean && clean.corruption_pct === 0, + }; + summary.acceptance.AC5_PASSES = summary.acceptance.known_corrupt_in_16_to_25_band + && summary.acceptance.known_clean_is_zero; + } + + console.log(JSON.stringify({ summary, results }, null, 2)); +} + +if (require.main === module) { + main().then(() => process.exit(0)).catch((e) => { + console.error('read-integrity FAILED:', e.message); + process.exit(1); + }); +} + +module.exports = { READERS, bindings }; diff --git a/scripts/reconstruct-game-environment.js b/scripts/reconstruct-game-environment.js index 663bd63..bbe3370 100644 --- a/scripts/reconstruct-game-environment.js +++ b/scripts/reconstruct-game-environment.js @@ -28,6 +28,8 @@ require('dotenv').config(); const axios = require('axios'); const { createClient } = require('@supabase/supabase-js'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -41,24 +43,35 @@ const ARCHIVE = (lat, lon, date) => const squash = (s) => String(s || '').toLowerCase().replace(/[^a-z]/g, ''); -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2/A2b) ──────────────────────────────────────────── +// This script used to walk pages with `.range()` and NO ORDER BY. Measured on +// production, that returned the correct row COUNT and the wrong ROWS: up to +// 33.6% of a read came back twice while an equal share never came back at all, +// so `rows.length` looked perfect while a fifth of the sample was missing. +// +// `pageSafe` routes every read through `src/utils/safePaginate`: a stable ORDER +// BY on the table's real UNIQUE key — single OR composite, looked up from +// `src/utils/tableKeys` rather than assumed — a runtime tuple-uniqueness check, +// and a THROWN error instead of a silent stop. The old unordered helper is gone +// rather than left beside it, because a dead broken helper is an invitation. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + ledger: (sb) => pageSafe(sb, 'ledger_entries', 'id, game_id, game_date, stat, outcome, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases']) + .in('outcome', ['hit', 'miss'])), +}; + const get = async (url) => (await axios.get(url, { timeout: 60_000 })).data; async function main() { const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const led = await page(sb, 'ledger_entries', 'game_id, game_date, stat, outcome, quarantine_reason', + const led = await pageSafe(sb, 'ledger_entries', 'id, game_id, game_date, stat, outcome, quarantine_reason', (q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases']) .in('outcome', ['hit', 'miss'])); const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')); @@ -167,4 +180,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS: (typeof READS !== 'undefined' ? READS : undefined), READS_LEDGER: (typeof READS_LEDGER !== 'undefined' ? READS_LEDGER : undefined) }; diff --git a/scripts/skill-v1-stagea.js b/scripts/skill-v1-stagea.js index 4df5a09..43a6fbf 100644 --- a/scripts/skill-v1-stagea.js +++ b/scripts/skill-v1-stagea.js @@ -43,6 +43,8 @@ const reg = require('../src/services/model/featureRegistry'); const mlb = require('../src/services/adapters/mlbStatsAdapter'); const { nameKey } = require('../src/utils/playerName'); const { knownRate } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); /** Team games played by the 2026-07-21 profile freeze — turns season PA into PA/game. */ const GAMES_SO_FAR = Number(process.env.STAGEA_GAMES_SO_FAR || 103); @@ -100,18 +102,30 @@ function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260803) { }; } -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2/A2b) ──────────────────────────────────────────── +// This script used to walk pages with `.range()` and NO ORDER BY. Measured on +// production, that returned the correct row COUNT and the wrong ROWS: up to +// 33.6% of a read came back twice while an equal share never came back at all, +// so `rows.length` looked perfect while a fifth of the sample was missing. +// +// `pageSafe` routes every read through `src/utils/safePaginate`: a stable ORDER +// BY on the table's real UNIQUE key — single OR composite, looked up from +// `src/utils/tableKeys` rather than assumed — a runtime tuple-uniqueness check, +// and a THROWN error instead of a silent stop. The old unordered helper is gone +// rather than left beside it, because a dead broken helper is an invitation. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits') + .in('outcome', ['hit', 'miss'])), +}; + /** * `date|team` → the starter that team FACED. * @@ -165,7 +179,7 @@ async function main() { const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); // Frozen skill profiles — by name (batters) and by source_id (pitchers). - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); const freeze = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), ''); const freezeDate = freeze.slice(0, 10); const batters = new Map(); @@ -182,8 +196,8 @@ async function main() { } } - const led = await page(sb, 'ledger_entries', - 'player_key, player_name, stat, line, side, outcome, game_date, p_win, team, opponent, quarantine_reason', + const led = await pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, stat, line, side, outcome, game_date, p_win, team, opponent, quarantine_reason', (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits') .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); @@ -276,4 +290,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/stagea-gate-run.js b/scripts/stagea-gate-run.js index 6cf5d0f..d3313ae 100644 --- a/scripts/stagea-gate-run.js +++ b/scripts/stagea-gate-run.js @@ -38,6 +38,8 @@ const sk = require('../src/services/model/skillProjection'); const reg = require('../src/services/model/featureRegistry'); const mlb = require('../src/services/adapters/mlbStatsAdapter'); const { knownRate, knownNumber } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -80,18 +82,29 @@ function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260803) { }; } -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── THE SAFE WALK (Fix A2/A2b) ──────────────────────────────────────────── +// This script used to walk pages with `.range()` and NO ORDER BY. Measured on +// production, that returned the correct row COUNT and the wrong ROWS: up to +// 33.6% of a read came back twice while an equal share never came back at all, +// so `rows.length` looked perfect while a fifth of the sample was missing. +// +// `pageSafe` routes every read through `src/utils/safePaginate`: a stable ORDER +// BY on the table's real UNIQUE key — single OR composite, looked up from +// `src/utils/tableKeys` rather than assumed — a runtime tuple-uniqueness check, +// and a THROWN error instead of a silent stop. The old unordered helper is gone +// rather than left beside it, because a dead broken helper is an invitation. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null)), +}; + async function opposingStarters(dates) { const m = new Map(); for (const d of dates) { @@ -126,7 +139,7 @@ async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10); const batters = new Map(); const pitchersById = new Map(); for (const r of statcast) { @@ -139,8 +152,8 @@ async function main() { } } - const led = await page(sb, 'ledger_entries', - 'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', + const led = await pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', (q) => q.eq('sport', 'mlb').is('user_id', null) .in('stat', ['hits', 'total_bases']) .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); @@ -244,4 +257,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/scripts/tb-solo-and-interactions.js b/scripts/tb-solo-and-interactions.js index 392bf6e..a20c56f 100644 --- a/scripts/tb-solo-and-interactions.js +++ b/scripts/tb-solo-and-interactions.js @@ -44,6 +44,8 @@ const sk = require('../src/services/model/skillProjection'); const reg = require('../src/services/model/featureRegistry'); const mlb = require('../src/services/adapters/mlbStatsAdapter'); const { knownRate, knownNumber } = require('../src/utils/known'); +const { paginate } = require('../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../src/utils/tableKeys'); const SB_URL = process.env.SUPABASE_URL; const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY; @@ -151,18 +153,35 @@ function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260804) { }; } -async function page(sb, table, select, apply) { - const out = []; - for (let from = 0; ; from += PAGE) { - const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1); - if (error) throw error; - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < PAGE) break; - } - return out; +// ── FIX A2 (2026-08-09) — THE SAFE WALK ─────────────────────────────────── +// `page()` above walks with no ORDER BY. Measured on production, that returned +// the correct row COUNT and the wrong ROWS: up to 33.6% of a read came back +// twice while an equal share never came back at all. `pageSafe` routes the same +// call through `src/utils/safePaginate`, which orders on a UNIQUE key on every +// page, verifies uniqueness at runtime, and THROWS on a query error instead of +// treating it as end-of-data. +// +// `page()` SURVIVES only for the context tables (statcast_aggregates, +// batter_spray, team_defense, platoon_splits, park_dimensions, game_context...). +// Those have COMPOSITE primary keys with no single unique column, so +// safePaginate cannot express them. They measure 0% corruption today; making +// them safe needs a composite-key ordering the helper does not yet have. Do not +// use `page()` for ledger_entries or model_snapshots. +async function pageSafe(sb, table, select, apply, key = uniqueKeyFor(table)) { + return paginate(() => apply(sb.from(table).select(select)), + { key, pageSize: PAGE, label: `${table}` }); } +/** THE MEASURED READS — main() and the harness call the same functions. */ +const READS = { + snaps: (sb) => pageSafe(sb, 'model_snapshots', 'id, player_key, game_date, archetype', + (q) => q.eq('sport', 'mlb').eq('stat', 'total_bases').not('archetype', 'is', null)), + ledger: (sb) => pageSafe(sb, 'ledger_entries', + 'id, player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', + (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases') + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)), +}; + async function opposingStarters(dates) { const m = new Map(); for (const d of dates) { @@ -233,7 +252,7 @@ async function main() { if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required'); const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } }); - const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); + const statcast = await pageSafe(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb')); const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10); const batters = new Map(); const pitchersById = new Map(); for (const r of statcast) { @@ -252,15 +271,11 @@ async function main() { // model_snapshots carries the actual classification per prop, so the // conditioning variable is now a genuine BOMBER indicator, which is // categorical and therefore not a transform of barrel at all. - const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype', - (q) => q.eq('sport', 'mlb').eq('stat', 'total_bases').not('archetype', 'is', null)); + const snaps = await READS.snaps(sb); const archetypeBy = new Map(); for (const r of snaps) if (r.player_key && r.game_date) archetypeBy.set(`${r.player_key}|${r.game_date}`, r.archetype); - const led = await page(sb, 'ledger_entries', - 'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason', - (q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases') - .in('outcome', ['hit', 'miss']).not('p_win', 'is', null)); + const led = await READS.ledger(sb); // POINT-IN-TIME IS NO LONGER AVAILABLE FROM THIS TABLE. // // `statcast_aggregates` is upserted in place and keeps one as-of date. The @@ -433,4 +448,9 @@ async function main() { process.exit(0); } -main().catch((e) => { console.error(e); process.exit(1); }); +if (require.main === module) { + main().catch((e) => { console.error(e); process.exit(1); }); +} + +// Exported so the read-integrity harness measures THE REAL FUNCTION. +module.exports = { READS }; diff --git a/specs/a8-shadow-factor-gate.md b/specs/a8-shadow-factor-gate.md new file mode 100644 index 0000000..5d755c4 --- /dev/null +++ b/specs/a8-shadow-factor-gate.md @@ -0,0 +1,140 @@ +# PRE-REGISTRATION — A8: do the shadow-fired hits factors improve the forecast? + +**Status: PRE-REGISTERED, NOT RUN.** Written 2026-08-11, before any accrual +exists, so the decision rule cannot be chosen after seeing the answer. + +--- + +## 1. WHAT IS BEING TESTED, AND WHY IT IS A NEW TEST + +A5 measured that all three hits factors were skipped on **596/596** real graded +props — `prop.opponent` and `prop.opposing_pitcher` were never set by anything, +so `positionOaa`, `throws` and `pitcherHardHit` resolved on **zero** rows. A6 +resolved those keys and fired the factors in **shadow**: 248 fires on 308 props, +median would-be multiplier 1.017, range 0.803–1.250. + +So `pitcher_contact_profile` and `defense_by_direction` were never "proven and +then disconnected." They were proven on an **audit reconstruction that the +production path never ran**, and then withdrawn (A0) because the reader that +produced that proof was 16–25% corrupt. + +**This is therefore a NEW hypothesis, not a re-test.** It counts as a fresh test +against the cumulative Bonferroni denominator (`testLedger`). The S87 rule that +re-asking the same question on more data is not a new shot on goal does not apply: +the question is different. Previously — "does this factor improve a reconstructed +audit?" Now — "does the multiplier this factor actually produces, on the rows +where it actually fires, improve the served forecast out of sample?" + +Three factors × one stat = **3 new hypotheses** to record before the run. + +--- + +## 2. THE UNIT OF EVIDENCE + +One testable pair = a `model_snapshots` row that carries **both**: + +- `factor_inputs.would_fire` — written at grade time, pre-game, never applied +- `outcome` — settled from a box score by `snapshotSettlementService` + +A `would_fire` with no settled outcome is **not yet evidence** and must not be +counted toward N. + +Rows are eligible only if they also satisfy the standing gates: +`model_version = engine1@2026-08-07-fullwindow` (A3) and +`quarantine_reason` not `nontakeable_book%`. + +--- + +## 3. THE TEST — TWO PARTS, BOTH REQUIRED + +Run per factor, on **rows where that factor fired** (`would_fire.per_factor` +contains it). Scoring a factor on rows it did not touch dilutes any real effect +toward zero (S77). + +**Baseline** = the served forecast, i.e. `p_win` as actually graded (the factors +were off). +**Conditioned** = `clamp(p_win × would_fire.multiplier, 0.01, 0.99)` — exactly +what would have been served had the factor been live. + +Adjudicated by `factorGate.adjudicate()`: + +1. **MOVEMENT** — mean `|Δp|` ≥ `MIN_MOVEMENT` (0.01). Movement alone is + **THEATER** and is named as such. +2. **IMPROVEMENT** — paired bootstrap on **Brier**, cluster-resampled, CI + excluding zero at `1 − 0.05/tests` where `tests` is the programme-lifetime + cumulative count from `testLedger`. + +`cluster = game_id`. Verified: for all three factors the treatment entity +(`player|opponent`, `starter_id`, `player_key`) is more numerous than the games, +so `factorGate`'s coarser-of rule resolves to the game. **`MIN_CLUSTERS = 40` +distinct settled games is the binding constraint**, not `MIN_N`. + +### Pre-registered outcomes + +| verdict | meaning | consequence | +|---|---|---| +| `PROVES` | moves AND corrected CI below zero | candidate for A9 live turn-on — still a separate order | +| `NOT_PROVEN_AT_CORRECTED_BAR` | point estimate improves, CI spans zero | keep accruing; do NOT turn on | +| `THEATER` | moves, Brier delta ≥ 0 | **do not turn on, ever, at this multiplier** | +| `INERT` | mean shift < 0.01 | wiring is pointless; drop the factor | +| `CANDIDATE_PENDING_SAMPLE` | below N or clusters | keep accruing | + +**Pre-registered fallback:** if a factor returns `THEATER`, that is a real result +and it must be recorded as such — a factor that moves ~80% of the board and does +not improve Brier is the arch-v1 failure repeating, and the correct action is to +leave it off, not to re-tune the multiplier until it passes. + +**A `PROVES` does not turn anything on.** It makes the case for A9, which needs +its own before/after on the served board. + +--- + +## 4. WHAT WOULD INVALIDATE THE RUN + +- **Fewer than 40 distinct settled games** for that factor → refuse, report N. +- **`would_fire` not re-derivable** from `would_fire.shadow_inputs` on any row + (A6's recompute check) → the evidence is not trustworthy; stop. +- **Any row whose `would_fire` was written after first pitch** — settlement + already refuses post-hoc rows, and the same rule applies here. +- **Keys resolved from anything other than the pre-game lineup + probable + pitcher.** A retrospectively-derived opponent would make this a reconstruction + again, which is the whole thing A5/A6 exist to avoid. + +--- + +## 5. WHEN IT CAN RUN + +Measured accrual (settled MLB hits, repaired-champion rows): + +| | 2026-08-07 | 2026-08-08 | +|---|---|---| +| settled hits rows | 758 | 1,096 | +| distinct settled players | 208 | 188 | +| **distinct settled games** | **11** | **10** | + +Shadow fire rates (A6, 2026-08-09): `pitcher_contact_profile` 80.5%, +`defense_by_direction` 60.1%, `platoon_severity` 47.7%. + +At ~925 settled hits rows and **~10.5 distinct settled games per slate**: + +| factor | rows/slate | slates to `MIN_N` 500 | **slates to `MIN_CLUSTERS` 40** | binding | +|---|---|---|---|---| +| `pitcher_contact_profile` | ~745 | 1 | **~4** | clusters | +| `defense_by_direction` | ~556 | 1 | **~4** | clusters | +| `platoon_severity` | ~441 | 2 | **~4** | clusters | + +**Accrual starts at DEPLOY, not today** — 0 rows carry `factor_inputs` now. Add +the settlement lag (a slate settles the following day), so: + +> **A8 is runnable ≈ 5 slates after A6/A7 deploy** — about **5 days**, for all +> three factors together. Nothing accrues until the deploy lands. + +--- + +## 6. WHAT THIS DOES NOT DECIDE + +It does not reinstate `defense_by_direction` or `pitcher_contact_profile`. Those +withdrawals stand on their own cause (drawn through a corrupt reader). A `PROVES` +here is evidence about **the live multiplier on rows where it fires** — a +different claim from the original verdicts, and it should be recorded as a new +finding rather than as a reinstatement. diff --git a/specs/read-integrity-harness.md b/specs/read-integrity-harness.md new file mode 100644 index 0000000..ccd3250 --- /dev/null +++ b/specs/read-integrity-harness.md @@ -0,0 +1,248 @@ +# SPEC — Read-Integrity Harness + Verdict Withdrawal (Fix A0) + +**Status:** built 2026-08-09. Measure-only instrument. Fixes no reader. + +--- + +## 1. WHY + +The 2026-08-09 probe established, by measurement, that the shared `page()` +pagination used by the learning loop — `.range(from, from+PAGE-1)` with **no +`ORDER BY`** — returns the correct row COUNT and the wrong ROWS. + +Measured on production, same query, minutes apart: + +| reader | true count | duplicates | missing | corruption | +|---|---|---|---|---| +| hits gate ledger read (`prove-hit-factors.js:136`) | 2,490 | 412 → 617 | same | **16.5% → 24.8%** | +| archetype join (`prove-hit-factors.js:131`) | 10,038 | 2,454 | 2,454 | **24.4%** | +| `calibrationService.fromLedger:110` (LIVE) | 2,490 | 617 | 617 | **24.8%** | +| `challenger-scoreboard.js:112` | 11,485 | 0 | 0 | **0%** | + +Two facts drive this spec: + +1. **Exposure cannot be reasoned from source.** `challenger-scoreboard` walks 12 + pages of a live table and is clean; `fromLedger` walks 3 pages of the *same* + table and is a quarter corrupt. The difference is the query PLAN, which shifts + with filters, table statistics and concurrent writes. Nothing in the code + predicts it. +2. **Duplicates are worse than a random subsample.** They inflate + `factorGate.movement.n`, which is the number checked against `MIN_N = 500` + (`factorGate.js:194`), AND they corrupt the leave-one-out per-player baseline + that is the gate's null hypothesis. Both the test statistic and the thing it is + tested against move. + +So: one instrument, measured per reader, per run. Not fifteen hand-written checks, +and never an inference from source text. + +--- + +## 2. THE META-SCAR RULE (load-bearing) + +> **A reader PASSES only on measured set identity against an ordered control. +> The presence of an `ORDER BY` clause is NEVER evidence of correctness.** + +A guard that reads its own source text for the string `order(` would be a guard +verifying its own intention rather than its effect — the exact failure class this +codebase has corrected repeatedly. `tests/unit/readIntegrity.test.js` therefore +asserts that a walk which HAS a stable order but still returns duplicates is +reported **FAIL**. + +Corollary: **the control must validate itself.** If the ordered walk's distinct +count does not equal the server's `count(*)`, the control is untrustworthy and the +result is `CONTROL_INVALID` — never `PASS`. An instrument that cannot tell a clean +reader from a broken control is not an instrument. + +--- + +## 3. ENDPOINTS / SURFACE + +None. This is a developer instrument with no HTTP surface, no cron, no product +behaviour. It reads; it writes nothing, anywhere. + +### Modules + +| file | role | +|---|---| +| `src/utils/readIntegrity.js` | PURE + injectable core. No Supabase import. | +| `scripts/read-integrity.js` | CLI runner + the declarative registry of reader specs | +| `src/services/model/withdrawnVerdicts.js` | append-only withdrawal record | + +### Data shapes + +A **reader spec** is declarative data, so the registry is auditable: + +```js +{ + id: 'prove-hit-factors:136', + source: 'scripts/prove-hit-factors.js:136', + table: 'ledger_entries', + select: 'id, player_key, stat, game_date', + key: ['id'], // identity columns for dedupe + order: 'id', // the control's stable order + filters: [ + ['eq', 'sport', 'mlb'], + ['is', 'user_id', null], + ['in', 'outcome', ['hit', 'miss']], + ['not', 'p_win', 'is', null], + ], +} +``` + +A **result**: + +```js +{ + id, source, table, + exact_count, // server count(*), the ground truth + pages, + unordered: { rows, distinct, duplicates, missing_vs_control, extra_vs_control }, + ordered: { rows, distinct, duplicates }, + corruption_pct, // missing / exact_count + key_unique, // ordered.distinct === ordered.rows + control_valid, // ordered.distinct === exact_count + verdict, // PASS | FAIL | CONTROL_INVALID | KEY_NOT_UNIQUE | ERROR + reason, +} +``` + +--- + +## 4. ACCEPTANCE CRITERIA + +**AC1** — `analyze()` returns `PASS` only when `duplicates === 0 AND +missing_vs_control === 0 AND control_valid`. Verified by unit test. + +**AC2** — An ordered walk that still returns duplicates reports `FAIL`. (The +meta-scar guard: a clause is not a result.) + +**AC3** — `ordered.distinct !== exact_count` reports `CONTROL_INVALID`, never +`PASS`. + +**AC4** — A non-unique key reports `KEY_NOT_UNIQUE` rather than counting +legitimate data duplicates as walk corruption. (`batter_spray` keyed on +`player_key|as_of_date` has 6 genuine duplicate keys; that is a data fact, not a +pagination fault.) + +**AC5** — Reproduces the probe on live data: the hits ledger read reports +corruption in the **16–25%** band; `challenger-scoreboard:112` reports **0%**. +If the harness cannot reproduce both, it is not trustworthy and must not be used. + +**AC6** — `walk()` is generic over a `fetchPage(from, to)` function, so the unit +tests never touch a network. + +**AC7** — The harness performs **zero writes**. No insert, update, upsert, delete, +or DDL anywhere in `readIntegrity.js` or `read-integrity.js`. + +--- + +## 5. TEST PLAN + +Unit (`tests/unit/readIntegrity.test.js`), all with fake page-fetchers: + +- clean walk → `PASS` +- walk whose last page re-emits earlier rows (the measured production shape) → + `FAIL` with the exact duplicate/missing counts +- **ordered-but-duplicating walk → `FAIL`** (AC2, the meta-scar) +- ordered walk short of `exact_count` → `CONTROL_INVALID` (AC3) +- non-unique key → `KEY_NOT_UNIQUE` (AC4) +- `walk()` stops on a short page and on an empty page +- `applyFilters` chains `eq/is/in/not/lt` in order onto a fake builder +- a query error propagates as `ERROR`, never as end-of-data (the + `fromLedger:117` failure class — an error swallowed as "no more rows") + +Live (`scripts/read-integrity.js`): AC5 above, run against prod, read-only. + +--- + +## 6. VERDICT WITHDRAWAL — WHAT IT IS AND IS NOT + +### Where standing verdicts actually live + +There is **no verdict table**. `featureRegistry.statVerdicts` is an in-memory +`Map` (`featureRegistry.js:145`) that resets every process; `mc_test_ledger` +stores hypothesis COUNTS for the Bonferroni denominator, not verdicts. The +standing record is the `specs/` markdown plus `CLAUDE.md`. + +So the withdrawal is **appended as a committed record** +(`src/services/model/withdrawnVerdicts.js` + this spec), never by editing or +deleting a prior finding. The original verdicts stay exactly as recorded; the +withdrawal sits beside them with its cause. + +### What is withdrawn + +| factor | stat | recorded verdict | now | +|---|---|---|---| +| `defense_by_direction` | hits | PROVES, n=528, Brier −0.0034, CI [−0.0059,−0.0009] @99.9% | **WITHDRAWN_PENDING_REAUDIT** | +| `pitcher_contact_profile` | hits | PROVES, n=741, Brier −0.0066, CI [−0.0114,−0.0016] @99.9% | **WITHDRAWN_PENDING_REAUDIT** | + +Cause, recorded verbatim on both: *drawn through an unordered-pagination reader; +16.5–24.8% set corruption measured on the identical query 2026-08-09; distinct-n +below the 500 floor for `defense_by_direction`; not retrospectively recoverable.* + +`defense_by_direction` additionally carries the eligibility note: at the measured +~20% duplication its distinct n is ≈422 against `MIN_N = 500`, so it was likely +**never eligible for adjudication at all**. This is arithmetic on the recorded n, +not a re-measurement. + +### What withdrawal does NOT do + +It does **not** stop the factors moving the served forecast. +`hitsFactors.js` hardcodes its three factors and consults no registry, so +withdrawal is a record of what we may CLAIM, not a change to what we SERVE. +Disarming or re-proving is a later order with its own blast radius. + +### Reinstatement + +A withdrawn verdict returns only by re-running its gate through readers that this +harness reports `PASS` for, on the day of the run. Nothing is reinstated by +argument. + +--- + +## 7. THE FIX — `safePaginate` (Fix A1, 2026-08-09) + +`src/utils/safePaginate.js` is the one way this codebase walks a paginated +PostgREST read. It closes both measured defects in a single call: + +```js +const rows = await paginate( + () => sb.from('ledger_entries').select('id, …').eq('sport', 'mlb') /* … */, + { key: 'id', pageSize: 1000, label: 'caller(name)' }, +); +``` + +**Requirements it enforces** + +1. **Stable ORDER BY on a UNIQUE key.** An order on a non-unique column is NOT a + fix — ties may be returned in any order, so pages still overlap. Default `id`. +2. **Uniqueness verified at runtime, not promised.** A repeated key mid-walk + means either a non-unique key or a scrambled read; both make the row set + unusable, so it THROWS rather than returning it. The harness PASS is + confirmation; this is the guard that runs every time. +3. **A query error THROWS.** Never end-of-data. `paginate` takes a query + *factory* (builders are single-use) and reuses `readIntegrity.walk`, which is + already tested for short-page / empty-page / error / runaway behaviour — it + does not hand-roll a second walk. + +**Refusal is not a swallow.** A caller that responds to a throw by producing +nothing (no calibrator, no map) is refusing honestly. A caller that continues +with partial rows is repeating the defect. `null` from `fromLedger` means "not +enough settled history"; a throw means "the read failed". Keeping those distinct +is the point. + +### Measuring a FIXED reader + +A registry spec that restates the corrected query would verify the restatement, +and a `fixed: true` flag would verify a comment. So a fixed reader supplies +`readerRows` instead — a call into the **real production function**, whose +actual output is compared to the ordered control. `arm_a` on the result reports +which was used (`real_reader_function` vs `query_as_written`). + +### AC7 — A1 acceptance (met 2026-08-09) + +- `calibrationService.fromLedger:110` — **FAIL 24.8% → PASS (0 dup / 0 missing, + 2,490 of 2,490)**, measured through `loadSettledRows`. +- `lowParamService.fromLedger:78` — the byte-identical twin, deliberately NOT + fixed in A1, still measures **24.8%**. The harness kept its teeth; the PASS is + not an artefact of editing the registry. diff --git a/src/config/modelVersion.js b/src/config/modelVersion.js new file mode 100644 index 0000000..851e281 --- /dev/null +++ b/src/config/modelVersion.js @@ -0,0 +1,54 @@ +'use strict'; + +/** + * MODEL VERSION — ONE source, because two were not the same. + * + * ── WHAT WENT WRONG ────────────────────────────────────────────────────── + * `ledgerService` and `retentionService` both read `process.env.MODEL_VERSION`, + * and both carried a HARDCODED DEFAULT — but different ones: + * + * ledgerService.js:57 … || 'engine1@2026-07-20' + * retentionService.js:59 … || 'engine1@2026-08-07-fullwindow' + * + * Production does not set the env var, so the two tables took different + * defaults. The 2026-08-07 champion repair bumped one and left the other behind. + * Result: `model_snapshots` carried the repaired marker on 34,128 rows while + * `ledger_entries` carried the OLD marker on all 15,739 — including rows graded + * by the repaired champion. + * + * That is not a cosmetic mismatch. `reAuditEligibility.isEligible` requires + * `model_version === REPAIRED_CHAMPION_VERSION`, so applied to the ledger it + * returned ZERO eligible rows and ZERO eligible dates, permanently, for all four + * pending measurements. The accrual clock read "blocked" when the real state was + * "mis-stamped". + * + * ── THE RULE ───────────────────────────────────────────────────────────── + * There is exactly one place a model version may be declared. Anything that + * stamps a row imports it from here. A default that lives next to its consumer + * will drift from the other consumer, and the drift is invisible because both + * sides look locally correct. + * + * ── WHAT THIS DOES NOT DO ──────────────────────────────────────────────── + * It does not re-stamp history. Rows already written keep the marker they were + * written with — rewriting them would destroy the one record of which forecast + * actually produced them, which is the thing the marker exists to preserve. Old + * rows stay honestly old; a backfill is a separate, explicit decision. + */ + +/** The version every NEW row is stamped with. Override with MODEL_VERSION. */ +const MODEL_VERSION = process.env.MODEL_VERSION || 'engine1@2026-08-07-fullwindow'; + +/** + * Rows at or after this marker were produced by the repaired champion and are + * eligible for forward re-audit. + * + * Deliberately NOT `MODEL_VERSION`: the eligibility bar is a fixed historical + * fact about a specific repair, and tying it to "whatever we stamp today" would + * make every future version silently re-qualify itself. + */ +const REPAIRED_CHAMPION_VERSION = 'engine1@2026-08-07-fullwindow'; + +/** The marker that preceded the repair — kept so old rows are nameable. */ +const PRE_REPAIR_VERSION = 'engine1@2026-07-20'; + +module.exports = { MODEL_VERSION, REPAIRED_CHAMPION_VERSION, PRE_REPAIR_VERSION }; diff --git a/src/services/gradeSlateService.js b/src/services/gradeSlateService.js index f54601f..379e4f4 100644 --- a/src/services/gradeSlateService.js +++ b/src/services/gradeSlateService.js @@ -101,8 +101,15 @@ async function gradeBestSide(grade, prop, sport, opts = {}) { // previously computed downstream of the grade it should inform. const factorContext = typeof opts.factorContext === 'function' ? opts.factorContext(prop, sport) : null; + // FIX A6 — the SHADOW pair. `factor_context` (above) is the LIVE context and + // is unchanged; these two let the engine compute what the factors WOULD do + // with the join keys, and freeze it. Neither reaches the served forecast. + const matchupKeys = typeof opts.matchupKeys === 'function' + ? opts.matchupKeys(prop, sport) : null; const base = { factor_context: factorContext, + factor_context_resolver: typeof opts.factorContext === 'function' ? opts.factorContext : null, + matchup_keys: matchupKeys, player: prop.player, stat_type: prop.stat_type, line: prop.line, diff --git a/src/services/intelligence/analyzeViaEngine1.js b/src/services/intelligence/analyzeViaEngine1.js index 43c654c..5122037 100644 --- a/src/services/intelligence/analyzeViaEngine1.js +++ b/src/services/intelligence/analyzeViaEngine1.js @@ -567,6 +567,27 @@ async function analyzeViaEngine1(rawProp = {}) { pOver = adj.p_adjusted; factorTrace = { multiplier: adj.multiplier, applied: adj.applied, skipped: adj.skipped, p_before: adj.p_base }; } + // ── FIX A5 — FREEZE THE INPUTS, NOT THE OUTPUT ────────────────── + // Captured from the SAME context object the factor was just handed, in + // the same pass, so the recorded value cannot drift from the used one. + // Purely additive: it reads `rawProp.factor_context` and writes a field. + // `pOver` above is untouched by this block, so the served grade cannot + // move — a freeze records, it does not compute. + try { + const ffz = require('../model/factorFreeze'); + // ── FIX A6 — SHADOW RESOLVE ────────────────────────────────── + // With the join keys the factors CAN fire. We compute what they + // WOULD apply and FREEZE it. `pOver` is not touched by this block: + // turning them on for real is a separate, gated decision. + let shadow = null; + const keys = rawProp.matchup_keys || null; + if (keys && (keys.opponent || keys.opposing_pitcher) + && typeof rawProp.factor_context_resolver === 'function') { + const shadowCtx = rawProp.factor_context_resolver(rawProp, rawProp.sport, keys); + if (shadowCtx) shadow = ffz.shadowBlock(shadowCtx, hf.hitsFactorMultiplier(shadowCtx), keys); + } + legacy.factor_inputs = ffz.freeze(rawProp.factor_context, rawProp, shadow); + } catch { /* recording must never break a grade either */ } } catch { /* a factor must never break the grade */ } } diff --git a/src/services/ledgerService.js b/src/services/ledgerService.js index 1257aed..01e40d9 100644 --- a/src/services/ledgerService.js +++ b/src/services/ledgerService.js @@ -54,7 +54,11 @@ const SETTLE_ATTEMPT_CAP = Number(process.env.SETTLE_ATTEMPT_CAP || 4); // from originals and the harness can filter by how a row was scored. const SETTLEMENT_VERSION = Number(process.env.SETTLEMENT_VERSION || 2); const settleSource = require('./settleSource'); -const MODEL_ERA_VERSION = process.env.MODEL_VERSION || 'engine1@2026-07-20'; +// FIX A3 — ONE SOURCE. This used to default to 'engine1@2026-07-20' while +// retentionService defaulted to the repaired version, so the same env var +// produced two different stamps and no ledger row ever carried the marker +// `reAuditEligibility` requires. Never re-declare a version default here. +const { MODEL_VERSION: MODEL_ERA_VERSION } = require('../config/modelVersion'); // Order Zero — which fair-probability RULER produced fair_prob_lock. It is the // denominator of every edge and CLV number, so a value computed under one ruler // is not the same measurement as one computed under another. Never pool across diff --git a/src/services/model/calibrationService.js b/src/services/model/calibrationService.js index 31f0bdc..0317075 100644 --- a/src/services/model/calibrationService.js +++ b/src/services/model/calibrationService.js @@ -29,6 +29,7 @@ */ const cal = require('./calibration'); +const { paginate } = require('../../utils/safePaginate'); const { knownNumber } = require('../../utils/known'); /** Default: hold out the most recent quarter of history to certify on. */ @@ -100,24 +101,54 @@ function build(settled, opts = {}) { * Injectable for tests; returns null rather than a permissive fallback, because * "no calibrator" must mean "nothing is stackable", not "pass the raw numbers * through". + * + * ── FIX A1 (2026-08-09) — THIS READ WAS 24.8% CORRUPT ──────────────────── + * It walked pages with `.range()` and NO ORDER BY, and swallowed a query error + * as end-of-data. Measured on production: 617 of 2,490 rows returned twice and + * an equal number never returned, while `rows.length` matched the server count + * exactly. A map fitted on that is fitted on a sample where a fifth of the + * history is double-weighted and another fifth is absent. + * + * Both defects are now `safePaginate`: a stable ORDER BY on the unique `id` + * (an order on a non-unique column is not a fix — ties still scramble), plus a + * thrown error instead of a silent stop. + * + * A THROW IS NOT A NULL HERE, AND THE DIFFERENCE IS LOAD-BEARING. `null` means + * "not enough settled history to fit" — a real, expected state. A throw means + * "the read failed", which used to be indistinguishable from the first. It + * propagates so a caller can never mistake a broken read for thin history and + * quietly serve a map built on a fragment. */ -async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) { - if (!sb) return null; - const cutoff = before || new Intl.DateTimeFormat('en-CA', { +function todayEt() { + return new Intl.DateTimeFormat('en-CA', { timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit', }).format(new Date()); - const rows = []; - for (let from = 0; ; from += 1000) { - const { data, error } = await sb.from('ledger_entries') - .select('p_win, outcome, game_date, quarantine_reason') +} + +/** + * The ROW LOAD, exported so the read-integrity harness can measure THE REAL + * FUNCTION rather than a re-declaration of its query. + * + * That distinction is the whole point of the harness: a spec that restates the + * query would verify the restatement, and a flag saying "this one is fixed now" + * would verify a comment. The harness calls this. + */ +async function loadSettledRows(sb, { sport = 'mlb', stat = 'hits', before = null } = {}) { + const cutoff = before || todayEt(); + return paginate( + () => sb.from('ledger_entries') + .select('id, p_win, outcome, game_date, quarantine_reason') .eq('sport', sport).is('user_id', null).eq('stat', stat) .in('outcome', ['hit', 'miss']).not('p_win', 'is', null) - .lt('game_date', cutoff) // STRICTLY before — the whole point - .range(from, from + 999); - if (error || !data || data.length === 0) break; - rows.push(...data); - if (data.length < 1000) break; - } + .lt('game_date', cutoff), // STRICTLY before — the whole point + { key: 'id', pageSize: 1000, label: `calibrationService.fromLedger(${sport}/${stat})` }, + ); +} + +async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) { + if (!sb) return null; + const cutoff = before || todayEt(); + const rows = await loadSettledRows(sb, { sport, stat, before: cutoff }); const clean = rows .filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')) .map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, date: String(r.game_date) })); @@ -125,4 +156,4 @@ async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ... return built ? { ...built, cutoff } : null; } -module.exports = { build, fromLedger, HOLDOUT_FRACTION, MIN_FIT }; +module.exports = { build, fromLedger, loadSettledRows, HOLDOUT_FRACTION, MIN_FIT }; diff --git a/src/services/model/factorFreeze.js b/src/services/model/factorFreeze.js new file mode 100644 index 0000000..e68d7b6 --- /dev/null +++ b/src/services/model/factorFreeze.js @@ -0,0 +1,164 @@ +'use strict'; + +/** + * factorFreeze — record WHAT THE FACTOR READ, at the moment it read it. + * + * Spec: specs/read-integrity-harness.md §10 + * + * ── WHY ────────────────────────────────────────────────────────────────── + * A4 established that a re-audit joining today's context to a past grade is + * contamination, and fixed it with an as-of cutoff. But an as-of join still + * RECONSTRUCTS: it asks "what would the context have been" rather than "what did + * the grader actually see". Those differ whenever a context table was refreshed + * late, backfilled, or corrected. The only way a row becomes independently + * checkable is if it carries its own inputs. + * + * ── INPUTS, NEVER OUTPUTS ──────────────────────────────────────────────── + * This freezes the RAW values the factor consumed — spray shares, the opposing + * team's positional OAA, the pitcher's hard-hit rate, the platoon split counts, + * handedness — and deliberately NOT the resulting multiplier. A stored + * multiplier is unfalsifiable: it can only be compared to itself. Stored inputs + * can be re-run through `hitsFactors` and the answer checked, which is what + * makes an audit an audit. `recompute()` below exists precisely so that check is + * one call. + * + * ── IT RECORDS ABSENCE AS ABSENCE ──────────────────────────────────────── + * Measured on 596 real graded hits props (2026-08-10): `positionOaa`, `throws` + * and `pitcherHardHit` resolved on ZERO of them, because nothing in the pipeline + * ever sets `prop.opponent` or `prop.opposing_pitcher` — the two join keys the + * resolver needs. All three factors were skipped on 596/596 and the multiplier + * was exactly 1 every time. + * + * So today this vector mostly freezes nulls, and that is the correct behaviour: + * it is the evidence that the factors are wired but inert. `available` records + * what COULD have been joined from the row, so a later order can measure what + * plumbing those keys would change before changing it. `available` is never fed + * to a factor — reading it here would silently alter the served grade. + */ + +const { knownNumber } = require('../../utils/known'); + +/** Version the shape, so a later change is distinguishable from a null. */ +const FREEZE_VERSION = 'hits-factors@1'; + +const num = (v) => knownNumber(v); + +/** The spray shares a defence read consumes — raw, not the multiplier. */ +function sprayInputs(spray) { + if (!spray) return null; + return { + as_of: spray.as_of_date || null, + pull_gb: num(spray.pull_gb), straight_gb: num(spray.straight_gb), oppo_gb: num(spray.oppo_gb), + pull_air: num(spray.pull_air), straight_air: num(spray.straight_air), oppo_air: num(spray.oppo_air), + }; +} + +/** Platoon split counts — the raw PAs and hits, so severity is re-derivable. */ +function platoonInputs(splits) { + if (!splits) return null; + const side = (s) => (s ? { pa: num(s.pa), ab: num(s.atBats), hits: num(s.hits) } : null); + return { vl: side(splits.vl), vr: side(splits.vr) }; +} + +/** + * Build the frozen vector from the SAME context object the factor was handed. + * + * @param {object} ctx the resolved factor context (exactly what hitsFactors got) + * @param {object} prop the graded prop, for the join keys as they stood + */ +/** + * FIX A6 — the SHADOW would-fire block. + * + * With the join keys resolved, the factors CAN compute. This records what they + * WOULD have applied, namespaced under `would_fire`, and it is never added to + * the served forecast. It is evidence for a later, gated decision — not a grade. + * + * `multiplier` is stored here (unlike the inputs rule) because the shadow inputs + * are stored ALONGSIDE it under `shadow_inputs`, so it stays recomputable and + * checkable. A number you cannot re-derive is the thing that rule forbids. + */ +function shadowBlock(shadowCtx, applied, keys) { + if (!shadowCtx || !applied) return null; + return { + keys: { + opponent: (keys && keys.opponent) || null, + opposing_pitcher: (keys && keys.opposing_pitcher) || null, + source: (keys && keys.source) || null, + refused: (keys && keys.refused) || null, + }, + multiplier: applied.multiplier, + factors_fired: applied.factors_fired, + per_factor: applied.applied.map((a) => ({ factor: a.factor, multiplier: a.multiplier })), + skipped: applied.skipped.map((sk) => sk.factor), + shadow_inputs: { + position_oaa: shadowCtx.positionOaa || null, + throws: shadowCtx.throws || null, + pitcher_hard_hit: knownNumber(shadowCtx.pitcherHardHit), + }, + }; +} + +function freeze(ctx, prop = {}, shadow = null) { + if (!ctx) return null; + return { + v: FREEZE_VERSION, + as_of: ctx.as_of || null, + + // ── the join keys, as the RESOLVER saw them ── + // Null here is the finding, not an omission: nothing sets these, so every + // opponent-side and pitcher-side input below is null as a consequence. + join: { + opponent: prop.opponent || prop.opp_team || null, + opposing_pitcher: prop.opposing_pitcher || null, + }, + + // ── the raw inputs each factor consumed ── + spray: sprayInputs(ctx.spray), + position_oaa: ctx.positionOaa || null, + bats: ctx.bats || null, + throws: ctx.throws || null, + pitcher_hard_hit: num(ctx.pitcherHardHit), + platoon: platoonInputs(ctx.platoonSplits), + + // Carried from the graded row, never invented (A4). + archetype: ctx.archetype || null, + + // ── WHAT WAS AVAILABLE BUT NOT USED ── + // Recorded so a later order can measure the effect of plumbing the join keys + // BEFORE it changes a served number. Never read by a factor. + available: { + team: prop.team || null, + game_id: prop.game_id || null, + game_date: prop.game_date || null, + }, + + // SHADOW ONLY (A6). Never read by the served grade. + would_fire: shadow || null, + }; +} + +/** + * Re-run the factors from a FROZEN vector. + * + * This is the whole point of freezing inputs rather than a multiplier: an audit + * can recompute and compare. Returns the same shape `hitsFactorMultiplier` does. + */ +function recompute(frozen, deps = {}) { + if (!frozen) return null; + const hf = deps.hitsFactors || require('./hitsFactors'); + return hf.hitsFactorMultiplier({ + spray: frozen.spray, + positionOaa: frozen.position_oaa, + bats: frozen.bats, + throws: frozen.throws, + pitcherHardHit: frozen.pitcher_hard_hit, + platoonSplits: frozen.platoon + ? { + vl: frozen.platoon.vl ? { pa: frozen.platoon.vl.pa, atBats: frozen.platoon.vl.ab, hits: frozen.platoon.vl.hits } : null, + vr: frozen.platoon.vr ? { pa: frozen.platoon.vr.pa, atBats: frozen.platoon.vr.ab, hits: frozen.platoon.vr.hits } : null, + } + : null, + }); +} + +module.exports = { freeze, recompute, shadowBlock, FREEZE_VERSION, sprayInputs, platoonInputs }; diff --git a/src/services/model/hitsFactorContext.js b/src/services/model/hitsFactorContext.js index 225c061..6d3ad45 100644 --- a/src/services/model/hitsFactorContext.js +++ b/src/services/model/hitsFactorContext.js @@ -14,10 +14,38 @@ * Every load is best-effort: a missing table yields an empty index, the factor * finds nothing readable, and the forecast is served unadjusted. A factor layer * must never be able to break the pipeline it rides in. + * + * ── AS-OF (Fix A4) ─────────────────────────────────────────────────────── + * These tables are dated snapshots — one row per entity per day. Without a + * cutoff, `latestBy` returns TODAY'S row, so re-auditing a 2026-08-07 grade + * would join a profile built from games played AFTER it. That is contamination + * through a perfectly clean read, and it is the reason a re-audit could not be + * trusted even once A1/A2b made every reader return the rows it thinks. + * + * `opts.asOf` bounds every dated read to `as_of_date <= asOf` and then takes the + * latest WITHIN that bound. + * + * REFUSAL OVER RECONSTRUCTION. If an entity has no row at or before `asOf`, its + * factor input is null and the factor does not apply — the base rate is left + * untouched. It NEVER falls back to the nearest or the latest row: a substituted + * row is a plausible wrong value wearing a date, which is worse than an absence + * because nothing downstream can see it. + * + * THE LIVE PATH IS UNTOUCHED. With no `asOf` the code below runs exactly as it + * did — same tables, same filters, same indexes — so the served grade cannot + * move. Only an explicitly dated call takes the audit path. + * + * STATCAST IS THE EXCEPTION THAT PROVES THE RULE. `statcast_aggregates` is + * upserted in place and keeps ONE as-of date, so it cannot answer an as-of + * question at all; the dated path reads `statcast_history` instead. Verified + * equivalent at the head: on 2026-08-09 the two agree on all 1,414 rows with 0 + * differences, so switching sources does not itself move a number. */ const { knownNumber } = require('../../utils/known'); const { nameKey } = require('../../utils/playerName'); +const { paginate } = require('../../utils/safePaginate'); +const { uniqueKeyFor } = require('../../utils/tableKeys'); /** Rows a factor table must have before we trust it at all. */ const MIN_ROWS = 1; @@ -32,20 +60,18 @@ const MIN_ROWS = 1; * wiring fault has worn the costume of an honest absence, so the error is now * surfaced rather than swallowed. */ -async function page(sb, table, select, orderBy, apply) { - const out = []; - for (let from = 0; ; from += 1000) { - const q = apply ? apply(sb.from(table).select(select)) : sb.from(table).select(select); - const { data, error } = await q.order(orderBy, { ascending: true }).range(from, from + 999); - if (error) throw new Error(`${table}: ${error.message}`); - if (!data || data.length === 0) break; - out.push(...data); - if (data.length < 1000) break; - } - return out; +async function page(sb, table, select, apply) { + return paginate(() => (apply ? apply(sb.from(table).select(select)) : sb.from(table).select(select)), + { key: uniqueKeyFor(table), pageSize: 1000, label: `hitsFactorContext:${table}` }); } -/** Keep the most recent dated row per key. */ +/** + * Keep the most recent dated row per key. + * + * With an as-of bound applied at the query, every row here is already at or + * before the cutoff, so "most recent" IS "the as-of row". An entity with no row + * in the bounded set simply never enters the map — which is the refusal. + */ function latestBy(rows, keyFn, dateFn) { const m = new Map(); for (const r of rows) { @@ -63,13 +89,25 @@ function latestBy(rows, keyFn, dateFn) { */ async function build(sb, opts = {}) { if (!sb) return null; + // null === LIVE. Only an explicitly dated call takes the audit path, so the + // served grade runs the same code it ran before this parameter existed. + const asOf = opts.asOf ? String(opts.asOf) : null; + const dated = (q) => (asOf ? q.lte('as_of_date', asOf) : q); + let spray; let defense; let platoon; let statcast; try { [spray, defense, platoon, statcast] = await Promise.all([ - page(sb, 'batter_spray', '*', 'player_key', (q) => q.eq('sport', 'mlb')), - page(sb, 'team_defense', '*', 'team', (q) => q.eq('sport', 'mlb')), - page(sb, 'platoon_splits', '*', 'player_key', (q) => q.eq('sport', 'mlb')), - page(sb, 'statcast_aggregates', 'player_key, source_id, role, bats, throws, hard_hit_pct', 'player_key', (q) => q.eq('sport', 'mlb')), + page(sb, 'batter_spray', '*', (q) => dated(q.eq('sport', 'mlb'))), + page(sb, 'team_defense', '*', (q) => dated(q.eq('sport', 'mlb'))), + page(sb, 'platoon_splits', '*', (q) => dated(q.eq('sport', 'mlb'))), + // `statcast_aggregates` holds ONE as-of date (upserted in place), so it + // cannot answer an as-of question; the dated path reads the retained + // history instead. + asOf + ? page(sb, 'statcast_history', 'player_key, source_id, role, bats, throws, hard_hit_pct, as_of_date', + (q) => q.eq('sport', 'mlb').lte('as_of_date', asOf)) + : page(sb, 'statcast_aggregates', 'player_key, source_id, role, bats, throws, hard_hit_pct', + (q) => q.eq('sport', 'mlb')), ]); } catch (e) { // Surfaced, not silent: a load failure must be distinguishable from a feed @@ -83,9 +121,16 @@ async function build(sb, opts = {}) { const defBy = latestBy(defense, (r) => r.team, (r) => r.as_of_date); const platBy = latestBy(platoon, (r) => r.player_key, (r) => r.as_of_date); + // On the dated path several as-of rows per player are in scope, so collapse to + // the latest WITHIN the bound before indexing. On the live path there is + // exactly one row per player and this is a no-op. + const statcastLatest = asOf + ? [...latestBy(statcast, (r) => `${r.source_id}|${r.role}`, (r) => r.as_of_date).values()] + : statcast; + const batBy = new Map(); const pitBy = new Map(); - for (const r of statcast) { + for (const r of statcastLatest) { if (!r.player_key) continue; if (r.role === 'pitcher') pitBy.set(r.player_key, r); else batBy.set(r.player_key, r); @@ -98,17 +143,28 @@ async function build(sb, opts = {}) { return n > 1 ? n / 100 : n; }; - const resolver = (prop) => { + /** + * @param {object} prop + * @param {string} [sport] + * @param {object} [keys] FIX A6 — externally resolved { opponent, + * opposing_pitcher }. Omitted on the LIVE path, so live behaviour is + * byte-identical to before this argument existed; supplied only by the + * SHADOW resolve, which never reaches the served grade. + */ + const resolver = (prop, sport, keys) => { const key = nameKey(prop && prop.player); if (!key) return null; const bat = batBy.get(key); const bats = bat && bat.bats ? String(bat.bats)[0] : null; - // The opposing team and its starter, from whatever the prop carries. - const oppName = prop && (prop.opponent || prop.opp_team || null); + // The opposing team and its starter. `keys` wins when supplied; otherwise + // this reads what the prop carries, which is what it has always done — and + // what nothing ever sets (A5). + const oppName = (keys && keys.opponent) || (prop && (prop.opponent || prop.opp_team)) || null; const def = oppName ? (defBy.get(oppName) || defBy.get(String(oppName).split(' ').pop())) : null; - const pitKey = prop && prop.opposing_pitcher ? nameKey(prop.opposing_pitcher) : null; + const pitName = (keys && keys.opposing_pitcher) || (prop && prop.opposing_pitcher) || null; + const pitKey = pitName ? nameKey(pitName) : null; const pit = pitKey ? pitBy.get(pitKey) : null; const sp = platBy.get(key); @@ -124,6 +180,12 @@ async function build(sb, opts = {}) { throws: pit && pit.throws ? String(pit.throws)[0] : null, pitcherHardHit: pit ? asFraction(pit.hard_hit_pct) : null, platoonSplits: splits, + // Carried, never invented. The archetype lives on the graded row + // (`model_snapshots.archetype`), not in any context table, so an audit + // caller supplies it and this passes it through for per-archetype + // conditioning. Absent stays absent. + archetype: (prop && prop.archetype) || null, + as_of: asOf, }; // Nothing readable at all -> null, so the engine skips the factor block // entirely rather than walking an empty context. @@ -131,7 +193,11 @@ async function build(sb, opts = {}) { return anything ? ctx : null; }; + resolver.asOf = asOf; resolver.__stats = { + as_of: asOf, + mode: asOf ? 'as-of (audit)' : 'live', + statcast_source: asOf ? 'statcast_history' : 'statcast_aggregates', spray_players: sprayBy.size, defense_teams: defBy.size, platoon_players: platBy.size, diff --git a/src/services/model/lowParamService.js b/src/services/model/lowParamService.js index 7c55d88..e471b03 100644 --- a/src/services/model/lowParamService.js +++ b/src/services/model/lowParamService.js @@ -21,6 +21,7 @@ const lp = require('./lowParamCalibrator'); const cal = require('./calibration'); +const { paginate } = require('../../utils/safePaginate'); const { knownNumber } = require('../../utils/known'); const MIN_FIT = 200; @@ -74,24 +75,44 @@ function build(rows, opts = {}) { }; } +function todayEt() { + return new Intl.DateTimeFormat('en-CA', { + timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit', + }).format(new Date()); +} + +/** + * The ROW LOAD — exported so the read-integrity harness measures THE REAL + * FUNCTION rather than a restatement of its query. + * + * ── FIX A2 (2026-08-09) — THIS READ WAS 24.8% CORRUPT ──────────────────── + * This is the PRIMARY calibrator (calibrationService is only its shadow), and it + * carried the byte-identical defect A1 fixed there: an unordered `.range()` walk, + * plus `if (error || !data) break` swallowing a failed read as end-of-data. + * Measured on production: 617 of 2,490 rows returned twice, an equal number never + * returned, with `rows.length` matching the server count exactly. + * + * Both are now `safePaginate` on the unique `id`. A throw means the read failed; + * `null` from `fromLedger` still means "not enough settled history to fit". Those + * are different states and collapsing them is what hid the defect. + */ +async function loadSettledRows(sb, { sport = 'mlb', stat = 'hits', before = null } = {}) { + const cutoff = before || todayEt(); + return paginate( + () => sb.from('ledger_entries') + .select('id, p_win, outcome, game_date, quarantine_reason') + .eq('sport', sport).is('user_id', null).eq('stat', stat) + .in('outcome', ['hit', 'miss']).not('p_win', 'is', null) + .lt('game_date', cutoff), + { key: 'id', pageSize: 1000, label: `lowParamService.fromLedger(${sport}/${stat})` }, + ); +} + /** Load settled history and build, POINT-IN-TIME (strictly before today). */ async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) { if (!sb) return null; - const cutoff = before || new Intl.DateTimeFormat('en-CA', { - timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit', - }).format(new Date()); - const rows = []; - for (let from = 0; ; from += 1000) { - const { data, error } = await sb.from('ledger_entries') - .select('p_win, outcome, game_date, quarantine_reason') - .eq('sport', sport).is('user_id', null).eq('stat', stat) - .in('outcome', ['hit', 'miss']).not('p_win', 'is', null) - .lt('game_date', cutoff) - .range(from, from + 999); - if (error || !data || data.length === 0) break; - rows.push(...data); - if (data.length < 1000) break; - } + const cutoff = before || todayEt(); + const rows = await loadSettledRows(sb, { sport, stat, before: cutoff }); const clean = rows .filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')) .map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, date: String(r.game_date) })); @@ -99,4 +120,4 @@ async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ... return built ? { ...built, cutoff } : null; } -module.exports = { build, fromLedger, MIN_FIT, HOLDOUT_FRACTION }; +module.exports = { build, fromLedger, loadSettledRows, MIN_FIT, HOLDOUT_FRACTION }; diff --git a/src/services/model/matchupKeys.js b/src/services/model/matchupKeys.js new file mode 100644 index 0000000..5f537aa --- /dev/null +++ b/src/services/model/matchupKeys.js @@ -0,0 +1,144 @@ +'use strict'; + +/** + * matchupKeys — resolve the two join keys the hits factors have always needed. + * + * Spec: specs/read-integrity-harness.md §11 + * + * ── WHY THIS EXISTS ────────────────────────────────────────────────────── + * A5 measured that all three hits factors are skipped on 596/596 real props, + * because `hitsFactorContext` reads `prop.opponent` and `prop.opposing_pitcher` + * and NOTHING in the pipeline ever sets them. The matchup model is built and + * disconnected; these are the two wires. + * + * ── AT GRADE TIME, THE GAME HAS NOT HAPPENED ───────────────────────────── + * `prove-hit-factors` resolves the opponent from each hitter's own GAME LOG, + * which is correct for auditing a settled row and useless before first pitch — + * tonight's game is not in the log yet. So the grade-time source is the pair the + * schedule already publishes: + * + * player -> team `lineup_context` (as-of dated, one row per posted lineup) + * team -> matchup `getScheduleWithPitchers(gameDate)` (probable pitchers) + * + * Both are as-of-correct by construction: the lineup read is bounded by + * `as_of_date <= asOf` exactly like A4's context reads, and a probable pitcher is + * a pre-game fact. + * + * ── REFUSAL OVER GUESSING ──────────────────────────────────────────────── + * A prop whose player has no posted lineup row, or whose game has no probable + * pitcher, resolves to NULL. It does NOT fall back to "the other team in the + * prop's game_id" — a prop carries `home_team` and `away_team`, so guessing + * which side a hitter bats for would be right about half the time and wrong + * invisibly. A guessed opponent would feed the defence factor a real team's + * fielders against the wrong hitter, which is worse than not firing. + */ + +const { paginate } = require('../../utils/safePaginate'); +const { uniqueKeyFor } = require('../../utils/tableKeys'); +const { nameKey } = require('../../utils/playerName'); + +/** Keep the latest dated row per key within an as-of bound. */ +function latestBy(rows, keyFn, dateFn) { + const m = new Map(); + for (const r of rows) { + const k = keyFn(r); + if (!k) continue; + const prev = m.get(k); + if (!prev || String(dateFn(r)) > String(dateFn(prev))) m.set(k, r); + } + return m; +} + +/** Normalise a team name for matching across feeds ("Chicago Cubs" vs "Cubs"). */ +function teamForms(name) { + const s = String(name || '').trim(); + if (!s) return []; + const last = s.split(' ').pop(); + return [...new Set([s, last])]; +} + +/** + * Build the per-slate index. + * + * @param {object} deps + * - sb supabase client (for lineup_context) + * - getSchedule async (date) => [{ home:{team,probablePitcher}, away:{...} }] + * - gameDate 'YYYY-MM-DD' + * - asOf as-of bound for the lineup read (default: gameDate) + * @returns {object|null} { resolve(prop), stats } — null when nothing loaded + */ +async function build(deps = {}) { + const { sb, getSchedule, gameDate } = deps; + if (!sb || !getSchedule || !gameDate) return null; + const asOf = deps.asOf || gameDate; + + let lineups = []; + let games = []; + try { + [lineups, games] = await Promise.all([ + paginate(() => sb.from('lineup_context') + .select('as_of_date, game_date, sport, game_pk, team, side, player_key') + .eq('sport', 'mlb').eq('game_date', gameDate).lte('as_of_date', asOf), + { key: uniqueKeyFor('lineup_context'), pageSize: 1000, label: 'matchupKeys:lineup_context' }), + getSchedule(gameDate), + ]); + } catch (e) { + // Surfaced, never swallowed as an empty feed (the recurring costume). + console.warn('[matchupKeys] load FAILED (not an empty feed):', e.message); + return null; + } + + // player -> the team he is posted to bat for, as of the cutoff + const teamByPlayer = latestBy(lineups, (r) => r.player_key, (r) => r.as_of_date); + + // team -> { opponent, opposing_pitcher } + const matchup = new Map(); + for (const g of games || []) { + const h = g && g.home; const a = g && g.away; + if (!h || !a || !h.team || !a.team) continue; + const put = (side, other) => { + const sp = other.probablePitcher && other.probablePitcher.name ? other.probablePitcher.name : null; + for (const form of teamForms(side.team)) { + if (!matchup.has(form)) matchup.set(form, { opponent: other.team, opposing_pitcher: sp }); + } + }; + put(h, a); + put(a, h); + } + + const stats = { + game_date: gameDate, + as_of: asOf, + lineup_players: teamByPlayer.size, + scheduled_games: (games || []).length, + teams_with_matchup: matchup.size, + games_with_both_probables: (games || []).filter((g) => g && g.home && g.away + && g.home.probablePitcher && g.away.probablePitcher).length, + }; + + /** + * @returns {object} { opponent, opposing_pitcher, source, refused } + * Nulls are refusals — never a guess from the prop's two teams. + */ + function resolve(prop) { + const key = nameKey(prop && (prop.player || prop.player_name)); + const empty = { opponent: null, opposing_pitcher: null, source: null, refused: 'no_lineup_row' }; + if (!key) return { ...empty, refused: 'no_player' }; + const lu = teamByPlayer.get(key); + if (!lu || !lu.team) return empty; + const m = matchup.get(lu.team) || matchup.get(String(lu.team).split(' ').pop()); + if (!m) return { opponent: null, opposing_pitcher: null, source: null, refused: 'team_not_in_schedule' }; + return { + opponent: m.opponent || null, + opposing_pitcher: m.opposing_pitcher || null, + team: lu.team, + source: 'lineup_context+schedule', + refused: m.opposing_pitcher ? null : 'no_probable_pitcher', + }; + } + + resolve.stats = stats; + return resolve; +} + +module.exports = { build, teamForms, latestBy }; diff --git a/src/services/model/withdrawnVerdicts.js b/src/services/model/withdrawnVerdicts.js new file mode 100644 index 0000000..0a6da9b --- /dev/null +++ b/src/services/model/withdrawnVerdicts.js @@ -0,0 +1,122 @@ +'use strict'; + +/** + * WITHDRAWN VERDICTS — the append-only record of findings we can no longer stand behind. + * + * Spec: specs/read-integrity-harness.md §6 + * + * WHY THIS FILE EXISTS AS A FILE: there is no verdict table. + * `featureRegistry.statVerdicts` is an in-memory Map that resets every process + * (featureRegistry.js:145), and `mc_test_ledger` stores hypothesis COUNTS for the + * Bonferroni denominator, not verdicts. The standing record has always been the + * specs plus CLAUDE.md. So a withdrawal is APPENDED here, in the repo, with its + * cause — the original findings are never edited or deleted. + * + * THE LEDGER TIGHTENING ITS OWN STANDARD IS THE RECORD WE WANT. A programme that + * only ever accumulates positive findings is not measuring; it is collecting. A + * withdrawal with a measured cause is worth more than the verdict it retracts. + * + * WHAT A WITHDRAWAL IS NOT: this module records what we may CLAIM. It does not + * change what the engine SERVES. `hitsFactors.js` hardcodes its factor list and + * consults no registry, so both withdrawn factors still move the served forecast. + * Disarming or re-proving them is a separate order with its own blast radius. + * + * REINSTATEMENT: a withdrawn verdict returns only by re-running its gate through + * readers that `src/utils/readIntegrity.js` reports PASS for ON THE DAY OF THE + * RUN. Never by argument, and never by re-reading the old numbers. + */ + +const STATUS = Object.freeze({ + WITHDRAWN_PENDING_REAUDIT: 'WITHDRAWN_PENDING_REAUDIT', +}); + +/** + * The measured cause, stated once so every entry cites the identical evidence + * rather than a paraphrase that can drift. + */ +const UNORDERED_PAGINATION_CAUSE = 'drawn through an unordered-pagination reader; ' + + '16.5–24.8% set corruption measured on the identical query 2026-08-09 ' + + '(412–617 of 2,490 rows duplicated, an equal number never returned); ' + + 'not retrospectively recoverable — the query plan varied per run and was never logged'; + +/** + * APPEND ONLY. Never edit an entry; never remove one. A superseded withdrawal + * gets a new entry that references the old one. + */ +const WITHDRAWALS = Object.freeze([ + Object.freeze({ + id: 'hits/defense_by_direction/2026-08-09', + factor: 'defense_by_direction', + sport: 'mlb', + stat: 'hits', + withdrawn_at: '2026-08-09', + status: STATUS.WITHDRAWN_PENDING_REAUDIT, + original_verdict: 'PROVES', + original_evidence: Object.freeze({ + n: 528, brier_delta: -0.0034, ci: Object.freeze([-0.0059, -0.0009]), ci_level: 0.999, + source: 'Session 93', + }), + cause: UNORDERED_PAGINATION_CAUSE, + // ARITHMETIC ON THE RECORDED n, not a re-measurement. Stated as such so it is + // never quoted back as a new result. + eligibility_note: 'at the measured ~20% duplication rate the distinct n is ≈422 against ' + + 'factorGate MIN_N = 500 — this verdict was likely never eligible for adjudication at all. ' + + 'This is arithmetic on the recorded n, not a re-measurement.', + readers_implicated: Object.freeze([ + 'scripts/prove-hit-factors.js:136 (ledger_entries walk)', + 'scripts/prove-hit-factors.js:131 (model_snapshots archetype join)', + ]), + still_served: true, + still_served_note: 'hitsFactors.js:63-78 applies it unconditionally; it consults no registry', + reinstatement: 're-run the two-part gate through readers measured PASS by src/utils/readIntegrity.js', + }), + Object.freeze({ + id: 'hits/pitcher_contact_profile/2026-08-09', + factor: 'pitcher_contact_profile', + sport: 'mlb', + stat: 'hits', + withdrawn_at: '2026-08-09', + status: STATUS.WITHDRAWN_PENDING_REAUDIT, + original_verdict: 'PROVES', + original_evidence: Object.freeze({ + n: 741, brier_delta: -0.0066, ci: Object.freeze([-0.0114, -0.0016]), ci_level: 0.999, + source: 'Session 92', + }), + cause: UNORDERED_PAGINATION_CAUSE, + eligibility_note: 'n=741 clears MIN_N even after ~20% duplication (≈593 distinct), so unlike ' + + 'defense_by_direction it was plausibly eligible — but the CI that decided PROVES was ' + + 'computed on a corrupted sample against a corrupted leave-one-out baseline, so the ' + + 'interval is not recoverable either.', + readers_implicated: Object.freeze([ + 'scripts/prove-hit-factors.js:136 (ledger_entries walk)', + 'scripts/prove-hit-factors.js:131 (model_snapshots archetype join)', + ]), + still_served: true, + still_served_note: 'hitsFactors.js:80-87 applies it unconditionally; it consults no registry', + reinstatement: 're-run the two-part gate through readers measured PASS by src/utils/readIntegrity.js', + }), +]); + +/** Is this factor's verdict withdrawn for this stat? */ +function isWithdrawn(sport, stat, factor) { + return WITHDRAWALS.some((w) => w.sport === String(sport || '').toLowerCase() + && w.stat === String(stat || '').toLowerCase() + && w.factor === factor); +} + +/** The withdrawal entries for a sport+stat (newest last — the list is append-only). */ +function withdrawalsFor(sport, stat) { + const sp = String(sport || '').toLowerCase(); + const st = String(stat || '').toLowerCase(); + return WITHDRAWALS.filter((w) => w.sport === sp && (!st || w.stat === st)); +} + +/** Factors still SERVED despite a withdrawn verdict — the honesty gap, listed. */ +function servedButWithdrawn() { + return WITHDRAWALS.filter((w) => w.still_served); +} + +module.exports = { + WITHDRAWALS, STATUS, UNORDERED_PAGINATION_CAUSE, + isWithdrawn, withdrawalsFor, servedButWithdrawn, +}; diff --git a/src/services/retentionService.js b/src/services/retentionService.js index d581293..e8326e6 100644 --- a/src/services/retentionService.js +++ b/src/services/retentionService.js @@ -56,9 +56,10 @@ function etDateOf(iso) { * forecast, and never on a mixture of the two, which is the trap that would * otherwise be invisible once both generations sit in the same table. */ -const MODEL_VERSION = process.env.MODEL_VERSION || 'engine1@2026-08-07-fullwindow'; -/** Rows at or after this marker are eligible for forward re-audit. */ -const REPAIRED_CHAMPION_VERSION = 'engine1@2026-08-07-fullwindow'; +// FIX A3 — ONE SOURCE (src/config/modelVersion.js). Re-exported so existing +// importers keep working, but never re-declared: the twin default in +// ledgerService is exactly how these two drifted apart. +const { MODEL_VERSION, REPAIRED_CHAMPION_VERSION } = require('../config/modelVersion'); function codeSha() { return process.env.SOURCE_COMMIT || process.env.GIT_SHA || process.env.COOLIFY_GIT_COMMIT_SHA || null; @@ -149,6 +150,12 @@ function rowsFromSides(base, sides, ctx = {}) { // The counterfactual enabler. Absent on pre-feature refusals (the juice // gate runs before features are computed) — honestly null, never faked. features: s._features && Object.keys(s._features).length ? s._features : null, + + // FIX A5 — the RAW inputs the hits factors read, frozen at grade time so a + // re-audit reads evidence instead of reconstructing context. Inputs only: + // the multiplier is recomputable from them and is deliberately not stored. + // Null on every non-hits row and on any row with no factor context. + factor_inputs: s.factor_inputs || null, }); } return rows; diff --git a/src/services/snapshotService.js b/src/services/snapshotService.js index bad6e35..108b925 100644 --- a/src/services/snapshotService.js +++ b/src/services/snapshotService.js @@ -483,8 +483,37 @@ async function runSnapshot(sport, opts = {}) { } } + // ── FIX A6 — THE JOIN KEYS, SHADOW ONLY ───────────────────────────────── + // A5 measured all three hits factors skipped on 596/596 because nothing sets + // `prop.opponent` / `prop.opposing_pitcher`. These resolve them from the + // posted lineup + the schedule's probable pitchers, as-of-correct. They feed a + // SHADOW resolve that is FROZEN onto the row — the served forecast is NOT + // adjusted by them. Turning them on live is a separate, gated decision. + let matchupKeys = null; + if (sp === 'mlb') { + try { + const mk = deps.matchupKeys || require('./model/matchupKeys'); + const sbm = require('../utils/supabase').getSupabaseServiceClient(); + const mlbAdapter = deps.mlbAdapter || require('./adapters/mlbStatsAdapter'); + const resolver = sbm ? await mk.build({ + sb: sbm, + getSchedule: (d) => mlbAdapter.getScheduleWithPitchers(d), + gameDate: retentionGameDate, + }) : null; + if (resolver) { + matchupKeys = (prop) => resolver(prop); + console.log(`[matchup] ${sp} keys loaded — ${JSON.stringify(resolver.stats)}`); + } else { + console.log(`[matchup] ${sp} — no key index; shadow factors will not fire`); + } + } catch (e) { + console.warn('[matchup] key resolve skipped:', e.message); + } + } + await deps.gradeAndCacheSlate(sp, props, { factorContext, + matchupKeys, // Bisect hook (2026-08-01): lets the internal trigger run a bounded slate // without a prod env change, so a cap regression can be isolated by // measurement instead of guessed at. Omitted => gradeSlateService's own diff --git a/src/services/snapshotSettlementService.js b/src/services/snapshotSettlementService.js new file mode 100644 index 0000000..ae0e160 --- /dev/null +++ b/src/services/snapshotSettlementService.js @@ -0,0 +1,311 @@ +'use strict'; + +/** + * snapshotSettlementService — settle `model_snapshots` ON A SCHEDULE. + * + * Spec: specs/read-integrity-harness.md §9 + * + * ── WHY THIS EXISTS ────────────────────────────────────────────────────── + * `model_snapshots` is the retention table built for replay — it holds the + * model's INPUTS (the frozen feature vector) alongside its prediction, which + * `ledger_entries` does not. It was settled exactly once, by hand, via + * `scripts/settle-model-snapshots.js`. Nothing ever settled it on a cron. + * + * Measured 2026-08-09: 104,834 rows were past-dated and settleable; 89,350 of + * them had no outcome. Worse, of the 34,128 rows carrying the repaired + * champion's marker — the only rows a forward re-audit may be measured on — + * ZERO were settled. The accrual clock could never advance. + * + * ── OUTCOMES ONLY. NO CONTEXT RECONSTRUCTION. ──────────────────────────── + * This joins each row to the realized stat from a box score, using the row's OWN + * keys (game_date + player_key + stat + line + side). It does not fetch, infer, + * or rebuild park, weather, defence, platoon, archetype or any other context. + * A settle pass that reconstructed context would be re-deciding what the model + * saw, which is precisely what the retention table exists to prevent. + * + * ── OUTCOME IS SIDE-ALIGNED ────────────────────────────────────────────── + * `p_win` is expressed for the GRADED SIDE, so `outcome` must be too. A raw + * `realized > line` indicator is the OVER perspective and would silently invert + * the target on every under row, making calibration measure the wrong thing. + * `actual_value` stores the raw realized stat; `outcome` stores whether the + * graded side won. + * + * ── A ROW LOGGED AFTER FIRST PITCH IS NOT A PREDICTION ─────────────────── + * The pipeline runs on UTC cron hours, so a 01:00-UTC cycle is 21:00 the + * previous evening ET — same game date, three hours into the slate. Those rows + * are refused rather than settled. Preserved verbatim from the hand-run script, + * because dropping it would quietly admit post-hoc rows into every measurement. + * + * Every dependency is injectable, so the unit tests never touch a network. + */ + +const { paginate } = require('../utils/safePaginate'); +const { uniqueKeyFor } = require('../utils/tableKeys'); +const { nameKey } = require('../utils/playerName'); +const { knownNumber } = require('../utils/known'); + +const STATS = ['hits', 'total_bases', 'rbi', 'runs']; +const PAGE = 1000; +/** Eastern first pitch, conservatively. At or after this is in-game. */ +const FIRST_PITCH_ET_HOUR = 19; +/** + * Dates fetched per run — bounded, because a cron that re-fetches 26 dates of + * box scores every five hours is a quota problem pretending to be thoroughness. + * + * NEWEST FIRST, and that ordering is a correctness property, not a preference. + * The first implementation drained oldest-first and DEADLOCKED: measured on + * production, 308 rows across the eight oldest dates are structurally + * unsettleable (184 logged after first pitch, 124 with no box-score line), so + * the window re-processed the same dead dates on every run and `dates_remaining` + * never moved. Newest-first also reaches the rows that matter — the repaired + * champion's, which are the only ones a forward re-audit may be measured on. + */ +const MAX_DATES_PER_RUN = Number(process.env.SNAPSHOT_SETTLE_MAX_DATES || 8); + +/** Realized value per stat, from the box-score batting line. */ +const FIELD = Object.freeze({ + hits: (b) => knownNumber(b.hits), + total_bases: (b) => knownNumber(b.totalBases), + rbi: (b) => knownNumber(b.rbi), + runs: (b) => knownNumber(b.runs), +}); + +function isPreGame(capturedAt, gameDate) { + if (!capturedAt || !gameDate) return false; + const cap = new Date(capturedAt); + if (Number.isNaN(cap.getTime())) return false; + const et = new Date(cap.getTime() - 4 * 3600 * 1000); // EDT + const etDate = et.toISOString().slice(0, 10); + if (etDate < String(gameDate)) return true; // day before, fine + if (etDate > String(gameDate)) return false; // day after, post-game + return et.getUTCHours() < FIRST_PITCH_ET_HOUR; +} + +function todayEt(now) { + return new Intl.DateTimeFormat('en-CA', { + timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit', + }).format(now || new Date()); +} + +async function pool(items, fn, n = 6) { + const out = []; let i = 0; + await Promise.all(Array.from({ length: n }, async () => { + while (i < items.length) { + const idx = i; i += 1; + try { out[idx] = await fn(items[idx]); } catch { out[idx] = null; } + } + })); + return out.filter(Boolean); +} + +/** + * Box-score batting lines for a set of dates, keyed `${date}|${nameKey}`. + * A doubleheader gives two lines; they are SUMMED, because the prop covers the + * day rather than a game. + */ +async function battingLines(dates, deps) { + const getJson = deps.getJson; + const games = []; + for (const d of dates) { + try { + const s = await getJson(`https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}`); + for (const day of (s && s.dates) || []) { + for (const g of day.games || []) { + if (String(g.status && g.status.detailedState) === 'Final') { + games.push({ pk: g.gamePk, date: g.officialDate || d }); + } + } + } + } catch { /* absent day — stays absent */ } + } + + const lines = {}; + const loaded = await pool(games, async (g) => { + const box = await getJson(`https://statsapi.mlb.com/api/v1/game/${g.pk}/boxscore`); + const out = []; + for (const side of ['home', 'away']) { + const t = box && box.teams && box.teams[side]; + if (!t) continue; + for (const id of t.batters || []) { + const pl = t.players[`ID${id}`]; + const b = pl && pl.stats && pl.stats.batting; + if (!b || b.atBats == null) continue; // did not bat → absent, never zero + out.push({ + date: g.date, + key: nameKey(pl.person && pl.person.fullName), + hits: b.hits, totalBases: b.totalBases, rbi: b.rbi, runs: b.runs, + }); + } + } + return out; + }, deps.concurrency || 6); + + for (const arr of loaded) { + for (const r of arr) { + const k = `${r.date}|${r.key}`; + if (!lines[k]) lines[k] = { ...r }; + else { + lines[k].hits += r.hits; lines[k].totalBases += r.totalBases; + lines[k].rbi += r.rbi; lines[k].runs += r.runs; + } + } + } + return lines; +} + +/** + * PURE — decide each row's outcome from the box-score index. + * Separated so the whole decision rule is unit-testable with no I/O. + */ +function decide(snaps, lines) { + const counts = { candidates: snaps.length, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 }; + const updates = []; + const seen = new Set(); + + for (const s of snaps) { + if (seen.has(s.id)) { + const e = new Error(`INTEGRITY: duplicate snapshot id ${s.id}`); + e.code = 'DUPLICATE_ROW'; + throw e; + } + seen.add(s.id); + + if (!isPreGame(s.captured_at, s.game_date)) { + counts.post_hoc_logged += 1; counts.unresolvable += 1; continue; + } + const line = knownNumber(s.line); + if (line === null || !s.side) { counts.unresolvable += 1; continue; } + + const b = lines[`${s.game_date}|${s.player_key}`]; + if (!b) { counts.orphaned += 1; continue; } + + const realized = FIELD[s.stat] ? FIELD[s.stat](b) : null; + if (realized === null) { counts.unresolvable += 1; continue; } + + const over = realized > line; + const won = String(s.side).toLowerCase() === 'under' ? !over : over; + updates.push({ id: s.id, outcome: won ? 'hit' : 'miss', actual_value: realized }); + counts.settled += 1; + } + + // CONSERVATION — hard fail. Every candidate lands in exactly one bucket. + const acc = counts.settled + counts.unresolvable + counts.orphaned; + if (acc !== counts.candidates) { + const e = new Error(`INTEGRITY: conservation violated ${acc} != ${counts.candidates}`); + e.code = 'CONSERVATION'; + throw e; + } + return { counts, updates }; +} + +/** + * Settle one pass. + * + * @param {object} deps + * - sb supabase service client (required to do anything) + * - getJson async (url) => json + * - now () => Date + * - write default true; false = dry run + * - maxDates dates fetched this run (oldest first) + * @returns {object} { skipped?, counts, dates, written, repaired_champion_settled } + */ +async function settleSnapshots(deps = {}) { + const sb = deps.sb; + if (!sb) return { skipped: 'supabase not configured', counts: null, written: 0 }; + const now = (deps.now || (() => new Date()))(); + const cutoff = todayEt(now); + const write = deps.write !== false; + + // Unsettled, PAST-DATED rows only — a game that has not finished cannot be + // settled, and asking would produce an orphan rather than an absence. + const snaps = await paginate( + () => sb.from('model_snapshots') + .select('id, game_date, captured_at, stat, player_key, line, side, model_version') + .eq('sport', 'mlb').in('stat', STATS).is('outcome', null) + .lt('game_date', cutoff), + { key: uniqueKeyFor('model_snapshots'), pageSize: PAGE, label: 'snapshotSettlement' }, + ); + if (!snaps.length) return { counts: { candidates: 0, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 }, dates: [], written: 0, repaired_champion_settled: 0 }; + + // CHOOSE DATES FROM ROWS THAT COULD ACTUALLY SETTLE. + // + // `isPreGame` is pure, so a row logged after first pitch is known-unsettleable + // WITHOUT fetching anything. Measured on production, 19,074 such rows sit in + // the recent dates; letting them pick the window meant the same dead dates + // were re-fetched on every run and the backlog never converged. Filtering + // first is what makes the drain terminate. + const settleable = snaps.filter((r) => isPreGame(r.captured_at, r.game_date)); + // NEWEST FIRST: the eligible rows — the repaired champion's — are the newest, + // and an older date whose remainder cannot settle must not block them. + const allDates = [...new Set(settleable.map((r) => r.game_date))].sort().reverse(); + const perWindow = deps.maxDates || MAX_DATES_PER_RUN; + const maxWindows = deps.maxWindows || Number(process.env.SNAPSHOT_SETTLE_MAX_WINDOWS || 4); + + // ADVANCE PAST A DRAINED WINDOW. The newest dates keep rows that can never + // settle (a player with no box-score line never gets one), so a fixed window + // would sit on them forever while older settleable dates were never reached. + // Move to the next window when this one produces nothing, bounded so a run + // still costs a predictable number of box-score fetches. + let dates = []; let counts = null; let updates = []; let windows = 0; + for (let w = 0; w < maxWindows; w += 1) { + const slice = allDates.slice(w * perWindow, (w + 1) * perWindow); + if (!slice.length) break; + windows = w + 1; + /* eslint-disable no-await-in-loop */ + const lines = await battingLines(slice, { getJson: deps.getJson, concurrency: deps.concurrency }); + const res = decide(snaps.filter((r) => slice.includes(r.game_date)), lines); + /* eslint-enable no-await-in-loop */ + dates = slice; counts = res.counts; updates = res.updates; + if (updates.length) break; // progress — stop here + } + if (!counts) return { counts: { candidates: 0, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 }, dates: [], written: 0, repaired_champion_settled: 0 }; + const inWindow = snaps.filter((r) => dates.includes(r.game_date)); + + let written = 0; + let repairedSettled = 0; + const byId = new Map(inWindow.map((r) => [r.id, r])); + if (write && updates.length) { + const settledAt = now.toISOString(); + for (let i = 0; i < updates.length; i += 500) { + const batch = updates.slice(i, i + 500); + /* eslint-disable no-await-in-loop */ + const results = await Promise.all(batch.map((u) => sb.from('model_snapshots') + .update({ + outcome: u.outcome, + actual_value: u.actual_value, + settled_at: settledAt, + settlement_source: 'statsapi_boxscore', + }) + // IDEMPOTENT: only an unsettled row is written, so a re-run can never + // overwrite an outcome that is already on the record. + .eq('id', u.id).is('outcome', null))); + /* eslint-enable no-await-in-loop */ + results.forEach((r, k) => { + if (r.error) return; + written += 1; + const row = byId.get(batch[k].id); + if (row && row.model_version === require('../config/modelVersion').REPAIRED_CHAMPION_VERSION) { + repairedSettled += 1; + } + }); + } + } + + return { + counts, + dates, + dates_remaining: Math.max(0, allDates.length - (windows * perWindow)), + windows_scanned: windows, + // Rows that can never settle, counted rather than hidden: a row logged after + // first pitch is not a prediction and will never become one. + permanently_unsettleable: snaps.length - settleable.length, + written, + repaired_champion_settled: repairedSettled, + mode: write ? 'write' : 'dry-run', + }; +} + +module.exports = { + settleSnapshots, decide, isPreGame, battingLines, + STATS, FIELD, MAX_DATES_PER_RUN, FIRST_PITCH_ET_HOUR, +}; diff --git a/src/snapshotScheduler.js b/src/snapshotScheduler.js index c58ab04..d3edfac 100644 --- a/src/snapshotScheduler.js +++ b/src/snapshotScheduler.js @@ -65,6 +65,11 @@ function startSnapshotScheduler(opts = {}) { // Session 58 — Phase 1 truth infrastructure: settle the persistent ledger // (outcome + actual + CLV) in the same pre-grade settle pass. const settleLedgers = opts.settleAllLedgers || require('./services/ledgerService').settleAllLedgers; + // FIX A3 — model_snapshots was settled ONCE, by hand, and never on a cron. + // 89,350 settleable rows sat unsettled and ZERO of the 34,128 repaired-champion + // rows had an outcome, so the re-audit accrual clock could not advance. + const settleSnaps = opts.settleSnapshots + || require('./services/snapshotSettlementService').settleSnapshots; const notify = opts.notify || require('./utils/opsNotify').notify; const cacheGet = opts.cacheGet || require('./utils/redis').cacheGet; const cacheSet = opts.cacheSet || require('./utils/redis').cacheSet; @@ -196,6 +201,28 @@ function startSnapshotScheduler(opts = {}) { title: 'VYNDR settlement', priority: 'high', tags: ['rotating_light'], }); } + // FIX A3 — settle the RETENTION table too, on the same pass that settles the + // ledger. Outcomes only: it joins each row to a box score by the row's own + // keys and never reconstructs context. Best-effort by design — the retention + // table is a measurement asset, not the public record, so a failure here must + // never take the ledger settle or the grade with it. + try { + const sbc = require('./utils/supabase').getSupabaseServiceClient(); + const axios = require('axios'); + const snapRes = await settleSnaps({ + sb: sbc, + getJson: async (url) => (await axios.get(url, { timeout: 45_000 })).data, + now: () => d, + }); + if (snapRes && snapRes.counts) { + console.log(`[snapshots] settle pass — ${snapRes.written} settled ` + + `(${snapRes.repaired_champion_settled} repaired-champion), ` + + `${snapRes.counts.orphaned} orphaned, ${snapRes.counts.post_hoc_logged} post-hoc, ` + + `${snapRes.dates_remaining} date(s) of backlog remaining`); + } + } catch (e) { + console.warn('[snapshots] settle run failed:', e.message); + } // Session 8 — zero-settle alarm, MORNING slot only (the book-closing pass). // Signal = the ledger settle's own Postgres-backed return values (see // opsWatch.zeroSettleAlarm for why not snapshot:{sport}:previous). Deduped diff --git a/src/utils/readIntegrity.js b/src/utils/readIntegrity.js new file mode 100644 index 0000000..8d11a35 Binary files /dev/null and b/src/utils/readIntegrity.js differ diff --git a/src/utils/safePaginate.js b/src/utils/safePaginate.js new file mode 100644 index 0000000..ae811c9 --- /dev/null +++ b/src/utils/safePaginate.js @@ -0,0 +1,124 @@ +'use strict'; + +/** + * safePaginate — the ONE way this codebase walks a paginated PostgREST read. + * + * Spec: specs/read-integrity-harness.md §7 + * + * It exists because the ad-hoc walk it replaces failed in two independent ways, + * both measured on production, both of which returned a plausible-looking result: + * + * 1. NO STABLE ORDER. `.range(from, from+PAGE-1)` without an ORDER BY returned + * the correct row COUNT and the wrong ROWS — 410-617 of 2,490 duplicated on + * the hits read, with an equal number never returned at all. Postgres + * guarantees no ordering without ORDER BY, and the planner can order + * differently between successive LIMIT/OFFSET queries. + * + * 2. ERROR SWALLOWED AS END-OF-DATA. `if (error || !data || !data.length) break` + * makes a failed read indistinguishable from a finished one, so a partial + * fit looks like a complete fit. This is the shape that silently killed + * settlement on 2026-08-01. + * + * ── THE KEY MUST BE UNIQUE ──────────────────────────────────────────────── + * An ORDER BY on a non-unique column is NOT a fix: ties may be returned in any + * order, so pages still overlap. The default is `id`. Because "unique" is a + * promise the caller makes, this module VERIFIES it at runtime — a repeated key + * during the walk means either a non-unique key or a scrambled read, and either + * way the row set is not trustworthy, so it THROWS rather than returning it. + * + * ── COMPOSITE KEYS (Fix A2b) ────────────────────────────────────────────── + * `key` accepts a string OR an ordered list of columns forming a unique TUPLE. + * The context tables (`batter_spray`, `team_defense`, `platoon_splits`, …) are + * dated-composite by design — one as-of snapshot per entity per day — and have + * no single unique column, so before this they could not be made safe at all. + * Every column is ordered, in the declared order, and the uniqueness guard + * checks the FULL tuple. `src/utils/tableKeys.js` holds the real constraints, + * pulled from the live schema rather than assumed. + * + * Ordering on a PREFIX of a unique key is not enough — the trailing columns are + * exactly where the ties live — so the caller passes the whole tuple or the + * guard will catch it. + * + * ── REFUSAL IS NOT A SWALLOW ────────────────────────────────────────────── + * Throwing here lets the caller decide, explicitly, between "no data" and "the + * read failed" — the distinction the old code destroyed. A caller that responds + * by producing NOTHING (no calibrator, no map) is refusing honestly. A caller + * that responds by continuing with partial rows is repeating the defect. + * + * The paging mechanics are `readIntegrity.walk`, which is already unit-tested + * for short-page, empty-page, error-propagation and runaway behaviour. This + * module adds the stable order and the uniqueness guard on top; it does not + * hand-roll a second walk. + */ + +const { walk, keyOf, DEFAULT_PAGE } = require('./readIntegrity'); + +/** + * Walk a filtered PostgREST query to exhaustion, safely. + * + * @param {function} makeQuery () => a FRESH filtered PostgrestFilterBuilder. + * A factory, not a builder: a builder is single-use, so each page needs + * its own. + * @param {object} opts + * - key {string|string[]} unique column, or an ordered list of columns + * forming a unique tuple (default 'id') + * - pageSize {number} default 1000 + * - ascending {boolean} default true + * - maxPages {number} runaway guard + * - label {string} included in thrown messages so a failure names itself + * @returns {Promise} every row, exactly once + * @throws on a query error, a duplicate key tuple, or a runaway walk — never silently + */ +async function paginate(makeQuery, opts = {}) { + const keyCols = normalizeKey(opts.key); + const pageSize = opts.pageSize || DEFAULT_PAGE; + const ascending = opts.ascending !== false; + const label = opts.label || 'safePaginate'; + const seen = new Set(); + + return walk(async (from, to) => { + // EVERY column is ordered, in the declared order. Ordering on a prefix + // leaves the trailing columns tied, which is precisely where pages overlap. + let q = makeQuery(); + for (const col of keyCols) q = q.order(col, { ascending }); + const { data, error } = await q.range(from, to); + + // AN ERROR IS AN ERROR. Never end-of-data. + if (error) { + const e = new Error(`${label}: read failed at range ${from}-${to} — ${error.message}`); + e.code = 'READ_FAILED'; + e.cause = error; + throw e; + } + const rows = data || []; + + // UNIQUENESS, VERIFIED RATHER THAN ASSUMED. A repeat means the ordering key + // is not unique or the walk scrambled; both make the row set unusable. + for (const r of rows) { + const k = keyOf(r, keyCols); + if (seen.has(k)) { + const e = new Error(`${label}: duplicate key (${keyCols.join(',')})=${k} at range ${from}-${to} — ` + + 'the ordering key is not unique or the read scrambled; refusing to return a corrupted row set'); + e.code = 'DUPLICATE_KEY'; + throw e; + } + seen.add(k); + } + return rows; + }, pageSize, opts.maxPages); +} + +/** A string key and a composite key are the same thing, one column long. */ +function normalizeKey(key) { + if (key == null) return ['id']; + const cols = Array.isArray(key) ? key : [key]; + const clean = cols.filter((c) => typeof c === 'string' && c.length); + if (!clean.length) { + const e = new Error('safePaginate: key must be a column name or a non-empty list of column names'); + e.code = 'BAD_KEY'; + throw e; + } + return clean; +} + +module.exports = { paginate, normalizeKey, DEFAULT_PAGE }; diff --git a/src/utils/tableKeys.js b/src/utils/tableKeys.js new file mode 100644 index 0000000..c6cfbcd --- /dev/null +++ b/src/utils/tableKeys.js @@ -0,0 +1,66 @@ +'use strict'; + +/** + * tableKeys — WHAT MAKES A ROW UNIQUE, per table. + * + * Spec: specs/read-integrity-harness.md §8 + * + * `safePaginate` needs a key whose tuple is unique, because an ORDER BY that + * leaves ties lets pages overlap and the whole fix evaporates. Guessing that key + * is exactly the kind of assumption this programme keeps getting burned by, so + * every entry below was PULLED FROM THE LIVE SCHEMA on 2026-08-09: + * + * SELECT c.relname, array_agg(a.attname ORDER BY k.ord) + * FROM pg_index ix ... WHERE ix.indisunique + * + * Note that only `ledger_entries` and `model_snapshots` have a single-column + * key. Every context table is dated-composite — which is the point of those + * tables (an as-of snapshot per entity per day), and the reason the A2 rollout + * could not reach them until `safePaginate` learned composite keys. + * + * IF A MIGRATION CHANGES A CONSTRAINT, CHANGE IT HERE. A stale entry does not + * fail loudly at the database — it fails as a duplicate tuple at runtime, which + * `safePaginate` throws on. That is the intended failure: loud, not silent. + */ + +const UNIQUE_KEY = Object.freeze({ + // single-column primary keys + ledger_entries: Object.freeze(['id']), + model_snapshots: Object.freeze(['id']), + game_context: Object.freeze(['game_id']), + + // dated composite keys — one as-of snapshot per entity per day + statcast_aggregates: Object.freeze(['sport', 'season', 'source_id', 'role']), + statcast_history: Object.freeze(['as_of_date', 'sport', 'season', 'source_id', 'role']), + batter_spray: Object.freeze(['as_of_date', 'sport', 'season', 'source_id']), + team_defense: Object.freeze(['as_of_date', 'sport', 'season', 'team']), + platoon_splits: Object.freeze(['as_of_date', 'sport', 'season', 'player_key']), + park_dimensions: Object.freeze(['as_of_date', 'sport', 'venue_id']), + hitter_opportunity: Object.freeze(['as_of_date', 'sport', 'season', 'player_key']), + lineup_context: Object.freeze(['as_of_date', 'sport', 'game_pk', 'player_key']), +}); + +/** + * The unique key for a table. + * + * THROWS for an unknown table rather than defaulting to `id`. A silent default + * would order by a column that may not exist or may not be unique, which is the + * failure this module exists to prevent — and it would do so while looking fixed. + */ +function uniqueKeyFor(table) { + const k = UNIQUE_KEY[table]; + if (!k) { + const e = new Error(`tableKeys: no unique key recorded for '${table}' — ` + + 'add it from the live schema (pg_index WHERE indisunique) rather than assuming id'); + e.code = 'UNKNOWN_TABLE'; + throw e; + } + return k; +} + +/** Is this table's key a single column? (Informational; both shapes work.) */ +function isSingleKey(table) { + return uniqueKeyFor(table).length === 1; +} + +module.exports = { UNIQUE_KEY, uniqueKeyFor, isSingleKey }; diff --git a/supabase/migrations/034_factor_inputs.sql b/supabase/migrations/034_factor_inputs.sql new file mode 100644 index 0000000..8f67fa5 --- /dev/null +++ b/supabase/migrations/034_factor_inputs.sql @@ -0,0 +1,23 @@ +-- Migration 034: model_snapshots.factor_inputs — FREEZE THE FACTOR INPUTS (Fix A5). +-- +-- A re-audit that joins context tables to a past grade RECONSTRUCTS what the +-- context probably was. This column carries what the grader ACTUALLY read, so a +-- row becomes independently checkable. +-- +-- INPUTS, NOT OUTPUTS. It stores the raw values (spray shares, positional OAA, +-- pitcher hard-hit rate, platoon split counts, handedness) and NOT the resulting +-- multiplier. A stored multiplier can only be compared to itself; stored inputs +-- can be re-run through hitsFactors and checked. +-- +-- A SEPARATE COLUMN, not extra keys inside `features`: champion-ablation.js +-- iterates every key of `features` for its residual scan, so widening it would +-- silently enlarge that multiple-comparisons denominator. +-- +-- NULLABLE and populated only on `hits` rows with a factor context, so the +-- storage cost is bounded. Adding a nullable column is metadata-only in +-- Postgres — no table rewrite — which matters while this database sits over its +-- free-tier size cap. +ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS factor_inputs jsonb; + +COMMENT ON COLUMN model_snapshots.factor_inputs IS + 'Fix A5: raw inputs the hits factors read at grade time (never the multiplier). Recompute via factorFreeze.recompute().'; diff --git a/tests/unit/calibrationServiceRead.test.js b/tests/unit/calibrationServiceRead.test.js new file mode 100644 index 0000000..0db939a --- /dev/null +++ b/tests/unit/calibrationServiceRead.test.js @@ -0,0 +1,144 @@ +'use strict'; + +/** + * calibrationService.loadSettledRows — the Fix A1 regression. + * + * This reader measured 24.8% corrupt on production (617 of 2,490 rows returned + * twice, an equal number never returned) via two independent defects: an + * unordered page walk, and a query error swallowed as end-of-data. Both are + * locked here against a fake `sb`, so the regression cannot return silently. + */ + +const calSvc = require('../../src/services/model/calibrationService'); + +/** + * A fake PostgREST client for `ledger_entries`. Every filter in the reader's + * chain is accepted and RECORDED, so a dropped filter fails the test rather than + * quietly changing the row set. + */ +function fakeSb(rows, opts = {}) { + const state = { rows, filters: [], selected: null, rangeCalls: 0, ordered: [] }; + const q = { + select(cols) { state.selected = cols; return q; }, + eq(c, v) { state.filters.push(['eq', c, v]); return q; }, + is(c, v) { state.filters.push(['is', c, v]); return q; }, + in(c, v) { state.filters.push(['in', c, v]); return q; }, + not(c, o, v) { state.filters.push(['not', c, o, v]); return q; }, + lt(c, v) { state.filters.push(['lt', c, v]); return q; }, + order(c, o) { state.ordered.push([c, o && o.ascending !== false]); return q; }, + async range(from, to) { + state.rangeCalls += 1; + if (opts.errorOnRange && opts.errorOnRange(from)) { + return { data: null, error: { message: 'connection reset by peer' } }; + } + if (opts.beforeRange) opts.beforeRange(from, state); + const key = state.ordered.length ? state.ordered[state.ordered.length - 1][0] : null; + const out = key + ? [...state.rows].sort((a, b) => (a[key] < b[key] ? -1 : a[key] > b[key] ? 1 : 0)) + : [...state.rows]; + return { data: out.slice(from, to + 1), error: null }; + }, + }; + return { state, from() { return q; } }; +} + +const led = (id, p, won, date) => ({ + id, p_win: p, outcome: won ? 'hit' : 'miss', game_date: date, quarantine_reason: null, +}); +const many = (n) => Array.from({ length: n }, (_, i) => led(i + 1, 0.5 + (i % 40) / 100, i % 2 === 0, '2026-08-01')); + +describe('loadSettledRows — orders on a UNIQUE key, every page', () => { + it('returns every row exactly once across a multi-page walk', async () => { + const sb = fakeSb(many(2490)); + const rows = await calSvc.loadSettledRows(sb, { sport: 'mlb', stat: 'hits', before: '2026-08-09' }); + expect(rows).toHaveLength(2490); + expect(new Set(rows.map((r) => r.id)).size).toBe(2490); + expect(sb.state.rangeCalls).toBe(3); + }); + + it('orders by id on EVERY page, not just the first', async () => { + const sb = fakeSb(many(2490)); + await calSvc.loadSettledRows(sb, { before: '2026-08-09' }); + expect(sb.state.ordered).toHaveLength(sb.state.rangeCalls); + for (const [col, asc] of sb.state.ordered) { + expect(col).toBe('id'); // a UNIQUE key — an order on a non-unique + expect(asc).toBe(true); // column still lets ties scramble + } + }); + + it('selects id — without it the ordering key is not in the payload', async () => { + const sb = fakeSb(many(10)); + await calSvc.loadSettledRows(sb, { before: '2026-08-09' }); + expect(sb.state.selected).toMatch(/\bid\b/); + }); + + it('keeps the point-in-time cut and every identity filter', async () => { + const sb = fakeSb(many(10)); + await calSvc.loadSettledRows(sb, { sport: 'mlb', stat: 'hits', before: '2026-08-09' }); + expect(sb.state.filters).toEqual(expect.arrayContaining([ + ['eq', 'sport', 'mlb'], + ['is', 'user_id', null], + ['eq', 'stat', 'hits'], + ['in', 'outcome', ['hit', 'miss']], + ['not', 'p_win', 'is', null], + ['lt', 'game_date', '2026-08-09'], // STRICTLY before — never the answer + ])); + }); +}); + +describe('loadSettledRows — a failed read THROWS, never a short fit', () => { + it('throws when a later page errors instead of returning the earlier pages', async () => { + // The :117 defect returned 1,000 rows here and `build()` fitted a map on + // them, indistinguishable from a complete fit. + const sb = fakeSb(many(2490), { errorOnRange: (from) => from === 1000 }); + await expect(calSvc.loadSettledRows(sb, { before: '2026-08-09' })) + .rejects.toThrow(/read failed at range 1000-1999/); + }); + + it('propagates through fromLedger — a read failure is not "thin history"', async () => { + // null means "not enough settled rows to fit"; a throw means "the read + // broke". Collapsing them is what let a fragment look like a fit. + const sb = fakeSb(many(2490), { errorOnRange: (from) => from === 1000 }); + await expect(calSvc.fromLedger(sb, { sport: 'mlb', stat: 'hits', before: '2026-08-09' })) + .rejects.toThrow(/read failed/); + }); + + it('still returns null (not a throw) when there is genuinely no client', async () => { + await expect(calSvc.fromLedger(null, {})).resolves.toBeNull(); + }); +}); + +describe('loadSettledRows — concurrent append (the production condition)', () => { + it('returns every pre-existing row exactly once while the cron appends mid-walk', async () => { + // ledger_entries is written at 14/19/22/1/3 UTC. With a stable ascending + // unique key, an appended row lands after the cursor: nothing already + // returned can repeat, nothing pending can be skipped. + const original = many(2500); + let nextId = 90000; + const sb = fakeSb([...original], { + beforeRange: (from, state) => { + if (from > 0) state.rows.push(led(nextId++, 0.6, true, '2026-08-08')); + }, + }); + const rows = await calSvc.loadSettledRows(sb, { before: '2026-08-09' }); + const seen = rows.map((r) => r.id); + + expect(new Set(seen).size).toBe(seen.length); // zero duplicates + for (const o of original) expect(seen).toContain(o.id); // zero dropped + }); + + it('THROWS rather than returning a corrupted set if the walk ever repeats a row', async () => { + // Belt and braces: if the ordering guarantee were ever lost, the reader + // refuses instead of handing a duplicated sample to the fit. + const sb = fakeSb([led(1, 0.5, true, '2026-08-01'), led(1, 0.5, true, '2026-08-01')]); + await expect(calSvc.loadSettledRows(sb, { before: '2026-08-09' })) + .rejects.toThrow(/duplicate key \(id\)=1/); + }); +}); + +describe('Fix A1 does not change what is SERVED', () => { + it('the deploy set is still empty — no stat serves a calibrated number', () => { + const snapshotService = require('../../src/services/snapshotService'); + expect(snapshotService.CALIBRATION_DEPLOYED).toEqual([]); + }); +}); diff --git a/tests/unit/factorFreeze.test.js b/tests/unit/factorFreeze.test.js new file mode 100644 index 0000000..488ea7a --- /dev/null +++ b/tests/unit/factorFreeze.test.js @@ -0,0 +1,167 @@ +'use strict'; + +/** + * factorFreeze — Fix A5. + * + * Freezes the RAW inputs a factor read, so a re-audit reads evidence instead of + * reconstructing context. The multiplier is deliberately NOT stored: a stored + * multiplier can only be compared to itself. + */ + +const ff = require('../../src/services/model/factorFreeze'); +const hf = require('../../src/services/model/hitsFactors'); + +const SPRAY = { + as_of_date: '2026-08-07', player_key: 'aaron judge', + pull_gb: 0.30, straight_gb: 0.20, oppo_gb: 0.10, + pull_air: 0.20, straight_air: 0.10, oppo_air: 0.10, +}; +const OAA = { + '3B': { oaa: 2 }, SS: { oaa: 1 }, LF: { oaa: 0 }, '1B': { oaa: 0 }, + '2B': { oaa: 1 }, RF: { oaa: -1 }, CF: { oaa: 0 }, +}; +const SPLITS = { + vl: { pa: 120, atBats: 110, hits: 33 }, + vr: { pa: 300, atBats: 280, hits: 70 }, +}; +const FULL_CTX = { + spray: SPRAY, positionOaa: OAA, bats: 'R', throws: 'L', + pitcherHardHit: 0.42, platoonSplits: SPLITS, archetype: 'BOMBER', as_of: '2026-08-07', +}; +const PROP = { + player: 'Aaron Judge', stat_type: 'hits', team: 'NYY', + opponent: 'Red Sox', opposing_pitcher: 'Chris Sale', + game_id: 'mlb:2026-08-07:BOS@NYY', game_date: '2026-08-07', +}; + +describe('INPUTS, NEVER OUTPUTS', () => { + it('stores the raw inputs', () => { + const f = ff.freeze(FULL_CTX, PROP); + expect(f.spray.pull_gb).toBe(0.30); + expect(f.position_oaa).toEqual(OAA); + expect(f.bats).toBe('R'); + expect(f.throws).toBe('L'); + expect(f.pitcher_hard_hit).toBe(0.42); + expect(f.platoon.vl).toEqual({ pa: 120, ab: 110, hits: 33 }); + }); + + it('does NOT store the multiplier — it must stay recomputable, not asserted', () => { + const f = ff.freeze(FULL_CTX, PROP); + const flat = JSON.stringify(f); + expect(f.multiplier).toBeUndefined(); + expect(f.factors_fired).toBeUndefined(); + expect(flat).not.toMatch(/multiplier/); + }); + + it('carries the join keys AS THE RESOLVER SAW THEM', () => { + const f = ff.freeze(FULL_CTX, PROP); + expect(f.join).toEqual({ opponent: 'Red Sox', opposing_pitcher: 'Chris Sale' }); + }); + + it('records the join keys as NULL when the resolver had none — the finding, not an omission', () => { + // This is the live case: nothing sets prop.opponent / prop.opposing_pitcher. + const f = ff.freeze({ spray: SPRAY, positionOaa: null, bats: 'R', throws: null, pitcherHardHit: null, platoonSplits: SPLITS }, { player: 'x', team: 'NYY', game_id: 'g1' }); + expect(f.join).toEqual({ opponent: null, opposing_pitcher: null }); + expect(f.position_oaa).toBeNull(); + expect(f.throws).toBeNull(); + expect(f.pitcher_hard_hit).toBeNull(); + // and what COULD have been joined is recorded separately, never fed to a factor + expect(f.available).toEqual({ team: 'NYY', game_id: 'g1', game_date: null }); + }); + + it('is versioned, so a shape change is distinguishable from a null', () => { + expect(ff.freeze(FULL_CTX, PROP).v).toBe(ff.FREEZE_VERSION); + }); + + it('returns null for no context rather than an empty husk', () => { + expect(ff.freeze(null, PROP)).toBeNull(); + }); + + it('carries the archetype through (A4), never invents it', () => { + expect(ff.freeze(FULL_CTX, PROP).archetype).toBe('BOMBER'); + expect(ff.freeze({ ...FULL_CTX, archetype: null }, PROP).archetype).toBeNull(); + }); +}); + +describe('RECOMPUTE — the reason inputs beat an output', () => { + it('a frozen vector reproduces the live multiplier EXACTLY', () => { + const live = hf.hitsFactorMultiplier(FULL_CTX); + const fromFrozen = ff.recompute(ff.freeze(FULL_CTX, PROP)); + expect(fromFrozen.multiplier).toBe(live.multiplier); + expect(fromFrozen.factors_fired).toBe(live.factors_fired); + expect(fromFrozen.applied.map((a) => a.factor).sort()) + .toEqual(live.applied.map((a) => a.factor).sort()); + }); + + it('reproduces the live multiplier on the DEGRADED live shape too', () => { + const degraded = { spray: SPRAY, positionOaa: null, bats: 'R', throws: null, pitcherHardHit: null, platoonSplits: SPLITS }; + const live = hf.hitsFactorMultiplier(degraded); + const fromFrozen = ff.recompute(ff.freeze(degraded, {})); + expect(fromFrozen.multiplier).toBe(live.multiplier); + expect(live.factors_fired).toBe(0); // the measured live state + expect(fromFrozen.factors_fired).toBe(0); + }); + + it('a frozen vector that fires would be CAUGHT if the factor code changed', () => { + // The audit property: recompute uses the CURRENT factor code against the + // FROZEN inputs, so a code change shows up as a difference. + const frozen = ff.freeze(FULL_CTX, PROP); + const fake = { hitsFactorMultiplier: () => ({ multiplier: 1.99, applied: [], skipped: [], factors_fired: 0 }) }; + expect(ff.recompute(frozen, { hitsFactors: fake }).multiplier).toBe(1.99); + expect(ff.recompute(frozen).multiplier).not.toBe(1.99); + }); + + it('recompute(null) is null', () => { + expect(ff.recompute(null)).toBeNull(); + }); +}); + +describe('RECORDING IS NOT COMPUTING — the served grade cannot move', () => { + it('freezing does not mutate the context the factor uses', () => { + const ctx = JSON.parse(JSON.stringify(FULL_CTX)); + const before = JSON.stringify(ctx); + ff.freeze(ctx, PROP); + expect(JSON.stringify(ctx)).toBe(before); + }); + + it('the grade path writes factor_inputs but leaves pOver alone', () => { + const src = require('fs').readFileSync( + require('path').join(__dirname, '..', '..', 'src/services/intelligence/analyzeViaEngine1.js'), 'utf8'); + // Window from the A5 marker to the end of the factor block. A fixed byte + // count was brittle — A6 inserted the shadow resolve inside it and pushed + // the assignment past the window, failing a test whose subject had not moved. + const start = src.indexOf('FIX A5 — FREEZE THE INPUTS'); + const block = src.slice(start, src.indexOf('const pWin = dir ===', start)); + expect(block).toMatch(/legacy\.factor_inputs = /); + // The recording path must never assign the served probability. + expect(block).not.toMatch(/pOver\s*=[^=]/); + }); + + it('retention persists it as inputs, not as a multiplier', () => { + const src = require('fs').readFileSync( + require('path').join(__dirname, '..', '..', 'src/services/retentionService.js'), 'utf8'); + expect(src).toMatch(/factor_inputs: s\.factor_inputs \|\| null/); + expect(src).not.toMatch(/factor_multiplier:/); + }); +}); + +describe('retentionService carries the frozen vector onto the row', () => { + const ret = require('../../src/services/retentionService'); + const ctx = { snapshotId: 's1', capturedAt: '2026-08-10T14:00:00Z', sport: 'mlb', gameDate: '2026-08-10' }; + + it('a graded hits side lands with its factor_inputs', () => { + const frozen = ff.freeze(FULL_CTX, PROP); + const rows = ret.__testing + ? [] + : ret.rowsFromSides({ player: 'Aaron Judge', stat_type: 'hits', line: 0.5, game_time: '2026-08-10T23:00:00Z' }, + [{ direction: 'over', grade: 'B', p_win: 0.55, factor_inputs: frozen }], ctx); + expect(rows[0].factor_inputs).toEqual(frozen); + expect(rows[0].factor_inputs.multiplier).toBeUndefined(); + }); + + it('is null when the grader had no factor context', () => { + const rows = ret.rowsFromSides({ player: 'A', stat_type: 'rbi', line: 0.5, game_time: '2026-08-10T23:00:00Z' }, + [{ direction: 'over', grade: 'C', p_win: 0.5 }], ctx); + expect(rows[0].factor_inputs).toBeNull(); + }); +}); diff --git a/tests/unit/hitsFactorContextAsOf.test.js b/tests/unit/hitsFactorContextAsOf.test.js new file mode 100644 index 0000000..b244038 --- /dev/null +++ b/tests/unit/hitsFactorContextAsOf.test.js @@ -0,0 +1,191 @@ +'use strict'; + +/** + * hitsFactorContext as-of — Fix A4. + * + * The contamination this locks out: without a cutoff, `latestBy` returns TODAY'S + * dated row, so auditing a 2026-08-07 grade would join a profile built from + * games played after it. A clean read of the wrong date is still contamination. + * + * Fake client throughout; nothing here touches a network. + */ + +const ctxSvc = require('../../src/services/model/hitsFactorContext'); +const hf = require('../../src/services/model/hitsFactors'); +const { nameKey } = require('../../src/utils/playerName'); + +// Derived, never hardcoded: nameKey resolves nicknames ('Chris' -> 'christopher'), +// so a literal key would silently miss and look like an honest absence. +const BATTER = nameKey('Aaron Judge'); +const PITCHER = nameKey('Chris Sale'); + +/** Records every filter so an unapplied as-of bound fails the test. */ +function fakeSb(tables) { + const calls = []; + return { + calls, + from(table) { + const state = { table, lte: null, eq: {} }; + calls.push(state); + const q = { + select() { return q; }, + eq(c, v) { state.eq[c] = v; return q; }, + lte(c, v) { if (c === 'as_of_date') state.lte = v; return q; }, + order() { return q; }, + async range(from, to) { + let rows = tables[table] || []; + if (state.lte) rows = rows.filter((r) => String(r.as_of_date) <= state.lte); + return { data: rows.slice(from, to + 1), error: null }; + }, + }; + return q; + }, + }; +} + +const spray = (d, key, id) => ({ + as_of_date: d, sport: 'mlb', season: 2026, source_id: id, player_key: key, + pull_gb: 0.3, straight_gb: 0.2, oppo_gb: 0.1, pull_air: 0.2, straight_air: 0.1, oppo_air: 0.1, +}); +// positionOaa is { position: { oaa, fielders } } — not a bare number. +const defense = (d, team) => ({ + as_of_date: d, sport: 'mlb', season: 2026, team, + position_oaa: { + '3B': { oaa: 2 }, SS: { oaa: 1 }, LF: { oaa: 0 }, '1B': { oaa: 0 }, + '2B': { oaa: 1 }, RF: { oaa: -1 }, CF: { oaa: 0 }, + }, +}); +const platoon = (d, key) => ({ + as_of_date: d, sport: 'mlb', season: 2026, player_key: key, + vl_pa: 120, vl_ab: 110, vl_hits: 33, vr_pa: 300, vr_ab: 280, vr_hits: 70, +}); +const sc = (d, key, id, role, extra = {}) => ({ + as_of_date: d, sport: 'mlb', season: 2026, source_id: id, role, player_key: key, + bats: 'R', throws: 'L', hard_hit_pct: 40, ...extra, +}); + +const TABLES = { + batter_spray: [spray('2026-08-05', BATTER, 1), spray('2026-08-09', BATTER, 1)], + team_defense: [defense('2026-08-05', 'Red Sox'), defense('2026-08-09', 'Red Sox')], + platoon_splits: [platoon('2026-08-05', BATTER), platoon('2026-08-09', BATTER)], + statcast_aggregates: [sc(null, BATTER, 1, 'batter'), sc(null, PITCHER, 2, 'pitcher', { hard_hit_pct: 30 })], + statcast_history: [ + sc('2026-08-05', BATTER, 1, 'batter'), sc('2026-08-09', BATTER, 1, 'batter'), + sc('2026-08-05', PITCHER, 2, 'pitcher', { hard_hit_pct: 30 }), + sc('2026-08-09', PITCHER, 2, 'pitcher', { hard_hit_pct: 99 }), // today's value + ], +}; + +const PROP = { player: 'Aaron Judge', opponent: 'Red Sox', opposing_pitcher: 'Chris Sale' }; + +describe('as-of bound is APPLIED, and only when asked', () => { + it('LIVE (no asOf) applies no as_of filter and reads statcast_aggregates', async () => { + const sb = fakeSb(TABLES); + const r = await ctxSvc.build(sb); + expect(r.asOf).toBeNull(); + expect(r.__stats.mode).toBe('live'); + expect(r.__stats.statcast_source).toBe('statcast_aggregates'); + expect(sb.calls.every((c) => c.lte === null)).toBe(true); + expect(sb.calls.map((c) => c.table)).toContain('statcast_aggregates'); + expect(sb.calls.map((c) => c.table)).not.toContain('statcast_history'); + }); + + it('AS-OF applies lte on EVERY dated table and reads statcast_history', async () => { + const sb = fakeSb(TABLES); + const r = await ctxSvc.build(sb, { asOf: '2026-08-07' }); + expect(r.__stats.mode).toBe('as-of (audit)'); + expect(r.__stats.statcast_source).toBe('statcast_history'); + for (const t of ['batter_spray', 'team_defense', 'platoon_splits', 'statcast_history']) { + const c = sb.calls.find((x) => x.table === t); + expect(c.lte).toBe('2026-08-07'); + } + }); +}); + +describe('THE CONTAMINATION, LOCKED OUT', () => { + it('never returns a row dated after asOf', async () => { + const r = await ctxSvc.build(fakeSb(TABLES), { asOf: '2026-08-07' }); + const c = r(PROP); + expect(c.spray.as_of_date).toBe('2026-08-05'); // NOT the 08-09 row + expect(c.as_of).toBe('2026-08-07'); + }); + + it("picks up TODAY's value on the live path — proving the test can tell them apart", async () => { + const live = await ctxSvc.build(fakeSb(TABLES)); + const past = await ctxSvc.build(fakeSb(TABLES), { asOf: '2026-08-07' }); + // The pitcher's hard-hit rate moved 30 -> 99 on 08-09. + expect(live(PROP).pitcherHardHit).toBeCloseTo(0.30, 5); + expect(past(PROP).pitcherHardHit).toBeCloseTo(0.30, 5); + const today = await ctxSvc.build(fakeSb(TABLES), { asOf: '2026-08-09' }); + expect(today(PROP).pitcherHardHit).toBeCloseTo(0.99, 5); + }); + + it('takes the LATEST row within the bound, not the earliest', async () => { + const tables = { ...TABLES, batter_spray: [spray('2026-08-01', BATTER, 1), spray('2026-08-05', BATTER, 1), spray('2026-08-09', BATTER, 1)] }; + const r = await ctxSvc.build(fakeSb(tables), { asOf: '2026-08-07' }); + expect(r(PROP).spray.as_of_date).toBe('2026-08-05'); + }); +}); + +describe('REFUSAL OVER RECONSTRUCTION', () => { + it('an entity with NO row at-or-before asOf yields null — never the nearest row', async () => { + // Every dated row is AFTER the cutoff. The honest answer is "no context". + const r = await ctxSvc.build(fakeSb(TABLES), { asOf: '2026-07-01' }); + expect(r).toBeNull(); // nothing loaded at all → serve unadjusted + }); + + it('refuses ONE entity while still serving another', async () => { + const tables = { + ...TABLES, + batter_spray: [spray('2026-08-09', BATTER, 1)], // only AFTER the cutoff + platoon_splits: [platoon('2026-08-05', BATTER)], // present at the cutoff + }; + const r = await ctxSvc.build(fakeSb(tables), { asOf: '2026-08-07' }); + const c = r(PROP); + expect(c.spray).toBeNull(); // refused + expect(c.platoonSplits).not.toBeNull(); + }); + + it('a refused input leaves the BASE RATE UNTOUCHED (factor does not apply)', () => { + // The end-to-end consequence: no spray → no defense_by_direction multiplier. + const withSpray = hf.hitsFactorMultiplier({ + spray: TABLES.batter_spray[0], positionOaa: TABLES.team_defense[0].position_oaa, bats: 'R', + }); + const refused = hf.hitsFactorMultiplier({ + spray: null, positionOaa: TABLES.team_defense[0].position_oaa, bats: 'R', + }); + expect(refused.applied.map((a) => a.factor)).not.toContain('defense_by_direction'); + expect(refused.skipped.map((s) => s.factor)).toContain('defense_by_direction'); + // and with nothing else readable the multiplier is exactly 1 — a no-op. + expect(refused.multiplier).toBe(1); + expect(withSpray.factors_fired).toBeGreaterThan(refused.factors_fired); + }); + + it('adjustProbability is unchanged when every input is refused', () => { + const out = hf.adjustProbability(0.42, { spray: null, positionOaa: null, bats: null }); + expect(out.p_adjusted).toBe(0.42); + }); +}); + +describe('archetype is CARRIED, never invented', () => { + it('passes the graded row archetype through', async () => { + const r = await ctxSvc.build(fakeSb(TABLES), { asOf: '2026-08-07' }); + expect(r({ ...PROP, archetype: 'BOMBER' }).archetype).toBe('BOMBER'); + }); + + it('is null when the caller supplies none — no context table holds it', async () => { + const r = await ctxSvc.build(fakeSb(TABLES), { asOf: '2026-08-07' }); + expect(r(PROP).archetype).toBeNull(); + }); +}); + +describe('the context read stays on safePaginate with the real composite key', () => { + it('orders by the full unique tuple of each table', () => { + const src = require('fs').readFileSync( + require('path').join(__dirname, '..', '..', 'src/services/model/hitsFactorContext.js'), 'utf8'); + expect(src).toMatch(/uniqueKeyFor\(table\)/); + expect(src).toMatch(/require\('\.\.\/\.\.\/utils\/safePaginate'\)/); + // No hand-rolled walk left behind. + expect(src).not.toMatch(/for \(let from = 0/); + }); +}); diff --git a/tests/unit/matchupKeysShadow.test.js b/tests/unit/matchupKeysShadow.test.js new file mode 100644 index 0000000..870a74f --- /dev/null +++ b/tests/unit/matchupKeysShadow.test.js @@ -0,0 +1,172 @@ +'use strict'; + +/** + * A6 — the join keys, resolved into a SHADOW resolve. + * + * A5 measured all three hits factors skipped on 596/596 because nothing sets + * `prop.opponent` / `prop.opposing_pitcher`. A6 resolves them and freezes what + * the factors WOULD do. The served grade must not move. + */ + +const mk = require('../../src/services/model/matchupKeys'); +const ctxSvc = require('../../src/services/model/hitsFactorContext'); +const hf = require('../../src/services/model/hitsFactors'); +const ff = require('../../src/services/model/factorFreeze'); +const { nameKey } = require('../../src/utils/playerName'); + +const JUDGE = nameKey('Aaron Judge'); + +const fakeSb = (rows) => ({ + from() { + const st = { lte: null }; + const q = { + select() { return q; }, eq() { return q; }, + lte(c, v) { if (c === 'as_of_date') st.lte = v; return q; }, + order() { return q; }, + async range(f, t2) { + const r = st.lte ? rows.filter((x) => String(x.as_of_date) <= st.lte) : rows; + return { data: r.slice(f, t2 + 1), error: null }; + }, + }; + return q; + }, +}); + +const LINEUP = [{ + as_of_date: '2026-08-09', game_date: '2026-08-09', sport: 'mlb', + game_pk: 1, team: 'New York Yankees', side: 'home', player_key: JUDGE, +}]; +const SCHED = [{ + home: { team: 'New York Yankees', probablePitcher: { id: 1, name: 'Gerrit Cole' } }, + away: { team: 'Boston Red Sox', probablePitcher: { id: 2, name: 'Chris Sale' } }, +}]; + +describe('key resolution', () => { + it('resolves opponent + opposing pitcher from lineup + schedule', async () => { + const r = await mk.build({ sb: fakeSb(LINEUP), getSchedule: async () => SCHED, gameDate: '2026-08-09' }); + const k = r({ player: 'Aaron Judge' }); + expect(k.opponent).toBe('Boston Red Sox'); + expect(k.opposing_pitcher).toBe('Chris Sale'); + expect(k.source).toBe('lineup_context+schedule'); + expect(k.refused).toBeNull(); + }); + + it('REFUSES a player with no posted lineup row — never guesses from the prop teams', async () => { + const r = await mk.build({ sb: fakeSb(LINEUP), getSchedule: async () => SCHED, gameDate: '2026-08-09' }); + // The prop knows both teams; the resolver still refuses, because which side + // the hitter bats for is not knowable from that. + const k = r({ player: 'Nobody Atall', home_team: 'NYY', away_team: 'BOS' }); + expect(k.opponent).toBeNull(); + expect(k.opposing_pitcher).toBeNull(); + expect(k.refused).toBe('no_lineup_row'); + }); + + it('refuses the pitcher when the game has no probable, keeping the opponent', async () => { + const noSP = [{ home: { team: 'New York Yankees' }, away: { team: 'Boston Red Sox' } }]; + const r = await mk.build({ sb: fakeSb(LINEUP), getSchedule: async () => noSP, gameDate: '2026-08-09' }); + const k = r({ player: 'Aaron Judge' }); + expect(k.opponent).toBe('Boston Red Sox'); + expect(k.opposing_pitcher).toBeNull(); + expect(k.refused).toBe('no_probable_pitcher'); + }); + + it('bounds the lineup read by as-of (A4 discipline)', async () => { + const future = [{ ...LINEUP[0], as_of_date: '2026-08-20' }]; + const r = await mk.build({ sb: fakeSb(future), getSchedule: async () => SCHED, gameDate: '2026-08-09', asOf: '2026-08-09' }); + expect(r({ player: 'Aaron Judge' }).refused).toBe('no_lineup_row'); + }); + + it('returns null (not a stub) when nothing loads', async () => { + expect(await mk.build({})).toBeNull(); + }); +}); + +describe('the resolver: LIVE unchanged, SHADOW fires', () => { + const TABLES = { + batter_spray: [{ as_of_date: '2026-08-09', sport: 'mlb', season: 2026, source_id: 1, player_key: JUDGE, pull_gb: 0.3, straight_gb: 0.2, oppo_gb: 0.1, pull_air: 0.2, straight_air: 0.1, oppo_air: 0.1 }], + team_defense: [{ as_of_date: '2026-08-09', sport: 'mlb', season: 2026, team: 'Boston Red Sox', position_oaa: { '3B': { oaa: 2 }, SS: { oaa: 1 }, LF: { oaa: 0 }, '1B': { oaa: 0 }, '2B': { oaa: 1 }, RF: { oaa: -1 }, CF: { oaa: 0 } } }], + platoon_splits: [{ as_of_date: '2026-08-09', sport: 'mlb', season: 2026, player_key: JUDGE, vl_pa: 120, vl_ab: 110, vl_hits: 33, vr_pa: 300, vr_ab: 280, vr_hits: 70 }], + statcast_aggregates: [ + { sport: 'mlb', season: 2026, source_id: 1, role: 'batter', player_key: JUDGE, bats: 'R', hard_hit_pct: 40 }, + { sport: 'mlb', season: 2026, source_id: 2, role: 'pitcher', player_key: nameKey('Chris Sale'), throws: 'L', hard_hit_pct: 45 }, + ], + }; + const sbCtx = () => ({ + from(table) { + const q = { select() { return q; }, eq() { return q; }, lte() { return q; }, order() { return q; }, + async range(f, t2) { return { data: (TABLES[table] || []).slice(f, t2 + 1), error: null }; } }; + return q; + }, + }); + const PROP = { player: 'Aaron Judge', stat_type: 'hits' }; + const KEYS = { opponent: 'Boston Red Sox', opposing_pitcher: 'Chris Sale', source: 'lineup_context+schedule', refused: null }; + + it('WITHOUT keys nothing fires — the measured live state', async () => { + const ctx = await ctxSvc.build(sbCtx()); + const c = ctx(PROP, 'mlb'); + const m = hf.hitsFactorMultiplier(c); + expect(c.positionOaa).toBeNull(); + expect(c.throws).toBeNull(); + expect(m.factors_fired).toBe(0); + expect(m.multiplier).toBe(1); + }); + + it('WITH keys the factors fire — same indexes, different join', async () => { + const ctx = await ctxSvc.build(sbCtx()); + const c = ctx(PROP, 'mlb', KEYS); + const m = hf.hitsFactorMultiplier(c); + expect(c.positionOaa).not.toBeNull(); + expect(c.throws).toBe('L'); + expect(m.factors_fired).toBeGreaterThan(0); + expect(m.multiplier).not.toBe(1); + }); + + it('the shadow block records the would-fire result and stays recomputable', async () => { + const ctx = await ctxSvc.build(sbCtx()); + const live = ctx(PROP, 'mlb'); + const shadowCtx = ctx(PROP, 'mlb', KEYS); + const shM = hf.hitsFactorMultiplier(shadowCtx); + const frozen = ff.freeze(live, PROP, ff.shadowBlock(shadowCtx, shM, KEYS)); + + expect(frozen.would_fire.multiplier).toBe(shM.multiplier); + expect(frozen.would_fire.factors_fired).toBe(shM.factors_fired); + expect(frozen.would_fire.keys.opponent).toBe('Boston Red Sox'); + // Recomputable from what was stored beside it. + const re = hf.hitsFactorMultiplier({ + spray: live.spray, bats: live.bats, platoonSplits: live.platoonSplits, + positionOaa: frozen.would_fire.shadow_inputs.position_oaa, + throws: frozen.would_fire.shadow_inputs.throws, + pitcherHardHit: frozen.would_fire.shadow_inputs.pitcher_hard_hit, + }); + expect(re.multiplier).toBe(frozen.would_fire.multiplier); + }); + + it('the LIVE frozen inputs still show the refusal — shadow does not overwrite them', async () => { + const ctx = await ctxSvc.build(sbCtx()); + const live = ctx(PROP, 'mlb'); + const shadowCtx = ctx(PROP, 'mlb', KEYS); + const frozen = ff.freeze(live, PROP, ff.shadowBlock(shadowCtx, hf.hitsFactorMultiplier(shadowCtx), KEYS)); + expect(frozen.position_oaa).toBeNull(); // what the SERVED grade saw + expect(frozen.throws).toBeNull(); + expect(frozen.would_fire.shadow_inputs.throws).toBe('L'); // what it WOULD have seen + }); + + it('no shadow block when the keys were refused', async () => { + const ctx = await ctxSvc.build(sbCtx()); + const live = ctx(PROP, 'mlb'); + expect(ff.freeze(live, PROP, null).would_fire).toBeNull(); + }); +}); + +describe('the served grade cannot read the shadow', () => { + it('the engine assigns pOver only from the LIVE adjust, never from would_fire', () => { + const src = require('fs').readFileSync( + require('path').join(__dirname, '..', '..', 'src/services/intelligence/analyzeViaEngine1.js'), 'utf8'); + const i = src.indexOf('FIX A6 — SHADOW RESOLVE'); + const block = src.slice(i, i + 900); + expect(block).toMatch(/shadow = ffz\.shadowBlock/); + expect(block).not.toMatch(/pOver\s*=/); + // and the only pOver assignment in the factor area is gated on the LIVE adj + expect(src).toMatch(/if \(adj\.p_adjusted != null && adj\.factors_fired > 0\) \{\s*\n\s*pOver = adj\.p_adjusted;/); + }); +}); diff --git a/tests/unit/readIntegrity.test.js b/tests/unit/readIntegrity.test.js new file mode 100644 index 0000000..97611d4 --- /dev/null +++ b/tests/unit/readIntegrity.test.js @@ -0,0 +1,203 @@ +'use strict'; + +/** + * Read-integrity harness — spec: specs/read-integrity-harness.md §5. + * Every case uses a FAKE page fetcher; nothing here touches a network. + */ + +const ri = require('../../src/utils/readIntegrity'); + +/** A fetcher that serves a fixed list of pages, then empties. */ +const pagesOf = (pages) => async (from, to) => { + const size = to - from + 1; + const idx = Math.floor(from / size); + return pages[idx] || []; +}; + +const row = (id) => ({ id }); +const ids = (a, b) => Array.from({ length: b - a + 1 }, (_, i) => row(a + i)); + +describe('walk', () => { + it('stops on a short page', async () => { + const rows = await ri.walk(pagesOf([ids(1, 10), ids(11, 15)]), 10); + expect(rows).toHaveLength(15); + }); + + it('stops on an empty page after a full one', async () => { + const rows = await ri.walk(pagesOf([ids(1, 10), []]), 10); + expect(rows).toHaveLength(10); + }); + + it('keeps duplicates — collapsing them here would erase the measurement', async () => { + const rows = await ri.walk(pagesOf([ids(1, 10), [row(3), row(4)]]), 10); + expect(rows).toHaveLength(12); + }); + + it('PROPAGATES a fetch error instead of treating it as end-of-data', async () => { + // The calibrationService.fromLedger:117 defect: `if (error || !data) break` + // makes a partial read indistinguishable from a complete one. + const boom = async () => { throw new Error('fetch failed'); }; + await expect(ri.walk(boom, 10)).rejects.toThrow('fetch failed'); + }); + + it('refuses to report a runaway walk as complete', async () => { + const never = async () => ids(1, 10); // always a full page + await expect(ri.walk(never, 10, 3)).rejects.toThrow(/maxPages/); + }); +}); + +describe('analyze — the verdict is measured, never inferred', () => { + it('PASSes a walk that is set-identical to the control', () => { + const r = ri.analyze({ exactCount: 15, unordered: ids(1, 15), ordered: ids(1, 15) }); + expect(r.verdict).toBe(ri.VERDICT.PASS); + expect(r.unordered.duplicates).toBe(0); + expect(r.corruption_pct).toBe(0); + }); + + it('FAILs the measured production shape: last page re-emits earlier rows', () => { + // 2,490-row read reproduced in miniature: the tail page repeats rows already + // returned, so an equal number of real rows are never seen. + const unordered = [...ids(1, 10), ...ids(11, 20), ...[row(3), row(4), row(5)]]; + const ordered = ids(1, 23); + const r = ri.analyze({ exactCount: 23, unordered, ordered }); + expect(r.verdict).toBe(ri.VERDICT.FAIL); + expect(r.unordered.rows).toBe(23); // the count looks perfect + expect(r.unordered.distinct).toBe(20); // the rows are not + expect(r.unordered.duplicates).toBe(3); + expect(r.unordered.missing_vs_control).toBe(3); + expect(r.corruption_pct).toBeCloseTo(13.0, 1); + }); + + it('META-SCAR GUARD: an ORDERED walk that still duplicates is FAIL, not PASS', () => { + // The clause is not the result. A harness that credited the presence of an + // ORDER BY would be verifying its own intention. + const orderedButBroken = [...ids(1, 20), row(20)]; + const r = ri.analyze({ exactCount: 21, unordered: orderedButBroken, ordered: ids(1, 21) }); + expect(r.verdict).toBe(ri.VERDICT.FAIL); + expect(r.unordered.duplicates).toBe(1); + }); + + it('reports CONTROL_INVALID when the control is short of the server count', () => { + const r = ri.analyze({ exactCount: 30, unordered: ids(1, 25), ordered: ids(1, 25) }); + expect(r.verdict).toBe(ri.VERDICT.CONTROL_INVALID); + expect(r.control_valid).toBe(false); + }); + + it('never issues PASS without a server count to validate against', () => { + const r = ri.analyze({ exactCount: null, unordered: ids(1, 5), ordered: ids(1, 5) }); + expect(r.verdict).not.toBe(ri.VERDICT.PASS); + expect(r.verdict).toBe(ri.VERDICT.CONTROL_INVALID); + }); + + it('reports KEY_NOT_UNIQUE rather than blaming the walk for duplicate data', () => { + // batter_spray keyed on player_key|as_of_date has genuine duplicate keys. + const dup = [{ k: 'a' }, { k: 'a' }, { k: 'b' }]; + const r = ri.analyze({ exactCount: 3, unordered: dup, ordered: dup, keyCols: ['k'] }); + expect(r.verdict).toBe(ri.VERDICT.KEY_NOT_UNIQUE); + expect(r.key_unique).toBe(false); + }); + + it('counts a composite key across columns', () => { + const rows = [{ a: 1, b: 'x' }, { a: 1, b: 'y' }]; + const r = ri.analyze({ exactCount: 2, unordered: rows, ordered: rows, keyCols: ['a', 'b'] }); + expect(r.verdict).toBe(ri.VERDICT.PASS); + }); + + it('distinguishes an EXTRA row from a missing one', () => { + const r = ri.analyze({ exactCount: 5, unordered: [...ids(1, 5), row(99)], ordered: ids(1, 5) }); + expect(r.unordered.extra_vs_control).toBe(1); + expect(r.unordered.missing_vs_control).toBe(0); + expect(r.verdict).toBe(ri.VERDICT.FAIL); + }); +}); + +describe('applyFilters', () => { + const fake = () => { + const calls = []; + const q = { + calls, + eq: (...a) => { calls.push(['eq', ...a]); return q; }, + is: (...a) => { calls.push(['is', ...a]); return q; }, + in: (...a) => { calls.push(['in', ...a]); return q; }, + not: (...a) => { calls.push(['not', ...a]); return q; }, + lt: (...a) => { calls.push(['lt', ...a]); return q; }, + gte: (...a) => { calls.push(['gte', ...a]); return q; }, + }; + return q; + }; + + it('chains ops in order', () => { + const q = ri.applyFilters(fake(), [ + ['eq', 'sport', 'mlb'], + ['is', 'user_id', null], + ['in', 'outcome', ['hit', 'miss']], + ['not', 'p_win', 'is', null], + ['lt', 'game_date', '2026-08-09'], + ]); + expect(q.calls).toEqual([ + ['eq', 'sport', 'mlb'], + ['is', 'user_id', null], + ['in', 'outcome', ['hit', 'miss']], + ['not', 'p_win', 'is', null], + ['lt', 'game_date', '2026-08-09'], + ]); + }); + + it('THROWS on an unknown op rather than skipping it — a dropped filter changes the row set', () => { + expect(() => ri.applyFilters(fake(), [['contains', 'x', 'y']])).toThrow(/unsupported filter op/); + }); + + it('handles an empty filter list', () => { + expect(ri.applyFilters(fake(), []).calls).toEqual([]); + }); +}); + +describe('measure', () => { + const spec = { id: 'r1', source: 'f.js:1', table: 't', key: ['id'] }; + + it('assembles a full result for a clean reader', async () => { + const r = await ri.measure(spec, { + pageSize: 10, + exactCount: async () => 15, + fetchUnordered: pagesOf([ids(1, 10), ids(11, 15)]), + fetchOrdered: pagesOf([ids(1, 10), ids(11, 15)]), + }); + expect(r.verdict).toBe(ri.VERDICT.PASS); + expect(r.pages).toBe(2); + expect(r.id).toBe('r1'); + }); + + it('reports ERROR (not PASS, not empty) when a page fetch throws', async () => { + const r = await ri.measure(spec, { + pageSize: 10, + exactCount: async () => 15, + fetchUnordered: async () => { throw new Error('boom'); }, + fetchOrdered: pagesOf([ids(1, 15)]), + }); + expect(r.verdict).toBe(ri.VERDICT.ERROR); + expect(r.reason).toMatch(/boom/); + }); + + it('reports ERROR when the count itself fails', async () => { + const r = await ri.measure(spec, { + pageSize: 10, + exactCount: async () => { throw new Error('count failed'); }, + fetchUnordered: pagesOf([ids(1, 5)]), + fetchOrdered: pagesOf([ids(1, 5)]), + }); + expect(r.verdict).toBe(ri.VERDICT.ERROR); + }); +}); + +describe('the harness writes nothing', () => { + it('has no insert/update/upsert/delete anywhere in its source', () => { + const fs = require('fs'); + const path = require('path'); + for (const f of ['../../src/utils/readIntegrity.js', '../../scripts/read-integrity.js']) { + const p = path.join(__dirname, f); + if (!fs.existsSync(p)) continue; + const src = fs.readFileSync(p, 'utf8'); + expect(src).not.toMatch(/\.insert\(|\.update\(|\.upsert\(|\.delete\(|\.rpc\(/); + } + }); +}); diff --git a/tests/unit/readerPagination.test.js b/tests/unit/readerPagination.test.js new file mode 100644 index 0000000..a9314e4 --- /dev/null +++ b/tests/unit/readerPagination.test.js @@ -0,0 +1,139 @@ +'use strict'; + +/** + * Fix A2 — the learning readers do not regress to the unordered walk. + * + * WHAT THIS TEST IS AND IS NOT. It does NOT assert that a reader is correct — + * correctness is measured by `scripts/read-integrity.js` against an ordered + * control on live data, and a source grep can never stand in for that (the + * meta-scar rule, specs/read-integrity-harness.md §2). + * + * What it DOES assert is structural: the two tables measured corrupt + * (`ledger_entries`, `model_snapshots`) are never read through the legacy + * `page()` helper again, the scripts stay importable without executing, and the + * named READS the harness measures still exist. Those are the properties that, + * if they silently regressed, would make the next harness run measure the wrong + * thing. + */ + +const fs = require('fs'); +const path = require('path'); + +const SCRIPTS = [ + 'prove-hit-factors', 'prove-tb-factors', 'cluster-prove', + 'tb-solo-and-interactions', 'build-grade-bands', 'proven-status', + 'champion-ablation', + // A2b — the context readers and the two write-scripts + 'prove-runs-rbi', 'pitcher-prove-k', 'prove-park-weather', 'skill-v1-stagea', + 'stagea-gate-run', 'backfill-context', 'reconstruct-game-environment', + 'calibrate-hits', +]; + +/** Every table these scripts page through must have a recorded unique key. */ +const KEYED_TABLES = ['ledger_entries', 'model_snapshots', 'statcast_aggregates', + 'batter_spray', 'team_defense', 'platoon_splits', 'park_dimensions', + 'hitter_opportunity', 'lineup_context', 'game_context']; + +const CORRUPTED_TABLES = ['ledger_entries', 'model_snapshots']; + +/** + * READS DELIBERATELY LEFT ON THE LEGACY WALK — named, with the reason. + * + * The A2 order scoped the fix to the eleven readers measured FAIL and said + * explicitly to leave the siblings that measured 0%. Those siblings are listed + * here rather than excluded by a loosened regex, so the gap is VISIBLE and + * tracked instead of invisible. + * + * They are not safe by construction — they are clean by plan. The same query on + * `ledger_entries` measured 16.5% one minute and 24.8% the next, so a 0% reading + * is a fact about today, not a property. These should move in a follow-up. + */ +const KNOWN_LEGACY_READS = { + // A2b closed the last exceptions: the three reads A2 left on the legacy walk + // are converted, and the unordered helper is DELETED from every script rather + // than parked beside the safe one — a dead broken helper is an invitation. +}; + +const read = (name) => fs.readFileSync(path.join(__dirname, '..', '..', 'scripts', `${name}.js`), 'utf8'); + +describe.each(SCRIPTS)('scripts/%s.js', (name) => { + const src = read(name); + const allowed = KNOWN_LEGACY_READS[name] || []; + + it('never reads a keyed table through an unordered walk', () => { + for (const table of KEYED_TABLES) { + // `page(sb, 'ledger_entries'` — the unordered walk. `pageSafe(...)` is fine. + expect(src).not.toMatch(new RegExp(`[^e]page\\(\\s*sb,\\s*'${table}'`)); + } + expect(allowed).toHaveLength(0); + }); + + it('carries NO unordered page() helper at all — the landmine is deleted', () => { + const code = src.replace(/\/\/.*/g, '').replace(/\/\*[\s\S]*?\*\//g, ''); + const defs = code.match(/async function page\(/g) || []; + for (const _ of defs) { + // A surviving helper must itself order; an unordered one must not exist. + expect(code).toMatch(/paginate\(/); + } + }); + + it('looks its key up from the schema rather than assuming id', () => { + expect(src).toMatch(/uniqueKeyFor/); + }); + + it('routes its fixed reads through safePaginate, not a hand-rolled walk', () => { + expect(src).toMatch(/require\('\.\.\/src\/utils\/safePaginate'\)/); + expect(src).toMatch(/async function pageSafe\(/); + // pageSafe must delegate — a second walk implementation is the thing the + // A2 order explicitly forbade. + expect(src).toMatch(/return paginate\(/); + }); + + it('exports the named READS the harness measures', () => { + const mod = require(path.join('..', '..', 'scripts', name)); // must not execute + expect(mod.READS).toBeDefined(); + for (const fn of Object.values(mod.READS)) expect(typeof fn).toBe('function'); + }); + + it('does NOT run main() on require — importing it must not hit the network or write', () => { + // prove-hit-factors and friends write to mc_test_ledger via + // tl.recordAndCount, which is the Bonferroni denominator. An unguarded + // entrypoint would let the harness inflate the correction just by measuring. + expect(src).toMatch(/if \(require\.main === module\)/); + expect(src).not.toMatch(/^main\(\)\.catch/m); + }); +}); + +describe('the calibration services load rows through safePaginate', () => { + it.each([ + ['calibrationService', '../../src/services/model/calibrationService'], + ['lowParamService', '../../src/services/model/lowParamService'], + ])('%s exports loadSettledRows', (label, modPath) => { + const mod = require(modPath); + expect(typeof mod.loadSettledRows).toBe('function'); + }); + + it.each([ + ['calibrationService', 'src/services/model/calibrationService.js'], + ['lowParamService', 'src/services/model/lowParamService.js'], + ])('%s no longer swallows a query error as end-of-data', (label, rel) => { + const src = fs.readFileSync(path.join(__dirname, '..', '..', rel), 'utf8'); + // The :117 defect, in CODE. Anchored to line-start so the comment that + // documents the defect does not count as the defect. + expect(src).not.toMatch(/^\s*if \(error \|\| !data/m); + expect(src).toMatch(/paginate\(/); + }); +}); + +describe('A2 did not open anything it was not meant to', () => { + it('CALIBRATION_DEPLOYED is still empty — no stat serves a calibrated number', () => { + const snapshotService = require('../../src/services/snapshotService'); + expect(snapshotService.CALIBRATION_DEPLOYED).toEqual([]); + }); + + it('both hits verdicts remain withdrawn', () => { + const wv = require('../../src/services/model/withdrawnVerdicts'); + expect(wv.isWithdrawn('mlb', 'hits', 'defense_by_direction')).toBe(true); + expect(wv.isWithdrawn('mlb', 'hits', 'pitcher_contact_profile')).toBe(true); + }); +}); diff --git a/tests/unit/safePaginate.test.js b/tests/unit/safePaginate.test.js new file mode 100644 index 0000000..bc9e694 --- /dev/null +++ b/tests/unit/safePaginate.test.js @@ -0,0 +1,294 @@ +'use strict'; + +/** + * safePaginate — spec: specs/read-integrity-harness.md §7. + * Fake query builders throughout; nothing here touches a network. + */ + +const { paginate } = require('../../src/utils/safePaginate'); + +/** + * A fake PostgREST table. `rows` may be mutated between pages to simulate + * concurrent writes. Honours order+range the way Postgres does WITH an ORDER BY. + */ +function fakeTable(rows, opts = {}) { + const state = { rows, orderCalls: 0, rangeCalls: 0 }; + state.make = () => { + let orderKey = null; let asc = true; + const q = { + order(k, o) { orderKey = k; asc = !o || o.ascending !== false; state.orderCalls += 1; return q; }, + async range(from, to) { + state.rangeCalls += 1; + if (opts.errorOnRange && opts.errorOnRange(from)) return { data: null, error: { message: 'connection reset' } }; + if (opts.beforeRange) opts.beforeRange(from, state); + // Unordered simulation: return physical order (the production defect). + let out = [...state.rows]; + if (orderKey) { + out.sort((a, b) => (a[orderKey] < b[orderKey] ? -1 : a[orderKey] > b[orderKey] ? 1 : 0)); + if (!asc) out.reverse(); + } + return { data: out.slice(from, to + 1), error: null }; + }, + }; + return q; + }; + return state; +} + +const row = (id) => ({ id, v: `r${id}` }); +const ids = (a, b) => Array.from({ length: b - a + 1 }, (_, i) => row(a + i)); + +describe('safePaginate — (a) returns the full set, in order', () => { + it('walks every page and returns each row exactly once, ascending by key', async () => { + const t = fakeTable(ids(1, 2490)); + const out = await paginate(t.make, { pageSize: 1000 }); + expect(out).toHaveLength(2490); + expect(new Set(out.map((r) => r.id)).size).toBe(2490); + expect(out.map((r) => r.id)).toEqual(ids(1, 2490).map((r) => r.id)); // in order + expect(t.rangeCalls).toBe(3); + }); + + it('always applies the order — every page is ordered, not just the first', async () => { + const t = fakeTable(ids(1, 2490)); + await paginate(t.make, { pageSize: 1000 }); + expect(t.orderCalls).toBe(t.rangeCalls); + }); + + it('orders by the configured key, and can descend', async () => { + const t = fakeTable(ids(1, 5)); + const out = await paginate(t.make, { pageSize: 10, ascending: false }); + expect(out.map((r) => r.id)).toEqual([5, 4, 3, 2, 1]); + }); + + it('returns [] for an empty table without throwing', async () => { + const out = await paginate(fakeTable([]).make, { pageSize: 10 }); + expect(out).toEqual([]); + }); +}); + +describe('safePaginate — (b) a mid-walk error THROWS, never a silent end', () => { + it('throws when page 2 fails instead of returning page 1 as the whole set', async () => { + // THE :117 DEFECT: `if (error || !data) break` would return 1,000 rows here + // and the caller would fit a map on them believing it had 2,490. + const t = fakeTable(ids(1, 2490), { errorOnRange: (from) => from === 1000 }); + await expect(paginate(t.make, { pageSize: 1000, label: 'cal' })) + .rejects.toThrow(/read failed at range 1000-1999 — connection reset/); + }); + + it('tags the error so a caller can tell a read failure from thin history', async () => { + const t = fakeTable(ids(1, 10), { errorOnRange: (from) => from === 0 }); + await expect(paginate(t.make, { pageSize: 5 })).rejects.toMatchObject({ code: 'READ_FAILED' }); + }); + + it('names itself in the message', async () => { + const t = fakeTable(ids(1, 10), { errorOnRange: () => true }); + await expect(paginate(t.make, { pageSize: 5, label: 'fromLedger(hits)' })) + .rejects.toThrow(/^fromLedger\(hits\):/); + }); + + it('refuses a runaway walk rather than reporting a truncated read as complete', async () => { + // Fresh unique ids every page, so the runaway guard is what fires — not the + // duplicate guard. + let n = 0; + const t = { + make: () => ({ + order() { return this; }, + async range() { const page = ids(n * 10 + 1, n * 10 + 10); n += 1; return { data: page, error: null }; }, + }), + }; + await expect(paginate(t.make, { pageSize: 10, maxPages: 3 })).rejects.toThrow(/maxPages/); + }); +}); + +describe('safePaginate — (c) set-identity vs an ordered full-table control', () => { + const control = (rows) => [...rows].sort((a, b) => a.id - b.id).map((r) => r.id); + + it('is set-identical to the control on a static table', async () => { + const rows = ids(1, 4321); + const out = await paginate(fakeTable(rows).make, { pageSize: 1000 }); + expect(out.map((r) => r.id)).toEqual(control(rows)); + }); + + it('is set-identical when the physical order is shuffled under it', async () => { + // The planner returning rows in an arbitrary physical order is exactly the + // production condition. A stable ORDER BY makes it irrelevant. + const rows = ids(1, 3000); + const t = fakeTable([...rows].reverse()); + const out = await paginate(t.make, { pageSize: 1000 }); + expect(out.map((r) => r.id)).toEqual(control(rows)); + }); + + it('THROWS on a duplicate key rather than returning a corrupted set', async () => { + // A non-unique ordering key: ties may come back in any order, so pages + // overlap. An .order() clause on such a column is not a fix. + const dup = [{ id: 1 }, { id: 2 }, { id: 2 }, { id: 3 }]; + await expect(paginate(fakeTable(dup).make, { pageSize: 2 })) + .rejects.toThrow(/duplicate key \(id\)=2/); + }); +}); + +describe('safePaginate — concurrent append (the regression this fix exists for)', () => { + it('returns every pre-existing row exactly once while rows are APPENDED mid-walk', async () => { + // The snapshot cron appends to ledger_entries at 14/19/22/1/3 UTC. With a + // stable ascending key, appended rows land after the cursor: nothing already + // returned can be returned again, and nothing pending can be skipped. + const original = ids(1, 2500); + let next = 10000; + const t = fakeTable([...original], { + beforeRange: (from, state) => { + if (from > 0) state.rows.push(row(next++)); // a write lands mid-walk + }, + }); + const out = await paginate(t.make, { pageSize: 1000 }); + const seen = out.map((r) => r.id); + + expect(new Set(seen).size).toBe(seen.length); // no duplicates + for (const r of original) expect(seen).toContain(r.id); // nothing dropped + }); + + it('the OLD unordered walk fails that same simulation — the test has teeth', async () => { + // Same table, paged over PHYSICAL order with no ORDER BY. An UPDATE in + // Postgres writes a new tuple version at the end of the heap, so a row that + // was already returned moves AFTER the cursor (returned twice) and every row + // behind it shifts back by one (one falls across the page boundary unseen). + // This is the production shape, and it is what the fix removes. + const original = ids(1, 2500); + const rows = [...original]; + const unorderedWalk = async () => { + const out = []; + for (let from = 0; ; from += 1000) { + if (from > 0) rows.push(rows.splice(0, 1)[0]); // a row is re-written mid-walk + const page = rows.slice(from, from + 1000); + if (page.length === 0) break; + out.push(...page); + if (page.length < 1000) break; + } + return out; + }; + const seen = (await unorderedWalk()).map((r) => r.id); + expect(new Set(seen).size).toBeLessThan(seen.length); // duplicates appear + const missing = original.filter((r) => !seen.includes(r.id)); + expect(missing.length).toBeGreaterThan(0); // and rows are lost + }); +}); + +// ── FIX A2b — COMPOSITE KEYS ──────────────────────────────────────────────── +const { normalizeKey } = require('../../src/utils/safePaginate'); +const { uniqueKeyFor } = require('../../src/utils/tableKeys'); + +/** A fake table keyed by a composite tuple, ordered the way Postgres would. */ +function fakeComposite(rows, opts = {}) { + const state = { rows, orderCols: [], rangeCalls: 0 }; + state.make = () => { + const cols = []; + const q = { + order(c) { cols.push(c); return q; }, + async range(from, to) { + state.rangeCalls += 1; + state.orderCols = [...cols]; + if (opts.errorOnRange && opts.errorOnRange(from)) return { data: null, error: { message: 'reset' } }; + const out = [...state.rows].sort((a, b) => { + for (const c of cols) { + if (a[c] < b[c]) return -1; + if (a[c] > b[c]) return 1; + } + return 0; + }); + return { data: out.slice(from, to + 1), error: null }; + }, + }; + return q; + }; + return state; +} + +const KEY = ['as_of_date', 'sport', 'season', 'player_key']; +const composite = (n) => Array.from({ length: n }, (_, i) => ({ + as_of_date: `2026-08-${String((i % 9) + 1).padStart(2, '0')}`, + sport: 'mlb', season: 2026, player_key: `p${i}`, v: i, +})); + +describe('safePaginate — composite keys (A2b)', () => { + it('normalizeKey treats a string and a one-column list identically', () => { + expect(normalizeKey('id')).toEqual(['id']); + expect(normalizeKey(['id'])).toEqual(['id']); + expect(normalizeKey(undefined)).toEqual(['id']); + expect(normalizeKey(KEY)).toEqual(KEY); + }); + + it('rejects an empty key rather than silently defaulting to id', () => { + expect(() => normalizeKey([])).toThrow(/must be a column name/); + expect(() => normalizeKey([null])).toThrow(/must be a column name/); + }); + + it('returns the full set exactly once, ordered by the whole tuple', async () => { + const t = fakeComposite(composite(2500)); + const rows = await paginate(t.make, { key: KEY, pageSize: 1000 }); + expect(rows).toHaveLength(2500); + expect(new Set(rows.map((r) => r.player_key)).size).toBe(2500); + // sorted by as_of_date first + for (let i = 1; i < rows.length; i += 1) { + expect(rows[i].as_of_date >= rows[i - 1].as_of_date).toBe(true); + } + }); + + it('orders EVERY column of the key, in the declared order', async () => { + const t = fakeComposite(composite(10)); + await paginate(t.make, { key: KEY, pageSize: 100 }); + expect(t.orderCols).toEqual(KEY); + }); + + it('a mid-walk error still THROWS with a composite key', async () => { + const t = fakeComposite(composite(2500), { errorOnRange: (f) => f === 1000 }); + await expect(paginate(t.make, { key: KEY, pageSize: 1000, label: 'ctx' })) + .rejects.toThrow(/ctx: read failed at range 1000-1999/); + }); + + it('THROWS on a duplicate TUPLE, naming every column', async () => { + const dup = [ + { as_of_date: '2026-08-01', sport: 'mlb', season: 2026, player_key: 'a' }, + { as_of_date: '2026-08-01', sport: 'mlb', season: 2026, player_key: 'a' }, + ]; + await expect(paginate(fakeComposite(dup).make, { key: KEY, pageSize: 10 })) + .rejects.toThrow(/duplicate key \(as_of_date,sport,season,player_key\)/); + }); + + it('does NOT throw when only a PREFIX of the tuple repeats', async () => { + // Same day, same sport, different player — a legitimate row, not a duplicate. + const rows = [ + { as_of_date: '2026-08-01', sport: 'mlb', season: 2026, player_key: 'a' }, + { as_of_date: '2026-08-01', sport: 'mlb', season: 2026, player_key: 'b' }, + ]; + await expect(paginate(fakeComposite(rows).make, { key: KEY, pageSize: 10 })) + .resolves.toHaveLength(2); + }); + + it('the single-key path is unchanged — backward compatible', async () => { + const t = fakeTable(ids(1, 2490)); + const out = await paginate(t.make, { pageSize: 1000 }); // no key => 'id' + expect(out).toHaveLength(2490); + expect(new Set(out.map((r) => r.id)).size).toBe(2490); + }); +}); + +describe('tableKeys — the constraints come from the schema, not a guess', () => { + it('gives the real composite key for each context table', () => { + expect(uniqueKeyFor('batter_spray')).toEqual(['as_of_date', 'sport', 'season', 'source_id']); + expect(uniqueKeyFor('team_defense')).toEqual(['as_of_date', 'sport', 'season', 'team']); + expect(uniqueKeyFor('statcast_aggregates')).toEqual(['sport', 'season', 'source_id', 'role']); + expect(uniqueKeyFor('platoon_splits')).toEqual(['as_of_date', 'sport', 'season', 'player_key']); + expect(uniqueKeyFor('park_dimensions')).toEqual(['as_of_date', 'sport', 'venue_id']); + expect(uniqueKeyFor('hitter_opportunity')).toEqual(['as_of_date', 'sport', 'season', 'player_key']); + expect(uniqueKeyFor('lineup_context')).toEqual(['as_of_date', 'sport', 'game_pk', 'player_key']); + }); + + it('keeps the single-column keys single', () => { + expect(uniqueKeyFor('ledger_entries')).toEqual(['id']); + expect(uniqueKeyFor('model_snapshots')).toEqual(['id']); + expect(uniqueKeyFor('game_context')).toEqual(['game_id']); + }); + + it('THROWS for an unknown table rather than defaulting to id', () => { + expect(() => uniqueKeyFor('some_new_table')).toThrow(/no unique key recorded/); + }); +}); diff --git a/tests/unit/shadowAccrual.test.js b/tests/unit/shadowAccrual.test.js new file mode 100644 index 0000000..848b733 --- /dev/null +++ b/tests/unit/shadowAccrual.test.js @@ -0,0 +1,172 @@ +'use strict'; + +/** + * A7 — does the shadow EVIDENCE actually accrue, on the scheduled path? + * + * A6 measured the shadow firing in a hand-run script. That proves the resolver + * works; it does not prove a scheduled snapshot writes anything. This drives the + * REAL `gradeAndCacheSlate` -> onGraded -> rowsFromSides chain with injected + * deps and asserts a `would_fire` block lands on the retention row. + * + * The distinction matters: A5 found factors that were built, correct, and never + * invoked. Evidence that is generated but never persisted is the same failure + * one layer along. + */ + +const gradeSlate = require('../../src/services/gradeSlateService'); +const retention = require('../../src/services/retentionService'); +const ff = require('../../src/services/model/factorFreeze'); +const hf = require('../../src/services/model/hitsFactors'); +const { nameKey } = require('../../src/utils/playerName'); + +const JUDGE = nameKey('Aaron Judge'); + +const SPRAY = { + as_of_date: '2026-08-10', player_key: JUDGE, + pull_gb: 0.30, straight_gb: 0.20, oppo_gb: 0.10, + pull_air: 0.20, straight_air: 0.10, oppo_air: 0.10, +}; +const OAA = { + '3B': { oaa: 2 }, SS: { oaa: 1 }, LF: { oaa: 0 }, '1B': { oaa: 0 }, + '2B': { oaa: 1 }, RF: { oaa: -1 }, CF: { oaa: 0 }, +}; +const SPLITS = { vl: { pa: 120, atBats: 110, hits: 33 }, vr: { pa: 300, atBats: 280, hits: 70 } }; + +/** The LIVE context: no opponent, no pitcher — the measured production state. */ +const liveCtx = () => ({ + spray: SPRAY, positionOaa: null, bats: 'R', throws: null, + pitcherHardHit: null, platoonSplits: SPLITS, archetype: 'BOMBER', as_of: null, +}); +/** The SHADOW context: same indexes, joined through the resolved keys. */ +const shadowCtx = () => ({ ...liveCtx(), positionOaa: OAA, throws: 'L', pitcherHardHit: 0.45 }); + +const KEYS = { + opponent: 'Boston Red Sox', opposing_pitcher: 'Chris Sale', + source: 'lineup_context+schedule', refused: null, +}; + +/** A resolver with A6's signature: keys omitted = live, keys supplied = shadow. */ +const resolver = (prop, sport, keys) => (keys ? shadowCtx() : liveCtx()); + +const PROP = { + player: 'Aaron Judge', stat_type: 'hits', line: 0.5, book: 'draftkings', + over_odds: -120, under_odds: 100, game_time: '2026-08-10T23:05:00Z', +}; + +/** A grader that runs the REAL engine freeze/shadow logic in miniature. */ +function fakeGrade(base) { + const legacy = { + player: base.player, stat_type: base.stat_type, line: base.line, + direction: base.direction, grade: 'B', confidence: 70, p_win: 0.55, + }; + // This mirrors analyzeViaEngine1's A5/A6 block exactly. + if (String(base.stat_type).toLowerCase() === 'hits' && base.factor_context) { + const live = hf.adjustProbability(0.55, base.factor_context); + if (live.p_adjusted != null && live.factors_fired > 0) legacy.p_win = live.p_adjusted; + let shadow = null; + const keys = base.matchup_keys || null; + if (keys && (keys.opponent || keys.opposing_pitcher) && typeof base.factor_context_resolver === 'function') { + const sc = base.factor_context_resolver(base, base.sport, keys); + if (sc) shadow = ff.shadowBlock(sc, hf.hitsFactorMultiplier(sc), keys); + } + legacy.factor_inputs = ff.freeze(base.factor_context, base, shadow); + } + return legacy; +} + +async function runSlate(opts = {}) { + const collected = []; + await gradeSlate.gradeAndCacheSlate('mlb', [PROP], { + grade: fakeGrade, + factorContext: resolver, + matchupKeys: opts.matchupKeys !== undefined ? opts.matchupKeys : (() => KEYS), + cacheSet: async () => {}, + onGraded: (base, sides) => { + collected.push(...retention.rowsFromSides(base, sides, { + snapshotId: 's1', capturedAt: '2026-08-10T14:00:00Z', sport: 'mlb', gameDate: '2026-08-10', + })); + }, + ...opts.slate, + }); + return collected; +} + +describe('the shadow accrues through the REAL slate path', () => { + it('a graded hits row lands with would_fire on it', async () => { + const rows = await runSlate(); + expect(rows.length).toBeGreaterThan(0); + const row = rows.find((r) => r.factor_inputs); + expect(row).toBeDefined(); + expect(row.factor_inputs.would_fire).not.toBeNull(); + expect(row.factor_inputs.would_fire.factors_fired).toBeGreaterThan(0); + expect(row.factor_inputs.would_fire.keys.opponent).toBe('Boston Red Sox'); + }); + + it('the row also carries the LIVE refusal — both states, side by side', async () => { + const rows = await runSlate(); + const fi = rows.find((r) => r.factor_inputs).factor_inputs; + expect(fi.position_oaa).toBeNull(); // what the grade saw + expect(fi.throws).toBeNull(); + expect(fi.would_fire.shadow_inputs.throws).toBe('L'); // what it would have + }); + + it('would_fire stays RE-DERIVABLE from what is stored beside it', async () => { + const rows = await runSlate(); + const fi = rows.find((r) => r.factor_inputs).factor_inputs; + const re = hf.hitsFactorMultiplier({ + spray: fi.spray, bats: fi.bats, + platoonSplits: fi.platoon + ? { vl: { pa: fi.platoon.vl.pa, atBats: fi.platoon.vl.ab, hits: fi.platoon.vl.hits }, + vr: { pa: fi.platoon.vr.pa, atBats: fi.platoon.vr.ab, hits: fi.platoon.vr.hits } } + : null, + positionOaa: fi.would_fire.shadow_inputs.position_oaa, + throws: fi.would_fire.shadow_inputs.throws, + pitcherHardHit: fi.would_fire.shadow_inputs.pitcher_hard_hit, + }); + expect(re.multiplier).toBe(fi.would_fire.multiplier); + }); + + it('THE SERVED GRADE IS UNTOUCHED — p_win is the unadjusted forecast', async () => { + const rows = await runSlate(); + for (const r of rows) expect(Number(r.p_win)).toBe(0.55); + }); + + it('no matchupKeys resolver => no would_fire, and the grade is the same', async () => { + const rows = await runSlate({ matchupKeys: null }); + const fi = rows.find((r) => r.factor_inputs).factor_inputs; + expect(fi.would_fire).toBeNull(); + for (const r of rows) expect(Number(r.p_win)).toBe(0.55); + }); + + it('refused keys => no shadow fire, never a guess', async () => { + const refused = { opponent: null, opposing_pitcher: null, source: null, refused: 'no_lineup_row' }; + const rows = await runSlate({ matchupKeys: () => refused }); + expect(rows.find((r) => r.factor_inputs).factor_inputs.would_fire).toBeNull(); + }); +}); + +describe('the scheduled path carries the shadow (not just a hand-run script)', () => { + const src = require('fs').readFileSync( + require('path').join(__dirname, '..', '..', 'src/services/snapshotService.js'), 'utf8'); + + it('runSnapshot builds the key index and passes it to the grader', () => { + expect(src).toMatch(/matchupKeys/); + expect(src).toMatch(/gradeAndCacheSlate\(sp, props, \{\s*\n\s*factorContext,\s*\n\s*matchupKeys,/); + }); + + it('the index is built BEFORE the grade call, not after', () => { + expect(src.indexOf('FIX A6 — THE JOIN KEYS')).toBeLessThan(src.indexOf('await deps.gradeAndCacheSlate(')); + }); + + it('runAllSnapshots (the cron entrypoint) routes through runSnapshot', () => { + expect(src).toMatch(/async function runAllSnapshots/); + expect(src).toMatch(/runSnapshot\(/); + }); + + it('a key-resolve failure degrades to no-shadow, never breaking the slate', () => { + const i = src.indexOf('FIX A6 — THE JOIN KEYS'); + const block = src.slice(i, i + 1600); + expect(block).toMatch(/catch \(e\)/); + expect(block).toMatch(/matchupKeys = null|no key index/); + }); +}); diff --git a/tests/unit/snapshotSettlementService.test.js b/tests/unit/snapshotSettlementService.test.js new file mode 100644 index 0000000..3ca6880 --- /dev/null +++ b/tests/unit/snapshotSettlementService.test.js @@ -0,0 +1,263 @@ +'use strict'; + +/** + * snapshotSettlementService — Fix A3. + * Every dependency injected; nothing here touches a network or a database. + */ + +const svc = require('../../src/services/snapshotSettlementService'); + +const line = (over) => ({ hits: over, totalBases: over, rbi: over, runs: over }); + +describe('decide — side-aligned outcomes', () => { + const base = { id: 1, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a' }; + const lines = { '2026-08-05|a': line(2) }; + + it('an OVER that cleared the line is a hit', () => { + const { updates } = svc.decide([{ ...base, line: 0.5, side: 'over' }], lines); + expect(updates).toEqual([{ id: 1, outcome: 'hit', actual_value: 2 }]); + }); + + it('an UNDER on the SAME row is the mirror — a miss, not a hit', () => { + // A raw `realized > line` indicator is the OVER perspective and would invert + // the target on every under row. + const { updates } = svc.decide([{ ...base, line: 0.5, side: 'under' }], lines); + expect(updates).toEqual([{ id: 1, outcome: 'miss', actual_value: 2 }]); + }); + + it('an UNDER that held is a hit', () => { + const { updates } = svc.decide([{ ...base, line: 3.5, side: 'under' }], lines); + expect(updates[0].outcome).toBe('hit'); + }); + + it('actual_value is the RAW realized stat, unopinionated by side', () => { + const o = svc.decide([{ ...base, line: 0.5, side: 'over' }], lines).updates[0]; + const u = svc.decide([{ ...base, line: 0.5, side: 'under' }], lines).updates[0]; + expect(o.actual_value).toBe(2); + expect(u.actual_value).toBe(2); + }); + + it('reads the right box-score field per stat', () => { + const l = { '2026-08-05|a': { hits: 1, totalBases: 4, rbi: 3, runs: 0 } }; + const at = (stat) => svc.decide([{ ...base, stat, line: 1.5, side: 'over' }], l).updates[0]; + expect(at('total_bases').actual_value).toBe(4); + expect(at('rbi').actual_value).toBe(3); + // hits=1 does not clear 1.5 — that is a settled MISS, not an absence. + expect(at('hits')).toEqual({ id: 1, outcome: 'miss', actual_value: 1 }); + }); +}); + +describe('decide — a row logged after first pitch is not a prediction', () => { + const lines = { '2026-08-05|a': line(2) }; + const row = (capturedAt) => ({ id: 1, game_date: '2026-08-05', captured_at: capturedAt, stat: 'hits', player_key: 'a', line: 0.5, side: 'over' }); + + it('accepts a capture from the morning of the game', () => { + // 14:00 UTC = 10:00 ET, well before first pitch. + expect(svc.decide([row('2026-08-05T14:00:00Z')], lines).counts.settled).toBe(1); + }); + + it('accepts a capture from the day before', () => { + expect(svc.decide([row('2026-08-04T22:00:00Z')], lines).counts.settled).toBe(1); + }); + + it('REFUSES a 01:00-UTC capture, which is 21:00 ET the same evening — in-game', () => { + // The trap: a UTC date compare would keep this row. + const r = svc.decide([{ ...row('2026-08-06T01:00:00Z'), game_date: '2026-08-05' }], lines); + expect(r.counts.settled).toBe(0); + expect(r.counts.post_hoc_logged).toBe(1); + }); + + it('REFUSES a capture from the following day', () => { + const r = svc.decide([row('2026-08-06T14:00:00Z')], lines); + expect(r.counts.post_hoc_logged).toBe(1); + }); + + it('isPreGame is absent-safe', () => { + expect(svc.isPreGame(null, '2026-08-05')).toBe(false); + expect(svc.isPreGame('2026-08-05T14:00:00Z', null)).toBe(false); + expect(svc.isPreGame('not-a-date', '2026-08-05')).toBe(false); + }); +}); + +describe('decide — integrity', () => { + it('every candidate lands in exactly one bucket (conservation)', () => { + const lines = { '2026-08-05|a': line(2) }; + const rows = [ + { id: 1, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over' }, + { id: 2, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'ghost', line: 0.5, side: 'over' }, + { id: 3, game_date: '2026-08-05', captured_at: '2026-08-06T01:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over' }, + { id: 4, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a', line: null, side: 'over' }, + ]; + const { counts } = svc.decide(rows, lines); + expect(counts.settled).toBe(1); + expect(counts.orphaned).toBe(1); + expect(counts.unresolvable).toBe(2); + expect(counts.settled + counts.orphaned + counts.unresolvable).toBe(counts.candidates); + }); + + it('THROWS on a duplicate row id — a doubled prediction poisons every measure', () => { + const r = { id: 7, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over' }; + expect(() => svc.decide([r, r], { '2026-08-05|a': line(2) })).toThrow(/duplicate snapshot id 7/); + }); + + it('a player with no box-score line is ORPHANED, never settled at zero', () => { + const r = svc.decide([{ id: 1, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'nobody', line: 0.5, side: 'over' }], {}); + expect(r.counts.orphaned).toBe(1); + expect(r.updates).toHaveLength(0); + }); +}); + +describe('battingLines — doubleheaders sum, non-batters are absent', () => { + const sched = (pks) => ({ dates: [{ games: pks.map((pk) => ({ gamePk: pk, officialDate: '2026-08-05', status: { detailedState: 'Final' } })) }] }); + const box = (rows) => ({ + teams: { + home: { batters: rows.map((_, i) => i + 1), players: Object.fromEntries(rows.map((r, i) => [`ID${i + 1}`, { person: { fullName: r.name }, stats: { batting: r.b } }])) }, + away: { batters: [], players: {} }, + }, + }); + + it('sums both games of a doubleheader — the prop covers the day', async () => { + const getJson = async (url) => { + if (url.includes('schedule')) return sched([1, 2]); + return box([{ name: 'Aaron Judge', b: { atBats: 4, hits: 1, totalBases: 2, rbi: 1, runs: 0 } }]); + }; + const lines = await svc.battingLines(['2026-08-05'], { getJson }); + const k = Object.keys(lines)[0]; + expect(lines[k].hits).toBe(2); // 1 + 1 + expect(lines[k].totalBases).toBe(4); + }); + + it('omits a player who did not bat (atBats null) rather than recording zero', async () => { + const getJson = async (url) => (url.includes('schedule') ? sched([1]) + : box([{ name: 'Benchy McBench', b: { atBats: null } }])); + const lines = await svc.battingLines(['2026-08-05'], { getJson }); + expect(Object.keys(lines)).toHaveLength(0); + }); + + it('ignores a game that is not Final', async () => { + const getJson = async (url) => (url.includes('schedule') + ? { dates: [{ games: [{ gamePk: 1, status: { detailedState: 'In Progress' } }] }] } : box([])); + expect(Object.keys(await svc.battingLines(['2026-08-05'], { getJson }))).toHaveLength(0); + }); + + it('a failed date is absent, not fatal', async () => { + const getJson = async () => { throw new Error('502'); }; + await expect(svc.battingLines(['2026-08-05'], { getJson })).resolves.toEqual({}); + }); +}); + +describe('settleSnapshots — orchestration', () => { + const fakeSb = (rows, captured = []) => ({ + from() { + const q = { + select() { return q; }, eq() { return q; }, in() { return q; }, + is() { return q; }, lt() { return q; }, order() { return q; }, + async range(from, to) { return { data: rows.slice(from, to + 1), error: null }; }, + update(patch) { return { eq: (_c, id) => ({ is: async () => { captured.push({ id, patch }); return { error: null }; } }) }; }, + }; + return q; + }, + }); + + it('no client → skipped, never a silent zero', async () => { + expect(await svc.settleSnapshots({})).toMatchObject({ skipped: 'supabase not configured' }); + }); + + it('settles and reports repaired-champion rows separately', async () => { + const rows = [ + { id: 1, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over', model_version: 'engine1@2026-08-07-fullwindow' }, + { id: 2, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over', model_version: 'engine1@2026-07-20' }, + ]; + const captured = []; + const res = await svc.settleSnapshots({ + sb: fakeSb(rows, captured), + now: () => new Date('2026-08-09T12:00:00Z'), + getJson: async (url) => (url.includes('schedule') + ? { dates: [{ games: [{ gamePk: 1, officialDate: '2026-08-05', status: { detailedState: 'Final' } }] }] } + : { teams: { home: { batters: [1], players: { ID1: { person: { fullName: 'A' }, stats: { batting: { atBats: 4, hits: 2, totalBases: 2, rbi: 0, runs: 0 } } } } }, away: { batters: [], players: {} } } }), + }); + expect(res.written).toBe(2); + expect(res.repaired_champion_settled).toBe(1); + // Real outcomes, not nulls. + for (const c of captured) { + expect(c.patch.outcome).toBe('hit'); + expect(c.patch.actual_value).toBe(2); + expect(c.patch.settlement_source).toBe('statsapi_boxscore'); + } + }); + + it('bounds the dates per run and drains NEWEST first (oldest-first deadlocks)', async () => { + const rows = Array.from({ length: 12 }, (_, i) => ({ + id: i + 1, game_date: `2026-07-${String(i + 10).padStart(2, '0')}`, + captured_at: `2026-07-${String(i + 10).padStart(2, '0')}T14:00:00Z`, + stat: 'hits', player_key: 'a', line: 0.5, side: 'over', + })); + const res = await svc.settleSnapshots({ + sb: fakeSb(rows), now: () => new Date('2026-08-09T12:00:00Z'), + maxDates: 3, maxWindows: 1, getJson: async () => ({ dates: [] }), + }); + expect(res.dates).toHaveLength(3); + expect(res.dates_remaining).toBe(9); + // Newest first: an old date whose remaining rows are all post-hoc or + // orphaned would otherwise re-process forever and block the window. + expect(res.dates[0]).toBe('2026-07-21'); + }); + + it('ADVANCES past a drained window so older settleable dates are reached', async () => { + // The newest dates keep rows that can never settle (no box-score line), so a + // fixed window would sit on them forever. Only the OLDEST date here has a + // matching box score; the run must walk to it. + const rows = Array.from({ length: 4 }, (_, i) => ({ + id: i + 1, game_date: `2026-08-0${i + 1}`, captured_at: `2026-08-0${i + 1}T14:00:00Z`, + stat: 'hits', player_key: i === 0 ? 'a' : `ghost${i}`, line: 0.5, side: 'over', + })); + const res = await svc.settleSnapshots({ + sb: fakeSb(rows), now: () => new Date('2026-08-09T12:00:00Z'), + maxDates: 1, maxWindows: 4, + getJson: async (url) => (url.includes('schedule') + ? { dates: [{ games: [{ gamePk: 1, officialDate: '2026-08-01', status: { detailedState: 'Final' } }] }] } + : { teams: { home: { batters: [1], players: { ID1: { person: { fullName: 'A' }, stats: { batting: { atBats: 4, hits: 2, totalBases: 2, rbi: 0, runs: 0 } } } } }, away: { batters: [], players: {} } } }), + }); + expect(res.windows_scanned).toBe(4); // walked past three drained windows + expect(res.written).toBe(1); + expect(res.dates).toEqual(['2026-08-01']); + }); + + it('counts permanently-unsettleable rows rather than hiding them', async () => { + const rows = [ + // logged at 01:00 UTC = 21:00 ET the same evening — in-game, never a prediction + { id: 1, game_date: '2026-08-05', captured_at: '2026-08-06T01:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over' }, + { id: 2, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over' }, + ]; + const res = await svc.settleSnapshots({ + sb: fakeSb(rows), now: () => new Date('2026-08-09T12:00:00Z'), + getJson: async () => ({ dates: [] }), + }); + expect(res.permanently_unsettleable).toBe(1); + }); + + it('dry run writes nothing', async () => { + const captured = []; + const rows = [{ id: 1, game_date: '2026-08-05', captured_at: '2026-08-05T14:00:00Z', stat: 'hits', player_key: 'a', line: 0.5, side: 'over' }]; + const res = await svc.settleSnapshots({ + sb: fakeSb(rows, captured), write: false, now: () => new Date('2026-08-09T12:00:00Z'), + getJson: async () => ({ dates: [] }), + }); + expect(res.mode).toBe('dry-run'); + expect(captured).toHaveLength(0); + }); +}); + +describe('settlement does NOT reconstruct context', () => { + it('touches no context table and no context column', () => { + const src = require('fs').readFileSync( + require('path').join(__dirname, '..', '..', 'src/services/snapshotSettlementService.js'), 'utf8'); + for (const t of ['batter_spray', 'team_defense', 'platoon_splits', 'park_dimensions', + 'game_context', 'hitter_opportunity', 'statcast_aggregates', 'lineup_context']) { + expect(src).not.toContain(`'${t}'`); + } + // It writes outcome columns only. + expect(src).toMatch(/outcome: u\.outcome/); + expect(src).not.toMatch(/features:/); + }); +}); diff --git a/tests/unit/withdrawnVerdicts.test.js b/tests/unit/withdrawnVerdicts.test.js new file mode 100644 index 0000000..543a2cf --- /dev/null +++ b/tests/unit/withdrawnVerdicts.test.js @@ -0,0 +1,67 @@ +'use strict'; + +/** The withdrawal record — spec: specs/read-integrity-harness.md §6. */ + +const wv = require('../../src/services/model/withdrawnVerdicts'); + +describe('withdrawnVerdicts', () => { + it('records BOTH hits PROVES as withdrawn', () => { + expect(wv.isWithdrawn('mlb', 'hits', 'defense_by_direction')).toBe(true); + expect(wv.isWithdrawn('mlb', 'hits', 'pitcher_contact_profile')).toBe(true); + }); + + it('does not withdraw a factor that was never proven', () => { + // platoon_severity never returned PROVES (n=452 < 500). Withdrawing it would + // imply it once stood, which is its own falsehood. + expect(wv.isWithdrawn('mlb', 'hits', 'platoon_severity')).toBe(false); + expect(wv.isWithdrawn('mlb', 'hits', 'park_hits')).toBe(false); + }); + + it('is scoped per stat — a hits withdrawal says nothing about total_bases', () => { + expect(wv.isWithdrawn('mlb', 'total_bases', 'defense_by_direction')).toBe(false); + }); + + it('APPENDS rather than replaces: the original verdict and evidence survive', () => { + for (const w of wv.WITHDRAWALS) { + expect(w.original_verdict).toBe('PROVES'); + expect(w.original_evidence.n).toBeGreaterThan(0); + expect(Array.isArray(w.original_evidence.ci)).toBe(true); + } + }); + + it('every withdrawal carries a measured cause and implicated readers', () => { + for (const w of wv.WITHDRAWALS) { + expect(w.cause).toMatch(/unordered-pagination/); + expect(w.cause).toMatch(/2026-08-09/); + expect(w.readers_implicated.length).toBeGreaterThan(0); + expect(w.status).toBe(wv.STATUS.WITHDRAWN_PENDING_REAUDIT); + } + }); + + it('names the eligibility arithmetic for defense_by_direction AS arithmetic', () => { + const d = wv.WITHDRAWALS.find((w) => w.factor === 'defense_by_direction'); + expect(d.eligibility_note).toMatch(/422/); + expect(d.eligibility_note).toMatch(/not a re-measurement/); + }); + + it('records the honesty gap: both are STILL SERVED despite withdrawal', () => { + const served = wv.servedButWithdrawn(); + expect(served).toHaveLength(2); + for (const w of served) expect(w.still_served_note).toMatch(/hitsFactors/); + }); + + it('states a reinstatement path that requires a measured clean read', () => { + for (const w of wv.WITHDRAWALS) expect(w.reinstatement).toMatch(/readIntegrity/); + }); + + it('is frozen — a withdrawal must not be mutated in place', () => { + expect(Object.isFrozen(wv.WITHDRAWALS)).toBe(true); + for (const w of wv.WITHDRAWALS) expect(Object.isFrozen(w)).toBe(true); + }); + + it('withdrawalsFor filters by sport and stat', () => { + expect(wv.withdrawalsFor('mlb', 'hits')).toHaveLength(2); + expect(wv.withdrawalsFor('mlb', 'runs')).toHaveLength(0); + expect(wv.withdrawalsFor('mlb')).toHaveLength(2); + }); +});