Read integrity, as-of context, and the shadow matchup resolve (A1-A7)

Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Kev
2026-08-11 22:49:56 -04:00
parent 387ae4d54e
commit f61ec6b391
49 changed files with 4874 additions and 308 deletions
+176 -1
View File
@@ -1,7 +1,182 @@
# VYNDR — Build State
## Last Updated
2026-08-03
2026-08-09
## Fix A0 (2026-08-09) — Withdraw unknowable verdicts + read-integrity harness ✅
4,569 tests / 363 suites green (repo config). **Measure-only: no production
reader edited, no eligibility opened, no delete, no settle, no schema.**
- **SHIPPED** `specs/read-integrity-harness.md` (spec-first),
`src/utils/readIntegrity.js` (pure + injectable core),
`scripts/read-integrity.js` (CLI + declarative 25-reader registry),
`src/services/model/withdrawnVerdicts.js` (append-only withdrawal record),
30 new unit tests across two suites.
- **AC5 verified on prod:** the harness reproduces the probe — known-corrupt
`prove-hit-factors:136` at 16.5%, known-clean `challenger-scoreboard:112` at 0%.
- **Baseline: 11 FAIL / 13 PASS / 1 KEY_NOT_UNIQUE of 25.** Worst
`champion-ablation:173` 33.6%, `proven-status:76` 32.6%,
`calibrationService.fromLedger:110` (the one LIVE reader) 24.8%.
- **WITHDRAWN:** `defense_by_direction` + `pitcher_contact_profile` on hits →
`WITHDRAWN_PENDING_REAUDIT`. Nothing is PROVEN on hits from the factor gate now.
- **OPEN, urgent, NOT in this order's scope:** (1) the live
`calibrationService.fromLedger` reader is 24.8% corrupt in production; (2)
`hitsFactors.js` still serves both withdrawn factors AND `platoon_severity`,
which never passed the gate at all. Both are A1/A2 decisions.
## Fix A1 (2026-08-09) — safePaginate + the first corrected reader ✅
4,592 tests / 365 suites green, web build exit 0. **Served grade FROZEN;
`CALIBRATION_DEPLOYED` still `[]`; nothing re-certified, re-deployed or
un-withdrawn.**
- **SHIPPED** `src/utils/safePaginate.js` (stable ORDER BY on a unique key +
runtime uniqueness guard + error propagation, built on the tested
`readIntegrity.walk`), 13 unit tests; `calibrationService.loadSettledRows`
extracted + converted, 11 unit tests.
- **HARNESS: FAIL 24.8% → PASS (0 dup / 0 missing of 2,490)**, measured through
the REAL function via the new `readerRows` arm (`arm_a: real_reader_function`).
The unfixed twin `lowParamService.fromLedger:78` still measures 24.8% — the
control that proves the PASS is not a registry edit.
- **CORRECTION to A0:** this reader was never live. `CALIBRATION_DEPLOYED = []`
(snapshotService.js:301) means the loop never runs; calibrationService is only
the SHADOW; `chainAcross` has no callers. **A1 changed no served number.**
- **MAP DELTA (reported, NOT deployed):** certified-band error −0.046 → +0.010,
ceiling 0.833 → 0.810, span unchanged at 0.50–0.70, `fitted_through` shifted a
day because duplicates moved the time split.
## Fix A2 (2026-08-09) — safePaginate rolled across every FAIL reader ✅
4,628 tests / 366 suites, web build exit 0. **Served grade FROZEN;
`CALIBRATION_DEPLOYED` still `[]`; both hits verdicts still withdrawn; no
eligibility opened, no delete/settle.**
- **11 FAIL → 0 FAIL.** Full 26-reader baseline: **25 PASS + 1 KEY_NOT_UNIQUE**
(the `batter_spray` composite-key data fact, unchanged and correctly named).
Worst corruption 33.6% → **0%**.
- Every fixed reader verified through the harness's `real_reader_function` arm,
i.e. the code that actually runs — not a restated query, never a `fixed:true`.
- **PRIMARY calibrator first:** `lowParamService.fromLedger:78` 24.8% → PASS.
- **7 scripts converted** via `pageSafe` + named `READS` exports + a
`require.main === module` guard (needed because the harness imports them and
several write to `mc_test_ledger`; verified unchanged at 168 rows).
- **RESISTED THE HELPER:** the context tables have composite PKs with no single
unique column, so `paginate()` cannot order them. They measure 0% and are left
on `page()`, documented in-file. Composite-key ordering is the open item.
- **OPEN:** an unresolved intermittent in `snapshotService.test.js` (2/30 on
branch, 0/29 at HEAD — not statistically distinguishable, assertion never
captured). Recorded, not dismissed.
## Fix A2b (2026-08-09) — every read clean by construction ✅
4,699 tests / 366 suites, web build exit 0. **Grade FROZEN; `CALIBRATION_DEPLOYED`
still `[]`; verdicts still withdrawn; no eligibility, no settlement, no
version-stamp, no delete. No context data altered.**
- **34 readers, 34 PASS, 0 FAIL, 0 KEY_NOT_UNIQUE** — up from 26 readers / 1
KEY_NOT_UNIQUE. 30 of 34 measured through the REAL reader function.
- **safePaginate takes composite keys**; `src/utils/tableKeys.js` holds the real
constraints from `pg_index`. All 7 context tables converted; `batter_spray`
resolved (wrong key, not bad data).
- **The 3 A2 legacy reads converted**; `KNOWN_LEGACY_READS` is now empty.
- **`calibrate-hits.js` found at 24.7%** during the sweep and converted.
- **The unordered `page()` helper is DELETED from all 15 scripts.**
- **Write-scripts assessed, data untouched:** omission-only failure mode; read is
0% today; historical exposure bounded at ~6.6 players / ~0.16 games.
## Fix A3 (2026-08-09) — eligibility can honestly count ✅
4,723 tests / 367 suites, web build exit 0. **Grade FROZEN; `CALIBRATION_DEPLOYED`
still `[]`; `hitsFactors`, `hitsFactorContext`, `withdrawnVerdicts`,
`reAuditEligibility`, `snapshotService` all EMPTY DIFF. No verdict re-run, no
calibration re-fit, no historical re-stamp.**
- **ONE version source:** `src/config/modelVersion.js` = `engine1@2026-08-07-fullwindow`.
Both `ledgerService` and `retentionService` import it; the twin defaults that
silently blocked eligibility are gone.
- **Snapshot settlement on the cron:** `snapshotSettlementService` wired into
`snapshotScheduler` beside the ledger settle. Outcomes only, no context
reconstruction. Settled **15,484 → 28,894**; repaired-champion **0 → 7,556**;
0 rows settled with a null actual_value. 24 new unit tests.
- **Eligibility (counted, not run):** 2 eligible dates against 10/10/14/14.
All four measurements remain BLOCKED — now for the honest reason.
## Fix A4 (2026-08-09) — as-of-correct context ✅
4,735 tests / 368 suites, web build exit 0, harness still 34/34 PASS.
**Grade output UNCHANGED (proven 400/400); inputs NOT frozen (that is A5); no
re-audit run, no verdict re-run, no calibration re-fit; `CALIBRATION_DEPLOYED`
still `[]`.**
- `hitsFactorContext.build(sb, {asOf})` — dated reads bounded `.lte('as_of_date')`,
latest-within-bound, **refusal on absence** (never nearest/latest).
- Dated statcast comes from `statcast_history` (aggregates keeps one date);
verified identical to aggregates at the head, 1,414 rows / 0 diffs.
- **Coverage on the two re-audit dates: spray/platoon/handedness 99% and 98.9%**,
2 players short per date. **Defence 0% from the row** — `opponent` is NULL on
every settled repaired-champion row; it needs the game-log join.
- 12 new unit tests incl. the contamination lock (never a row dated after asOf).
## Fix A5 (2026-08-10) — input freeze, and a finding that reframes it ✅
4,751 tests / 369 suites, web build exit 0. **Grade UNCHANGED (496/496); no
re-audit, no verdict re-run, no calibration; `CALIBRATION_DEPLOYED` still `[]`.**
- **FINDING: the three hits factors NEVER FIRE.** 596/596 real props skipped,
multiplier 1 every time, because `prop.opponent` / `prop.opposing_pitcher` are
never set by anything. They are wired and inert. Plumbing them is a MODEL
CHANGE and needs its own order.
- **`model_snapshots.factor_inputs` jsonb** (migration 034, APPLIED) freezes the
raw inputs at grade time — never the multiplier, which stays recomputable via
`factorFreeze.recompute()`. Proven: frozen == read (496/496), recompute == live
(496/496).
- **Forward-only.** Existing 7,556 settled rows keep NULL `factor_inputs` and
NULL `opponent`; defence re-audit on them still needs the game-log join.
## Fix A6 (2026-08-10) — join keys plumbed into a SHADOW resolve ✅
4,762 tests / 370 suites, web build exit 0. **Served grade FROZEN; no re-audit,
no verdict reinstated, no calibration; `CALIBRATION_DEPLOYED` still `[]`.**
- `matchupKeys` resolves opponent + opposing pitcher from `lineup_context`
(as-of) + schedule probables. Refuses rather than guessing.
- **Key resolution 248/308 (80.5%)** on 2026-08-09; 60 refused (no lineup row).
- **Shadow fire: pitcher_contact 248, defense_by_direction 185,
platoon_severity 147** — up from 0/0/0. Multiplier median 1.017, range
0.803–1.250; 245/248 would move the served number.
- Served p_win identical on every evaluated row; shadow recompute-check 248/248.
## Fix A7 (2026-08-11) — shadow accrual proven + A8 pre-registered ✅
4,772 tests / 371 suites, web build exit 0. **Nothing turned on. Served grade
frozen; `CALIBRATION_DEPLOYED` still `[]`; no verdict reinstated; gate NOT run.**
- **Accrual is automatic on the scheduled pass** — proven end-to-end through the
real `gradeAndCacheSlate → onGraded → rowsFromSides` chain (10 new tests),
plus assertions that `runSnapshot` builds the key index pre-grade and degrades
safely.
- **Complete (would_fire, outcome) pairs today: 0.** `factor_inputs` is
0/122,276 — A5/A6 are not deployed. Accrual begins at deploy.
- **`specs/a8-shadow-factor-gate.md`** pre-registers the two-part gate (movement
AND Brier, cluster-resampled, cumulative-Bonferroni, new-test α) with named
fallbacks. Not run.
- **Binding constraint is CLUSTERS: ~10.5 settled games/slate ⇒ `MIN_CLUSTERS`
40 in ~4 slates**, while `MIN_N` 500 is met in 1–2. A8 runnable ≈5 slates
after deploy.
### Next
- **A8 — run the pre-registered gate** once ≈5 slates have accrued.
- **A9 — turn the factors on LIVE**, only if A8 returns PROVES, and with its own
before/after on the served board.
- **(superseded) A7 — turn the factors on LIVE.** This is the biggest served-grade change in
VYNDR's history: ~80% of hits props would move, median +1.7%, tails ±20-25%.
Gate it on (a) accrued shadow evidence, (b) a re-proof of the withdrawn
verdicts on rows where the factors actually fire.
- **(superseded) A6 — plumb `opponent` + `opposing_pitcher` into the graded prop.** This is a
MODEL CHANGE (grades will move): measure the delta against the frozen
`available` fields first, then decide.
- **(superseded) A5 — freeze the factor inputs onto the graded row** so a re-audit does not
depend on context tables at all (and so defence stops needing a game-log join).
- **(done in A4) as-of-correct context** (`hitsFactorContext.build` has no as-of cutoff;
it also orders context walks on a non-unique prefix). **The re-audit RUN is
gated on this, not on clean reads.**
- **A2c (was A2b)** — composite-key ordering in `safePaginate` so the context tables can be
made safe by construction rather than clean by luck; then convert the 3 known
legacy reads and register `backfill-context` / `reconstruct-game-environment`.
- **A2c** — chase the `snapshotService.test.js` intermittent to an assertion.
- **(superseded)** roll `safePaginate` across the remaining FAIL readers, starting with
`lowParamService.fromLedger:78` (the PRIMARY calibrator, 24.8%). Each must move
FAIL → PASS on a measured re-run through `readerRows`, never on the presence of
an `.order()` clause. Remaining: `champion-ablation:173` 33.6%,
`proven-status:76` 32.6%, `prove-hit-factors:131/136`, `prove-tb-factors:149/154`,
`cluster-prove:337`, `tb-solo-and-interactions:260`, `build-grade-bands:49/54`.
- **A3** — decide what to do about `hitsFactors.js` serving two withdrawn factors
plus `platoon_severity`, which never passed the gate.
- **Re-certification of calibration is a separate gated order** on the accrual
clock — not unlocked by A1.
## Session 94 (2026-08-04) — Causally-correct platoon + park inputs ✅
4,307 tests / 344 suites green, build exit 0. Counter + frozen clusters