Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move. A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS: 410-617 of 2,490 duplicated with an equal number never returned, while rows.length matched the server exactly. safePaginate orders on a real unique key, verifies the tuple at runtime, and THROWS on a query error instead of treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn through that reader, and defense_by_direction's distinct-n was likely below the gate floor all along. A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from pg_index (the context tables are dated-composite and had no single unique column). The unordered helper is deleted, not parked. A3 — ledgerService and retentionService defaulted the SAME env var to DIFFERENT versions, so no ledger row ever carried the marker eligibility requires. One source now. model_snapshots settlement moved onto the cron: 15,484 -> 28,894 settled, repaired-champion 0 -> 7,556. A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no row at-or-before the date means the factor does not apply, never the nearest row. Live path unchanged, proven 400/400 on real rows. A5 — factor_inputs freezes what the factor READ, never the multiplier, so an audit can recompute and check. It also recorded the finding: the three hits factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by the resolver and written by nothing. A6/A7 — matchupKeys resolves those keys from the posted lineup plus the schedule's probable pitchers, and fires the factors into a SHADOW freeze: 248 fires on 308 props, 245 of which would move the grade. The served forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test that decides whether they ever go live. Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity harness 34/34. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+176
-1
@@ -1,7 +1,182 @@
|
||||
# VYNDR — Build State
|
||||
|
||||
## Last Updated
|
||||
2026-08-03
|
||||
2026-08-09
|
||||
|
||||
## Fix A0 (2026-08-09) — Withdraw unknowable verdicts + read-integrity harness ✅
|
||||
4,569 tests / 363 suites green (repo config). **Measure-only: no production
|
||||
reader edited, no eligibility opened, no delete, no settle, no schema.**
|
||||
- **SHIPPED** `specs/read-integrity-harness.md` (spec-first),
|
||||
`src/utils/readIntegrity.js` (pure + injectable core),
|
||||
`scripts/read-integrity.js` (CLI + declarative 25-reader registry),
|
||||
`src/services/model/withdrawnVerdicts.js` (append-only withdrawal record),
|
||||
30 new unit tests across two suites.
|
||||
- **AC5 verified on prod:** the harness reproduces the probe — known-corrupt
|
||||
`prove-hit-factors:136` at 16.5%, known-clean `challenger-scoreboard:112` at 0%.
|
||||
- **Baseline: 11 FAIL / 13 PASS / 1 KEY_NOT_UNIQUE of 25.** Worst
|
||||
`champion-ablation:173` 33.6%, `proven-status:76` 32.6%,
|
||||
`calibrationService.fromLedger:110` (the one LIVE reader) 24.8%.
|
||||
- **WITHDRAWN:** `defense_by_direction` + `pitcher_contact_profile` on hits →
|
||||
`WITHDRAWN_PENDING_REAUDIT`. Nothing is PROVEN on hits from the factor gate now.
|
||||
- **OPEN, urgent, NOT in this order's scope:** (1) the live
|
||||
`calibrationService.fromLedger` reader is 24.8% corrupt in production; (2)
|
||||
`hitsFactors.js` still serves both withdrawn factors AND `platoon_severity`,
|
||||
which never passed the gate at all. Both are A1/A2 decisions.
|
||||
|
||||
## Fix A1 (2026-08-09) — safePaginate + the first corrected reader ✅
|
||||
4,592 tests / 365 suites green, web build exit 0. **Served grade FROZEN;
|
||||
`CALIBRATION_DEPLOYED` still `[]`; nothing re-certified, re-deployed or
|
||||
un-withdrawn.**
|
||||
- **SHIPPED** `src/utils/safePaginate.js` (stable ORDER BY on a unique key +
|
||||
runtime uniqueness guard + error propagation, built on the tested
|
||||
`readIntegrity.walk`), 13 unit tests; `calibrationService.loadSettledRows`
|
||||
extracted + converted, 11 unit tests.
|
||||
- **HARNESS: FAIL 24.8% → PASS (0 dup / 0 missing of 2,490)**, measured through
|
||||
the REAL function via the new `readerRows` arm (`arm_a: real_reader_function`).
|
||||
The unfixed twin `lowParamService.fromLedger:78` still measures 24.8% — the
|
||||
control that proves the PASS is not a registry edit.
|
||||
- **CORRECTION to A0:** this reader was never live. `CALIBRATION_DEPLOYED = []`
|
||||
(snapshotService.js:301) means the loop never runs; calibrationService is only
|
||||
the SHADOW; `chainAcross` has no callers. **A1 changed no served number.**
|
||||
- **MAP DELTA (reported, NOT deployed):** certified-band error −0.046 → +0.010,
|
||||
ceiling 0.833 → 0.810, span unchanged at 0.50–0.70, `fitted_through` shifted a
|
||||
day because duplicates moved the time split.
|
||||
|
||||
## Fix A2 (2026-08-09) — safePaginate rolled across every FAIL reader ✅
|
||||
4,628 tests / 366 suites, web build exit 0. **Served grade FROZEN;
|
||||
`CALIBRATION_DEPLOYED` still `[]`; both hits verdicts still withdrawn; no
|
||||
eligibility opened, no delete/settle.**
|
||||
- **11 FAIL → 0 FAIL.** Full 26-reader baseline: **25 PASS + 1 KEY_NOT_UNIQUE**
|
||||
(the `batter_spray` composite-key data fact, unchanged and correctly named).
|
||||
Worst corruption 33.6% → **0%**.
|
||||
- Every fixed reader verified through the harness's `real_reader_function` arm,
|
||||
i.e. the code that actually runs — not a restated query, never a `fixed:true`.
|
||||
- **PRIMARY calibrator first:** `lowParamService.fromLedger:78` 24.8% → PASS.
|
||||
- **7 scripts converted** via `pageSafe` + named `READS` exports + a
|
||||
`require.main === module` guard (needed because the harness imports them and
|
||||
several write to `mc_test_ledger`; verified unchanged at 168 rows).
|
||||
- **RESISTED THE HELPER:** the context tables have composite PKs with no single
|
||||
unique column, so `paginate()` cannot order them. They measure 0% and are left
|
||||
on `page()`, documented in-file. Composite-key ordering is the open item.
|
||||
- **OPEN:** an unresolved intermittent in `snapshotService.test.js` (2/30 on
|
||||
branch, 0/29 at HEAD — not statistically distinguishable, assertion never
|
||||
captured). Recorded, not dismissed.
|
||||
|
||||
## Fix A2b (2026-08-09) — every read clean by construction ✅
|
||||
4,699 tests / 366 suites, web build exit 0. **Grade FROZEN; `CALIBRATION_DEPLOYED`
|
||||
still `[]`; verdicts still withdrawn; no eligibility, no settlement, no
|
||||
version-stamp, no delete. No context data altered.**
|
||||
- **34 readers, 34 PASS, 0 FAIL, 0 KEY_NOT_UNIQUE** — up from 26 readers / 1
|
||||
KEY_NOT_UNIQUE. 30 of 34 measured through the REAL reader function.
|
||||
- **safePaginate takes composite keys**; `src/utils/tableKeys.js` holds the real
|
||||
constraints from `pg_index`. All 7 context tables converted; `batter_spray`
|
||||
resolved (wrong key, not bad data).
|
||||
- **The 3 A2 legacy reads converted**; `KNOWN_LEGACY_READS` is now empty.
|
||||
- **`calibrate-hits.js` found at 24.7%** during the sweep and converted.
|
||||
- **The unordered `page()` helper is DELETED from all 15 scripts.**
|
||||
- **Write-scripts assessed, data untouched:** omission-only failure mode; read is
|
||||
0% today; historical exposure bounded at ~6.6 players / ~0.16 games.
|
||||
|
||||
## Fix A3 (2026-08-09) — eligibility can honestly count ✅
|
||||
4,723 tests / 367 suites, web build exit 0. **Grade FROZEN; `CALIBRATION_DEPLOYED`
|
||||
still `[]`; `hitsFactors`, `hitsFactorContext`, `withdrawnVerdicts`,
|
||||
`reAuditEligibility`, `snapshotService` all EMPTY DIFF. No verdict re-run, no
|
||||
calibration re-fit, no historical re-stamp.**
|
||||
- **ONE version source:** `src/config/modelVersion.js` = `engine1@2026-08-07-fullwindow`.
|
||||
Both `ledgerService` and `retentionService` import it; the twin defaults that
|
||||
silently blocked eligibility are gone.
|
||||
- **Snapshot settlement on the cron:** `snapshotSettlementService` wired into
|
||||
`snapshotScheduler` beside the ledger settle. Outcomes only, no context
|
||||
reconstruction. Settled **15,484 → 28,894**; repaired-champion **0 → 7,556**;
|
||||
0 rows settled with a null actual_value. 24 new unit tests.
|
||||
- **Eligibility (counted, not run):** 2 eligible dates against 10/10/14/14.
|
||||
All four measurements remain BLOCKED — now for the honest reason.
|
||||
|
||||
## Fix A4 (2026-08-09) — as-of-correct context ✅
|
||||
4,735 tests / 368 suites, web build exit 0, harness still 34/34 PASS.
|
||||
**Grade output UNCHANGED (proven 400/400); inputs NOT frozen (that is A5); no
|
||||
re-audit run, no verdict re-run, no calibration re-fit; `CALIBRATION_DEPLOYED`
|
||||
still `[]`.**
|
||||
- `hitsFactorContext.build(sb, {asOf})` — dated reads bounded `.lte('as_of_date')`,
|
||||
latest-within-bound, **refusal on absence** (never nearest/latest).
|
||||
- Dated statcast comes from `statcast_history` (aggregates keeps one date);
|
||||
verified identical to aggregates at the head, 1,414 rows / 0 diffs.
|
||||
- **Coverage on the two re-audit dates: spray/platoon/handedness 99% and 98.9%**,
|
||||
2 players short per date. **Defence 0% from the row** — `opponent` is NULL on
|
||||
every settled repaired-champion row; it needs the game-log join.
|
||||
- 12 new unit tests incl. the contamination lock (never a row dated after asOf).
|
||||
|
||||
## Fix A5 (2026-08-10) — input freeze, and a finding that reframes it ✅
|
||||
4,751 tests / 369 suites, web build exit 0. **Grade UNCHANGED (496/496); no
|
||||
re-audit, no verdict re-run, no calibration; `CALIBRATION_DEPLOYED` still `[]`.**
|
||||
- **FINDING: the three hits factors NEVER FIRE.** 596/596 real props skipped,
|
||||
multiplier 1 every time, because `prop.opponent` / `prop.opposing_pitcher` are
|
||||
never set by anything. They are wired and inert. Plumbing them is a MODEL
|
||||
CHANGE and needs its own order.
|
||||
- **`model_snapshots.factor_inputs` jsonb** (migration 034, APPLIED) freezes the
|
||||
raw inputs at grade time — never the multiplier, which stays recomputable via
|
||||
`factorFreeze.recompute()`. Proven: frozen == read (496/496), recompute == live
|
||||
(496/496).
|
||||
- **Forward-only.** Existing 7,556 settled rows keep NULL `factor_inputs` and
|
||||
NULL `opponent`; defence re-audit on them still needs the game-log join.
|
||||
|
||||
## Fix A6 (2026-08-10) — join keys plumbed into a SHADOW resolve ✅
|
||||
4,762 tests / 370 suites, web build exit 0. **Served grade FROZEN; no re-audit,
|
||||
no verdict reinstated, no calibration; `CALIBRATION_DEPLOYED` still `[]`.**
|
||||
- `matchupKeys` resolves opponent + opposing pitcher from `lineup_context`
|
||||
(as-of) + schedule probables. Refuses rather than guessing.
|
||||
- **Key resolution 248/308 (80.5%)** on 2026-08-09; 60 refused (no lineup row).
|
||||
- **Shadow fire: pitcher_contact 248, defense_by_direction 185,
|
||||
platoon_severity 147** — up from 0/0/0. Multiplier median 1.017, range
|
||||
0.803–1.250; 245/248 would move the served number.
|
||||
- Served p_win identical on every evaluated row; shadow recompute-check 248/248.
|
||||
|
||||
## Fix A7 (2026-08-11) — shadow accrual proven + A8 pre-registered ✅
|
||||
4,772 tests / 371 suites, web build exit 0. **Nothing turned on. Served grade
|
||||
frozen; `CALIBRATION_DEPLOYED` still `[]`; no verdict reinstated; gate NOT run.**
|
||||
- **Accrual is automatic on the scheduled pass** — proven end-to-end through the
|
||||
real `gradeAndCacheSlate → onGraded → rowsFromSides` chain (10 new tests),
|
||||
plus assertions that `runSnapshot` builds the key index pre-grade and degrades
|
||||
safely.
|
||||
- **Complete (would_fire, outcome) pairs today: 0.** `factor_inputs` is
|
||||
0/122,276 — A5/A6 are not deployed. Accrual begins at deploy.
|
||||
- **`specs/a8-shadow-factor-gate.md`** pre-registers the two-part gate (movement
|
||||
AND Brier, cluster-resampled, cumulative-Bonferroni, new-test α) with named
|
||||
fallbacks. Not run.
|
||||
- **Binding constraint is CLUSTERS: ~10.5 settled games/slate ⇒ `MIN_CLUSTERS`
|
||||
40 in ~4 slates**, while `MIN_N` 500 is met in 1–2. A8 runnable ≈5 slates
|
||||
after deploy.
|
||||
|
||||
### Next
|
||||
- **A8 — run the pre-registered gate** once ≈5 slates have accrued.
|
||||
- **A9 — turn the factors on LIVE**, only if A8 returns PROVES, and with its own
|
||||
before/after on the served board.
|
||||
- **(superseded) A7 — turn the factors on LIVE.** This is the biggest served-grade change in
|
||||
VYNDR's history: ~80% of hits props would move, median +1.7%, tails ±20-25%.
|
||||
Gate it on (a) accrued shadow evidence, (b) a re-proof of the withdrawn
|
||||
verdicts on rows where the factors actually fire.
|
||||
- **(superseded) A6 — plumb `opponent` + `opposing_pitcher` into the graded prop.** This is a
|
||||
MODEL CHANGE (grades will move): measure the delta against the frozen
|
||||
`available` fields first, then decide.
|
||||
- **(superseded) A5 — freeze the factor inputs onto the graded row** so a re-audit does not
|
||||
depend on context tables at all (and so defence stops needing a game-log join).
|
||||
- **(done in A4) as-of-correct context** (`hitsFactorContext.build` has no as-of cutoff;
|
||||
it also orders context walks on a non-unique prefix). **The re-audit RUN is
|
||||
gated on this, not on clean reads.**
|
||||
- **A2c (was A2b)** — composite-key ordering in `safePaginate` so the context tables can be
|
||||
made safe by construction rather than clean by luck; then convert the 3 known
|
||||
legacy reads and register `backfill-context` / `reconstruct-game-environment`.
|
||||
- **A2c** — chase the `snapshotService.test.js` intermittent to an assertion.
|
||||
- **(superseded)** roll `safePaginate` across the remaining FAIL readers, starting with
|
||||
`lowParamService.fromLedger:78` (the PRIMARY calibrator, 24.8%). Each must move
|
||||
FAIL → PASS on a measured re-run through `readerRows`, never on the presence of
|
||||
an `.order()` clause. Remaining: `champion-ablation:173` 33.6%,
|
||||
`proven-status:76` 32.6%, `prove-hit-factors:131/136`, `prove-tb-factors:149/154`,
|
||||
`cluster-prove:337`, `tb-solo-and-interactions:260`, `build-grade-bands:49/54`.
|
||||
- **A3** — decide what to do about `hitsFactors.js` serving two withdrawn factors
|
||||
plus `platoon_severity`, which never passed the gate.
|
||||
- **Re-certification of calibration is a separate gated order** on the accrual
|
||||
clock — not unlocked by A1.
|
||||
|
||||
## Session 94 (2026-08-04) — Causally-correct platoon + park inputs ✅
|
||||
4,307 tests / 344 suites green, build exit 0. Counter + frozen clusters
|
||||
|
||||
Reference in New Issue
Block a user