Files
vyndr/specs/read-integrity-harness.md
builtbykev f61ec6b391 Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 22:49:56 -04:00

10 KiB
Raw Permalink Blame History

SPEC — Read-Integrity Harness + Verdict Withdrawal (Fix A0)

Status: built 2026-08-09. Measure-only instrument. Fixes no reader.


1. WHY

The 2026-08-09 probe established, by measurement, that the shared page() pagination used by the learning loop — .range(from, from+PAGE-1) with no ORDER BY — returns the correct row COUNT and the wrong ROWS.

Measured on production, same query, minutes apart:

reader true count duplicates missing corruption
hits gate ledger read (prove-hit-factors.js:136) 2,490 412 → 617 same 16.5% → 24.8%
archetype join (prove-hit-factors.js:131) 10,038 2,454 2,454 24.4%
calibrationService.fromLedger:110 (LIVE) 2,490 617 617 24.8%
challenger-scoreboard.js:112 11,485 0 0 0%

Two facts drive this spec:

  1. Exposure cannot be reasoned from source. challenger-scoreboard walks 12 pages of a live table and is clean; fromLedger walks 3 pages of the same table and is a quarter corrupt. The difference is the query PLAN, which shifts with filters, table statistics and concurrent writes. Nothing in the code predicts it.
  2. Duplicates are worse than a random subsample. They inflate factorGate.movement.n, which is the number checked against MIN_N = 500 (factorGate.js:194), AND they corrupt the leave-one-out per-player baseline that is the gate's null hypothesis. Both the test statistic and the thing it is tested against move.

So: one instrument, measured per reader, per run. Not fifteen hand-written checks, and never an inference from source text.


2. THE META-SCAR RULE (load-bearing)

A reader PASSES only on measured set identity against an ordered control. The presence of an ORDER BY clause is NEVER evidence of correctness.

A guard that reads its own source text for the string order( would be a guard verifying its own intention rather than its effect — the exact failure class this codebase has corrected repeatedly. tests/unit/readIntegrity.test.js therefore asserts that a walk which HAS a stable order but still returns duplicates is reported FAIL.

Corollary: the control must validate itself. If the ordered walk's distinct count does not equal the server's count(*), the control is untrustworthy and the result is CONTROL_INVALID — never PASS. An instrument that cannot tell a clean reader from a broken control is not an instrument.


3. ENDPOINTS / SURFACE

None. This is a developer instrument with no HTTP surface, no cron, no product behaviour. It reads; it writes nothing, anywhere.

Modules

file role
src/utils/readIntegrity.js PURE + injectable core. No Supabase import.
scripts/read-integrity.js CLI runner + the declarative registry of reader specs
src/services/model/withdrawnVerdicts.js append-only withdrawal record

Data shapes

A reader spec is declarative data, so the registry is auditable:

{
  id: 'prove-hit-factors:136',
  source: 'scripts/prove-hit-factors.js:136',
  table: 'ledger_entries',
  select: 'id, player_key, stat, game_date',
  key: ['id'],                       // identity columns for dedupe
  order: 'id',                       // the control's stable order
  filters: [
    ['eq', 'sport', 'mlb'],
    ['is', 'user_id', null],
    ['in', 'outcome', ['hit', 'miss']],
    ['not', 'p_win', 'is', null],
  ],
}

A result:

{
  id, source, table,
  exact_count,                       // server count(*), the ground truth
  pages,
  unordered: { rows, distinct, duplicates, missing_vs_control, extra_vs_control },
  ordered:   { rows, distinct, duplicates },
  corruption_pct,                    // missing / exact_count
  key_unique,                        // ordered.distinct === ordered.rows
  control_valid,                     // ordered.distinct === exact_count
  verdict,                           // PASS | FAIL | CONTROL_INVALID | KEY_NOT_UNIQUE | ERROR
  reason,
}

4. ACCEPTANCE CRITERIA

AC1 — analyze() returns PASS only when duplicates === 0 AND missing_vs_control === 0 AND control_valid. Verified by unit test.

AC2 — An ordered walk that still returns duplicates reports FAIL. (The meta-scar guard: a clause is not a result.)

AC3 — ordered.distinct !== exact_count reports CONTROL_INVALID, never PASS.

AC4 — A non-unique key reports KEY_NOT_UNIQUE rather than counting legitimate data duplicates as walk corruption. (batter_spray keyed on player_key|as_of_date has 6 genuine duplicate keys; that is a data fact, not a pagination fault.)

AC5 — Reproduces the probe on live data: the hits ledger read reports corruption in the 16–25% band; challenger-scoreboard:112 reports 0%. If the harness cannot reproduce both, it is not trustworthy and must not be used.

AC6 — walk() is generic over a fetchPage(from, to) function, so the unit tests never touch a network.

AC7 — The harness performs zero writes. No insert, update, upsert, delete, or DDL anywhere in readIntegrity.js or read-integrity.js.


5. TEST PLAN

Unit (tests/unit/readIntegrity.test.js), all with fake page-fetchers:

  • clean walk → PASS
  • walk whose last page re-emits earlier rows (the measured production shape) → FAIL with the exact duplicate/missing counts
  • ordered-but-duplicating walk → FAIL (AC2, the meta-scar)
  • ordered walk short of exact_count → CONTROL_INVALID (AC3)
  • non-unique key → KEY_NOT_UNIQUE (AC4)
  • walk() stops on a short page and on an empty page
  • applyFilters chains eq/is/in/not/lt in order onto a fake builder
  • a query error propagates as ERROR, never as end-of-data (the fromLedger:117 failure class — an error swallowed as "no more rows")

Live (scripts/read-integrity.js): AC5 above, run against prod, read-only.


6. VERDICT WITHDRAWAL — WHAT IT IS AND IS NOT

Where standing verdicts actually live

There is no verdict table. featureRegistry.statVerdicts is an in-memory Map (featureRegistry.js:145) that resets every process; mc_test_ledger stores hypothesis COUNTS for the Bonferroni denominator, not verdicts. The standing record is the specs/ markdown plus CLAUDE.md.

So the withdrawal is appended as a committed record (src/services/model/withdrawnVerdicts.js + this spec), never by editing or deleting a prior finding. The original verdicts stay exactly as recorded; the withdrawal sits beside them with its cause.

What is withdrawn

factor stat recorded verdict now
defense_by_direction hits PROVES, n=528, Brier −0.0034, CI [−0.0059,−0.0009] @99.9% WITHDRAWN_PENDING_REAUDIT
pitcher_contact_profile hits PROVES, n=741, Brier −0.0066, CI [−0.0114,−0.0016] @99.9% WITHDRAWN_PENDING_REAUDIT

Cause, recorded verbatim on both: drawn through an unordered-pagination reader; 16.5–24.8% set corruption measured on the identical query 2026-08-09; distinct-n below the 500 floor for defense_by_direction; not retrospectively recoverable.

defense_by_direction additionally carries the eligibility note: at the measured ~20% duplication its distinct n is ≈422 against MIN_N = 500, so it was likely never eligible for adjudication at all. This is arithmetic on the recorded n, not a re-measurement.

What withdrawal does NOT do

It does not stop the factors moving the served forecast. hitsFactors.js hardcodes its three factors and consults no registry, so withdrawal is a record of what we may CLAIM, not a change to what we SERVE. Disarming or re-proving is a later order with its own blast radius.

Reinstatement

A withdrawn verdict returns only by re-running its gate through readers that this harness reports PASS for, on the day of the run. Nothing is reinstated by argument.


7. THE FIX — safePaginate (Fix A1, 2026-08-09)

src/utils/safePaginate.js is the one way this codebase walks a paginated PostgREST read. It closes both measured defects in a single call:

const rows = await paginate(
  () => sb.from('ledger_entries').select('id, …').eq('sport', 'mlb') /* … */,
  { key: 'id', pageSize: 1000, label: 'caller(name)' },
);

Requirements it enforces

  1. Stable ORDER BY on a UNIQUE key. An order on a non-unique column is NOT a fix — ties may be returned in any order, so pages still overlap. Default id.
  2. Uniqueness verified at runtime, not promised. A repeated key mid-walk means either a non-unique key or a scrambled read; both make the row set unusable, so it THROWS rather than returning it. The harness PASS is confirmation; this is the guard that runs every time.
  3. A query error THROWS. Never end-of-data. paginate takes a query factory (builders are single-use) and reuses readIntegrity.walk, which is already tested for short-page / empty-page / error / runaway behaviour — it does not hand-roll a second walk.

Refusal is not a swallow. A caller that responds to a throw by producing nothing (no calibrator, no map) is refusing honestly. A caller that continues with partial rows is repeating the defect. null from fromLedger means "not enough settled history"; a throw means "the read failed". Keeping those distinct is the point.

Measuring a FIXED reader

A registry spec that restates the corrected query would verify the restatement, and a fixed: true flag would verify a comment. So a fixed reader supplies readerRows instead — a call into the real production function, whose actual output is compared to the ordered control. arm_a on the result reports which was used (real_reader_function vs query_as_written).

AC7 — A1 acceptance (met 2026-08-09)

  • calibrationService.fromLedger:110 — FAIL 24.8% → PASS (0 dup / 0 missing, 2,490 of 2,490), measured through loadSettledRows.
  • lowParamService.fromLedger:78 — the byte-identical twin, deliberately NOT fixed in A1, still measures 24.8%. The harness kept its teeth; the PASS is not an artefact of editing the registry.