Files
vyndr/specs/read-integrity-harness.md
builtbykev f61ec6b391 Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 22:49:56 -04:00

249 lines
10 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SPEC — Read-Integrity Harness + Verdict Withdrawal (Fix A0)
**Status:** built 2026-08-09. Measure-only instrument. Fixes no reader.
---
## 1. WHY
The 2026-08-09 probe established, by measurement, that the shared `page()`
pagination used by the learning loop — `.range(from, from+PAGE-1)` with **no
`ORDER BY`** — returns the correct row COUNT and the wrong ROWS.
Measured on production, same query, minutes apart:
| reader | true count | duplicates | missing | corruption |
|---|---|---|---|---|
| hits gate ledger read (`prove-hit-factors.js:136`) | 2,490 | 412 → 617 | same | **16.5% → 24.8%** |
| archetype join (`prove-hit-factors.js:131`) | 10,038 | 2,454 | 2,454 | **24.4%** |
| `calibrationService.fromLedger:110` (LIVE) | 2,490 | 617 | 617 | **24.8%** |
| `challenger-scoreboard.js:112` | 11,485 | 0 | 0 | **0%** |
Two facts drive this spec:
1. **Exposure cannot be reasoned from source.** `challenger-scoreboard` walks 12
pages of a live table and is clean; `fromLedger` walks 3 pages of the *same*
table and is a quarter corrupt. The difference is the query PLAN, which shifts
with filters, table statistics and concurrent writes. Nothing in the code
predicts it.
2. **Duplicates are worse than a random subsample.** They inflate
`factorGate.movement.n`, which is the number checked against `MIN_N = 500`
(`factorGate.js:194`), AND they corrupt the leave-one-out per-player baseline
that is the gate's null hypothesis. Both the test statistic and the thing it is
tested against move.
So: one instrument, measured per reader, per run. Not fifteen hand-written checks,
and never an inference from source text.
---
## 2. THE META-SCAR RULE (load-bearing)
> **A reader PASSES only on measured set identity against an ordered control.
> The presence of an `ORDER BY` clause is NEVER evidence of correctness.**
A guard that reads its own source text for the string `order(` would be a guard
verifying its own intention rather than its effect — the exact failure class this
codebase has corrected repeatedly. `tests/unit/readIntegrity.test.js` therefore
asserts that a walk which HAS a stable order but still returns duplicates is
reported **FAIL**.
Corollary: **the control must validate itself.** If the ordered walk's distinct
count does not equal the server's `count(*)`, the control is untrustworthy and the
result is `CONTROL_INVALID` — never `PASS`. An instrument that cannot tell a clean
reader from a broken control is not an instrument.
---
## 3. ENDPOINTS / SURFACE
None. This is a developer instrument with no HTTP surface, no cron, no product
behaviour. It reads; it writes nothing, anywhere.
### Modules
| file | role |
|---|---|
| `src/utils/readIntegrity.js` | PURE + injectable core. No Supabase import. |
| `scripts/read-integrity.js` | CLI runner + the declarative registry of reader specs |
| `src/services/model/withdrawnVerdicts.js` | append-only withdrawal record |
### Data shapes
A **reader spec** is declarative data, so the registry is auditable:
```js
{
id: 'prove-hit-factors:136',
source: 'scripts/prove-hit-factors.js:136',
table: 'ledger_entries',
select: 'id, player_key, stat, game_date',
key: ['id'], // identity columns for dedupe
order: 'id', // the control's stable order
filters: [
['eq', 'sport', 'mlb'],
['is', 'user_id', null],
['in', 'outcome', ['hit', 'miss']],
['not', 'p_win', 'is', null],
],
}
```
A **result**:
```js
{
id, source, table,
exact_count, // server count(*), the ground truth
pages,
unordered: { rows, distinct, duplicates, missing_vs_control, extra_vs_control },
ordered: { rows, distinct, duplicates },
corruption_pct, // missing / exact_count
key_unique, // ordered.distinct === ordered.rows
control_valid, // ordered.distinct === exact_count
verdict, // PASS | FAIL | CONTROL_INVALID | KEY_NOT_UNIQUE | ERROR
reason,
}
```
---
## 4. ACCEPTANCE CRITERIA
**AC1** — `analyze()` returns `PASS` only when `duplicates === 0 AND
missing_vs_control === 0 AND control_valid`. Verified by unit test.
**AC2** — An ordered walk that still returns duplicates reports `FAIL`. (The
meta-scar guard: a clause is not a result.)
**AC3** — `ordered.distinct !== exact_count` reports `CONTROL_INVALID`, never
`PASS`.
**AC4** — A non-unique key reports `KEY_NOT_UNIQUE` rather than counting
legitimate data duplicates as walk corruption. (`batter_spray` keyed on
`player_key|as_of_date` has 6 genuine duplicate keys; that is a data fact, not a
pagination fault.)
**AC5** — Reproduces the probe on live data: the hits ledger read reports
corruption in the **16–25%** band; `challenger-scoreboard:112` reports **0%**.
If the harness cannot reproduce both, it is not trustworthy and must not be used.
**AC6** — `walk()` is generic over a `fetchPage(from, to)` function, so the unit
tests never touch a network.
**AC7** — The harness performs **zero writes**. No insert, update, upsert, delete,
or DDL anywhere in `readIntegrity.js` or `read-integrity.js`.
---
## 5. TEST PLAN
Unit (`tests/unit/readIntegrity.test.js`), all with fake page-fetchers:
- clean walk → `PASS`
- walk whose last page re-emits earlier rows (the measured production shape) →
`FAIL` with the exact duplicate/missing counts
- **ordered-but-duplicating walk → `FAIL`** (AC2, the meta-scar)
- ordered walk short of `exact_count` → `CONTROL_INVALID` (AC3)
- non-unique key → `KEY_NOT_UNIQUE` (AC4)
- `walk()` stops on a short page and on an empty page
- `applyFilters` chains `eq/is/in/not/lt` in order onto a fake builder
- a query error propagates as `ERROR`, never as end-of-data (the
`fromLedger:117` failure class — an error swallowed as "no more rows")
Live (`scripts/read-integrity.js`): AC5 above, run against prod, read-only.
---
## 6. VERDICT WITHDRAWAL — WHAT IT IS AND IS NOT
### Where standing verdicts actually live
There is **no verdict table**. `featureRegistry.statVerdicts` is an in-memory
`Map` (`featureRegistry.js:145`) that resets every process; `mc_test_ledger`
stores hypothesis COUNTS for the Bonferroni denominator, not verdicts. The
standing record is the `specs/` markdown plus `CLAUDE.md`.
So the withdrawal is **appended as a committed record**
(`src/services/model/withdrawnVerdicts.js` + this spec), never by editing or
deleting a prior finding. The original verdicts stay exactly as recorded; the
withdrawal sits beside them with its cause.
### What is withdrawn
| factor | stat | recorded verdict | now |
|---|---|---|---|
| `defense_by_direction` | hits | PROVES, n=528, Brier −0.0034, CI [−0.0059,−0.0009] @99.9% | **WITHDRAWN_PENDING_REAUDIT** |
| `pitcher_contact_profile` | hits | PROVES, n=741, Brier −0.0066, CI [−0.0114,−0.0016] @99.9% | **WITHDRAWN_PENDING_REAUDIT** |
Cause, recorded verbatim on both: *drawn through an unordered-pagination reader;
16.5–24.8% set corruption measured on the identical query 2026-08-09; distinct-n
below the 500 floor for `defense_by_direction`; not retrospectively recoverable.*
`defense_by_direction` additionally carries the eligibility note: at the measured
~20% duplication its distinct n is ≈422 against `MIN_N = 500`, so it was likely
**never eligible for adjudication at all**. This is arithmetic on the recorded n,
not a re-measurement.
### What withdrawal does NOT do
It does **not** stop the factors moving the served forecast.
`hitsFactors.js` hardcodes its three factors and consults no registry, so
withdrawal is a record of what we may CLAIM, not a change to what we SERVE.
Disarming or re-proving is a later order with its own blast radius.
### Reinstatement
A withdrawn verdict returns only by re-running its gate through readers that this
harness reports `PASS` for, on the day of the run. Nothing is reinstated by
argument.
---
## 7. THE FIX — `safePaginate` (Fix A1, 2026-08-09)
`src/utils/safePaginate.js` is the one way this codebase walks a paginated
PostgREST read. It closes both measured defects in a single call:
```js
const rows = await paginate(
() => sb.from('ledger_entries').select('id, …').eq('sport', 'mlb') /* … */,
{ key: 'id', pageSize: 1000, label: 'caller(name)' },
);
```
**Requirements it enforces**
1. **Stable ORDER BY on a UNIQUE key.** An order on a non-unique column is NOT a
fix — ties may be returned in any order, so pages still overlap. Default `id`.
2. **Uniqueness verified at runtime, not promised.** A repeated key mid-walk
means either a non-unique key or a scrambled read; both make the row set
unusable, so it THROWS rather than returning it. The harness PASS is
confirmation; this is the guard that runs every time.
3. **A query error THROWS.** Never end-of-data. `paginate` takes a query
*factory* (builders are single-use) and reuses `readIntegrity.walk`, which is
already tested for short-page / empty-page / error / runaway behaviour — it
does not hand-roll a second walk.
**Refusal is not a swallow.** A caller that responds to a throw by producing
nothing (no calibrator, no map) is refusing honestly. A caller that continues
with partial rows is repeating the defect. `null` from `fromLedger` means "not
enough settled history"; a throw means "the read failed". Keeping those distinct
is the point.
### Measuring a FIXED reader
A registry spec that restates the corrected query would verify the restatement,
and a `fixed: true` flag would verify a comment. So a fixed reader supplies
`readerRows` instead — a call into the **real production function**, whose
actual output is compared to the ordered control. `arm_a` on the result reports
which was used (`real_reader_function` vs `query_as_written`).
### AC7 — A1 acceptance (met 2026-08-09)
- `calibrationService.fromLedger:110` — **FAIL 24.8% → PASS (0 dup / 0 missing,
2,490 of 2,490)**, measured through `loadSettledRows`.
- `lowParamService.fromLedger:78` — the byte-identical twin, deliberately NOT
fixed in A1, still measures **24.8%**. The harness kept its teeth; the PASS is
not an artefact of editing the registry.