Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move. A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS: 410-617 of 2,490 duplicated with an equal number never returned, while rows.length matched the server exactly. safePaginate orders on a real unique key, verifies the tuple at runtime, and THROWS on a query error instead of treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn through that reader, and defense_by_direction's distinct-n was likely below the gate floor all along. A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from pg_index (the context tables are dated-composite and had no single unique column). The unordered helper is deleted, not parked. A3 — ledgerService and retentionService defaulted the SAME env var to DIFFERENT versions, so no ledger row ever carried the marker eligibility requires. One source now. model_snapshots settlement moved onto the cron: 15,484 -> 28,894 settled, repaired-champion 0 -> 7,556. A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no row at-or-before the date means the factor does not apply, never the nearest row. Live path unchanged, proven 400/400 on real rows. A5 — factor_inputs freezes what the factor READ, never the multiplier, so an audit can recompute and check. It also recorded the finding: the three hits factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by the resolver and written by nothing. A6/A7 — matchupKeys resolves those keys from the posted lineup plus the schedule's probable pitchers, and fires the factors into a SHADOW freeze: 248 fires on 308 props, 245 of which would move the grade. The served forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test that decides whether they ever go live. Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity harness 34/34. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,140 @@
|
||||
# PRE-REGISTRATION — A8: do the shadow-fired hits factors improve the forecast?
|
||||
|
||||
**Status: PRE-REGISTERED, NOT RUN.** Written 2026-08-11, before any accrual
|
||||
exists, so the decision rule cannot be chosen after seeing the answer.
|
||||
|
||||
---
|
||||
|
||||
## 1. WHAT IS BEING TESTED, AND WHY IT IS A NEW TEST
|
||||
|
||||
A5 measured that all three hits factors were skipped on **596/596** real graded
|
||||
props — `prop.opponent` and `prop.opposing_pitcher` were never set by anything,
|
||||
so `positionOaa`, `throws` and `pitcherHardHit` resolved on **zero** rows. A6
|
||||
resolved those keys and fired the factors in **shadow**: 248 fires on 308 props,
|
||||
median would-be multiplier 1.017, range 0.803–1.250.
|
||||
|
||||
So `pitcher_contact_profile` and `defense_by_direction` were never "proven and
|
||||
then disconnected." They were proven on an **audit reconstruction that the
|
||||
production path never ran**, and then withdrawn (A0) because the reader that
|
||||
produced that proof was 16–25% corrupt.
|
||||
|
||||
**This is therefore a NEW hypothesis, not a re-test.** It counts as a fresh test
|
||||
against the cumulative Bonferroni denominator (`testLedger`). The S87 rule that
|
||||
re-asking the same question on more data is not a new shot on goal does not apply:
|
||||
the question is different. Previously — "does this factor improve a reconstructed
|
||||
audit?" Now — "does the multiplier this factor actually produces, on the rows
|
||||
where it actually fires, improve the served forecast out of sample?"
|
||||
|
||||
Three factors × one stat = **3 new hypotheses** to record before the run.
|
||||
|
||||
---
|
||||
|
||||
## 2. THE UNIT OF EVIDENCE
|
||||
|
||||
One testable pair = a `model_snapshots` row that carries **both**:
|
||||
|
||||
- `factor_inputs.would_fire` — written at grade time, pre-game, never applied
|
||||
- `outcome` — settled from a box score by `snapshotSettlementService`
|
||||
|
||||
A `would_fire` with no settled outcome is **not yet evidence** and must not be
|
||||
counted toward N.
|
||||
|
||||
Rows are eligible only if they also satisfy the standing gates:
|
||||
`model_version = engine1@2026-08-07-fullwindow` (A3) and
|
||||
`quarantine_reason` not `nontakeable_book%`.
|
||||
|
||||
---
|
||||
|
||||
## 3. THE TEST — TWO PARTS, BOTH REQUIRED
|
||||
|
||||
Run per factor, on **rows where that factor fired** (`would_fire.per_factor`
|
||||
contains it). Scoring a factor on rows it did not touch dilutes any real effect
|
||||
toward zero (S77).
|
||||
|
||||
**Baseline** = the served forecast, i.e. `p_win` as actually graded (the factors
|
||||
were off).
|
||||
**Conditioned** = `clamp(p_win × would_fire.multiplier, 0.01, 0.99)` — exactly
|
||||
what would have been served had the factor been live.
|
||||
|
||||
Adjudicated by `factorGate.adjudicate()`:
|
||||
|
||||
1. **MOVEMENT** — mean `|Δp|` ≥ `MIN_MOVEMENT` (0.01). Movement alone is
|
||||
**THEATER** and is named as such.
|
||||
2. **IMPROVEMENT** — paired bootstrap on **Brier**, cluster-resampled, CI
|
||||
excluding zero at `1 − 0.05/tests` where `tests` is the programme-lifetime
|
||||
cumulative count from `testLedger`.
|
||||
|
||||
`cluster = game_id`. Verified: for all three factors the treatment entity
|
||||
(`player|opponent`, `starter_id`, `player_key`) is more numerous than the games,
|
||||
so `factorGate`'s coarser-of rule resolves to the game. **`MIN_CLUSTERS = 40`
|
||||
distinct settled games is the binding constraint**, not `MIN_N`.
|
||||
|
||||
### Pre-registered outcomes
|
||||
|
||||
| verdict | meaning | consequence |
|
||||
|---|---|---|
|
||||
| `PROVES` | moves AND corrected CI below zero | candidate for A9 live turn-on — still a separate order |
|
||||
| `NOT_PROVEN_AT_CORRECTED_BAR` | point estimate improves, CI spans zero | keep accruing; do NOT turn on |
|
||||
| `THEATER` | moves, Brier delta ≥ 0 | **do not turn on, ever, at this multiplier** |
|
||||
| `INERT` | mean shift < 0.01 | wiring is pointless; drop the factor |
|
||||
| `CANDIDATE_PENDING_SAMPLE` | below N or clusters | keep accruing |
|
||||
|
||||
**Pre-registered fallback:** if a factor returns `THEATER`, that is a real result
|
||||
and it must be recorded as such — a factor that moves ~80% of the board and does
|
||||
not improve Brier is the arch-v1 failure repeating, and the correct action is to
|
||||
leave it off, not to re-tune the multiplier until it passes.
|
||||
|
||||
**A `PROVES` does not turn anything on.** It makes the case for A9, which needs
|
||||
its own before/after on the served board.
|
||||
|
||||
---
|
||||
|
||||
## 4. WHAT WOULD INVALIDATE THE RUN
|
||||
|
||||
- **Fewer than 40 distinct settled games** for that factor → refuse, report N.
|
||||
- **`would_fire` not re-derivable** from `would_fire.shadow_inputs` on any row
|
||||
(A6's recompute check) → the evidence is not trustworthy; stop.
|
||||
- **Any row whose `would_fire` was written after first pitch** — settlement
|
||||
already refuses post-hoc rows, and the same rule applies here.
|
||||
- **Keys resolved from anything other than the pre-game lineup + probable
|
||||
pitcher.** A retrospectively-derived opponent would make this a reconstruction
|
||||
again, which is the whole thing A5/A6 exist to avoid.
|
||||
|
||||
---
|
||||
|
||||
## 5. WHEN IT CAN RUN
|
||||
|
||||
Measured accrual (settled MLB hits, repaired-champion rows):
|
||||
|
||||
| | 2026-08-07 | 2026-08-08 |
|
||||
|---|---|---|
|
||||
| settled hits rows | 758 | 1,096 |
|
||||
| distinct settled players | 208 | 188 |
|
||||
| **distinct settled games** | **11** | **10** |
|
||||
|
||||
Shadow fire rates (A6, 2026-08-09): `pitcher_contact_profile` 80.5%,
|
||||
`defense_by_direction` 60.1%, `platoon_severity` 47.7%.
|
||||
|
||||
At ~925 settled hits rows and **~10.5 distinct settled games per slate**:
|
||||
|
||||
| factor | rows/slate | slates to `MIN_N` 500 | **slates to `MIN_CLUSTERS` 40** | binding |
|
||||
|---|---|---|---|---|
|
||||
| `pitcher_contact_profile` | ~745 | 1 | **~4** | clusters |
|
||||
| `defense_by_direction` | ~556 | 1 | **~4** | clusters |
|
||||
| `platoon_severity` | ~441 | 2 | **~4** | clusters |
|
||||
|
||||
**Accrual starts at DEPLOY, not today** — 0 rows carry `factor_inputs` now. Add
|
||||
the settlement lag (a slate settles the following day), so:
|
||||
|
||||
> **A8 is runnable ≈ 5 slates after A6/A7 deploy** — about **5 days**, for all
|
||||
> three factors together. Nothing accrues until the deploy lands.
|
||||
|
||||
---
|
||||
|
||||
## 6. WHAT THIS DOES NOT DECIDE
|
||||
|
||||
It does not reinstate `defense_by_direction` or `pitcher_contact_profile`. Those
|
||||
withdrawals stand on their own cause (drawn through a corrupt reader). A `PROVES`
|
||||
here is evidence about **the live multiplier on rows where it fires** — a
|
||||
different claim from the original verdicts, and it should be recorded as a new
|
||||
finding rather than as a reinstatement.
|
||||
@@ -0,0 +1,248 @@
|
||||
# SPEC — Read-Integrity Harness + Verdict Withdrawal (Fix A0)
|
||||
|
||||
**Status:** built 2026-08-09. Measure-only instrument. Fixes no reader.
|
||||
|
||||
---
|
||||
|
||||
## 1. WHY
|
||||
|
||||
The 2026-08-09 probe established, by measurement, that the shared `page()`
|
||||
pagination used by the learning loop — `.range(from, from+PAGE-1)` with **no
|
||||
`ORDER BY`** — returns the correct row COUNT and the wrong ROWS.
|
||||
|
||||
Measured on production, same query, minutes apart:
|
||||
|
||||
| reader | true count | duplicates | missing | corruption |
|
||||
|---|---|---|---|---|
|
||||
| hits gate ledger read (`prove-hit-factors.js:136`) | 2,490 | 412 → 617 | same | **16.5% → 24.8%** |
|
||||
| archetype join (`prove-hit-factors.js:131`) | 10,038 | 2,454 | 2,454 | **24.4%** |
|
||||
| `calibrationService.fromLedger:110` (LIVE) | 2,490 | 617 | 617 | **24.8%** |
|
||||
| `challenger-scoreboard.js:112` | 11,485 | 0 | 0 | **0%** |
|
||||
|
||||
Two facts drive this spec:
|
||||
|
||||
1. **Exposure cannot be reasoned from source.** `challenger-scoreboard` walks 12
|
||||
pages of a live table and is clean; `fromLedger` walks 3 pages of the *same*
|
||||
table and is a quarter corrupt. The difference is the query PLAN, which shifts
|
||||
with filters, table statistics and concurrent writes. Nothing in the code
|
||||
predicts it.
|
||||
2. **Duplicates are worse than a random subsample.** They inflate
|
||||
`factorGate.movement.n`, which is the number checked against `MIN_N = 500`
|
||||
(`factorGate.js:194`), AND they corrupt the leave-one-out per-player baseline
|
||||
that is the gate's null hypothesis. Both the test statistic and the thing it is
|
||||
tested against move.
|
||||
|
||||
So: one instrument, measured per reader, per run. Not fifteen hand-written checks,
|
||||
and never an inference from source text.
|
||||
|
||||
---
|
||||
|
||||
## 2. THE META-SCAR RULE (load-bearing)
|
||||
|
||||
> **A reader PASSES only on measured set identity against an ordered control.
|
||||
> The presence of an `ORDER BY` clause is NEVER evidence of correctness.**
|
||||
|
||||
A guard that reads its own source text for the string `order(` would be a guard
|
||||
verifying its own intention rather than its effect — the exact failure class this
|
||||
codebase has corrected repeatedly. `tests/unit/readIntegrity.test.js` therefore
|
||||
asserts that a walk which HAS a stable order but still returns duplicates is
|
||||
reported **FAIL**.
|
||||
|
||||
Corollary: **the control must validate itself.** If the ordered walk's distinct
|
||||
count does not equal the server's `count(*)`, the control is untrustworthy and the
|
||||
result is `CONTROL_INVALID` — never `PASS`. An instrument that cannot tell a clean
|
||||
reader from a broken control is not an instrument.
|
||||
|
||||
---
|
||||
|
||||
## 3. ENDPOINTS / SURFACE
|
||||
|
||||
None. This is a developer instrument with no HTTP surface, no cron, no product
|
||||
behaviour. It reads; it writes nothing, anywhere.
|
||||
|
||||
### Modules
|
||||
|
||||
| file | role |
|
||||
|---|---|
|
||||
| `src/utils/readIntegrity.js` | PURE + injectable core. No Supabase import. |
|
||||
| `scripts/read-integrity.js` | CLI runner + the declarative registry of reader specs |
|
||||
| `src/services/model/withdrawnVerdicts.js` | append-only withdrawal record |
|
||||
|
||||
### Data shapes
|
||||
|
||||
A **reader spec** is declarative data, so the registry is auditable:
|
||||
|
||||
```js
|
||||
{
|
||||
id: 'prove-hit-factors:136',
|
||||
source: 'scripts/prove-hit-factors.js:136',
|
||||
table: 'ledger_entries',
|
||||
select: 'id, player_key, stat, game_date',
|
||||
key: ['id'], // identity columns for dedupe
|
||||
order: 'id', // the control's stable order
|
||||
filters: [
|
||||
['eq', 'sport', 'mlb'],
|
||||
['is', 'user_id', null],
|
||||
['in', 'outcome', ['hit', 'miss']],
|
||||
['not', 'p_win', 'is', null],
|
||||
],
|
||||
}
|
||||
```
|
||||
|
||||
A **result**:
|
||||
|
||||
```js
|
||||
{
|
||||
id, source, table,
|
||||
exact_count, // server count(*), the ground truth
|
||||
pages,
|
||||
unordered: { rows, distinct, duplicates, missing_vs_control, extra_vs_control },
|
||||
ordered: { rows, distinct, duplicates },
|
||||
corruption_pct, // missing / exact_count
|
||||
key_unique, // ordered.distinct === ordered.rows
|
||||
control_valid, // ordered.distinct === exact_count
|
||||
verdict, // PASS | FAIL | CONTROL_INVALID | KEY_NOT_UNIQUE | ERROR
|
||||
reason,
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. ACCEPTANCE CRITERIA
|
||||
|
||||
**AC1** — `analyze()` returns `PASS` only when `duplicates === 0 AND
|
||||
missing_vs_control === 0 AND control_valid`. Verified by unit test.
|
||||
|
||||
**AC2** — An ordered walk that still returns duplicates reports `FAIL`. (The
|
||||
meta-scar guard: a clause is not a result.)
|
||||
|
||||
**AC3** — `ordered.distinct !== exact_count` reports `CONTROL_INVALID`, never
|
||||
`PASS`.
|
||||
|
||||
**AC4** — A non-unique key reports `KEY_NOT_UNIQUE` rather than counting
|
||||
legitimate data duplicates as walk corruption. (`batter_spray` keyed on
|
||||
`player_key|as_of_date` has 6 genuine duplicate keys; that is a data fact, not a
|
||||
pagination fault.)
|
||||
|
||||
**AC5** — Reproduces the probe on live data: the hits ledger read reports
|
||||
corruption in the **16–25%** band; `challenger-scoreboard:112` reports **0%**.
|
||||
If the harness cannot reproduce both, it is not trustworthy and must not be used.
|
||||
|
||||
**AC6** — `walk()` is generic over a `fetchPage(from, to)` function, so the unit
|
||||
tests never touch a network.
|
||||
|
||||
**AC7** — The harness performs **zero writes**. No insert, update, upsert, delete,
|
||||
or DDL anywhere in `readIntegrity.js` or `read-integrity.js`.
|
||||
|
||||
---
|
||||
|
||||
## 5. TEST PLAN
|
||||
|
||||
Unit (`tests/unit/readIntegrity.test.js`), all with fake page-fetchers:
|
||||
|
||||
- clean walk → `PASS`
|
||||
- walk whose last page re-emits earlier rows (the measured production shape) →
|
||||
`FAIL` with the exact duplicate/missing counts
|
||||
- **ordered-but-duplicating walk → `FAIL`** (AC2, the meta-scar)
|
||||
- ordered walk short of `exact_count` → `CONTROL_INVALID` (AC3)
|
||||
- non-unique key → `KEY_NOT_UNIQUE` (AC4)
|
||||
- `walk()` stops on a short page and on an empty page
|
||||
- `applyFilters` chains `eq/is/in/not/lt` in order onto a fake builder
|
||||
- a query error propagates as `ERROR`, never as end-of-data (the
|
||||
`fromLedger:117` failure class — an error swallowed as "no more rows")
|
||||
|
||||
Live (`scripts/read-integrity.js`): AC5 above, run against prod, read-only.
|
||||
|
||||
---
|
||||
|
||||
## 6. VERDICT WITHDRAWAL — WHAT IT IS AND IS NOT
|
||||
|
||||
### Where standing verdicts actually live
|
||||
|
||||
There is **no verdict table**. `featureRegistry.statVerdicts` is an in-memory
|
||||
`Map` (`featureRegistry.js:145`) that resets every process; `mc_test_ledger`
|
||||
stores hypothesis COUNTS for the Bonferroni denominator, not verdicts. The
|
||||
standing record is the `specs/` markdown plus `CLAUDE.md`.
|
||||
|
||||
So the withdrawal is **appended as a committed record**
|
||||
(`src/services/model/withdrawnVerdicts.js` + this spec), never by editing or
|
||||
deleting a prior finding. The original verdicts stay exactly as recorded; the
|
||||
withdrawal sits beside them with its cause.
|
||||
|
||||
### What is withdrawn
|
||||
|
||||
| factor | stat | recorded verdict | now |
|
||||
|---|---|---|---|
|
||||
| `defense_by_direction` | hits | PROVES, n=528, Brier −0.0034, CI [−0.0059,−0.0009] @99.9% | **WITHDRAWN_PENDING_REAUDIT** |
|
||||
| `pitcher_contact_profile` | hits | PROVES, n=741, Brier −0.0066, CI [−0.0114,−0.0016] @99.9% | **WITHDRAWN_PENDING_REAUDIT** |
|
||||
|
||||
Cause, recorded verbatim on both: *drawn through an unordered-pagination reader;
|
||||
16.5–24.8% set corruption measured on the identical query 2026-08-09; distinct-n
|
||||
below the 500 floor for `defense_by_direction`; not retrospectively recoverable.*
|
||||
|
||||
`defense_by_direction` additionally carries the eligibility note: at the measured
|
||||
~20% duplication its distinct n is ≈422 against `MIN_N = 500`, so it was likely
|
||||
**never eligible for adjudication at all**. This is arithmetic on the recorded n,
|
||||
not a re-measurement.
|
||||
|
||||
### What withdrawal does NOT do
|
||||
|
||||
It does **not** stop the factors moving the served forecast.
|
||||
`hitsFactors.js` hardcodes its three factors and consults no registry, so
|
||||
withdrawal is a record of what we may CLAIM, not a change to what we SERVE.
|
||||
Disarming or re-proving is a later order with its own blast radius.
|
||||
|
||||
### Reinstatement
|
||||
|
||||
A withdrawn verdict returns only by re-running its gate through readers that this
|
||||
harness reports `PASS` for, on the day of the run. Nothing is reinstated by
|
||||
argument.
|
||||
|
||||
---
|
||||
|
||||
## 7. THE FIX — `safePaginate` (Fix A1, 2026-08-09)
|
||||
|
||||
`src/utils/safePaginate.js` is the one way this codebase walks a paginated
|
||||
PostgREST read. It closes both measured defects in a single call:
|
||||
|
||||
```js
|
||||
const rows = await paginate(
|
||||
() => sb.from('ledger_entries').select('id, …').eq('sport', 'mlb') /* … */,
|
||||
{ key: 'id', pageSize: 1000, label: 'caller(name)' },
|
||||
);
|
||||
```
|
||||
|
||||
**Requirements it enforces**
|
||||
|
||||
1. **Stable ORDER BY on a UNIQUE key.** An order on a non-unique column is NOT a
|
||||
fix — ties may be returned in any order, so pages still overlap. Default `id`.
|
||||
2. **Uniqueness verified at runtime, not promised.** A repeated key mid-walk
|
||||
means either a non-unique key or a scrambled read; both make the row set
|
||||
unusable, so it THROWS rather than returning it. The harness PASS is
|
||||
confirmation; this is the guard that runs every time.
|
||||
3. **A query error THROWS.** Never end-of-data. `paginate` takes a query
|
||||
*factory* (builders are single-use) and reuses `readIntegrity.walk`, which is
|
||||
already tested for short-page / empty-page / error / runaway behaviour — it
|
||||
does not hand-roll a second walk.
|
||||
|
||||
**Refusal is not a swallow.** A caller that responds to a throw by producing
|
||||
nothing (no calibrator, no map) is refusing honestly. A caller that continues
|
||||
with partial rows is repeating the defect. `null` from `fromLedger` means "not
|
||||
enough settled history"; a throw means "the read failed". Keeping those distinct
|
||||
is the point.
|
||||
|
||||
### Measuring a FIXED reader
|
||||
|
||||
A registry spec that restates the corrected query would verify the restatement,
|
||||
and a `fixed: true` flag would verify a comment. So a fixed reader supplies
|
||||
`readerRows` instead — a call into the **real production function**, whose
|
||||
actual output is compared to the ordered control. `arm_a` on the result reports
|
||||
which was used (`real_reader_function` vs `query_as_written`).
|
||||
|
||||
### AC7 — A1 acceptance (met 2026-08-09)
|
||||
|
||||
- `calibrationService.fromLedger:110` — **FAIL 24.8% → PASS (0 dup / 0 missing,
|
||||
2,490 of 2,490)**, measured through `loadSettledRows`.
|
||||
- `lowParamService.fromLedger:78` — the byte-identical twin, deliberately NOT
|
||||
fixed in A1, still measures **24.8%**. The harness kept its teeth; the PASS is
|
||||
not an artefact of editing the registry.
|
||||
Reference in New Issue
Block a user