Files
vyndr/specs/a8-shadow-factor-gate.md
builtbykev f61ec6b391 Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 22:49:56 -04:00

141 lines
6.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PRE-REGISTRATION — A8: do the shadow-fired hits factors improve the forecast?
**Status: PRE-REGISTERED, NOT RUN.** Written 2026-08-11, before any accrual
exists, so the decision rule cannot be chosen after seeing the answer.
---
## 1. WHAT IS BEING TESTED, AND WHY IT IS A NEW TEST
A5 measured that all three hits factors were skipped on **596/596** real graded
props — `prop.opponent` and `prop.opposing_pitcher` were never set by anything,
so `positionOaa`, `throws` and `pitcherHardHit` resolved on **zero** rows. A6
resolved those keys and fired the factors in **shadow**: 248 fires on 308 props,
median would-be multiplier 1.017, range 0.803–1.250.
So `pitcher_contact_profile` and `defense_by_direction` were never "proven and
then disconnected." They were proven on an **audit reconstruction that the
production path never ran**, and then withdrawn (A0) because the reader that
produced that proof was 16–25% corrupt.
**This is therefore a NEW hypothesis, not a re-test.** It counts as a fresh test
against the cumulative Bonferroni denominator (`testLedger`). The S87 rule that
re-asking the same question on more data is not a new shot on goal does not apply:
the question is different. Previously — "does this factor improve a reconstructed
audit?" Now — "does the multiplier this factor actually produces, on the rows
where it actually fires, improve the served forecast out of sample?"
Three factors × one stat = **3 new hypotheses** to record before the run.
---
## 2. THE UNIT OF EVIDENCE
One testable pair = a `model_snapshots` row that carries **both**:
- `factor_inputs.would_fire` — written at grade time, pre-game, never applied
- `outcome` — settled from a box score by `snapshotSettlementService`
A `would_fire` with no settled outcome is **not yet evidence** and must not be
counted toward N.
Rows are eligible only if they also satisfy the standing gates:
`model_version = engine1@2026-08-07-fullwindow` (A3) and
`quarantine_reason` not `nontakeable_book%`.
---
## 3. THE TEST — TWO PARTS, BOTH REQUIRED
Run per factor, on **rows where that factor fired** (`would_fire.per_factor`
contains it). Scoring a factor on rows it did not touch dilutes any real effect
toward zero (S77).
**Baseline** = the served forecast, i.e. `p_win` as actually graded (the factors
were off).
**Conditioned** = `clamp(p_win × would_fire.multiplier, 0.01, 0.99)` — exactly
what would have been served had the factor been live.
Adjudicated by `factorGate.adjudicate()`:
1. **MOVEMENT** — mean `|Δp|` ≥ `MIN_MOVEMENT` (0.01). Movement alone is
**THEATER** and is named as such.
2. **IMPROVEMENT** — paired bootstrap on **Brier**, cluster-resampled, CI
excluding zero at `1 − 0.05/tests` where `tests` is the programme-lifetime
cumulative count from `testLedger`.
`cluster = game_id`. Verified: for all three factors the treatment entity
(`player|opponent`, `starter_id`, `player_key`) is more numerous than the games,
so `factorGate`'s coarser-of rule resolves to the game. **`MIN_CLUSTERS = 40`
distinct settled games is the binding constraint**, not `MIN_N`.
### Pre-registered outcomes
| verdict | meaning | consequence |
|---|---|---|
| `PROVES` | moves AND corrected CI below zero | candidate for A9 live turn-on — still a separate order |
| `NOT_PROVEN_AT_CORRECTED_BAR` | point estimate improves, CI spans zero | keep accruing; do NOT turn on |
| `THEATER` | moves, Brier delta ≥ 0 | **do not turn on, ever, at this multiplier** |
| `INERT` | mean shift < 0.01 | wiring is pointless; drop the factor |
| `CANDIDATE_PENDING_SAMPLE` | below N or clusters | keep accruing |
**Pre-registered fallback:** if a factor returns `THEATER`, that is a real result
and it must be recorded as such — a factor that moves ~80% of the board and does
not improve Brier is the arch-v1 failure repeating, and the correct action is to
leave it off, not to re-tune the multiplier until it passes.
**A `PROVES` does not turn anything on.** It makes the case for A9, which needs
its own before/after on the served board.
---
## 4. WHAT WOULD INVALIDATE THE RUN
- **Fewer than 40 distinct settled games** for that factor → refuse, report N.
- **`would_fire` not re-derivable** from `would_fire.shadow_inputs` on any row
(A6's recompute check) → the evidence is not trustworthy; stop.
- **Any row whose `would_fire` was written after first pitch** — settlement
already refuses post-hoc rows, and the same rule applies here.
- **Keys resolved from anything other than the pre-game lineup + probable
pitcher.** A retrospectively-derived opponent would make this a reconstruction
again, which is the whole thing A5/A6 exist to avoid.
---
## 5. WHEN IT CAN RUN
Measured accrual (settled MLB hits, repaired-champion rows):
| | 2026-08-07 | 2026-08-08 |
|---|---|---|
| settled hits rows | 758 | 1,096 |
| distinct settled players | 208 | 188 |
| **distinct settled games** | **11** | **10** |
Shadow fire rates (A6, 2026-08-09): `pitcher_contact_profile` 80.5%,
`defense_by_direction` 60.1%, `platoon_severity` 47.7%.
At ~925 settled hits rows and **~10.5 distinct settled games per slate**:
| factor | rows/slate | slates to `MIN_N` 500 | **slates to `MIN_CLUSTERS` 40** | binding |
|---|---|---|---|---|
| `pitcher_contact_profile` | ~745 | 1 | **~4** | clusters |
| `defense_by_direction` | ~556 | 1 | **~4** | clusters |
| `platoon_severity` | ~441 | 2 | **~4** | clusters |
**Accrual starts at DEPLOY, not today** — 0 rows carry `factor_inputs` now. Add
the settlement lag (a slate settles the following day), so:
> **A8 is runnable ≈ 5 slates after A6/A7 deploy** — about **5 days**, for all
> three factors together. Nothing accrues until the deploy lands.
---
## 6. WHAT THIS DOES NOT DECIDE
It does not reinstate `defense_by_direction` or `pitcher_contact_profile`. Those
withdrawals stand on their own cause (drawn through a corrupt reader). A `PROVES`
here is evidence about **the live multiplier on rows where it fires** — a
different claim from the original verdicts, and it should be recorded as a new
finding rather than as a reinstatement.