f61ec6b391
Seven orders of measurement-first repair. The served grade does not move. A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS: 410-617 of 2,490 duplicated with an equal number never returned, while rows.length matched the server exactly. safePaginate orders on a real unique key, verifies the tuple at runtime, and THROWS on a query error instead of treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn through that reader, and defense_by_direction's distinct-n was likely below the gate floor all along. A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from pg_index (the context tables are dated-composite and had no single unique column). The unordered helper is deleted, not parked. A3 — ledgerService and retentionService defaulted the SAME env var to DIFFERENT versions, so no ledger row ever carried the marker eligibility requires. One source now. model_snapshots settlement moved onto the cron: 15,484 -> 28,894 settled, repaired-champion 0 -> 7,556. A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no row at-or-before the date means the factor does not apply, never the nearest row. Live path unchanged, proven 400/400 on real rows. A5 — factor_inputs freezes what the factor READ, never the multiplier, so an audit can recompute and check. It also recorded the finding: the three hits factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by the resolver and written by nothing. A6/A7 — matchupKeys resolves those keys from the posted lineup plus the schedule's probable pitchers, and fires the factors into a SHADOW freeze: 248 fires on 308 props, 245 of which would move the grade. The served forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test that decides whether they ever go live. Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity harness 34/34. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
141 lines
6.0 KiB
Markdown
141 lines
6.0 KiB
Markdown
# PRE-REGISTRATION — A8: do the shadow-fired hits factors improve the forecast?
|
||
|
||
**Status: PRE-REGISTERED, NOT RUN.** Written 2026-08-11, before any accrual
|
||
exists, so the decision rule cannot be chosen after seeing the answer.
|
||
|
||
---
|
||
|
||
## 1. WHAT IS BEING TESTED, AND WHY IT IS A NEW TEST
|
||
|
||
A5 measured that all three hits factors were skipped on **596/596** real graded
|
||
props — `prop.opponent` and `prop.opposing_pitcher` were never set by anything,
|
||
so `positionOaa`, `throws` and `pitcherHardHit` resolved on **zero** rows. A6
|
||
resolved those keys and fired the factors in **shadow**: 248 fires on 308 props,
|
||
median would-be multiplier 1.017, range 0.803–1.250.
|
||
|
||
So `pitcher_contact_profile` and `defense_by_direction` were never "proven and
|
||
then disconnected." They were proven on an **audit reconstruction that the
|
||
production path never ran**, and then withdrawn (A0) because the reader that
|
||
produced that proof was 16–25% corrupt.
|
||
|
||
**This is therefore a NEW hypothesis, not a re-test.** It counts as a fresh test
|
||
against the cumulative Bonferroni denominator (`testLedger`). The S87 rule that
|
||
re-asking the same question on more data is not a new shot on goal does not apply:
|
||
the question is different. Previously — "does this factor improve a reconstructed
|
||
audit?" Now — "does the multiplier this factor actually produces, on the rows
|
||
where it actually fires, improve the served forecast out of sample?"
|
||
|
||
Three factors × one stat = **3 new hypotheses** to record before the run.
|
||
|
||
---
|
||
|
||
## 2. THE UNIT OF EVIDENCE
|
||
|
||
One testable pair = a `model_snapshots` row that carries **both**:
|
||
|
||
- `factor_inputs.would_fire` — written at grade time, pre-game, never applied
|
||
- `outcome` — settled from a box score by `snapshotSettlementService`
|
||
|
||
A `would_fire` with no settled outcome is **not yet evidence** and must not be
|
||
counted toward N.
|
||
|
||
Rows are eligible only if they also satisfy the standing gates:
|
||
`model_version = engine1@2026-08-07-fullwindow` (A3) and
|
||
`quarantine_reason` not `nontakeable_book%`.
|
||
|
||
---
|
||
|
||
## 3. THE TEST — TWO PARTS, BOTH REQUIRED
|
||
|
||
Run per factor, on **rows where that factor fired** (`would_fire.per_factor`
|
||
contains it). Scoring a factor on rows it did not touch dilutes any real effect
|
||
toward zero (S77).
|
||
|
||
**Baseline** = the served forecast, i.e. `p_win` as actually graded (the factors
|
||
were off).
|
||
**Conditioned** = `clamp(p_win × would_fire.multiplier, 0.01, 0.99)` — exactly
|
||
what would have been served had the factor been live.
|
||
|
||
Adjudicated by `factorGate.adjudicate()`:
|
||
|
||
1. **MOVEMENT** — mean `|Δp|` ≥ `MIN_MOVEMENT` (0.01). Movement alone is
|
||
**THEATER** and is named as such.
|
||
2. **IMPROVEMENT** — paired bootstrap on **Brier**, cluster-resampled, CI
|
||
excluding zero at `1 − 0.05/tests` where `tests` is the programme-lifetime
|
||
cumulative count from `testLedger`.
|
||
|
||
`cluster = game_id`. Verified: for all three factors the treatment entity
|
||
(`player|opponent`, `starter_id`, `player_key`) is more numerous than the games,
|
||
so `factorGate`'s coarser-of rule resolves to the game. **`MIN_CLUSTERS = 40`
|
||
distinct settled games is the binding constraint**, not `MIN_N`.
|
||
|
||
### Pre-registered outcomes
|
||
|
||
| verdict | meaning | consequence |
|
||
|---|---|---|
|
||
| `PROVES` | moves AND corrected CI below zero | candidate for A9 live turn-on — still a separate order |
|
||
| `NOT_PROVEN_AT_CORRECTED_BAR` | point estimate improves, CI spans zero | keep accruing; do NOT turn on |
|
||
| `THEATER` | moves, Brier delta ≥ 0 | **do not turn on, ever, at this multiplier** |
|
||
| `INERT` | mean shift < 0.01 | wiring is pointless; drop the factor |
|
||
| `CANDIDATE_PENDING_SAMPLE` | below N or clusters | keep accruing |
|
||
|
||
**Pre-registered fallback:** if a factor returns `THEATER`, that is a real result
|
||
and it must be recorded as such — a factor that moves ~80% of the board and does
|
||
not improve Brier is the arch-v1 failure repeating, and the correct action is to
|
||
leave it off, not to re-tune the multiplier until it passes.
|
||
|
||
**A `PROVES` does not turn anything on.** It makes the case for A9, which needs
|
||
its own before/after on the served board.
|
||
|
||
---
|
||
|
||
## 4. WHAT WOULD INVALIDATE THE RUN
|
||
|
||
- **Fewer than 40 distinct settled games** for that factor → refuse, report N.
|
||
- **`would_fire` not re-derivable** from `would_fire.shadow_inputs` on any row
|
||
(A6's recompute check) → the evidence is not trustworthy; stop.
|
||
- **Any row whose `would_fire` was written after first pitch** — settlement
|
||
already refuses post-hoc rows, and the same rule applies here.
|
||
- **Keys resolved from anything other than the pre-game lineup + probable
|
||
pitcher.** A retrospectively-derived opponent would make this a reconstruction
|
||
again, which is the whole thing A5/A6 exist to avoid.
|
||
|
||
---
|
||
|
||
## 5. WHEN IT CAN RUN
|
||
|
||
Measured accrual (settled MLB hits, repaired-champion rows):
|
||
|
||
| | 2026-08-07 | 2026-08-08 |
|
||
|---|---|---|
|
||
| settled hits rows | 758 | 1,096 |
|
||
| distinct settled players | 208 | 188 |
|
||
| **distinct settled games** | **11** | **10** |
|
||
|
||
Shadow fire rates (A6, 2026-08-09): `pitcher_contact_profile` 80.5%,
|
||
`defense_by_direction` 60.1%, `platoon_severity` 47.7%.
|
||
|
||
At ~925 settled hits rows and **~10.5 distinct settled games per slate**:
|
||
|
||
| factor | rows/slate | slates to `MIN_N` 500 | **slates to `MIN_CLUSTERS` 40** | binding |
|
||
|---|---|---|---|---|
|
||
| `pitcher_contact_profile` | ~745 | 1 | **~4** | clusters |
|
||
| `defense_by_direction` | ~556 | 1 | **~4** | clusters |
|
||
| `platoon_severity` | ~441 | 2 | **~4** | clusters |
|
||
|
||
**Accrual starts at DEPLOY, not today** — 0 rows carry `factor_inputs` now. Add
|
||
the settlement lag (a slate settles the following day), so:
|
||
|
||
> **A8 is runnable ≈ 5 slates after A6/A7 deploy** — about **5 days**, for all
|
||
> three factors together. Nothing accrues until the deploy lands.
|
||
|
||
---
|
||
|
||
## 6. WHAT THIS DOES NOT DECIDE
|
||
|
||
It does not reinstate `defense_by_direction` or `pitcher_contact_profile`. Those
|
||
withdrawals stand on their own cause (drawn through a corrupt reader). A `PROVES`
|
||
here is evidence about **the live multiplier on rows where it fires** — a
|
||
different claim from the original verdicts, and it should be recorded as a new
|
||
finding rather than as a reinstatement.
|