Files
vyndr/specs/a8-shadow-factor-gate.md
T
builtbykev f61ec6b391 Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 22:49:56 -04:00

6.0 KiB
Raw Blame History

PRE-REGISTRATION — A8: do the shadow-fired hits factors improve the forecast?

Status: PRE-REGISTERED, NOT RUN. Written 2026-08-11, before any accrual exists, so the decision rule cannot be chosen after seeing the answer.


1. WHAT IS BEING TESTED, AND WHY IT IS A NEW TEST

A5 measured that all three hits factors were skipped on 596/596 real graded props — prop.opponent and prop.opposing_pitcher were never set by anything, so positionOaa, throws and pitcherHardHit resolved on zero rows. A6 resolved those keys and fired the factors in shadow: 248 fires on 308 props, median would-be multiplier 1.017, range 0.803–1.250.

So pitcher_contact_profile and defense_by_direction were never "proven and then disconnected." They were proven on an audit reconstruction that the production path never ran, and then withdrawn (A0) because the reader that produced that proof was 16–25% corrupt.

This is therefore a NEW hypothesis, not a re-test. It counts as a fresh test against the cumulative Bonferroni denominator (testLedger). The S87 rule that re-asking the same question on more data is not a new shot on goal does not apply: the question is different. Previously — "does this factor improve a reconstructed audit?" Now — "does the multiplier this factor actually produces, on the rows where it actually fires, improve the served forecast out of sample?"

Three factors × one stat = 3 new hypotheses to record before the run.


2. THE UNIT OF EVIDENCE

One testable pair = a model_snapshots row that carries both:

  • factor_inputs.would_fire — written at grade time, pre-game, never applied
  • outcome — settled from a box score by snapshotSettlementService

A would_fire with no settled outcome is not yet evidence and must not be counted toward N.

Rows are eligible only if they also satisfy the standing gates: model_version = engine1@2026-08-07-fullwindow (A3) and quarantine_reason not nontakeable_book%.


3. THE TEST — TWO PARTS, BOTH REQUIRED

Run per factor, on rows where that factor fired (would_fire.per_factor contains it). Scoring a factor on rows it did not touch dilutes any real effect toward zero (S77).

Baseline = the served forecast, i.e. p_win as actually graded (the factors were off). Conditioned = clamp(p_win × would_fire.multiplier, 0.01, 0.99) — exactly what would have been served had the factor been live.

Adjudicated by factorGate.adjudicate():

  1. MOVEMENT — mean |Δp| ≥ MIN_MOVEMENT (0.01). Movement alone is THEATER and is named as such.
  2. IMPROVEMENT — paired bootstrap on Brier, cluster-resampled, CI excluding zero at 1 − 0.05/tests where tests is the programme-lifetime cumulative count from testLedger.

cluster = game_id. Verified: for all three factors the treatment entity (player|opponent, starter_id, player_key) is more numerous than the games, so factorGate's coarser-of rule resolves to the game. MIN_CLUSTERS = 40 distinct settled games is the binding constraint, not MIN_N.

Pre-registered outcomes

verdict meaning consequence
PROVES moves AND corrected CI below zero candidate for A9 live turn-on — still a separate order
NOT_PROVEN_AT_CORRECTED_BAR point estimate improves, CI spans zero keep accruing; do NOT turn on
THEATER moves, Brier delta ≥ 0 do not turn on, ever, at this multiplier
INERT mean shift < 0.01 wiring is pointless; drop the factor
CANDIDATE_PENDING_SAMPLE below N or clusters keep accruing

Pre-registered fallback: if a factor returns THEATER, that is a real result and it must be recorded as such — a factor that moves ~80% of the board and does not improve Brier is the arch-v1 failure repeating, and the correct action is to leave it off, not to re-tune the multiplier until it passes.

A PROVES does not turn anything on. It makes the case for A9, which needs its own before/after on the served board.


4. WHAT WOULD INVALIDATE THE RUN

  • Fewer than 40 distinct settled games for that factor → refuse, report N.
  • would_fire not re-derivable from would_fire.shadow_inputs on any row (A6's recompute check) → the evidence is not trustworthy; stop.
  • Any row whose would_fire was written after first pitch — settlement already refuses post-hoc rows, and the same rule applies here.
  • Keys resolved from anything other than the pre-game lineup + probable pitcher. A retrospectively-derived opponent would make this a reconstruction again, which is the whole thing A5/A6 exist to avoid.

5. WHEN IT CAN RUN

Measured accrual (settled MLB hits, repaired-champion rows):

2026-08-07 2026-08-08
settled hits rows 758 1,096
distinct settled players 208 188
distinct settled games 11 10

Shadow fire rates (A6, 2026-08-09): pitcher_contact_profile 80.5%, defense_by_direction 60.1%, platoon_severity 47.7%.

At ~925 settled hits rows and ~10.5 distinct settled games per slate:

factor rows/slate slates to MIN_N 500 slates to MIN_CLUSTERS 40 binding
pitcher_contact_profile ~745 1 ~4 clusters
defense_by_direction ~556 1 ~4 clusters
platoon_severity ~441 2 ~4 clusters

Accrual starts at DEPLOY, not today — 0 rows carry factor_inputs now. Add the settlement lag (a slate settles the following day), so:

A8 is runnable ≈ 5 slates after A6/A7 deploy — about 5 days, for all three factors together. Nothing accrues until the deploy lands.


6. WHAT THIS DOES NOT DECIDE

It does not reinstate defense_by_direction or pitcher_contact_profile. Those withdrawals stand on their own cause (drawn through a corrupt reader). A PROVES here is evidence about the live multiplier on rows where it fires — a different claim from the original verdicts, and it should be recorded as a new finding rather than as a reinstatement.