Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move. A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS: 410-617 of 2,490 duplicated with an equal number never returned, while rows.length matched the server exactly. safePaginate orders on a real unique key, verifies the tuple at runtime, and THROWS on a query error instead of treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn through that reader, and defense_by_direction's distinct-n was likely below the gate floor all along. A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from pg_index (the context tables are dated-composite and had no single unique column). The unordered helper is deleted, not parked. A3 — ledgerService and retentionService defaulted the SAME env var to DIFFERENT versions, so no ledger row ever carried the marker eligibility requires. One source now. model_snapshots settlement moved onto the cron: 15,484 -> 28,894 settled, repaired-champion 0 -> 7,556. A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no row at-or-before the date means the factor does not apply, never the nearest row. Live path unchanged, proven 400/400 on real rows. A5 — factor_inputs freezes what the factor READ, never the multiplier, so an audit can recompute and check. It also recorded the finding: the three hits factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by the resolver and written by nothing. A6/A7 — matchupKeys resolves those keys from the posted lineup plus the schedule's probable pitchers, and fires the factors into a SHADOW freeze: 248 fires on 308 props, 245 of which would move the grade. The served forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test that decides whether they ever go live. Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity harness 34/34. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1749,6 +1749,342 @@ phased plan in the Session-57 conversation / BUILD-STATE Next section).
|
||||
- **Proven factors for hits remain: `pitcher_contact_profile`,
|
||||
`defense_by_direction`.** Crude `defense` and `platoon` both still NOT_PROVEN.
|
||||
|
||||
## Read integrity + withdrawn verdicts (Fix A0 — non-obvious)
|
||||
- **THE SHARED `page()` HELPER RETURNS THE RIGHT COUNT AND THE WRONG ROWS.**
|
||||
`.range(from, from+PAGE-1)` with no `ORDER BY` was measured on prod returning
|
||||
410–617 duplicates out of 2,490 on the hits gate's ledger read, with an equal
|
||||
number of real rows NEVER returned. `rows.length` matches the server count
|
||||
exactly, which is why it hid for months. `settle-model-snapshots.js:92-98`
|
||||
found and fixed this in ONE script and it was never carried across.
|
||||
- **EXPOSURE CANNOT BE REASONED FROM SOURCE — only measured, per reader, per
|
||||
run.** `challenger-scoreboard` walks 12 pages of `ledger_entries` clean;
|
||||
`calibrationService.fromLedger` walks 3 pages of the SAME table at 24.8%
|
||||
corrupt. The difference is the query PLAN (filters × table stats × concurrent
|
||||
writes), which nothing in the code predicts. Same query re-measured minutes
|
||||
apart moved 16.5% → 24.8%. **Never infer a reader is safe; run the harness.**
|
||||
- **`src/utils/readIntegrity.js` + `scripts/read-integrity.js` are the
|
||||
instrument.** Spec `specs/read-integrity-harness.md`. A reader PASSES only on
|
||||
MEASURED set identity vs an ordered control — **the presence of an `.order()`
|
||||
clause is never evidence** (a grep-for-the-clause guard would verify its own
|
||||
intention, the meta-scar). The control validates itself: ordered-distinct ≠
|
||||
server `count(*)` ⇒ `CONTROL_INVALID`, never PASS. A non-unique key ⇒
|
||||
`KEY_NOT_UNIQUE`, so duplicate DATA is never blamed on the walk.
|
||||
`walk()` PROPAGATES a fetch error — swallowing it as end-of-data is the
|
||||
`calibrationService.fromLedger:117` defect and the 2026-08-01 settlement shape.
|
||||
- **Baseline 2026-08-09: 11 of 25 registered readers FAIL, worst 33.6%**
|
||||
(`champion-ablation:173`), `proven-status:76` 32.6%. Duplicates are WORSE than
|
||||
a random subsample: they inflate `factorGate.movement.n` against `MIN_N=500`
|
||||
AND corrupt the leave-one-out per-player baseline that is the gate's null.
|
||||
- **BOTH hits PROVES are WITHDRAWN_PENDING_REAUDIT**
|
||||
(`src/services/model/withdrawnVerdicts.js`, append-only — originals never
|
||||
edited). `defense_by_direction` proved at n=528; at ~20% duplication distinct-n
|
||||
≈422 < MIN_N 500, so it was likely never eligible (arithmetic on the recorded
|
||||
n, NOT a re-measurement). `pitcher_contact_profile` n=741 was plausibly
|
||||
eligible but its CI was computed on a corrupt sample against a corrupt
|
||||
baseline. Reinstatement requires re-running the gate through readers the
|
||||
harness reports PASS for ON THE DAY.
|
||||
- **WITHDRAWAL ≠ DISARMAMENT.** `hitsFactors.js` HARDCODES its three factors and
|
||||
consults no registry, so both withdrawn factors still move served `p_win`.
|
||||
Worse: its header calls all three "the three PROVEN hits factors" but
|
||||
**`platoon_severity` never passed the gate** (n=452 < 500) — a factor that was
|
||||
never proven is live in the served forecast.
|
||||
- **There is NO verdict table.** `featureRegistry.statVerdicts` is an in-memory
|
||||
Map that resets per process; `mc_test_ledger` holds hypothesis COUNTS for the
|
||||
Bonferroni denominator, not verdicts. Standing verdicts live in `specs/` +
|
||||
CLAUDE.md, which is why the withdrawal is a committed file.
|
||||
|
||||
## safePaginate + the calibration fence (Fix A1 — non-obvious)
|
||||
- **`src/utils/safePaginate.js` is the ONE way to walk a paginated PostgREST
|
||||
read.** `paginate(makeQuery, {key:'id', label})` — takes a query FACTORY (a
|
||||
builder is single-use), applies a stable ORDER BY on a UNIQUE key on every
|
||||
page, VERIFIES uniqueness at runtime (a repeated key throws — the guard that
|
||||
runs every time, vs the harness which runs on demand), and THROWS on a query
|
||||
error. It reuses `readIntegrity.walk`; never hand-roll a second walk.
|
||||
**An `.order()` on a NON-UNIQUE column is not a fix** — ties still scramble.
|
||||
- **`null` and `throw` mean different things and must stay different.**
|
||||
`fromLedger` returns `null` for "not enough settled history" and PROPAGATES a
|
||||
throw for "the read failed". Collapsing them (`if (error || !data) break`) is
|
||||
what let a 1,000-row fragment look like a complete 2,490-row fit.
|
||||
- **A fixed reader is measured through the REAL FUNCTION.** The harness spec
|
||||
takes `readerRows` (a call into production code), not a restatement of the
|
||||
corrected query and never a `fixed:true` flag — those verify a restatement and
|
||||
a comment respectively. `arm_a` on the result says which was used.
|
||||
- **CORRECTION to the A0 report: `calibrationService.fromLedger` was NOT live.**
|
||||
`snapshotService.js:301` is `CALIBRATION_DEPLOYED = Object.freeze([])` with no
|
||||
env override, so `for (const stat of CALIBRATION_DEPLOYED)` never executes.
|
||||
It is fenced THREE deep: empty deploy list; `calibrationService` is only the
|
||||
SHADOW (the primary is `lowParamService`); and `chain.chainAcross` has ZERO
|
||||
callers anywhere. `p_win_calibrated` appears in no route and no web file.
|
||||
**Nothing serves a calibrated number.** `calibrationDeployGate.test.js` locks it.
|
||||
- **The corrupt fit was wrong in a way row-count could not show.** Corrupt vs
|
||||
corrected on the same 2,490 rows: certified-band observed error **−0.046 →
|
||||
+0.010** (4.6x closer), ceiling 0.833 → 0.810, map output moving up to ±0.043.
|
||||
The certified span stayed 0.50–0.70 either way. And the duplicates shifted the
|
||||
TIME split: `fitted_through` 2026-08-06 → 2026-08-05, so the corrupt read
|
||||
fitted a different window than intended.
|
||||
- **`lowParamService.fromLedger:78` carries the byte-identical defect and is the
|
||||
PRIMARY calibrator** (calibrationService is its shadow). Left broken on purpose
|
||||
in A1 as the harness's control — it still measures 24.8%, which is how we know
|
||||
the PASS is real and not a registry edit. First target for A2.
|
||||
|
||||
## A2 rollout — 0 FAIL readers (non-obvious)
|
||||
- **`pageSafe()` is the per-script wrapper**, present in the 7 gate scripts:
|
||||
`pageSafe(sb, table, select, apply, key='id')` → `safePaginate.paginate`. The
|
||||
legacy `page()` SURVIVES in the same files **on purpose** — the context tables
|
||||
(`statcast_aggregates`, `batter_spray`, `team_defense`, `platoon_splits`,
|
||||
`park_dimensions`, `hitter_opportunity`, `lineup_context`) have **COMPOSITE
|
||||
primary keys with no single unique column**, which `paginate()` cannot express.
|
||||
They measure 0% today. **Never use `page()` for `ledger_entries` or
|
||||
`model_snapshots`** — `tests/unit/readerPagination.test.js` fails if you do.
|
||||
- **`ledger_entries` and `model_snapshots` are the only two tables with a
|
||||
single-column PK (`id`)** — which is why all 11 corrupt readers were fixable
|
||||
and the context tables were not. Composite-key ordering is the open item.
|
||||
- **Every fixed read is a named `READS` export**, and `main()` is wrapped in
|
||||
`if (require.main === module)`. Both are load-bearing: the harness `require()`s
|
||||
these scripts to call the REAL function, and several of them WRITE to
|
||||
`mc_test_ledger` (the Bonferroni denominator) via `tl.recordAndCount`. Without
|
||||
the guard, measuring a reader would inflate the multiple-comparisons
|
||||
correction. Verified: `mc_test_ledger` stayed 168 rows / last write 2026-08-06
|
||||
across the full harness run.
|
||||
- **3 reads deliberately left on the legacy walk** (`proven-status:66`, `:85`,
|
||||
`champion-ablation:176`) — they measured 0% and the A2 order scoped to the FAIL
|
||||
set. They are listed in `KNOWN_LEGACY_READS` in the test so the gap is tracked,
|
||||
not invisible. They are clean **by plan, not by construction**.
|
||||
- **`backfill-context.js` and `reconstruct-game-environment.js` still walk
|
||||
`ledger_entries` unordered and are NOT in the harness registry** — unmeasured,
|
||||
not proven clean.
|
||||
- **SUITE FLAKE, UNRESOLVED:** `snapshotService.test.js` "Wave 2A — threads the
|
||||
resolved athlete id" failed 2 times in 30 full-suite runs on this branch and
|
||||
0 in 29 at HEAD. That difference is NOT statistically significant (Fisher
|
||||
≈ p 0.5) and the assertion could not be captured in 24 instrumented runs. The
|
||||
plausible mechanism is that +89 tests reshuffles jest worker sharding and
|
||||
exposes a latent cross-suite interference. Isolated: 5/5 clean. **Do not treat
|
||||
it as resolved.**
|
||||
|
||||
## A2b — clean by construction (non-obvious)
|
||||
- **`safePaginate` takes a COMPOSITE key**: `key: ['as_of_date','sport','season','player_key']`.
|
||||
Every column is ordered, in order, and the uniqueness guard checks the FULL
|
||||
tuple. Ordering on a PREFIX is not a fix — the trailing columns are exactly
|
||||
where the ties live. String keys still work; the 12 A2 readers were unaffected.
|
||||
- **`src/utils/tableKeys.js` is the single source of truth for row identity**,
|
||||
pulled from `pg_index WHERE indisunique` on 2026-08-09 — never assumed. It
|
||||
THROWS for an unknown table rather than defaulting to `id`, because a silent
|
||||
`id` default would order by a column that may not exist and would look fixed
|
||||
while doing nothing. `pageSafe(sb, table, ...)` defaults its key from here.
|
||||
**If a migration changes a constraint, change it here** — a stale entry
|
||||
surfaces as a runtime duplicate-tuple throw, which is the intended loud failure.
|
||||
- **Only `ledger_entries`, `model_snapshots` and `game_context` have a
|
||||
single-column key.** Every other context table is dated-composite
|
||||
(`as_of_date, sport, season, <entity>`) by design — one as-of snapshot per
|
||||
entity per day. That is why A2 could not reach them.
|
||||
- **`batter_spray`'s KEY_NOT_UNIQUE was a wrong key, not bad data** — 6 "duplicate"
|
||||
`player_key|as_of_date` pairs resolve cleanly on the real
|
||||
`as_of_date,sport,season,source_id`. The harness naming it rather than
|
||||
swallowing it is what made that diagnosable.
|
||||
- **THE UNORDERED `page()` HELPER IS DELETED FROM ALL 15 SCRIPTS**, not parked
|
||||
next to `pageSafe`. It had become dead code, and a dead broken helper is an
|
||||
invitation. `readerPagination.test.js` fails if one reappears.
|
||||
- **A converted read MUST select its key columns.** 10 reads selected a subset
|
||||
that omitted `id` or the composite columns; the uniqueness guard would have
|
||||
thrown at runtime. A checker script verified every converted select covers its
|
||||
key — do that after any conversion.
|
||||
- **`calibrate-hits.js:40` was an unregistered 24.7%-corrupt ledger walker**
|
||||
found during the sweep — it fits the calibration reliability curve, so a fifth
|
||||
of the history was double-weighted and another fifth absent from every bin.
|
||||
Converted.
|
||||
- **WRITE-SCRIPT VERDICT (data untouched):** `backfill-context` and
|
||||
`reconstruct-game-environment` build DISTINCT SETS from their read, so
|
||||
corruption causes OMISSION, never a wrong value — the values come from
|
||||
statsapi/Open-Meteo. Their shared read measures **0% today** (5,107 rows), and
|
||||
a random-drop sensitivity at 24.8% bounds historical exposure at **~6.6 of 419
|
||||
players** (18 have a single row) and **~0.16 of 153 games** (median 37 rows).
|
||||
Small, bounded, and NOT a reason to rewrite context data.
|
||||
- **STILL OPEN — a clean read is not an as-of-correct read.**
|
||||
`hitsFactorContext.build` has no as-of cutoff (it takes `latestBy(as_of_date)`)
|
||||
AND orders its context walks by a NON-UNIQUE prefix (`player_key`/`team`). Same
|
||||
prefix problem in `factor-wiring-audit.js:62-65`. Reading today's context for a
|
||||
past row is contamination even when the read is perfect. Next order.
|
||||
|
||||
## A3 — version alignment + snapshot settlement (non-obvious)
|
||||
- **`src/config/modelVersion.js` is the ONLY place a model version is declared.**
|
||||
`ledgerService` and `retentionService` both read `process.env.MODEL_VERSION`
|
||||
but carried DIFFERENT hardcoded defaults (`engine1@2026-07-20` vs
|
||||
`engine1@2026-08-07-fullwindow`). Prod sets no env var, so the two tables took
|
||||
different defaults and **no ledger row ever carried the marker
|
||||
`reAuditEligibility` requires** — the accrual clock read "blocked" when the
|
||||
real state was "mis-stamped". Never re-declare a version default beside its
|
||||
consumer; it will drift from the other consumer and both sides look locally
|
||||
correct.
|
||||
- **`REPAIRED_CHAMPION_VERSION` is deliberately NOT `MODEL_VERSION`.** The
|
||||
eligibility bar is a fixed fact about one repair; tying it to "whatever we
|
||||
stamp today" would make every future version silently re-qualify itself.
|
||||
- **History is NOT re-stamped.** Rows keep the marker they were written with —
|
||||
rewriting them destroys the only record of which forecast produced them. Old
|
||||
rows stay honestly old.
|
||||
- **`snapshotSettlementService` settles `model_snapshots` on the cron** that
|
||||
already settles the ledger, best-effort (the retention table is a measurement
|
||||
asset, not the public record — a failure must not take the ledger settle or the
|
||||
grade with it). **Outcomes only: it never reconstructs context** (a test greps
|
||||
for every context table name).
|
||||
- **THE DRAIN ORDER IS A CORRECTNESS PROPERTY, and oldest-first is wrong.**
|
||||
Measured: 308 rows across the eight oldest dates are structurally unsettleable
|
||||
(post-hoc-logged or orphaned), so an oldest-first window re-processed dead
|
||||
dates forever and `dates_remaining` never moved. It is now (1) NEWEST-first,
|
||||
(2) date selection filtered by `isPreGame` — pure, so unsettleable rows are
|
||||
known without fetching — and (3) window-advancing when a window produces
|
||||
nothing. All three were needed to make it terminate.
|
||||
- **A row logged after first pitch is not a prediction.** A 01:00-UTC cycle is
|
||||
21:00 ET the *previous* evening — same game date, three hours into the slate.
|
||||
A UTC date compare would keep those rows. **20,064 rows are permanently
|
||||
unsettleable** for this reason and are counted, not hidden.
|
||||
- **`outcome` is SIDE-ALIGNED, `actual_value` is raw.** `p_win` is expressed for
|
||||
the graded side, so a raw `realized > line` indicator would invert the target
|
||||
on every under row.
|
||||
- **RESULT:** settled 15,484 → **28,894**; repaired-champion **0 → 7,556**;
|
||||
`settled_but_null_actual = 0` on every row.
|
||||
- **ELIGIBILITY NOW COUNTS BUT IS STILL SHORT: 2 eligible dates vs 10/10/14/14
|
||||
needed.** And the re-audit must not RUN on clean reads alone — it is gated on
|
||||
the **as-of-correct context fix** (`hitsFactorContext.build` has no as-of
|
||||
cutoff), which is a separate order.
|
||||
|
||||
## A4 — as-of-correct context (non-obvious)
|
||||
- **`hitsFactorContext.build(sb, { asOf })`.** No `asOf` = LIVE = the exact code
|
||||
that ran before, so the served grade cannot move. Only an explicitly dated call
|
||||
takes the audit path. **Proven, not asserted: 400 real graded rows resolved
|
||||
both ways gave identical context, identical multiplier and identical
|
||||
p_adjusted — 400/400.**
|
||||
- **REFUSAL OVER RECONSTRUCTION.** No row at-or-before `asOf` for an entity ⇒ its
|
||||
factor input is null ⇒ the factor does not apply and the base rate is
|
||||
untouched. NEVER the nearest or latest row: a substituted row is a plausible
|
||||
wrong value wearing a date, and nothing downstream can see it.
|
||||
- **`statcast_aggregates` CANNOT answer an as-of question** — it is upserted in
|
||||
place and keeps one date. The dated path reads **`statcast_history`**. Verified
|
||||
equivalent at the head before switching: on 2026-08-09 the two agree on all
|
||||
1,414 rows with **0 differences**, so changing source does not itself move a
|
||||
number. History starts **2026-08-03**; an as-of before that refuses.
|
||||
- **Dated-table coverage starts 2026-08-04** (`batter_spray`, `team_defense`,
|
||||
`platoon_splits`). An as-of earlier than that refuses everything — correct, and
|
||||
it means a re-audit cannot reach back past early August no matter how many
|
||||
dates accrue.
|
||||
- **`model_snapshots.opponent` is NULL on 100% of settled repaired-champion rows**
|
||||
(0 of 3,064 and 0 of 4,492). `team` is present on ~98%, `archetype` on ~63%.
|
||||
So `defense_by_direction` — one of the two withdrawn PROVES — **cannot be
|
||||
re-audited from the snapshot row alone**; the opponent must come from the
|
||||
player's own game log, as `prove-hit-factors` already does. Spray/platoon/
|
||||
handedness reach 99%; defence reaches 0% without that join.
|
||||
- **Archetype is CARRIED, never invented** — no context table holds it; it lives
|
||||
on `model_snapshots.archetype`, so an audit caller supplies it and the resolver
|
||||
passes it through for per-archetype conditioning.
|
||||
- `factor-wiring-audit.js` gets the same `asOf` (env `FWA_AS_OF`) and its context
|
||||
walks moved off a NON-UNIQUE prefix order (`player_key`/`team`) onto the real
|
||||
composite key.
|
||||
|
||||
## A5 — THE HITS FACTORS NEVER FIRE (the finding) + input freeze
|
||||
- **MEASURED 2026-08-10, 596 real graded hits props: ALL THREE hits factors are
|
||||
skipped on 596/596 and the multiplier is exactly 1 every time.** The context
|
||||
loads fine (618 spray players, 31 defence teams, 789 pitcher profiles) and
|
||||
resolves batter-side inputs — spray 590/596, bats 594, platoon 594 — but
|
||||
**`positionOaa` 0, `throws` 0, `pitcherHardHit` 0.**
|
||||
- **CAUSE: the two join keys are never set.** `prop.opponent` and
|
||||
`prop.opposing_pitcher` are read by `hitsFactorContext` and written by NOTHING
|
||||
— `oddsNormalizer` never sets them, and `snapshotService` attaches team/opponent
|
||||
AFTER grading (`:500`). `grep opposing_pitcher src/` returns exactly one hit:
|
||||
the line that reads it.
|
||||
- **CONSEQUENCE, and it corrects several earlier notes:** no hits factor has ever
|
||||
moved a served `p_win`. So (a) A0/A2's "hitsFactors serves two withdrawn
|
||||
factors plus platoon_severity" is WRONG — they are wired but inert; (b) S94's
|
||||
"PROVEN FACTORS, LOADED BEFORE THE GRADE — this ordering IS the fix" was
|
||||
necessary but NOT sufficient; the ordering was fixed and the join keys were
|
||||
never plumbed; (c) `defense_by_direction` did not "do the thing without
|
||||
recording it" — it never did the thing.
|
||||
- **Plumbing the keys WOULD change served grades**, so it is not a bug-fix, it is
|
||||
a model change and needs its own order with a before/after.
|
||||
- **`factorFreeze.freeze(ctx, prop)` records INPUTS, never the multiplier**
|
||||
(`model_snapshots.factor_inputs` jsonb, migration 034). A stored multiplier can
|
||||
only be compared to itself; stored inputs re-run through `hitsFactors` via
|
||||
`factorFreeze.recompute()` and get CHECKED. Captured in the same pass that
|
||||
reads the context (`analyzeViaEngine1`), so recorded cannot drift from used.
|
||||
- **A separate column, not extra keys in `features`** — `champion-ablation.js`
|
||||
iterates every `features` key for its residual scan, so widening it would
|
||||
silently enlarge that multiple-comparisons denominator.
|
||||
- **`available` (team/game_id/game_date) is recorded but NEVER fed to a factor.**
|
||||
It is the evidence a later order needs to measure what plumbing the join keys
|
||||
would change, before changing it.
|
||||
- Proven on 496 real rows: frozen == what the factor read (496/496), recompute
|
||||
from frozen == live multiplier (496/496), `p_win` unchanged (496/496).
|
||||
- **FORWARD-ONLY.** Rows graded before deploy keep `factor_inputs` NULL and
|
||||
`opponent` NULL; `defense_by_direction` re-audit on them still needs the
|
||||
game-log opponent join. No backfill — that would reconstruct.
|
||||
|
||||
## A6 — the join keys, SHADOW (non-obvious)
|
||||
- **`matchupKeys.build({sb,getSchedule,gameDate,asOf})`** resolves the two keys
|
||||
the hits factors have always needed: `player -> team` from **`lineup_context`**
|
||||
(as-of dated) and `team -> {opponent, opposing_pitcher}` from
|
||||
`getScheduleWithPitchers`. **At grade time the game has not happened**, so
|
||||
`prove-hit-factors`' game-log source is useless here — the probable pitcher is
|
||||
the only honest pre-game source.
|
||||
- **NEVER guess the opponent from the prop.** A prop carries `home_team` and
|
||||
`away_team`, so inferring which side a hitter bats for would be right about
|
||||
half the time and wrong invisibly — and a wrong opponent feeds the defence
|
||||
factor a real team's fielders against the wrong hitter. No lineup row ⇒
|
||||
`refused: 'no_lineup_row'`, no fire.
|
||||
- **`hitsFactorContext`'s resolver takes an OPTIONAL third arg `keys`.** Omitted
|
||||
on the live path ⇒ byte-identical to before it existed. Supplied only by the
|
||||
shadow resolve. That is what keeps the served grade frozen while the factors
|
||||
compute.
|
||||
- **MEASURED 2026-08-09, 308 hits props:** keys resolved on **248 (80.5%)**, 60
|
||||
refused for no lineup row. Live fire **0/0/0** (unchanged). **Shadow fire:
|
||||
`pitcher_contact_profile` 248, `defense_by_direction` 185,
|
||||
`platoon_severity` 147** — from 0/596 in A5. Would-be multiplier: min 0.803,
|
||||
p25 0.966, **median 1.017**, p75 1.064, max 1.250; **245 of 248 would move the
|
||||
number**. Served p_win identical on every evaluated row; shadow recompute-check
|
||||
248/248.
|
||||
- **`factor_inputs.would_fire` stores a multiplier — deliberately breaking A5's
|
||||
inputs-only rule, and only because `shadow_inputs` is stored beside it**, so it
|
||||
stays re-derivable. A number you cannot re-derive is what that rule forbids.
|
||||
- **The live frozen inputs still record the refusal** (`position_oaa: null`,
|
||||
`throws: null`) while `would_fire.shadow_inputs` records what it WOULD have
|
||||
seen. Both, side by side, on the same row.
|
||||
- A6 is EVIDENCE, not a verdict. The withdrawn factors are still withdrawn; A7
|
||||
(turning them on live) is a model change gated on this evidence plus a re-proof
|
||||
on rows where they actually fire.
|
||||
- **Test-window gotcha:** the A5 test sliced a fixed 700 bytes from its marker
|
||||
and A6's insertion pushed the assertion out of the window — a test failing
|
||||
because a neighbour grew. Window to a syntactic landmark, never a byte count.
|
||||
|
||||
## A7 — shadow accrual + the pre-registered A8 gate (non-obvious)
|
||||
- **`specs/a8-shadow-factor-gate.md` is PRE-REGISTERED and NOT RUN.** Written
|
||||
before any accrual exists, so the decision rule cannot be picked after seeing
|
||||
the answer. It names the fallback out loud: a factor that moves ~80% of the
|
||||
board and does not improve Brier is THEATER, and the correct action is to leave
|
||||
it off — not to re-tune the multiplier until it passes.
|
||||
- **A8 is a NEW hypothesis, not a re-test**, and counts fresh against the
|
||||
cumulative Bonferroni denominator. The S87 "re-asking on more data is not a new
|
||||
shot on goal" rule does NOT apply: the old question was "does this improve a
|
||||
reconstructed audit"; the new one is "does the multiplier this factor actually
|
||||
produces, on rows where it actually fires, improve the served forecast".
|
||||
- **THE BINDING CONSTRAINT IS CLUSTERS, NOT ROWS.** For all three factors the
|
||||
treatment entity (`player|opponent`, `starter_id`, `player_key`) outnumbers the
|
||||
games, so `factorGate`'s coarser-of rule resolves to `game_id`. Measured:
|
||||
~925 settled hits rows but only **~10.5 distinct settled GAMES per slate**, so
|
||||
`MIN_N` 500 is met in 1–2 slates and `MIN_CLUSTERS` 40 takes **~4**. Reading
|
||||
the row count alone would have said "ready tomorrow" and been wrong.
|
||||
- **A `would_fire` with no settled outcome is NOT evidence.** Complete pairs
|
||||
today: **0** — `factor_inputs` is 0/122,276 because A5/A6 are not deployed.
|
||||
Accrual starts at deploy, not at merge; add a one-day settle lag. **A8 is
|
||||
runnable ≈5 slates after deploy.**
|
||||
- **`tests/unit/shadowAccrual.test.js` drives the REAL
|
||||
`gradeAndCacheSlate → onGraded → rowsFromSides` chain** rather than reading
|
||||
code, because A5's whole finding was a factor that was built, correct, and
|
||||
never invoked — evidence that is generated but never persisted is that same
|
||||
failure one layer along. It also asserts the SCHEDULED path (`runSnapshot`
|
||||
builds the key index before the grade call, `runAllSnapshots` routes through
|
||||
it, a key-resolve failure degrades to no-shadow).
|
||||
- **Settlement lags the slate**: 2026-08-09/10/11 had 1,438–1,804 graded hits
|
||||
rows and **0 settled** at the time of writing, because snapshot settlement runs
|
||||
on the cron and had last drained through 08-08. Accrual counts must be read off
|
||||
SETTLED rows, never graded ones.
|
||||
|
||||
## Active Skills
|
||||
- vyndr-voice (all user-facing output)
|
||||
- prop-analysis (grading methodology)
|
||||
|
||||
Reference in New Issue
Block a user