5cad851922f31ff844ef009e80f4fa90b5129686
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f61ec6b391 |
Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move. A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS: 410-617 of 2,490 duplicated with an equal number never returned, while rows.length matched the server exactly. safePaginate orders on a real unique key, verifies the tuple at runtime, and THROWS on a query error instead of treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn through that reader, and defense_by_direction's distinct-n was likely below the gate floor all along. A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from pg_index (the context tables are dated-composite and had no single unique column). The unordered helper is deleted, not parked. A3 — ledgerService and retentionService defaulted the SAME env var to DIFFERENT versions, so no ledger row ever carried the marker eligibility requires. One source now. model_snapshots settlement moved onto the cron: 15,484 -> 28,894 settled, repaired-champion 0 -> 7,556. A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no row at-or-before the date means the factor does not apply, never the nearest row. Live path unchanged, proven 400/400 on real rows. A5 — factor_inputs freezes what the factor READ, never the multiplier, so an audit can recompute and check. It also recorded the finding: the three hits factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by the resolver and written by nothing. A6/A7 — matchupKeys resolves those keys from the posted lineup plus the schedule's probable pitchers, and fires the factors into a SHADOW freeze: 248 fires on 308 props, 245 of which would move the grade. The served forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test that decides whether they ever go live. Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity harness 34/34. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ff037e40c2 |
Re-adjudicate: nothing to demote, and close the hole that would have mattered
There is nothing to re-adjudicate. The proven set is empty and always has
been -- verified three ways: proven-status reports EMPTY, validatedSkills()
returns {} for every archetype, and zero conditioning entries have ever
reached PROVEN. The one PROVEN feature is recent_frequency_prior, which is the
incumbent counter itself, proven by the S78 ablation as ~100% of the
champion's resolution. It is the baseline every challenger is measured
against, not a conditioning interaction, and demoting it would leave the model
with nothing to grade from.
A correction to the premise: the cumulative gate did NOT catch a false
positive last session. It caught nothing, because there was nothing in the
proven set to catch. What it did was tighten alpha from 0.0026 to 0.0013
within one session, which demonstrated the mechanism working rather than a
demotion. So steps 3 and 4 -- demote, recalibrate -- are vacuous here, and
readjudicateAll says so plainly rather than glossing a no-op.
But the worry behind the order was well founded, and the audit found the real
exposure: promote() did not require the cumulative denominator. It checked n,
lift and CI, and nothing stopped a future session from testing eight
hypotheses, correcting by eight, and promoting on a p-value that would not
survive the programme's real denominator. That is precisely the hole that
makes a retroactive re-adjudication pass necessary later, so it is closed at
promotion time instead. isSufficient now refuses evidence carrying no
correction, evidence corrected against fewer tests than the cumulative count,
and any p-value that does not clear 0.05 over its own test count. The same
rule guards a PROVEN conditioning entry.
The second audit found two of four analysis scripts still correcting
per-session; pitcher-prove-k and tb-solo-and-interactions now use the
cumulative ledger, so the correction is native on every path.
reAblation.js is the standing second line: pure and injectable, so the
decision rule cannot drift from the gate's, and every verdict records both
p-values and both test counts so a demotion is re-derivable by anyone. A
feature promoted at alpha 0.05/20 can demote on the same p-value once the bar
is 0.05/60 -- correct, because the bar rose only after the programme had more
chances to get lucky. No fresh measurement is PENDING_RETEST and never a
demotion: absence of a re-test is not evidence, and demoting on it would
punish whichever stat happens to be off-season.
Net effect on the proven set is zero. No demotions, no recalibrations, and no
public ledger event -- announcing "recalibrated after re-adjudication" when
nothing changed would itself be a false signal of rigour.
4,238 tests green (337 suites); web build exit 0; counter byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
4aab18096f |
Prove both on total bases -- and find that my own fix destroyed the backtest
Nothing passed. Nothing promoted. Counter byte-identical. THE BLOCKER, which is the real finding. statcast_aggregates is upserted in place and holds exactly one as-of date. Yesterday's skill backtest was honest only by accident: the nightly refresh was unreachable code, so the profiles sat frozen at 2026-07-21 -- before the settled window. Repairing that cron was right for production and it refreshed them to today, destroying every prior version. Scoring a 2026-07-25 game now uses a season aggregate that contains that game. Point-in-time validation is structurally impossible from that table, so every number in this run is contaminated and directional, and none of it is a gate verdict. Fixed forward: statcast_history retains a dated snapshot on every refresh, so point-in-time becomes "as_of_date < game_date, most recent". Retention is best-effort and cannot fail the refresh; both properties are unit-tested. It has one day of data, which is not yet a window. SOLO BASELINE, n=383, Bonferroni across 12 tests (alpha 0.00417): nothing passes. hard_hit_pct is closest at marginal r 0.135 with p 0.0080, failing both the 0.15 effect bar and the corrected alpha. And it drifted DOWN from 0.153 at n=295 -- an estimate regressing as noise averages out, not an effect firming up. I called that number encouraging yesterday; on 88 more rows it is fading, and it should not keep being quoted at its best value. INTERACTIONS, each scored by partial correlation against the counter residual controlling for both of its own components: none pass. Only barrel x power archetype has an incremental exceeding its parts (-0.101 against 0.019) at n=260 -- the shape Discipline 2 predicts, but a lead, not a finding. A methodological catch worth keeping. The archetype conditioner was first built as barrel_pct over league barrel -- a monotone transform of one of its own components -- so the "interaction" was barrel squared, measuring nonlinearity in barrel rate rather than any archetype effect, and it produced this run's only positive result. A Gauss-Jordan pivot test does not catch that, because the two columns differ by a scale factor. Fixed with a scale-free collinearity check plus real archetype labels joined from model_snapshots. Without it this document would have reported a fabricated interaction as the session's finding. COMBINED vs COUNTER on total bases: 0.2718 against 0.2647, delta +0.0071, CI [-0.065, +0.079] -- inconclusive, and the first time a challenger has not lost. The same engine on hits was -0.116 with a CI excluding zero. That contrast is the whole argument for total bases, and it is what the physics said: contact quality governs extra bases, not whether a grounder finds a hole. Also built: the compound TB projection. skillProjection no longer refuses total bases -- a deterministic bases-per-hit multiplier had made P(TB>=2) exactly P(hits>=1), a relabelled hits curve. It is now a convolution over per-PA base outcomes with hit-type shares shifted by skill. Non-degeneracy is locked by test. 4,204 tests green (334 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |