Commit Graph

3 Commits

Author SHA1 Message Date
builtbykev 981b26010a The first real cohort found the leak: one artifact was calibrating every stat
Shadow converged at 18:09:07Z and the 19:00Z slot produced cohort 0353c551.
It immediately falsified something no test had asked: 513 PUBLISHED NON-HITS
rows came back CERTIFIED_CALIBRATED with a served probability drawn from the
mlb-hits curve — total_bases 246, runs 149, rbi 125, walks 109, outs 28,
strikeouts 24, hits_allowed 20, earned_runs 17.

Two bugs, one on top of the other. `mergeProbabilityContract` passed only
{model_version, p_win}, dropping the row's identity; and the service's resolve
then stamped the {sport, stat} it had been BUILT with onto every read. So all
3,000 rows in the batch resolved as mlb hits.

The governance tests could not see it. They asked "does build() refuse another
stat?" — it does, and always did — and then exercised the merge with
hits-only rows. Production sends one mixed batch. The regression test now drives
the REAL collector with hits, total_bases, rbi, runs, walks, strikeouts and
home_runs at the same p_win and requires hits certified and every other stat
neither certified nor numeric.

Fixed in three layers, because one would have been the same single point that
just failed:
  1. the service no longer substitutes its own identity — the row's decides,
     and a read naming no stat resolves to no contract, which is UNSUPPORTED;
  2. the merge carries the row's sport and stat;
  3. probabilityContract refuses an artifact whose own sport/stat disagree with
     the contract it is being used under, independent of plumbing.

NO USER IMPACT. Shadow only: every block carries servable:false, live serving is
OFF, CALIBRATION_DEPLOYED is [], and the anonymous payload showed zero
calibration fields before and after. But this is exactly the defect that would
have served a hits calibration curve for strikeouts on the day live was enabled,
and only a real cohort surfaced it.

Two teeth were themselves wrong. Both runners checked "retention identity
changes" by grepping the diff for `stat:`, which fired on `stat: r.stat` — a
line that READS identity to hand it to a reader, not one that changes what
identifies a row. A guard that cannot tell those apart blocks the fix for the
defect it exists to protect against. Both are now behavioural: build a row
through the real collector and compare the identity tuple.

Artifact unchanged: mlb-hits-isotonic@2026-09-03, knot 5ae940ea163b7da2.
Suite 405/405, 5,659 passed. Teeth 34/34 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 16:07:39 -04:00
builtbykev 044df7a426 The monitoring contract gets a callsite
Traced 2026-09-03: `calibrationRegistry.reverify` had ZERO production callers,
and the only reference to the artifact machinery outside its own directory was
the shadow builder, behind a flag that is off. The previous tranche declared
that future settled outcomes are forward evaluation evidence and shipped no
evaluator — the contract lived in a comment. That is the same shape as
statModel.js and correlateValidator.js, cited for months and never present.

`forwardMonitor` scores the FROZEN artifact on outcomes settled strictly after
its training_cutoff, and `snapshotScheduler.calibrationMonitorTick` runs it on
the existing per-minute cadence, throttled 6h, every failure swallowed. A test
asserts the tick is defined, invoked AND exported, and drives it with a fake to
prove it passes the promoted artifact and a forward-only window.

It cannot change what it watches: the module imports no fitter and no registry,
and a test greps the stripped source for fitIsotonic, fitPlatt, writeFileSync,
upsert, update, promote( , artifactRegistry and PROMOTED.

LOW N IS ITS OWN ANSWER. Below the floor it reports INSUFFICIENT_SAMPLE with
`healthy: null` — never false, never true. The floor is DERIVED, not chosen:
resolving an error of the certified tolerance at two standard errors needs
n >= 0.25/(0.05/2)^2 = 400, and a test recomputes it from TOLERANCE so the two
cannot drift apart. TOLERANCE is 0.05, the same number certifyBands used, so the
monitor can be neither stricter nor laxer than the thing it watches.

Wrong-era forward rows are INVALID, not scored — scoring them would measure a
different forecaster, which is the defect this line of work removed. An
unreadable read is AUDIT_UNAVAILABLE, which is not a health verdict. A drift
alert says in its own text that the artifact is frozen and unchanged, because
the alert is not a demotion.

`currentEraSource.loadRows` gains an optional `after` bound, strictly greater
so the cutoff date itself can never be scored as forward evidence.

SHADOW REMAINS BLOCKED: the production variable is still absent
(configuration_source "default") after a restart at 05:09:02Z, so no shadow
cohort was obtained and no replay was substituted for one.

Artifact unchanged: mlb-hits-isotonic@2026-09-03, knots 5ae940ea163b7da2.
Live OFF. Suite 405/405, 5,654 passed. Teeth 31/31 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:29:47 -04:00
builtbykev be8e16aca9 Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state
anyway. A validator that changes nothing is decoration, so `servable:false` is
now load-bearing: an artifact that fails its policy returns
ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes
with it. The gate sits inside the resolution, not beside the flag that turns the
shadow on, so no environment variable can reach past it — a test asserts
`resolve` never reads process.env at all. Shadow and live consume the SAME
decision, differing only in which promotion stage they demand.

Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This
artifact belongs to a different forecaster" is more precise than "policy
blocked", and the existing state already says it exactly.

THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the
era from config/modelVersion so the query, the artifact and the validator all
read one identity. Measured on the actual fitted set, not a second count:
6,069 current-era rows, 0 wrong-era.

The procedure was then certified on current-era rows ONLY — four walk-forward
folds, training strictly before each evaluation block, 0 future rows in train on
every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323,
delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold.
Mapping spread inside support is 0.001-0.018. The prior mixed-era certification
did not substitute for this.

Policy B selected. A (era-filtered 65/35) and B (all current-era) are
statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has
the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the
PROCEDURE — it is not permanently withheld from the artifact that ships.
withheld_from_fit is 0.

FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01
(distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the
eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…,
served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON.

The runtime no longer fits. It loads. A test greps the service for fitIsotonic,
fromLedger and loadRows and requires all three absent, because the old behaviour
meant a user's number could move with no version, no review and no rollback, and
a past Read could not be reconstructed because its curve no longer existed.
New settled outcomes are forward evidence now; they cannot touch this curve.

Independent reconstruction from the declared training contract alone — fresh
read, fresh digest, fresh fit — reproduces every digest and the curve byte for
byte. Calling the builder twice would only have proven the builder deterministic.

Promotion is a frozen source constant. A snapshot cannot promote, a settlement
cannot promote, a successful fit cannot promote, and dropping a file into the
artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is
explicitly false.

Two coverage holes found by their own teeth. The promotion guard could be
deleted with every test still green, because the promoted file naturally agrees
with itself — extracted as `acceptFile` and tested on the case `load()` cannot
reach. And `validate(null)` returned no `servable` field at all, which is falsy
at a call site and so would have read as correct while asserting nothing.

Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change.
Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:06:18 -04:00