84f1fc075c2ebb07cdfd41783a2e23b93ab3a27e
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
044df7a426 |
The monitoring contract gets a callsite
Traced 2026-09-03: `calibrationRegistry.reverify` had ZERO production callers, and the only reference to the artifact machinery outside its own directory was the shadow builder, behind a flag that is off. The previous tranche declared that future settled outcomes are forward evaluation evidence and shipped no evaluator — the contract lived in a comment. That is the same shape as statModel.js and correlateValidator.js, cited for months and never present. `forwardMonitor` scores the FROZEN artifact on outcomes settled strictly after its training_cutoff, and `snapshotScheduler.calibrationMonitorTick` runs it on the existing per-minute cadence, throttled 6h, every failure swallowed. A test asserts the tick is defined, invoked AND exported, and drives it with a fake to prove it passes the promoted artifact and a forward-only window. It cannot change what it watches: the module imports no fitter and no registry, and a test greps the stripped source for fitIsotonic, fitPlatt, writeFileSync, upsert, update, promote( , artifactRegistry and PROMOTED. LOW N IS ITS OWN ANSWER. Below the floor it reports INSUFFICIENT_SAMPLE with `healthy: null` — never false, never true. The floor is DERIVED, not chosen: resolving an error of the certified tolerance at two standard errors needs n >= 0.25/(0.05/2)^2 = 400, and a test recomputes it from TOLERANCE so the two cannot drift apart. TOLERANCE is 0.05, the same number certifyBands used, so the monitor can be neither stricter nor laxer than the thing it watches. Wrong-era forward rows are INVALID, not scored — scoring them would measure a different forecaster, which is the defect this line of work removed. An unreadable read is AUDIT_UNAVAILABLE, which is not a health verdict. A drift alert says in its own text that the artifact is frozen and unchanged, because the alert is not a demotion. `currentEraSource.loadRows` gains an optional `after` bound, strictly greater so the cutoff date itself can never be scored as forward evidence. SHADOW REMAINS BLOCKED: the production variable is still absent (configuration_source "default") after a restart at 05:09:02Z, so no shadow cohort was obtained and no replay was substituted for one. Artifact unchanged: mlb-hits-isotonic@2026-09-03, knots 5ae940ea163b7da2. Live OFF. Suite 405/405, 5,654 passed. Teeth 31/31 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
be8e16aca9 |
Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state anyway. A validator that changes nothing is decoration, so `servable:false` is now load-bearing: an artifact that fails its policy returns ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes with it. The gate sits inside the resolution, not beside the flag that turns the shadow on, so no environment variable can reach past it — a test asserts `resolve` never reads process.env at all. Shadow and live consume the SAME decision, differing only in which promotion stage they demand. Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This artifact belongs to a different forecaster" is more precise than "policy blocked", and the existing state already says it exactly. THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the era from config/modelVersion so the query, the artifact and the validator all read one identity. Measured on the actual fitted set, not a second count: 6,069 current-era rows, 0 wrong-era. The procedure was then certified on current-era rows ONLY — four walk-forward folds, training strictly before each evaluation block, 0 future rows in train on every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323, delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold. Mapping spread inside support is 0.001-0.018. The prior mixed-era certification did not substitute for this. Policy B selected. A (era-filtered 65/35) and B (all current-era) are statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the PROCEDURE — it is not permanently withheld from the artifact that ships. withheld_from_fit is 0. FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01 (distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…, served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON. The runtime no longer fits. It loads. A test greps the service for fitIsotonic, fromLedger and loadRows and requires all three absent, because the old behaviour meant a user's number could move with no version, no review and no rollback, and a past Read could not be reconstructed because its curve no longer existed. New settled outcomes are forward evidence now; they cannot touch this curve. Independent reconstruction from the declared training contract alone — fresh read, fresh digest, fresh fit — reproduces every digest and the curve byte for byte. Calling the builder twice would only have proven the builder deterministic. Promotion is a frozen source constant. A snapshot cannot promote, a settlement cannot promote, a successful fit cannot promote, and dropping a file into the artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is explicitly false. Two coverage holes found by their own teeth. The promotion guard could be deleted with every test still green, because the promoted file naturally agrees with itself — extracted as `acceptFile` and tested on the case `load()` cannot reach. And `validate(null)` returned no `servable` field at all, which is falsy at a call site and so would have read as correct while asserting nothing. Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change. Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |