2 Commits

Author SHA1 Message Date
builtbykev 84f1fc075c The live machinery ships dark, behind two gates that cannot substitute for each other
ATTRIBUTION RECEIPT (cohort 27ce152f, writer 94f7c3c, 01:02:19Z): 2,316 non-hits
rows — certified 0, numeric 0, and artifact_id / estimator / certification
version / source / knot / curve ALL ZERO. The hits slice held: 209 published,
201 certified, curve mismatches 0/201, artifact identity deviations 0, support
violations 0, raw erased 0. Shadow tranche frozen.

AUTHENTICATED NON-LEAK: obtained through the owner magic-link flow — the real
auth system, no bypass, no stored password, token never printed, session
discarded. /api/ledger, /api/ledger/accuracy, /api/preferences, /api/accuracy
and the authenticated /api/snapshot/mlb (677KB, full model fields) carry zero
shadow keys. One scanner hit adjudicated: `served_grade.calibrated` is false on
all 504 rows — servedGrade's hardcoded constant from before this work, a name
collision with my keyword list, not the shadow.

A REAL MONITOR DEFECT, asked for and found. The row floor alone let 400
observations from ONE slate reach HEALTHY or DRIFT — one correlated draw wearing
the costume of four hundred. The monitor now also requires settled DATES, and
the floor is not invented: it reads
`fitPolicy.POLICY_V1.certification.eval_block_dates`, the 3-date fold the
walk-forward was actually certified with. One date and two dates now return
INSUFFICIENT_SAMPLE regardless of row count.

That fix had a bug of its own that a test caught: `scored` never carried `date`,
so the distinct-date count read `undefined` for every row and always returned 1.
The gate looked correct while measuring nothing.

THE SEAM. EV, Kelly and VALUE are produced in exactly one place at grade time,
all from p_win, and every served row passes exactly one boundary on the way out
(`snapshotGating.stripModelPrice`, used by the snapshot route, hero route,
topGraded and props). So `servedProbability.applyToRows` sits there — one place,
not four call sites and four chances to miss one. Consumers never see the
artifact, the curve, the support region or the environment; a test forbids
`applyCurve`, `applyIsotonic`, `artifactRegistry` and `served_curve` in all three
serving consumers.

Placed BEFORE the tier strip deliberately: calibration decides what the number
IS, entitlement decides who may see it. Reversed, it would calibrate fields that
had already been removed.

TWO GATES, NEITHER SUFFICIENT. The artifact is promoted APPROVED_FOR_LIVE — stage
only, same id, same source/knot/curve digests, same cutoff, no refit. Behaviour
is still OFF because PROBABILITY_CONTRACT_LIVE is unset. Flag without approval:
OFF. Approval without flag: OFF. Both: ON. Live OFF returns rows byte-identical
and consults no artifact.

Two teeth had to be rewritten rather than satisfied: they pinned "artifact
unapproved" as the safety property, which would have blocked the deliberate
promotion. The property is that live BEHAVIOUR is off, and they now test that.
And my own splice while fixing one of them silently deleted ten newly-added
teeth — caught by counting ids, not by the runner going green.

Suite 406/406, 5,681 passed. Teeth 41/41 + 10/10 + 23/23. Live OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-04 06:36:11 -04:00
builtbykev 044df7a426 The monitoring contract gets a callsite
Traced 2026-09-03: `calibrationRegistry.reverify` had ZERO production callers,
and the only reference to the artifact machinery outside its own directory was
the shadow builder, behind a flag that is off. The previous tranche declared
that future settled outcomes are forward evaluation evidence and shipped no
evaluator — the contract lived in a comment. That is the same shape as
statModel.js and correlateValidator.js, cited for months and never present.

`forwardMonitor` scores the FROZEN artifact on outcomes settled strictly after
its training_cutoff, and `snapshotScheduler.calibrationMonitorTick` runs it on
the existing per-minute cadence, throttled 6h, every failure swallowed. A test
asserts the tick is defined, invoked AND exported, and drives it with a fake to
prove it passes the promoted artifact and a forward-only window.

It cannot change what it watches: the module imports no fitter and no registry,
and a test greps the stripped source for fitIsotonic, fitPlatt, writeFileSync,
upsert, update, promote( , artifactRegistry and PROMOTED.

LOW N IS ITS OWN ANSWER. Below the floor it reports INSUFFICIENT_SAMPLE with
`healthy: null` — never false, never true. The floor is DERIVED, not chosen:
resolving an error of the certified tolerance at two standard errors needs
n >= 0.25/(0.05/2)^2 = 400, and a test recomputes it from TOLERANCE so the two
cannot drift apart. TOLERANCE is 0.05, the same number certifyBands used, so the
monitor can be neither stricter nor laxer than the thing it watches.

Wrong-era forward rows are INVALID, not scored — scoring them would measure a
different forecaster, which is the defect this line of work removed. An
unreadable read is AUDIT_UNAVAILABLE, which is not a health verdict. A drift
alert says in its own text that the artifact is frozen and unchanged, because
the alert is not a demotion.

`currentEraSource.loadRows` gains an optional `after` bound, strictly greater
so the cutoff date itself can never be scored as forward evidence.

SHADOW REMAINS BLOCKED: the production variable is still absent
(configuration_source "default") after a restart at 05:09:02Z, so no shadow
cohort was obtained and no replay was substituted for one.

Artifact unchanged: mlb-hits-isotonic@2026-09-03, knots 5ae940ea163b7da2.
Live OFF. Suite 405/405, 5,654 passed. Teeth 31/31 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:29:47 -04:00