84f1fc075c2ebb07cdfd41783a2e23b93ab3a27e
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
84f1fc075c |
The live machinery ships dark, behind two gates that cannot substitute for each other
ATTRIBUTION RECEIPT (cohort 27ce152f, writer
|
||
|
|
be8e16aca9 |
Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state anyway. A validator that changes nothing is decoration, so `servable:false` is now load-bearing: an artifact that fails its policy returns ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes with it. The gate sits inside the resolution, not beside the flag that turns the shadow on, so no environment variable can reach past it — a test asserts `resolve` never reads process.env at all. Shadow and live consume the SAME decision, differing only in which promotion stage they demand. Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This artifact belongs to a different forecaster" is more precise than "policy blocked", and the existing state already says it exactly. THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the era from config/modelVersion so the query, the artifact and the validator all read one identity. Measured on the actual fitted set, not a second count: 6,069 current-era rows, 0 wrong-era. The procedure was then certified on current-era rows ONLY — four walk-forward folds, training strictly before each evaluation block, 0 future rows in train on every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323, delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold. Mapping spread inside support is 0.001-0.018. The prior mixed-era certification did not substitute for this. Policy B selected. A (era-filtered 65/35) and B (all current-era) are statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the PROCEDURE — it is not permanently withheld from the artifact that ships. withheld_from_fit is 0. FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01 (distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…, served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON. The runtime no longer fits. It loads. A test greps the service for fitIsotonic, fromLedger and loadRows and requires all three absent, because the old behaviour meant a user's number could move with no version, no review and no rollback, and a past Read could not be reconstructed because its curve no longer existed. New settled outcomes are forward evidence now; they cannot touch this curve. Independent reconstruction from the declared training contract alone — fresh read, fresh digest, fresh fit — reproduces every digest and the curve byte for byte. Calling the builder twice would only have proven the builder deterministic. Promotion is a frozen source constant. A snapshot cannot promote, a settlement cannot promote, a successful fit cannot promote, and dropping a file into the artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is explicitly false. Two coverage holes found by their own teeth. The promotion guard could be deleted with every test still green, because the promoted file naturally agrees with itself — extracted as `acceptFile` and tested on the case `load()` cannot reach. And `validate(null)` returned no `servable` field at all, which is falsy at a call site and so would have read as correct while asserting nothing. Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change. Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
5cad851922 |
A digest names the curve; it does not vouch for the procedure that makes the next one
Artifact identity was the previous tranche's answer. It is not certification. `knot_digest` says WHICH mapping ran and nothing about whether tomorrow's refit deserves the same trust — that one would get its own digest and be equally "identified" while fitted on anything at all. So the object being certified is named: VYNDR certifies a PROCEDURE, not a frozen curve. A frozen curve goes stale against a live model and has to be replaced by hand on no schedule; "fit past, apply forward" is procedural by construction. `fitPolicy.POLICY_V1` declares it — data selection, horizon, algorithm and version, minimum rows, model-version restriction, sport, stat, and a support contract that a refit may NOT widen. Each artifact still carries its own digest. `fitPolicy.validate` is the Step-22 gate: an artifact does not become servable because the algorithm ran. It refuses a widened support, a wrong era, a wrong estimator, a thin fit, a missing identity or a missing training cutoff, and `servable` is false whenever any violation stands, with no override argument. The statistical bars stay where they already live in calibrationRegistry — this is not a second governance system. MEASURED, AND THE REASON LIVE SERVING STAYS BLOCKED: production does not match the declaration. `loadSettledRows` applies no model_version filter, so at fit_as_of 2026-09-02 the fit drew 6,084 rows from 9,361 settled — all 3,292 from the superseded engine1@2026-07-20 plus 2,792 current-era. 54.1% of the served map is fitted on a forecaster it was never certified for, while the artifact declares the current era. That is a provenance contradiction, not a performance claim: era-filtered scores 0.24065 against pooled 0.24068 on 1,120 out-of-sample rows and both intervals span zero. It is blocked because nothing prevents the next era change from repeating it, and because the freshness lag grows. `era_restricted` is answered STRUCTURALLY, not by an extra read — the query is in this service and applies no filter, so the artifact records ERA_NOT_RESTRICTED rather than claiming a restriction that did not hold. An unverified restriction is recorded as a violation, because "we did not check" is exactly the state production is in. The violation does NOT distort the shadow. Flipping every row to UNCERTIFIED would make the shadow measure the violation instead of the contract, so the policy state rides beside the resolution and a test asserts the shadow still reads CERTIFIED_CALIBRATED at 0.65 and UNCERTIFIED at 0.91. Nothing serves. CALIBRATION_DEPLOYED still []. Shadow still OFF (the production variable remains absent — the probe reads configuration_source "default"). Suite 402/402, 5,607 passed, 4 skipped. Teeth 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
556d186ff1 |
One evaluator for the shadow, so the probe cannot report a mode the pipeline is not in
"Is the shadow effective?" was only answerable by waiting for a snapshot to write a row. That leaves a blind spot with real cost: a variable SET IN COOLIFY BUT NOT YET APPLIED to the running process is indistinguishable from an unset one, and the runtime probe already proves the distinction matters — code_sha |
||
|
|
22cf51c4b0 |
The mapping that ran now has a name, and the row carries it
Runtime probes say the fleet is on
|