3f7aa368c4c5466467de6252ab5dd29097912d90
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
84f1fc075c |
The live machinery ships dark, behind two gates that cannot substitute for each other
ATTRIBUTION RECEIPT (cohort 27ce152f, writer
|
||
|
|
94f7c3c3ef |
The probability never leaked across stats; the attribution did
Post-fix cohort df4ec562 closed the primary question: eleven non-hits stats, 1,910 rows, ZERO certified and ZERO numeric served probabilities. The cross-stat repair holds. But every one of those 1,910 rows still recorded `artifact_id: mlb-hits-isotonic@2026-09-03` beside state UNSUPPORTED. A `doubles` row named the hits artifact. Nothing was calibrated by it, so no number leaked — but a later query for "rows this artifact produced" would have returned 2,186 instead of 133, and that is the shape of footgun this programme keeps finding. The contract check runs BEFORE any artifact is relevant: with no certified contract for the sport/stat, no artifact applies, and naming one asserts a relationship that does not exist. UNSUPPORTED now carries null artifact, artifact_id, estimator_type, estimator_version, certification_version and procedure_version. Attribution is KEPT where the artifact is genuinely the thing that declined — UNCERTIFIED (out of support) and VERSION_MISMATCH both still name it. A test holds both directions so this does not over-correct into erasing real provenance. One existing test called resolve() without naming a stat and relied on the service substituting one. That substitution was the original defect, so the test now names its stat, as production does. Artifact unchanged. Live OFF. Suite 405/405, 5,662 passed. Teeth 35/35 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
981b26010a |
The first real cohort found the leak: one artifact was calibrating every stat
Shadow converged at 18:09:07Z and the 19:00Z slot produced cohort 0353c551.
It immediately falsified something no test had asked: 513 PUBLISHED NON-HITS
rows came back CERTIFIED_CALIBRATED with a served probability drawn from the
mlb-hits curve — total_bases 246, runs 149, rbi 125, walks 109, outs 28,
strikeouts 24, hits_allowed 20, earned_runs 17.
Two bugs, one on top of the other. `mergeProbabilityContract` passed only
{model_version, p_win}, dropping the row's identity; and the service's resolve
then stamped the {sport, stat} it had been BUILT with onto every read. So all
3,000 rows in the batch resolved as mlb hits.
The governance tests could not see it. They asked "does build() refuse another
stat?" — it does, and always did — and then exercised the merge with
hits-only rows. Production sends one mixed batch. The regression test now drives
the REAL collector with hits, total_bases, rbi, runs, walks, strikeouts and
home_runs at the same p_win and requires hits certified and every other stat
neither certified nor numeric.
Fixed in three layers, because one would have been the same single point that
just failed:
1. the service no longer substitutes its own identity — the row's decides,
and a read naming no stat resolves to no contract, which is UNSUPPORTED;
2. the merge carries the row's sport and stat;
3. probabilityContract refuses an artifact whose own sport/stat disagree with
the contract it is being used under, independent of plumbing.
NO USER IMPACT. Shadow only: every block carries servable:false, live serving is
OFF, CALIBRATION_DEPLOYED is [], and the anonymous payload showed zero
calibration fields before and after. But this is exactly the defect that would
have served a hits calibration curve for strikeouts on the day live was enabled,
and only a real cohort surfaced it.
Two teeth were themselves wrong. Both runners checked "retention identity
changes" by grepping the diff for `stat:`, which fired on `stat: r.stat` — a
line that READS identity to hand it to a reader, not one that changes what
identifies a row. A guard that cannot tell those apart blocks the fix for the
defect it exists to protect against. Both are now behavioural: build a row
through the real collector and compare the identity tuple.
Artifact unchanged: mlb-hits-isotonic@2026-09-03, knot 5ae940ea163b7da2.
Suite 405/405, 5,659 passed. Teeth 34/34 + 10/10 + 23/23.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
f4f51ed7ca |
The runtime names its artifact before anything is switched on
Step 35 asks the deployed process to prove WHICH frozen curve it resolves — artifact id, procedure, era, source digest, knot digest, curve digest, support, servable, stage — with the shadow still OFF. The probe could not answer, so the same shadowState() evaluator now resolves the promoted artifact and reports its identity. Identity only: the curve is committed and addressable by artifact_id, and repeating it on a status page would put a second copy of the truth there. null means nothing is promoted, which is also the honest answer when the file is missing or fails its policy at load. Suite 404/404. Shadow OFF. Live OFF. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
be8e16aca9 |
Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state anyway. A validator that changes nothing is decoration, so `servable:false` is now load-bearing: an artifact that fails its policy returns ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes with it. The gate sits inside the resolution, not beside the flag that turns the shadow on, so no environment variable can reach past it — a test asserts `resolve` never reads process.env at all. Shadow and live consume the SAME decision, differing only in which promotion stage they demand. Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This artifact belongs to a different forecaster" is more precise than "policy blocked", and the existing state already says it exactly. THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the era from config/modelVersion so the query, the artifact and the validator all read one identity. Measured on the actual fitted set, not a second count: 6,069 current-era rows, 0 wrong-era. The procedure was then certified on current-era rows ONLY — four walk-forward folds, training strictly before each evaluation block, 0 future rows in train on every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323, delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold. Mapping spread inside support is 0.001-0.018. The prior mixed-era certification did not substitute for this. Policy B selected. A (era-filtered 65/35) and B (all current-era) are statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the PROCEDURE — it is not permanently withheld from the artifact that ships. withheld_from_fit is 0. FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01 (distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…, served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON. The runtime no longer fits. It loads. A test greps the service for fitIsotonic, fromLedger and loadRows and requires all three absent, because the old behaviour meant a user's number could move with no version, no review and no rollback, and a past Read could not be reconstructed because its curve no longer existed. New settled outcomes are forward evidence now; they cannot touch this curve. Independent reconstruction from the declared training contract alone — fresh read, fresh digest, fresh fit — reproduces every digest and the curve byte for byte. Calling the builder twice would only have proven the builder deterministic. Promotion is a frozen source constant. A snapshot cannot promote, a settlement cannot promote, a successful fit cannot promote, and dropping a file into the artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is explicitly false. Two coverage holes found by their own teeth. The promotion guard could be deleted with every test still green, because the promoted file naturally agrees with itself — extracted as `acceptFile` and tested on the case `load()` cannot reach. And `validate(null)` returned no `servable` field at all, which is falsy at a call site and so would have read as correct while asserting nothing. Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change. Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |