main
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3f7aa368c4 |
The seam never ran: an undeclared variable, and a bare catch that hid it
LIVE was switched on in production and nothing happened, because
`snapshotGating.stripModelPrice` contained
try { rows = require(...).applyToRows(rows); } catch { }
`rows` is not declared in that function — the parameter is `grades`. Under
'use strict' that is a ReferenceError on every call, and the empty catch
swallowed it. For an entire release the seam was dead, the status surface
reported LIVE ON, and every row was served raw.
Two failures, and the second is the one that mattered: a silent catch turned a
hard crash into nothing at all. It now names the failure in the log and degrades
to uncalibrated rows explicitly.
The entitled branch also returned `grades` — the original array — so even a
working seam would have had its output discarded for exactly the tier meant to
receive it. Both fixed; `grades` is threaded through.
MY TESTS COULD NOT SEE IT. They called `applyToRows` directly and grepped the
source for the require line. Neither exercises the boundary, and a source grep
is not proof a line runs: the dead wiring contained that exact require. The new
tests call `stripModelPrice` and assert on its RETURN VALUE, entitled and
unentitled, plus a throwing-seam case that requires the warning.
Verified on real production rows through the real entitled path:
0.706 -> 0.625 CERTIFIED, EV recomputed 8.7 -> -3.7, grade B unchanged
0.858 -> unavailable UNCERTIFIED, its value:true WITHDRAWN, grade B+ unchanged
rbi -> untouched
free tier -> model fields still stripped
Found only because the live acceptance measured real rows instead of trusting a
green suite.
Suite 407/407, 5,697 passed. Teeth 48/48 + 10/10 + 23/23.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
355615d27e |
The interface was about to restore the lie: null confidence rendered as 0%
The backend already refuses to state a probability it cannot certify. The card
did not. `gradeAdapter.js:59` read
input.confidence != null ? Math.round(Number(input.confidence)) : 0
and `GradeResultCard` printed `{d.confidence}% confidence` unconditionally, so
an UNCERTIFIED Read — which by contract carries no exact probability — would
have rendered "0% CONFIDENCE" to a user. That is `Number(null) === 0`, the
standing fabrication bug in this codebase, arriving at the last surface before
the eye.
Found by running two REAL production rows (model_snapshots 453021 and 453033
from proven cohort 27ce152f) through the real serving seam and the real adapter,
before rendering anything.
Repaired at the two places it lives, presentation only:
- the adapter carries null through instead of coercing;
- the card renders "Confidence unavailable" as TEXT when there is no number.
Wording follows the house pattern already in `readHistory.js` ("History
unavailable") — no dash, no dimmed zero, no icon-only state, and no engineering
language: a test forbids isotonic / artifact / probability_contract /
UNCERTIFIED / certification_version reaching the card.
A second, narrower correction: the derived-field withdrawal was
`if (f in out) out[f] = null`, so it depended on whether the input happened to
carry the key and a field added by a later merge would have arrived
un-withdrawn. An uncertified row now DECLARES each probability-derived claim
null rather than merely lacking it.
The uncertified fixture is the strong case, not a convenient one: Luis Robert
carries value:true and ev_pct 38.5 on the raw row, so the contract has a live
VALUE badge to withdraw rather than an absent one to leave absent. Verified
withdrawn, while grade B+, the side, the Read and the market (book -160,
fair -135) all survive.
The certified case moves the number materially and visibly: raw 0.798 -> served
0.638, and EV is recomputed from the served probability, not carried from raw.
PriceTriplet needed no change — it already renders an absent leg as a dash and
was documented never to print a zero.
Suite 407/407, 5,691 passed. Teeth 45/45 + 10/10 + 23/23. Live OFF.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
84f1fc075c |
The live machinery ships dark, behind two gates that cannot substitute for each other
ATTRIBUTION RECEIPT (cohort 27ce152f, writer
|
||
|
|
94f7c3c3ef |
The probability never leaked across stats; the attribution did
Post-fix cohort df4ec562 closed the primary question: eleven non-hits stats, 1,910 rows, ZERO certified and ZERO numeric served probabilities. The cross-stat repair holds. But every one of those 1,910 rows still recorded `artifact_id: mlb-hits-isotonic@2026-09-03` beside state UNSUPPORTED. A `doubles` row named the hits artifact. Nothing was calibrated by it, so no number leaked — but a later query for "rows this artifact produced" would have returned 2,186 instead of 133, and that is the shape of footgun this programme keeps finding. The contract check runs BEFORE any artifact is relevant: with no certified contract for the sport/stat, no artifact applies, and naming one asserts a relationship that does not exist. UNSUPPORTED now carries null artifact, artifact_id, estimator_type, estimator_version, certification_version and procedure_version. Attribution is KEPT where the artifact is genuinely the thing that declined — UNCERTIFIED (out of support) and VERSION_MISMATCH both still name it. A test holds both directions so this does not over-correct into erasing real provenance. One existing test called resolve() without naming a stat and relied on the service substituting one. That substitution was the original defect, so the test now names its stat, as production does. Artifact unchanged. Live OFF. Suite 405/405, 5,662 passed. Teeth 35/35 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
981b26010a |
The first real cohort found the leak: one artifact was calibrating every stat
Shadow converged at 18:09:07Z and the 19:00Z slot produced cohort 0353c551.
It immediately falsified something no test had asked: 513 PUBLISHED NON-HITS
rows came back CERTIFIED_CALIBRATED with a served probability drawn from the
mlb-hits curve — total_bases 246, runs 149, rbi 125, walks 109, outs 28,
strikeouts 24, hits_allowed 20, earned_runs 17.
Two bugs, one on top of the other. `mergeProbabilityContract` passed only
{model_version, p_win}, dropping the row's identity; and the service's resolve
then stamped the {sport, stat} it had been BUILT with onto every read. So all
3,000 rows in the batch resolved as mlb hits.
The governance tests could not see it. They asked "does build() refuse another
stat?" — it does, and always did — and then exercised the merge with
hits-only rows. Production sends one mixed batch. The regression test now drives
the REAL collector with hits, total_bases, rbi, runs, walks, strikeouts and
home_runs at the same p_win and requires hits certified and every other stat
neither certified nor numeric.
Fixed in three layers, because one would have been the same single point that
just failed:
1. the service no longer substitutes its own identity — the row's decides,
and a read naming no stat resolves to no contract, which is UNSUPPORTED;
2. the merge carries the row's sport and stat;
3. probabilityContract refuses an artifact whose own sport/stat disagree with
the contract it is being used under, independent of plumbing.
NO USER IMPACT. Shadow only: every block carries servable:false, live serving is
OFF, CALIBRATION_DEPLOYED is [], and the anonymous payload showed zero
calibration fields before and after. But this is exactly the defect that would
have served a hits calibration curve for strikeouts on the day live was enabled,
and only a real cohort surfaced it.
Two teeth were themselves wrong. Both runners checked "retention identity
changes" by grepping the diff for `stat:`, which fired on `stat: r.stat` — a
line that READS identity to hand it to a reader, not one that changes what
identifies a row. A guard that cannot tell those apart blocks the fix for the
defect it exists to protect against. Both are now behavioural: build a row
through the real collector and compare the identity tuple.
Artifact unchanged: mlb-hits-isotonic@2026-09-03, knot 5ae940ea163b7da2.
Suite 405/405, 5,659 passed. Teeth 34/34 + 10/10 + 23/23.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
044df7a426 |
The monitoring contract gets a callsite
Traced 2026-09-03: `calibrationRegistry.reverify` had ZERO production callers, and the only reference to the artifact machinery outside its own directory was the shadow builder, behind a flag that is off. The previous tranche declared that future settled outcomes are forward evaluation evidence and shipped no evaluator — the contract lived in a comment. That is the same shape as statModel.js and correlateValidator.js, cited for months and never present. `forwardMonitor` scores the FROZEN artifact on outcomes settled strictly after its training_cutoff, and `snapshotScheduler.calibrationMonitorTick` runs it on the existing per-minute cadence, throttled 6h, every failure swallowed. A test asserts the tick is defined, invoked AND exported, and drives it with a fake to prove it passes the promoted artifact and a forward-only window. It cannot change what it watches: the module imports no fitter and no registry, and a test greps the stripped source for fitIsotonic, fitPlatt, writeFileSync, upsert, update, promote( , artifactRegistry and PROMOTED. LOW N IS ITS OWN ANSWER. Below the floor it reports INSUFFICIENT_SAMPLE with `healthy: null` — never false, never true. The floor is DERIVED, not chosen: resolving an error of the certified tolerance at two standard errors needs n >= 0.25/(0.05/2)^2 = 400, and a test recomputes it from TOLERANCE so the two cannot drift apart. TOLERANCE is 0.05, the same number certifyBands used, so the monitor can be neither stricter nor laxer than the thing it watches. Wrong-era forward rows are INVALID, not scored — scoring them would measure a different forecaster, which is the defect this line of work removed. An unreadable read is AUDIT_UNAVAILABLE, which is not a health verdict. A drift alert says in its own text that the artifact is frozen and unchanged, because the alert is not a demotion. `currentEraSource.loadRows` gains an optional `after` bound, strictly greater so the cutoff date itself can never be scored as forward evidence. SHADOW REMAINS BLOCKED: the production variable is still absent (configuration_source "default") after a restart at 05:09:02Z, so no shadow cohort was obtained and no replay was substituted for one. Artifact unchanged: mlb-hits-isotonic@2026-09-03, knots 5ae940ea163b7da2. Live OFF. Suite 405/405, 5,654 passed. Teeth 31/31 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
be8e16aca9 |
Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state anyway. A validator that changes nothing is decoration, so `servable:false` is now load-bearing: an artifact that fails its policy returns ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes with it. The gate sits inside the resolution, not beside the flag that turns the shadow on, so no environment variable can reach past it — a test asserts `resolve` never reads process.env at all. Shadow and live consume the SAME decision, differing only in which promotion stage they demand. Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This artifact belongs to a different forecaster" is more precise than "policy blocked", and the existing state already says it exactly. THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the era from config/modelVersion so the query, the artifact and the validator all read one identity. Measured on the actual fitted set, not a second count: 6,069 current-era rows, 0 wrong-era. The procedure was then certified on current-era rows ONLY — four walk-forward folds, training strictly before each evaluation block, 0 future rows in train on every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323, delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold. Mapping spread inside support is 0.001-0.018. The prior mixed-era certification did not substitute for this. Policy B selected. A (era-filtered 65/35) and B (all current-era) are statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the PROCEDURE — it is not permanently withheld from the artifact that ships. withheld_from_fit is 0. FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01 (distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…, served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON. The runtime no longer fits. It loads. A test greps the service for fitIsotonic, fromLedger and loadRows and requires all three absent, because the old behaviour meant a user's number could move with no version, no review and no rollback, and a past Read could not be reconstructed because its curve no longer existed. New settled outcomes are forward evidence now; they cannot touch this curve. Independent reconstruction from the declared training contract alone — fresh read, fresh digest, fresh fit — reproduces every digest and the curve byte for byte. Calling the builder twice would only have proven the builder deterministic. Promotion is a frozen source constant. A snapshot cannot promote, a settlement cannot promote, a successful fit cannot promote, and dropping a file into the artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is explicitly false. Two coverage holes found by their own teeth. The promotion guard could be deleted with every test still green, because the promoted file naturally agrees with itself — extracted as `acceptFile` and tested on the case `load()` cannot reach. And `validate(null)` returned no `servable` field at all, which is falsy at a call site and so would have read as correct while asserting nothing. Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change. Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |