4 Commits

Author SHA1 Message Date
builtbykev 94f7c3c3ef The probability never leaked across stats; the attribution did
Post-fix cohort df4ec562 closed the primary question: eleven non-hits stats,
1,910 rows, ZERO certified and ZERO numeric served probabilities. The cross-stat
repair holds.

But every one of those 1,910 rows still recorded
`artifact_id: mlb-hits-isotonic@2026-09-03` beside state UNSUPPORTED. A `doubles`
row named the hits artifact. Nothing was calibrated by it, so no number leaked —
but a later query for "rows this artifact produced" would have returned 2,186
instead of 133, and that is the shape of footgun this programme keeps finding.

The contract check runs BEFORE any artifact is relevant: with no certified
contract for the sport/stat, no artifact applies, and naming one asserts a
relationship that does not exist. UNSUPPORTED now carries null artifact,
artifact_id, estimator_type, estimator_version, certification_version and
procedure_version.

Attribution is KEPT where the artifact is genuinely the thing that declined —
UNCERTIFIED (out of support) and VERSION_MISMATCH both still name it. A test
holds both directions so this does not over-correct into erasing real provenance.

One existing test called resolve() without naming a stat and relied on the
service substituting one. That substitution was the original defect, so the test
now names its stat, as production does.

Artifact unchanged. Live OFF. Suite 405/405, 5,662 passed. Teeth 35/35 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 18:33:31 -04:00
builtbykev be8e16aca9 Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state
anyway. A validator that changes nothing is decoration, so `servable:false` is
now load-bearing: an artifact that fails its policy returns
ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes
with it. The gate sits inside the resolution, not beside the flag that turns the
shadow on, so no environment variable can reach past it — a test asserts
`resolve` never reads process.env at all. Shadow and live consume the SAME
decision, differing only in which promotion stage they demand.

Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This
artifact belongs to a different forecaster" is more precise than "policy
blocked", and the existing state already says it exactly.

THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the
era from config/modelVersion so the query, the artifact and the validator all
read one identity. Measured on the actual fitted set, not a second count:
6,069 current-era rows, 0 wrong-era.

The procedure was then certified on current-era rows ONLY — four walk-forward
folds, training strictly before each evaluation block, 0 future rows in train on
every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323,
delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold.
Mapping spread inside support is 0.001-0.018. The prior mixed-era certification
did not substitute for this.

Policy B selected. A (era-filtered 65/35) and B (all current-era) are
statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has
the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the
PROCEDURE — it is not permanently withheld from the artifact that ships.
withheld_from_fit is 0.

FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01
(distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the
eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…,
served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON.

The runtime no longer fits. It loads. A test greps the service for fitIsotonic,
fromLedger and loadRows and requires all three absent, because the old behaviour
meant a user's number could move with no version, no review and no rollback, and
a past Read could not be reconstructed because its curve no longer existed.
New settled outcomes are forward evidence now; they cannot touch this curve.

Independent reconstruction from the declared training contract alone — fresh
read, fresh digest, fresh fit — reproduces every digest and the curve byte for
byte. Calling the builder twice would only have proven the builder deterministic.

Promotion is a frozen source constant. A snapshot cannot promote, a settlement
cannot promote, a successful fit cannot promote, and dropping a file into the
artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is
explicitly false.

Two coverage holes found by their own teeth. The promotion guard could be
deleted with every test still green, because the promoted file naturally agrees
with itself — extracted as `acceptFile` and tested on the case `load()` cannot
reach. And `validate(null)` returned no `servable` field at all, which is falsy
at a call site and so would have read as correct while asserting nothing.

Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change.
Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:06:18 -04:00
builtbykev 22cf51c4b0 The mapping that ran now has a name, and the row carries it
Runtime probes say the fleet is on a8de676, one process generation. But
"isotonic" was a label, not a claim: probabilityContractService refits per
snapshot against `game_date < todayEt()`, so the mapping changes as outcomes
settle, and nothing on a row could say WHICH mapping produced its number.

The artifact now has an identity:

  estimator_type / estimator_version / certification_version / model_version
  fit_as_of         the exact lt(game_date) bound   2026-09-02
  training_cutoff   last date INSIDE the fit        2026-08-21
  fit_n / knot_count                                6,084 / 28
  knot_digest       d9d571d728ba76de
  served_curve      the COMPLETE served function over [0.50,0.80)
  served_curve_digest

The served curve is not a sample. p_win is quantised to three decimals at the
source, so a step table at 0.001 granularity is the mapping itself for every
input that can occur — six steps, ~200 bytes. Storing it makes a Read
reconstructable WITHOUT re-deriving a training set that may since have been
re-settled, and a claim you can only verify when the inputs happen not to have
moved is not a reconstructable claim.

Proven, not asserted: the production construction path run twice gives an
identical digest, and an INDEPENDENT reconstruction — re-walk 9,361 settled
ledger rows at the declared bound, refit from scratch — reproduces
d9d571d728ba76de exactly, 28 knots for 28.

A teeth injection found a real defect behind a coverage hole. `resolve` checked
the CONTRACT's model era and never the ARTIFACT's, so a mapping fitted for a
different era could be recorded beside a served number with every test green.
Both the era and the estimator type are now checked, and a mismatch serves
nothing rather than serving quietly.

OBSERVED AND NOT CHANGED: calibrationService splits 65/35 to certify its own
bands, a step this contract does not consume because support comes from the
frozen artifact. So the served map is fitted through 2026-08-21 while 3,277
more recent settled rows sit unused, and that lag grows with history. Changing
it would change the fitted function, which this tranche froze.

Shadow still defaults OFF. CALIBRATION_DEPLOYED still []. served_probability is
referenced by nothing outside the contract layer — asserted by a tooth.

Suite 401/401, 5,593 passed, 4 skipped. Teeth 23/23 (prior) + 7/7 (new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:44:52 -04:00
builtbykev a8de676756 A probability is served because evidence supports it, not because nothing else answered
The band gate was blocked for its `else` branch. It read:

    candidate = F(raw)
    served    = inCertifiedBand(candidate) ? candidate : RAW

and above raw 0.60 the model is measured overconfident — holdout raw 0.80-0.90
predicts 0.843 and realizes 0.639. So "the calibrator is not supported here" was
being answered with a number already proven wrong. Unsupported calibration does
not make raw true.

Four candidates were adjudicated on ONE split — fit on the earliest 60% of
train, decide support on the last 40%, evaluate on a holdout that saw neither:

  A low-param      80.2% coverage  0.24374  REFUTED — its extra region
                   (raw 0.80-0.90) certified on cert (err +0.040, n=55) and
                   refuted on holdout (served 0.754 vs observed 0.639), and it
                   leaves a hole at 0.70-0.80 while serving the island above it
  B isotonic       91.3% coverage  0.24337  CERTIFIED, contiguous raw [0.50,0.80)
  C empirical band 91.3% coverage  0.24335  REFUTED — refitted point-in-time on
                   current-model hits the realized rates INVERT in grade order
                   (B+ 0.593 < B 0.614 < C+ 0.623), so the served function steps
                   down at raw 0.78. Its shipped constants come from 3,417 props
                   pooled across four batter stats and do not reproduce here
  D raw identity   43.1% coverage  0.24866  certifies raw 0.50-0.60 and only there

Raw is candidate D, not a fallback. It earns exactly one region (holdout error
+0.010 on n=1,316), which is why the law is "raw must earn its region" rather
than "raw is never true". B already covers that region, so no hybrid is built.

Above raw 0.80 nothing is certified and nothing is served. That is the region
where raw is most wrong, isotonic over-corrects (cert err -0.093) and its LODO
mapping at 0.95 has spread 0.180. 8.7% of holdout rows land there.

The registry did not need changing. `serves(stat, p)` already tested certified
bands against the RAW p_win — support in the input domain, the correct question —
and returned {serve:false, reason}. It never said "serve raw". The output-space
gate and the raw fallback were both invented downstream in calibrationService.

ACTIVATION IS OFF. PROBABILITY_CONTRACT_SHADOW defaults to 0, CALIBRATION_DEPLOYED
stays frozen empty, and every served field is byte-identical. This releases the
support first, which is the required order. The shadow records raw belief, the
candidate served value, the state, the estimator identity, and what EV/Kelly/VALUE
would be under the actionability law — into its own column, read by nothing.

Migration 051 was applied to production BEFORE retentionService named the column.
PostgREST builds a bulk insert from the first row's shape, so a key whose column
does not exist 400s the whole batch silently — that is how migration 038 took
retention down for three days.

The user-facing contradiction is NOT fixed here. A B+ still says "realized about
66%" beside a confidence of 84. Fixing that is activation, and activation costs
32% of VALUE flags and 46% of Kelly recommendations on the holdout.

Suite 401/401, 5,580 passed, 4 skipped, deterministic across three runs.
Teeth 23/23, each independently injected and restored byte-identically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:03:16 -04:00