94f7c3c3ef134874cbb6dcceb5e9b0b0ca8baa91
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
981b26010a |
The first real cohort found the leak: one artifact was calibrating every stat
Shadow converged at 18:09:07Z and the 19:00Z slot produced cohort 0353c551.
It immediately falsified something no test had asked: 513 PUBLISHED NON-HITS
rows came back CERTIFIED_CALIBRATED with a served probability drawn from the
mlb-hits curve — total_bases 246, runs 149, rbi 125, walks 109, outs 28,
strikeouts 24, hits_allowed 20, earned_runs 17.
Two bugs, one on top of the other. `mergeProbabilityContract` passed only
{model_version, p_win}, dropping the row's identity; and the service's resolve
then stamped the {sport, stat} it had been BUILT with onto every read. So all
3,000 rows in the batch resolved as mlb hits.
The governance tests could not see it. They asked "does build() refuse another
stat?" — it does, and always did — and then exercised the merge with
hits-only rows. Production sends one mixed batch. The regression test now drives
the REAL collector with hits, total_bases, rbi, runs, walks, strikeouts and
home_runs at the same p_win and requires hits certified and every other stat
neither certified nor numeric.
Fixed in three layers, because one would have been the same single point that
just failed:
1. the service no longer substitutes its own identity — the row's decides,
and a read naming no stat resolves to no contract, which is UNSUPPORTED;
2. the merge carries the row's sport and stat;
3. probabilityContract refuses an artifact whose own sport/stat disagree with
the contract it is being used under, independent of plumbing.
NO USER IMPACT. Shadow only: every block carries servable:false, live serving is
OFF, CALIBRATION_DEPLOYED is [], and the anonymous payload showed zero
calibration fields before and after. But this is exactly the defect that would
have served a hits calibration curve for strikeouts on the day live was enabled,
and only a real cohort surfaced it.
Two teeth were themselves wrong. Both runners checked "retention identity
changes" by grepping the diff for `stat:`, which fired on `stat: r.stat` — a
line that READS identity to hand it to a reader, not one that changes what
identifies a row. A guard that cannot tell those apart blocks the fix for the
defect it exists to protect against. Both are now behavioural: build a row
through the real collector and compare the identity tuple.
Artifact unchanged: mlb-hits-isotonic@2026-09-03, knot 5ae940ea163b7da2.
Suite 405/405, 5,659 passed. Teeth 34/34 + 10/10 + 23/23.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
a8de676756 |
A probability is served because evidence supports it, not because nothing else answered
The band gate was blocked for its `else` branch. It read:
candidate = F(raw)
served = inCertifiedBand(candidate) ? candidate : RAW
and above raw 0.60 the model is measured overconfident — holdout raw 0.80-0.90
predicts 0.843 and realizes 0.639. So "the calibrator is not supported here" was
being answered with a number already proven wrong. Unsupported calibration does
not make raw true.
Four candidates were adjudicated on ONE split — fit on the earliest 60% of
train, decide support on the last 40%, evaluate on a holdout that saw neither:
A low-param 80.2% coverage 0.24374 REFUTED — its extra region
(raw 0.80-0.90) certified on cert (err +0.040, n=55) and
refuted on holdout (served 0.754 vs observed 0.639), and it
leaves a hole at 0.70-0.80 while serving the island above it
B isotonic 91.3% coverage 0.24337 CERTIFIED, contiguous raw [0.50,0.80)
C empirical band 91.3% coverage 0.24335 REFUTED — refitted point-in-time on
current-model hits the realized rates INVERT in grade order
(B+ 0.593 < B 0.614 < C+ 0.623), so the served function steps
down at raw 0.78. Its shipped constants come from 3,417 props
pooled across four batter stats and do not reproduce here
D raw identity 43.1% coverage 0.24866 certifies raw 0.50-0.60 and only there
Raw is candidate D, not a fallback. It earns exactly one region (holdout error
+0.010 on n=1,316), which is why the law is "raw must earn its region" rather
than "raw is never true". B already covers that region, so no hybrid is built.
Above raw 0.80 nothing is certified and nothing is served. That is the region
where raw is most wrong, isotonic over-corrects (cert err -0.093) and its LODO
mapping at 0.95 has spread 0.180. 8.7% of holdout rows land there.
The registry did not need changing. `serves(stat, p)` already tested certified
bands against the RAW p_win — support in the input domain, the correct question —
and returned {serve:false, reason}. It never said "serve raw". The output-space
gate and the raw fallback were both invented downstream in calibrationService.
ACTIVATION IS OFF. PROBABILITY_CONTRACT_SHADOW defaults to 0, CALIBRATION_DEPLOYED
stays frozen empty, and every served field is byte-identical. This releases the
support first, which is the required order. The shadow records raw belief, the
candidate served value, the state, the estimator identity, and what EV/Kelly/VALUE
would be under the actionability law — into its own column, read by nothing.
Migration 051 was applied to production BEFORE retentionService named the column.
PostgREST builds a bulk insert from the first row's shape, so a key whose column
does not exist 400s the whole batch silently — that is how migration 038 took
retention down for three days.
The user-facing contradiction is NOT fixed here. A B+ still says "realized about
66%" beside a confidence of 84. Fixing that is activation, and activation costs
32% of VALUE flags and 46% of Kelly recommendations on the holdout.
Suite 401/401, 5,580 passed, 4 skipped, deterministic across three runs.
Teeth 23/23, each independently injected and restored byte-identically.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|