Commit Graph

3 Commits

Author SHA1 Message Date
builtbykev 556d186ff1 One evaluator for the shadow, so the probe cannot report a mode the pipeline is not in
"Is the shadow effective?" was only answerable by waiting for a snapshot to
write a row. That leaves a blind spot with real cost: a variable SET IN COOLIFY
BUT NOT YET APPLIED to the running process is indistinguishable from an unset
one, and the runtime probe already proves the distinction matters — code_sha
22cf51c with started_at 01:45:44Z means anything set after that is not in this
process's environment.

probabilityContract.shadowState() is now the single evaluator. snapshotService
calls it and the protected status probe calls it, and a test asserts NEITHER
reads process.env directly — the same rule that keeps lineage_write_mode honest.
Reading the env in two places is how a status page and a gate come to disagree.

Strict by construction: only the exact string '1' enables it. 'true', 'yes',
'on', '01', ' 1 ' and '' are all OFF, because a loose parse turns a typo into an
activation. `configuration_source` separates an unset variable from one
explicitly set to '0', and `live_serving` is reported as its own switch so the
shadow can never be read as implying serving.

No behaviour changes. The shadow still defaults OFF, CALIBRATION_DEPLOYED is
still [], and served fields are untouched.

Frontend byte-identical to the last green build (git reports zero changes under
web/), so the build from 22cf51c stands.

Suite 401/401, 5,597 passed, 4 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 22:53:21 -04:00
builtbykev 22cf51c4b0 The mapping that ran now has a name, and the row carries it
Runtime probes say the fleet is on a8de676, one process generation. But
"isotonic" was a label, not a claim: probabilityContractService refits per
snapshot against `game_date < todayEt()`, so the mapping changes as outcomes
settle, and nothing on a row could say WHICH mapping produced its number.

The artifact now has an identity:

  estimator_type / estimator_version / certification_version / model_version
  fit_as_of         the exact lt(game_date) bound   2026-09-02
  training_cutoff   last date INSIDE the fit        2026-08-21
  fit_n / knot_count                                6,084 / 28
  knot_digest       d9d571d728ba76de
  served_curve      the COMPLETE served function over [0.50,0.80)
  served_curve_digest

The served curve is not a sample. p_win is quantised to three decimals at the
source, so a step table at 0.001 granularity is the mapping itself for every
input that can occur — six steps, ~200 bytes. Storing it makes a Read
reconstructable WITHOUT re-deriving a training set that may since have been
re-settled, and a claim you can only verify when the inputs happen not to have
moved is not a reconstructable claim.

Proven, not asserted: the production construction path run twice gives an
identical digest, and an INDEPENDENT reconstruction — re-walk 9,361 settled
ledger rows at the declared bound, refit from scratch — reproduces
d9d571d728ba76de exactly, 28 knots for 28.

A teeth injection found a real defect behind a coverage hole. `resolve` checked
the CONTRACT's model era and never the ARTIFACT's, so a mapping fitted for a
different era could be recorded beside a served number with every test green.
Both the era and the estimator type are now checked, and a mismatch serves
nothing rather than serving quietly.

OBSERVED AND NOT CHANGED: calibrationService splits 65/35 to certify its own
bands, a step this contract does not consume because support comes from the
frozen artifact. So the served map is fitted through 2026-08-21 while 3,277
more recent settled rows sit unused, and that lag grows with history. Changing
it would change the fitted function, which this tranche froze.

Shadow still defaults OFF. CALIBRATION_DEPLOYED still []. served_probability is
referenced by nothing outside the contract layer — asserted by a tooth.

Suite 401/401, 5,593 passed, 4 skipped. Teeth 23/23 (prior) + 7/7 (new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:44:52 -04:00
builtbykev a8de676756 A probability is served because evidence supports it, not because nothing else answered
The band gate was blocked for its `else` branch. It read:

    candidate = F(raw)
    served    = inCertifiedBand(candidate) ? candidate : RAW

and above raw 0.60 the model is measured overconfident — holdout raw 0.80-0.90
predicts 0.843 and realizes 0.639. So "the calibrator is not supported here" was
being answered with a number already proven wrong. Unsupported calibration does
not make raw true.

Four candidates were adjudicated on ONE split — fit on the earliest 60% of
train, decide support on the last 40%, evaluate on a holdout that saw neither:

  A low-param      80.2% coverage  0.24374  REFUTED — its extra region
                   (raw 0.80-0.90) certified on cert (err +0.040, n=55) and
                   refuted on holdout (served 0.754 vs observed 0.639), and it
                   leaves a hole at 0.70-0.80 while serving the island above it
  B isotonic       91.3% coverage  0.24337  CERTIFIED, contiguous raw [0.50,0.80)
  C empirical band 91.3% coverage  0.24335  REFUTED — refitted point-in-time on
                   current-model hits the realized rates INVERT in grade order
                   (B+ 0.593 < B 0.614 < C+ 0.623), so the served function steps
                   down at raw 0.78. Its shipped constants come from 3,417 props
                   pooled across four batter stats and do not reproduce here
  D raw identity   43.1% coverage  0.24866  certifies raw 0.50-0.60 and only there

Raw is candidate D, not a fallback. It earns exactly one region (holdout error
+0.010 on n=1,316), which is why the law is "raw must earn its region" rather
than "raw is never true". B already covers that region, so no hybrid is built.

Above raw 0.80 nothing is certified and nothing is served. That is the region
where raw is most wrong, isotonic over-corrects (cert err -0.093) and its LODO
mapping at 0.95 has spread 0.180. 8.7% of holdout rows land there.

The registry did not need changing. `serves(stat, p)` already tested certified
bands against the RAW p_win — support in the input domain, the correct question —
and returned {serve:false, reason}. It never said "serve raw". The output-space
gate and the raw fallback were both invented downstream in calibrationService.

ACTIVATION IS OFF. PROBABILITY_CONTRACT_SHADOW defaults to 0, CALIBRATION_DEPLOYED
stays frozen empty, and every served field is byte-identical. This releases the
support first, which is the required order. The shadow records raw belief, the
candidate served value, the state, the estimator identity, and what EV/Kelly/VALUE
would be under the actionability law — into its own column, read by nothing.

Migration 051 was applied to production BEFORE retentionService named the column.
PostgREST builds a bulk insert from the first row's shape, so a key whose column
does not exist 400s the whole batch silently — that is how migration 038 took
retention down for three days.

The user-facing contradiction is NOT fixed here. A B+ still says "realized about
66%" beside a confidence of 84. Fixing that is activation, and activation costs
32% of VALUE flags and 46% of Kelly recommendations on the holdout.

Suite 401/401, 5,580 passed, 4 skipped, deterministic across three runs.
Teeth 23/23, each independently injected and restored byte-identically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:03:16 -04:00