Commit Graph

5 Commits

Author SHA1 Message Date
builtbykev 84f1fc075c The live machinery ships dark, behind two gates that cannot substitute for each other
ATTRIBUTION RECEIPT (cohort 27ce152f, writer 94f7c3c, 01:02:19Z): 2,316 non-hits
rows — certified 0, numeric 0, and artifact_id / estimator / certification
version / source / knot / curve ALL ZERO. The hits slice held: 209 published,
201 certified, curve mismatches 0/201, artifact identity deviations 0, support
violations 0, raw erased 0. Shadow tranche frozen.

AUTHENTICATED NON-LEAK: obtained through the owner magic-link flow — the real
auth system, no bypass, no stored password, token never printed, session
discarded. /api/ledger, /api/ledger/accuracy, /api/preferences, /api/accuracy
and the authenticated /api/snapshot/mlb (677KB, full model fields) carry zero
shadow keys. One scanner hit adjudicated: `served_grade.calibrated` is false on
all 504 rows — servedGrade's hardcoded constant from before this work, a name
collision with my keyword list, not the shadow.

A REAL MONITOR DEFECT, asked for and found. The row floor alone let 400
observations from ONE slate reach HEALTHY or DRIFT — one correlated draw wearing
the costume of four hundred. The monitor now also requires settled DATES, and
the floor is not invented: it reads
`fitPolicy.POLICY_V1.certification.eval_block_dates`, the 3-date fold the
walk-forward was actually certified with. One date and two dates now return
INSUFFICIENT_SAMPLE regardless of row count.

That fix had a bug of its own that a test caught: `scored` never carried `date`,
so the distinct-date count read `undefined` for every row and always returned 1.
The gate looked correct while measuring nothing.

THE SEAM. EV, Kelly and VALUE are produced in exactly one place at grade time,
all from p_win, and every served row passes exactly one boundary on the way out
(`snapshotGating.stripModelPrice`, used by the snapshot route, hero route,
topGraded and props). So `servedProbability.applyToRows` sits there — one place,
not four call sites and four chances to miss one. Consumers never see the
artifact, the curve, the support region or the environment; a test forbids
`applyCurve`, `applyIsotonic`, `artifactRegistry` and `served_curve` in all three
serving consumers.

Placed BEFORE the tier strip deliberately: calibration decides what the number
IS, entitlement decides who may see it. Reversed, it would calibrate fields that
had already been removed.

TWO GATES, NEITHER SUFFICIENT. The artifact is promoted APPROVED_FOR_LIVE — stage
only, same id, same source/knot/curve digests, same cutoff, no refit. Behaviour
is still OFF because PROBABILITY_CONTRACT_LIVE is unset. Flag without approval:
OFF. Approval without flag: OFF. Both: ON. Live OFF returns rows byte-identical
and consults no artifact.

Two teeth had to be rewritten rather than satisfied: they pinned "artifact
unapproved" as the safety property, which would have blocked the deliberate
promotion. The property is that live BEHAVIOUR is off, and they now test that.
And my own splice while fixing one of them silently deleted ten newly-added
teeth — caught by counting ids, not by the runner going green.

Suite 406/406, 5,681 passed. Teeth 41/41 + 10/10 + 23/23. Live OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-04 06:36:11 -04:00
builtbykev be8e16aca9 Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state
anyway. A validator that changes nothing is decoration, so `servable:false` is
now load-bearing: an artifact that fails its policy returns
ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes
with it. The gate sits inside the resolution, not beside the flag that turns the
shadow on, so no environment variable can reach past it — a test asserts
`resolve` never reads process.env at all. Shadow and live consume the SAME
decision, differing only in which promotion stage they demand.

Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This
artifact belongs to a different forecaster" is more precise than "policy
blocked", and the existing state already says it exactly.

THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the
era from config/modelVersion so the query, the artifact and the validator all
read one identity. Measured on the actual fitted set, not a second count:
6,069 current-era rows, 0 wrong-era.

The procedure was then certified on current-era rows ONLY — four walk-forward
folds, training strictly before each evaluation block, 0 future rows in train on
every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323,
delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold.
Mapping spread inside support is 0.001-0.018. The prior mixed-era certification
did not substitute for this.

Policy B selected. A (era-filtered 65/35) and B (all current-era) are
statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has
the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the
PROCEDURE — it is not permanently withheld from the artifact that ships.
withheld_from_fit is 0.

FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01
(distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the
eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…,
served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON.

The runtime no longer fits. It loads. A test greps the service for fitIsotonic,
fromLedger and loadRows and requires all three absent, because the old behaviour
meant a user's number could move with no version, no review and no rollback, and
a past Read could not be reconstructed because its curve no longer existed.
New settled outcomes are forward evidence now; they cannot touch this curve.

Independent reconstruction from the declared training contract alone — fresh
read, fresh digest, fresh fit — reproduces every digest and the curve byte for
byte. Calling the builder twice would only have proven the builder deterministic.

Promotion is a frozen source constant. A snapshot cannot promote, a settlement
cannot promote, a successful fit cannot promote, and dropping a file into the
artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is
explicitly false.

Two coverage holes found by their own teeth. The promotion guard could be
deleted with every test still green, because the promoted file naturally agrees
with itself — extracted as `acceptFile` and tested on the case `load()` cannot
reach. And `validate(null)` returned no `servable` field at all, which is falsy
at a call site and so would have read as correct while asserting nothing.

Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change.
Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:06:18 -04:00
builtbykev 5cad851922 A digest names the curve; it does not vouch for the procedure that makes the next one
Artifact identity was the previous tranche's answer. It is not certification.
`knot_digest` says WHICH mapping ran and nothing about whether tomorrow's refit
deserves the same trust — that one would get its own digest and be equally
"identified" while fitted on anything at all.

So the object being certified is named: VYNDR certifies a PROCEDURE, not a
frozen curve. A frozen curve goes stale against a live model and has to be
replaced by hand on no schedule; "fit past, apply forward" is procedural by
construction. `fitPolicy.POLICY_V1` declares it — data selection, horizon,
algorithm and version, minimum rows, model-version restriction, sport, stat, and
a support contract that a refit may NOT widen. Each artifact still carries its
own digest.

`fitPolicy.validate` is the Step-22 gate: an artifact does not become servable
because the algorithm ran. It refuses a widened support, a wrong era, a wrong
estimator, a thin fit, a missing identity or a missing training cutoff, and
`servable` is false whenever any violation stands, with no override argument.
The statistical bars stay where they already live in calibrationRegistry — this
is not a second governance system.

MEASURED, AND THE REASON LIVE SERVING STAYS BLOCKED: production does not match
the declaration. `loadSettledRows` applies no model_version filter, so at
fit_as_of 2026-09-02 the fit drew 6,084 rows from 9,361 settled — all 3,292 from
the superseded engine1@2026-07-20 plus 2,792 current-era. 54.1% of the served
map is fitted on a forecaster it was never certified for, while the artifact
declares the current era.

That is a provenance contradiction, not a performance claim: era-filtered scores
0.24065 against pooled 0.24068 on 1,120 out-of-sample rows and both intervals
span zero. It is blocked because nothing prevents the next era change from
repeating it, and because the freshness lag grows.

`era_restricted` is answered STRUCTURALLY, not by an extra read — the query is
in this service and applies no filter, so the artifact records
ERA_NOT_RESTRICTED rather than claiming a restriction that did not hold. An
unverified restriction is recorded as a violation, because "we did not check" is
exactly the state production is in.

The violation does NOT distort the shadow. Flipping every row to UNCERTIFIED
would make the shadow measure the violation instead of the contract, so the
policy state rides beside the resolution and a test asserts the shadow still
reads CERTIFIED_CALIBRATED at 0.65 and UNCERTIFIED at 0.91.

Nothing serves. CALIBRATION_DEPLOYED still []. Shadow still OFF (the production
variable remains absent — the probe reads configuration_source "default").

Suite 402/402, 5,607 passed, 4 skipped. Teeth 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 23:43:31 -04:00
builtbykev 556d186ff1 One evaluator for the shadow, so the probe cannot report a mode the pipeline is not in
"Is the shadow effective?" was only answerable by waiting for a snapshot to
write a row. That leaves a blind spot with real cost: a variable SET IN COOLIFY
BUT NOT YET APPLIED to the running process is indistinguishable from an unset
one, and the runtime probe already proves the distinction matters — code_sha
22cf51c with started_at 01:45:44Z means anything set after that is not in this
process's environment.

probabilityContract.shadowState() is now the single evaluator. snapshotService
calls it and the protected status probe calls it, and a test asserts NEITHER
reads process.env directly — the same rule that keeps lineage_write_mode honest.
Reading the env in two places is how a status page and a gate come to disagree.

Strict by construction: only the exact string '1' enables it. 'true', 'yes',
'on', '01', ' 1 ' and '' are all OFF, because a loose parse turns a typo into an
activation. `configuration_source` separates an unset variable from one
explicitly set to '0', and `live_serving` is reported as its own switch so the
shadow can never be read as implying serving.

No behaviour changes. The shadow still defaults OFF, CALIBRATION_DEPLOYED is
still [], and served fields are untouched.

Frontend byte-identical to the last green build (git reports zero changes under
web/), so the build from 22cf51c stands.

Suite 401/401, 5,597 passed, 4 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 22:53:21 -04:00
builtbykev 22cf51c4b0 The mapping that ran now has a name, and the row carries it
Runtime probes say the fleet is on a8de676, one process generation. But
"isotonic" was a label, not a claim: probabilityContractService refits per
snapshot against `game_date < todayEt()`, so the mapping changes as outcomes
settle, and nothing on a row could say WHICH mapping produced its number.

The artifact now has an identity:

  estimator_type / estimator_version / certification_version / model_version
  fit_as_of         the exact lt(game_date) bound   2026-09-02
  training_cutoff   last date INSIDE the fit        2026-08-21
  fit_n / knot_count                                6,084 / 28
  knot_digest       d9d571d728ba76de
  served_curve      the COMPLETE served function over [0.50,0.80)
  served_curve_digest

The served curve is not a sample. p_win is quantised to three decimals at the
source, so a step table at 0.001 granularity is the mapping itself for every
input that can occur — six steps, ~200 bytes. Storing it makes a Read
reconstructable WITHOUT re-deriving a training set that may since have been
re-settled, and a claim you can only verify when the inputs happen not to have
moved is not a reconstructable claim.

Proven, not asserted: the production construction path run twice gives an
identical digest, and an INDEPENDENT reconstruction — re-walk 9,361 settled
ledger rows at the declared bound, refit from scratch — reproduces
d9d571d728ba76de exactly, 28 knots for 28.

A teeth injection found a real defect behind a coverage hole. `resolve` checked
the CONTRACT's model era and never the ARTIFACT's, so a mapping fitted for a
different era could be recorded beside a served number with every test green.
Both the era and the estimator type are now checked, and a mismatch serves
nothing rather than serving quietly.

OBSERVED AND NOT CHANGED: calibrationService splits 65/35 to certify its own
bands, a step this contract does not consume because support comes from the
frozen artifact. So the served map is fitted through 2026-08-21 while 3,277
more recent settled rows sit unused, and that lag grows with history. Changing
it would change the fitted function, which this tranche froze.

Shadow still defaults OFF. CALIBRATION_DEPLOYED still []. served_probability is
referenced by nothing outside the contract layer — asserted by a tooth.

Suite 401/401, 5,593 passed, 4 skipped. Teeth 23/23 (prior) + 7/7 (new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:44:52 -04:00