The live machinery ships dark, behind two gates that cannot substitute for each other

ATTRIBUTION RECEIPT (cohort 27ce152f, writer 94f7c3c, 01:02:19Z): 2,316 non-hits
rows — certified 0, numeric 0, and artifact_id / estimator / certification
version / source / knot / curve ALL ZERO. The hits slice held: 209 published,
201 certified, curve mismatches 0/201, artifact identity deviations 0, support
violations 0, raw erased 0. Shadow tranche frozen.

AUTHENTICATED NON-LEAK: obtained through the owner magic-link flow — the real
auth system, no bypass, no stored password, token never printed, session
discarded. /api/ledger, /api/ledger/accuracy, /api/preferences, /api/accuracy
and the authenticated /api/snapshot/mlb (677KB, full model fields) carry zero
shadow keys. One scanner hit adjudicated: `served_grade.calibrated` is false on
all 504 rows — servedGrade's hardcoded constant from before this work, a name
collision with my keyword list, not the shadow.

A REAL MONITOR DEFECT, asked for and found. The row floor alone let 400
observations from ONE slate reach HEALTHY or DRIFT — one correlated draw wearing
the costume of four hundred. The monitor now also requires settled DATES, and
the floor is not invented: it reads
`fitPolicy.POLICY_V1.certification.eval_block_dates`, the 3-date fold the
walk-forward was actually certified with. One date and two dates now return
INSUFFICIENT_SAMPLE regardless of row count.

That fix had a bug of its own that a test caught: `scored` never carried `date`,
so the distinct-date count read `undefined` for every row and always returned 1.
The gate looked correct while measuring nothing.

THE SEAM. EV, Kelly and VALUE are produced in exactly one place at grade time,
all from p_win, and every served row passes exactly one boundary on the way out
(`snapshotGating.stripModelPrice`, used by the snapshot route, hero route,
topGraded and props). So `servedProbability.applyToRows` sits there — one place,
not four call sites and four chances to miss one. Consumers never see the
artifact, the curve, the support region or the environment; a test forbids
`applyCurve`, `applyIsotonic`, `artifactRegistry` and `served_curve` in all three
serving consumers.

Placed BEFORE the tier strip deliberately: calibration decides what the number
IS, entitlement decides who may see it. Reversed, it would calibrate fields that
had already been removed.

TWO GATES, NEITHER SUFFICIENT. The artifact is promoted APPROVED_FOR_LIVE — stage
only, same id, same source/knot/curve digests, same cutoff, no refit. Behaviour
is still OFF because PROBABILITY_CONTRACT_LIVE is unset. Flag without approval:
OFF. Approval without flag: OFF. Both: ON. Live OFF returns rows byte-identical
and consults no artifact.

Two teeth had to be rewritten rather than satisfied: they pinned "artifact
unapproved" as the safety property, which would have blocked the deliberate
promotion. The property is that live BEHAVIOUR is off, and they now test that.
And my own splice while fixing one of them silently deleted ten newly-added
teeth — caught by counting ids, not by the runner going green.

Suite 406/406, 5,681 passed. Teeth 41/41 + 10/10 + 23/23. Live OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
Kev
2026-09-04 06:36:11 -04:00
parent 94f7c3c3ef
commit 84f1fc075c
12 changed files with 470 additions and 55 deletions
+16 -6
View File
@@ -81,16 +81,24 @@ describe('no environment variable can override the gate', () => {
});
describe('one validity decision for shadow AND live', () => {
it('a shadow-approved artifact is refused for live use', () => {
expect(good.approved_for_shadow).toBe(true);
expect(good.approved_for_live).toBe(false);
const live = pc.resolve(read(0.65), { estimate: est, artifact: good, usage: 'live' });
it('an artifact promoted only for shadow is refused for live use', () => {
// the SHIPPED artifact is now APPROVED_FOR_LIVE (promoted 2026-09-04), so
// the stage gate is exercised against a shadow-only artifact explicitly.
const shadowOnly = { ...good, approved_for_live: false, approved_for_shadow: true };
const live = pc.resolve(read(0.65), { estimate: est, artifact: shadowOnly, usage: 'live' });
expect(live.probability_state).toBe(pc.STATE.ARTIFACT_POLICY_BLOCKED);
expect(live.reason).toContain('not promoted for live');
const shadow = pc.resolve(read(0.65), { estimate: est, artifact: good, usage: 'shadow' });
const shadow = pc.resolve(read(0.65), { estimate: est, artifact: shadowOnly, usage: 'shadow' });
expect(shadow.probability_state).toBe(pc.STATE.CERTIFIED_CALIBRATED);
});
it('the shipped artifact IS promoted for live — and that alone changes nothing', () => {
expect(good.approved_for_shadow).toBe(true);
expect(good.approved_for_live).toBe(true);
// approval is one gate; the runtime flag is the other, and it is unset
expect(pc.liveState({}).live).toBe('OFF');
});
it('live consumes the SAME servable decision, not a parallel one', () => {
const bad = { ...good, servable: false, approved_for_live: true, fit_policy_violations: ['X'] };
expect(pc.resolve(read(0.65), { estimate: est, artifact: bad, usage: 'live' }).probability_state)
@@ -128,6 +136,7 @@ describe('promotion is a deliberate act', () => {
it('the promotion table is a frozen source constant, not runtime state', () => {
expect(Object.isFrozen(registry.PROMOTED)).toBe(true);
expect(Object.isFrozen(registry.PROMOTED['mlb:hits'])).toBe(true);
expect(registry.PROMOTED['mlb:hits'].stage).toBe(registry.STAGE.APPROVED_FOR_LIVE);
// strict mode makes the write throw rather than fail silently — either way
// the table is unchanged, which is the property under test
expect(() => { registry.PROMOTED['mlb:rbi'] = { artifact_id: 'x', stage: 'APPROVED_FOR_LIVE' }; }).toThrow();
@@ -184,7 +193,8 @@ describe('the runtime can name its artifact with the shadow OFF', () => {
'certified_bands', 'stage']) expect(a[k]).toBeTruthy();
expect(a.servable).toBe(true);
expect(a.approved_for_shadow).toBe(true);
expect(a.approved_for_live).toBe(false);
// promoted 2026-09-04; the runtime flag remains the second, unset gate
expect(a.approved_for_live).toBe(true);
expect(a.wrong_era_rows).toBe(0);
expect(a.withheld_from_fit).toBe(0);
});
+55 -2
View File
@@ -10,7 +10,11 @@ const fm = require('../../src/services/model/forwardMonitor');
const registry = require('../../src/services/model/artifactRegistry');
const A = registry.load('mlb', 'hits');
const AFTER = '2026-09-02'; // strictly after training_cutoff 2026-09-01
// Strictly after training_cutoff 2026-09-01. THREE dates, because the monitor
// requires one certification-equivalent block of date evidence — a single-date
// fixture can never reach a verdict, by design.
const AFTER_DATES = ['2026-09-02', '2026-09-03', '2026-09-04'];
const AFTER = AFTER_DATES[0];
/** Rows whose outcomes track the frozen curve, optionally shifted to force drift. */
function rows(n, shift = 0, over = {}) {
@@ -18,7 +22,7 @@ function rows(n, shift = 0, over = {}) {
const raw = Math.round((0.50 + (i % 29) / 100) * 1000) / 1000;
const served = registry.applyCurve(A, raw) ?? 0.6;
return { p: raw, won: ((i * 2654435761) % 1000) / 1000 < served + shift ? 1 : 0,
date: AFTER, model_version: A.model_version, ...over };
date: AFTER_DATES[i % AFTER_DATES.length], model_version: A.model_version, ...over };
});
}
@@ -158,3 +162,52 @@ describe('THE CALLSITE — production actually runs it', () => {
}
});
});
describe('row count must not masquerade as temporal evidence', () => {
const fp = require('../../src/services/model/fitPolicy');
const many = (n, dates) => Array.from({ length: n }, (_, i) => {
const raw = Math.round((0.50 + (i % 29) / 100) * 1000) / 1000;
const served = registry.applyCurve(A, raw) ?? 0.6;
return { p: raw, won: ((i * 2654435761) % 1000) / 1000 < served ? 1 : 0,
date: dates[i % dates.length], model_version: A.model_version };
});
it('the date floor is READ FROM the procedure, not invented here', () => {
expect(fm.MIN_FORWARD_DATES).toBe(fp.POLICY_V1.certification.eval_block_dates);
expect(fm.MIN_FORWARD_DATES).toBe(3); // the walk-forward's fold width
});
it('1,200 rows from ONE slate is not a verdict', () => {
const r = fm.evaluate(A, many(1200, ['2026-09-02']), registry.applyCurve);
expect(r.n).toBeGreaterThanOrEqual(fm.MIN_FORWARD_ROWS); // rows are ample
expect(r.health).toBe(fm.HEALTH.INSUFFICIENT_SAMPLE); // dates are not
expect(r.healthy).toBeNull();
expect(r.settled_date_count).toBe(1);
expect(r.reason).toContain('settled date');
});
it('two dates is still not enough', () => {
const r = fm.evaluate(A, many(1200, ['2026-09-02', '2026-09-03']), registry.applyCurve);
expect(r.health).toBe(fm.HEALTH.INSUFFICIENT_SAMPLE);
expect(r.settled_date_count).toBe(2);
});
it('one certification-equivalent block of dates unlocks a verdict', () => {
const r = fm.evaluate(A, many(1200, ['2026-09-02', '2026-09-03', '2026-09-04']), registry.applyCurve);
expect(r.settled_date_count).toBe(3);
expect([fm.HEALTH.HEALTHY, fm.HEALTH.DRIFT_WARNING]).toContain(r.health);
});
it('enough dates does NOT rescue a thin row count', () => {
const r = fm.evaluate(A, many(50, ['2026-09-02', '2026-09-03', '2026-09-04']), registry.applyCurve);
expect(r.health).toBe(fm.HEALTH.INSUFFICIENT_SAMPLE);
expect(r.healthy).toBeNull();
});
it('the date count is computed from the SCORED rows, not from undefined', () => {
// Without `date` on the scored row every row hashes to undefined and the
// count is always 1 — the gate would look right while measuring nothing.
const r = fm.evaluate(A, many(900, ['2026-09-02', '2026-09-03', '2026-09-04']), registry.applyCurve);
expect(r.settled_dates).toEqual(['2026-09-02', '2026-09-03', '2026-09-04']);
});
});
+125
View File
@@ -0,0 +1,125 @@
'use strict';
/**
* ONE SEAM, TWO GATES, AND A DARK DEFAULT.
*/
const sp = require('../../src/services/model/servedProbability');
const pc = require('../../src/services/model/probabilityContract');
const registry = require('../../src/services/model/artifactRegistry');
const gating = require('../../src/utils/snapshotGating');
const ERA = 'engine1@2026-08-07-fullwindow';
const A = registry.load('mlb', 'hits');
const row = (over = {}) => ({ sport: 'mlb', stat_type: 'hits', model_version: ERA,
p_win: 0.65, confidence: 65, ev_pct: 12.3, value: true, takeable: true, model_odds: -186,
book_odds: -110, grade: 'C+', side: 'over', player: 'P', line: 0.5, ...over });
const LIVE_ON = { live: 'ON' };
describe('live is OFF by default and that is load-bearing', () => {
it('the default state is OFF with the flag named as the reason', () => {
const st = pc.liveState({});
expect(st.live).toBe('OFF');
expect(st.flag_set).toBe(false);
expect(st.blocked_reason).toBe('runtime flag not set');
});
it('rows pass through BYTE-IDENTICAL with live off', () => {
const rows = [row(), row({ p_win: 0.91 }), row({ stat_type: 'total_bases' })];
const before = JSON.stringify(rows);
expect(JSON.stringify(sp.applyToRows(rows))).toBe(before);
expect(JSON.stringify(gating.stripModelPrice(rows, 'desk'))).toBe(before);
});
it('only the exact string 1 sets the flag', () => {
for (const v of ['true', 'yes', 'on', '0', ' 1 ', '', '01']) {
expect(pc.liveState({ PROBABILITY_CONTRACT_LIVE: v }).flag_set).toBe(false);
}
expect(pc.liveState({ PROBABILITY_CONTRACT_LIVE: '1' }).flag_set).toBe(true);
});
});
describe('BOTH gates are required — neither activates alone', () => {
it('the runtime flag alone does NOT serve an unapproved artifact', () => {
const st = pc.liveState({ PROBABILITY_CONTRACT_LIVE: '1' },
{ load: () => ({ ...A, approved_for_live: false }) });
expect(st.flag_set).toBe(true);
expect(st.live).toBe('OFF');
expect(st.blocked_reason).toMatch(/not approved_for_live|not servable/);
});
it('the runtime flag alone does NOT serve an unservable artifact', () => {
const st = pc.liveState({ PROBABILITY_CONTRACT_LIVE: '1' },
{ load: () => ({ ...A, approved_for_live: true, servable: false }) });
expect(st.live).toBe('OFF');
});
it('artifact approval alone does NOT activate behaviour', () => {
expect(A.approved_for_live).toBe(true); // it IS promoted
expect(pc.liveState({}).live).toBe('OFF'); // and behaviour is still off
});
it('both together turn it on', () => {
expect(pc.liveState({ PROBABILITY_CONTRACT_LIVE: '1' }).live).toBe('ON');
});
});
describe('when live is on', () => {
it('a CERTIFIED row serves the calibrated number and derives everything from it', () => {
const [out] = sp.applyToRows([row()], { liveState: LIVE_ON });
const expected = registry.applyCurve(A, 0.65);
expect(out.p_win).toBe(Math.round(expected * 1000) / 1000);
expect(out.confidence).toBe(Math.round(out.p_win * 100));
expect(out.confidence_basis).toBe('served_probability');
expect(out.probability_state).toBe(pc.STATE.CERTIFIED_CALIBRATED);
const { evPct } = require('../../src/utils/devig');
expect(out.ev_pct).toBe(evPct(out.p_win, -110));
expect(out.ev_pct).not.toBe(evPct(0.65, -110)); // NOT from raw
expect(out.raw_model_probability).toBe(0.65); // raw survives
expect(out.grade).toBe('C+'); // grade untouched
expect(out.side).toBe('over'); // side untouched
});
it('an UNCERTIFIED row loses the number AND every claim derived from it', () => {
const [out] = sp.applyToRows([row({ p_win: 0.91, grade: 'B+' })], { liveState: LIVE_ON });
expect(out.p_win).toBeNull();
expect(out.confidence).toBeNull();
expect(out.ev_pct).toBeNull();
expect(out.value).toBeNull();
expect(out.model_odds).toBeNull();
expect(out.kelly).toBeNull();
// and what must survive
expect(out.raw_model_probability).toBe(0.91);
expect(out.grade).toBe('B+');
expect(out.side).toBe('over');
expect(out.book_odds).toBe(-110);
expect(out.player).toBe('P'); // the Read still exists
});
it('a stat outside the contract is returned untouched', () => {
const tb = row({ stat_type: 'total_bases' });
const [out] = sp.applyToRows([tb], { liveState: LIVE_ON });
expect(out).toEqual(tb);
});
it('a different model era is returned untouched', () => {
const old = row({ model_version: 'engine1@2026-07-20' });
const [out] = sp.applyToRows([old], { liveState: LIVE_ON });
expect(out).toEqual(old);
});
});
describe('one seam', () => {
it('serving routes reach it through stripModelPrice, not by duplicating logic', () => {
const fs = require('fs'); const path = require('path');
const gsrc = fs.readFileSync(path.join(__dirname, '../../src/utils/snapshotGating.js'), 'utf8');
expect(gsrc).toContain("require('../services/model/servedProbability').applyToRows(rows)");
// no consumer may hold calibration logic of its own
for (const f of ['src/routes/snapshot.js', 'src/routes/heroProp.js', 'src/services/topGradedService.js']) {
const src = fs.readFileSync(path.join(__dirname, '../../', f), 'utf8')
.replace(/\/\*[\s\S]*?\*\//g, '').replace(/^\s*\/\/.*$/gm, '');
for (const forbidden of ['applyCurve', 'applyIsotonic', 'artifactRegistry', 'served_curve']) {
expect(src).not.toContain(forbidden);
}
}
});
});