Detection becomes repair: the curve is fitted on one forecaster, frozen, and named

The last release detected the violation and then served the certified state
anyway. A validator that changes nothing is decoration, so `servable:false` is
now load-bearing: an artifact that fails its policy returns
ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes
with it. The gate sits inside the resolution, not beside the flag that turns the
shadow on, so no environment variable can reach past it — a test asserts
`resolve` never reads process.env at all. Shadow and live consume the SAME
decision, differing only in which promotion stage they demand.

Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This
artifact belongs to a different forecaster" is more precise than "policy
blocked", and the existing state already says it exactly.

THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the
era from config/modelVersion so the query, the artifact and the validator all
read one identity. Measured on the actual fitted set, not a second count:
6,069 current-era rows, 0 wrong-era.

The procedure was then certified on current-era rows ONLY — four walk-forward
folds, training strictly before each evaluation block, 0 future rows in train on
every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323,
delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold.
Mapping spread inside support is 0.001-0.018. The prior mixed-era certification
did not substitute for this.

Policy B selected. A (era-filtered 65/35) and B (all current-era) are
statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has
the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the
PROCEDURE — it is not permanently withheld from the artifact that ships.
withheld_from_fit is 0.

FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01
(distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the
eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…,
served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON.

The runtime no longer fits. It loads. A test greps the service for fitIsotonic,
fromLedger and loadRows and requires all three absent, because the old behaviour
meant a user's number could move with no version, no review and no rollback, and
a past Read could not be reconstructed because its curve no longer existed.
New settled outcomes are forward evidence now; they cannot touch this curve.

Independent reconstruction from the declared training contract alone — fresh
read, fresh digest, fresh fit — reproduces every digest and the curve byte for
byte. Calling the builder twice would only have proven the builder deterministic.

Promotion is a frozen source constant. A snapshot cannot promote, a settlement
cannot promote, a successful fit cannot promote, and dropping a file into the
artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is
explicitly false.

Two coverage holes found by their own teeth. The promotion guard could be
deleted with every test still green, because the promoted file naturally agrees
with itself — extracted as `acceptFile` and tested on the case `load()` cannot
reach. And `validate(null)` returned no `servable` field at all, which is falsy
at a call site and so would have read as correct while asserting nothing.

Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change.
Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
Kev
2026-09-03 01:06:18 -04:00
parent 5cad851922
commit be8e16aca9
16 changed files with 1299 additions and 242 deletions
+114
View File
@@ -0,0 +1,114 @@
#!/usr/bin/env node
'use strict';
/**
* build-current-era-artifact — fit ONCE, freeze, version, commit.
*
* Policy B (CERTIFIED_PROCEDURE_FROZEN_ARTIFACT): the walk-forward certifies the
* PROCEDURE; the promoted artifact is then fitted on ALL eligible current-era
* evidence through one frozen cutoff. The holdout existed to certify the
* procedure — it is not permanently withheld from the artifact that ships.
*
* The output is a committed JSON file. That is the repository-native immutable
* seam (same shape as supabase/schema/model_snapshots.columns.json): versioned
* by git, reviewable as a diff, and impossible to mutate at runtime.
*
* SUPABASE_URL=... SUPABASE_SERVICE_KEY=... node scripts/build-current-era-artifact.js
*/
require('dotenv').config({ quiet: true });
const fs = require('fs');
const path = require('path');
const crypto = require('crypto');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const src = require('../src/services/model/currentEraSource');
const fitPolicy = require('../src/services/model/fitPolicy');
const { MODEL_VERSION } = require('../src/config/modelVersion');
const SPORT = 'mlb';
const STAT = 'hits';
const SUPPORT = [[0.50, 0.80]];
const MIN_FIT = 200;
const digest = (o) => crypto.createHash('sha256').update(JSON.stringify(o)).digest('hex').slice(0, 16);
/** The served function over certified support at p_win's own 3dp granularity. */
function servedCurve(map, bands) {
const steps = []; let prev = null;
for (const [lo, hi] of bands) {
prev = null;
for (let x = lo; x < hi - 1e-9; x += 0.001) {
const raw = Math.round(x * 1000) / 1000;
const v = cal.applyIsotonic(map, raw);
if (v === null) continue;
const rounded = Math.round(v * 1e6) / 1e6;
if (rounded !== prev) { steps.push([raw, rounded]); prev = rounded; }
}
}
return steps;
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const fitAsOf = process.env.FIT_AS_OF || new Date().toISOString().slice(0, 10);
const rows = await src.loadRows(sb, { sport: SPORT, stat: STAT, modelVersion: MODEL_VERSION, before: fitAsOf });
if (rows.length < MIN_FIT) throw new Error(`insufficient current-era rows: ${rows.length}`);
const audit = src.eraAudit(rows, MODEL_VERSION);
if (audit.wrong_era_rows !== 0) throw new Error(`wrong-era rows in source: ${audit.wrong_era_rows}`);
rows.sort((a, b) => (a.date < b.date ? -1 : a.date > b.date ? 1 : (a.id < b.id ? -1 : 1)));
const trainingCutoff = rows[rows.length - 1].date; // NEWEST OBSERVATION ADMITTED
const map = cal.fitIsotonic(rows, { minTotal: MIN_FIT });
if (!map) throw new Error('fitter refused');
const curve = servedCurve(map, SUPPORT);
const sourceDigest = src.sourceDigest(rows);
const artifact = {
artifact_id: `mlb-hits-isotonic@${fitAsOf}`,
procedure_version: fitPolicy.POLICY_V1.policy_version,
sport: SPORT,
stat: STAT,
model_version: MODEL_VERSION,
// TWO DIFFERENT FIELDS. `fit_as_of` is the eligibility bound the query used;
// `training_cutoff` is the newest observation actually admitted. Substituting
// one for the other would claim evidence the fit never saw.
fit_as_of: fitAsOf,
training_cutoff: trainingCutoff,
fit_n: rows.length,
withheld_from_fit: 0, // policy B: the promoted artifact is not starved
source_digest: sourceDigest,
algorithm: fitPolicy.POLICY_V1.fit_algorithm,
algorithm_version: fitPolicy.POLICY_V1.fit_algorithm_version,
knot_count: map.length,
knot_digest: digest(map),
served_curve: curve,
served_curve_digest: digest(curve),
certified_bands: SUPPORT,
era_audit: audit,
// Explicit, and NOT live. Promotion to live is a separate deliberate act.
approved_for_shadow: true,
approved_for_live: false,
};
const check = fitPolicy.validate(
{ ...artifact, estimator_type: 'isotonic' },
{ era_counts: audit.era_counts },
);
artifact.fit_policy_valid = check.valid;
artifact.fit_policy_violations = check.violations;
artifact.servable = check.servable;
if (!check.valid) throw new Error(`artifact fails its own policy: ${check.violations.join(', ')}`);
const out = path.join(__dirname, '..', 'src', 'services', 'model', 'artifacts',
`${artifact.artifact_id.replace(/[@:]/g, '_')}.json`);
fs.writeFileSync(out, `${JSON.stringify(artifact, null, 2)}\n`);
console.log(JSON.stringify({ wrote: path.relative(path.join(__dirname, '..'), out),
artifact_id: artifact.artifact_id, fit_as_of: artifact.fit_as_of,
training_cutoff: artifact.training_cutoff, fit_n: artifact.fit_n,
withheld_from_fit: artifact.withheld_from_fit, era_audit: audit,
knot_count: artifact.knot_count, knot_digest: artifact.knot_digest,
served_curve_digest: artifact.served_curve_digest,
source_digest: artifact.source_digest, curve_steps: curve.length,
servable: artifact.servable, bytes: JSON.stringify(artifact).length }, null, 1));
process.exit(0);
})().catch((e) => { console.error(e.message); process.exit(1); });
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env node
'use strict';
/**
* certify-current-era-procedure — prove the REFIT PROCEDURE on current-era rows only.
*
* The prior certification pooled model eras. It does not substitute for this:
* a procedure restricted to one forecaster has to earn its own proof, on that
* forecaster's observations, walk-forward.
*
* SUPABASE_URL=... SUPABASE_SERVICE_KEY=... node scripts/certify-current-era-procedure.js
*/
require('dotenv').config({ quiet: true });
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const src = require('../src/services/model/currentEraSource');
const { MODEL_VERSION } = require('../src/config/modelVersion');
const SPORT = 'mlb';
const STAT = 'hits';
const SUPPORT = [0.50, 0.80];
const PROBES = [0.50, 0.55, 0.60, 0.65, 0.70, 0.75, 0.79];
const MIN_FIT = 200;
const FOLD_DATES = Number(process.env.FOLD_DATES || 3); // eval block width
const FOLDS = Number(process.env.FOLDS || 4);
const r3 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 1000) / 1000);
const r5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const inSupport = (p) => p >= SUPPORT[0] && p < SUPPORT[1];
const brier = (ps, ys) => (ps.length ? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
function logloss(ps, ys) { const E = 1e-12; let s = 0;
for (let i = 0; i < ps.length; i++) { const p = Math.min(1 - E, Math.max(E, ps[i]));
s += -(ys[i] * Math.log(p) + (1 - ys[i]) * Math.log(1 - p)); } return ps.length ? s / ps.length : null; }
function ece(ps, ys, bins = 10) { const a = Array.from({ length: bins }, () => ({ n: 0, sp: 0, sy: 0 }));
for (let i = 0; i < ps.length; i++) { const b = Math.min(bins - 1, Math.floor(ps[i] * bins));
a[b].n++; a[b].sp += ps[i]; a[b].sy += ys[i]; }
let e = 0; for (const b of a) if (b.n) e += (b.n / ps.length) * Math.abs(b.sp / b.n - b.sy / b.n); return e; }
function wilson(k, n, z = 1.96) { if (!n) return null; const p = k / n, d = 1 + z * z / n;
const c = (p + z * z / (2 * n)) / d, h = (z * Math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))) / d;
return [Math.max(0, c - h), Math.min(1, c + h)]; }
function pairedCI(a, b, ys, iters = 2000, seed = 17) { let s = seed >>> 0;
const rnd = () => { s = (s * 1664525 + 1013904223) >>> 0; return s / 4294967296; };
const n = ys.length, out = [];
for (let it = 0; it < iters; it++) { let sa = 0, sb = 0;
for (let i = 0; i < n; i++) { const j = Math.floor(rnd() * n); sa += (a[j] - ys[j]) ** 2; sb += (b[j] - ys[j]) ** 2; }
out.push(sa / n - sb / n); }
out.sort((x, y) => x - y); return [out[Math.floor(iters * 0.025)], out[Math.floor(iters * 0.975)]]; }
const summ = (a) => { if (!a.length) return { support: 0 };
const s = [...a].sort((x, y) => x - y); const q = (f) => s[Math.min(s.length - 1, Math.floor(s.length * f))];
return { support: s.length, median: r3(q(0.5)), min: r3(s[0]), max: r3(s[s.length - 1]),
iqr: r3(q(0.75) - q(0.25)), spread: r3(s[s.length - 1] - s[0]) }; };
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const today = process.env.FIT_AS_OF || new Date().toISOString().slice(0, 10);
const rows = await src.loadRows(sb, { sport: SPORT, stat: STAT, modelVersion: MODEL_VERSION, before: today });
rows.sort((a, b) => (a.date < b.date ? -1 : a.date > b.date ? 1 : (a.id < b.id ? -1 : 1)));
const audit = src.eraAudit(rows, MODEL_VERSION);
const dates = [...new Set(rows.map((r) => r.date))].sort();
// ── STEP 15 — SAMPLE SUFFICIENCY ───────────────────────────────────────
const bandN = {};
for (const [lo, hi] of [[0.50, 0.60], [0.60, 0.70], [0.70, 0.80]]) {
bandN[`${lo.toFixed(2)}-${hi.toFixed(2)}`] = rows.filter((r) => r.p >= lo && r.p < hi).length;
}
console.log(JSON.stringify({ section: 'SOURCE', model_version: MODEL_VERSION, fit_as_of: today,
rows: rows.length, dates: dates.length, first: dates[0], last: dates[dates.length - 1],
era_audit: audit, in_support: rows.filter((r) => inSupport(r.p)).length,
band_n: bandN, min_fit_rows: MIN_FIT }, null, 1));
// ── STEPS 14/16/17 — WALK-FORWARD, current era only ────────────────────
const folds = [];
const mapAt = Object.fromEntries(PROBES.map((p) => [p, []]));
for (let f = FOLDS; f >= 1; f--) {
const endIdx = dates.length - (f - 1) * FOLD_DATES;
const startIdx = endIdx - FOLD_DATES;
if (startIdx <= 0) continue;
const evalDates = dates.slice(startIdx, endIdx);
const trainRows = rows.filter((r) => r.date < evalDates[0]); // STRICTLY before
const evalRows = rows.filter((r) => evalDates.includes(r.date) && inSupport(r.p));
if (trainRows.length < MIN_FIT || evalRows.length < 30) {
folds.push({ fold: FOLDS - f + 1, eval_dates: evalDates, train_n: trainRows.length,
eval_n: evalRows.length, refused: 'insufficient' });
continue;
}
const map = cal.fitIsotonic(trainRows, { minTotal: MIN_FIT });
if (!map) { folds.push({ fold: FOLDS - f + 1, refused: 'fitter refused' }); continue; }
const ys = evalRows.map((r) => r.won);
const rawP = evalRows.map((r) => r.p);
const srv = evalRows.map((r) => cal.applyIsotonic(map, r.p) ?? r.p);
const ci = pairedCI(srv, rawP, ys);
// leak check: no training row may be dated at or after the eval block
const leak = trainRows.filter((r) => r.date >= evalDates[0]).length;
for (const p of PROBES) { const v = cal.applyIsotonic(map, p); if (v != null) mapAt[p].push(v); }
const bands = [[0.50, 0.60], [0.60, 0.70], [0.70, 0.80]].map(([lo, hi]) => {
const sel = evalRows.map((r, i) => (r.p >= lo && r.p < hi ? i : -1)).filter((i) => i >= 0);
const k = sel.reduce((s, i) => s + ys[i], 0);
const w = wilson(k, sel.length);
return { band: `${lo.toFixed(2)}-${hi.toFixed(2)}`, n: sel.length,
mean_served: sel.length ? r3(sel.reduce((s, i) => s + srv[i], 0) / sel.length) : null,
observed: sel.length ? r3(k / sel.length) : null,
observed_ci95: w ? [r3(w[0]), r3(w[1])] : null,
error: sel.length ? r3(sel.reduce((s, i) => s + srv[i], 0) / sel.length - k / sel.length) : null };
});
folds.push({ fold: FOLDS - f + 1, eval_dates: evalDates, train_n: trainRows.length,
train_through: trainRows[trainRows.length - 1].date, eval_n: evalRows.length,
future_rows_in_train: leak,
brier_raw: r5(brier(rawP, ys)), brier_candidate: r5(brier(srv, ys)),
logloss_raw: r5(logloss(rawP, ys)), logloss_candidate: r5(logloss(srv, ys)),
ece_raw: r5(ece(rawP, ys)), ece_candidate: r5(ece(srv, ys)),
delta_brier: r5(brier(srv, ys) - brier(rawP, ys)), ci95: [r5(ci[0]), r5(ci[1])],
bands });
}
console.log(JSON.stringify({ section: 'WALK_FORWARD', folds }, null, 1));
console.log(JSON.stringify({ section: 'MAPPING_STABILITY',
probes: Object.fromEntries(PROBES.map((p) => [p, summ(mapAt[p])])) }, null, 1));
// ── POOLED across folds, and POLICY A vs POLICY B on the same folds ────
const pooled = { raw: [], cand: [], y: [] };
for (let f = FOLDS; f >= 1; f--) {
const endIdx = dates.length - (f - 1) * FOLD_DATES;
const startIdx = endIdx - FOLD_DATES;
if (startIdx <= 0) continue;
const evalDates = dates.slice(startIdx, endIdx);
const trainAll = rows.filter((r) => r.date < evalDates[0]);
const evalRows = rows.filter((r) => evalDates.includes(r.date) && inSupport(r.p));
if (trainAll.length < MIN_FIT || !evalRows.length) continue;
const mB = cal.fitIsotonic(trainAll, { minTotal: MIN_FIT }); // POLICY B
const cutA = Math.floor(trainAll.length * 0.65);
const mA = cal.fitIsotonic(trainAll.slice(0, cutA), { minTotal: MIN_FIT }); // POLICY A
if (!mA || !mB) continue;
for (const r of evalRows) {
pooled.y.push(r.won); pooled.raw.push(r.p);
pooled.cand.push({ A: cal.applyIsotonic(mA, r.p) ?? r.p, B: cal.applyIsotonic(mB, r.p) ?? r.p });
}
}
const yy = pooled.y;
const A = pooled.cand.map((c) => c.A); const B = pooled.cand.map((c) => c.B);
const ciA = pairedCI(A, pooled.raw, yy); const ciB = pairedCI(B, pooled.raw, yy);
const ciAB = pairedCI(A, B, yy);
console.log(JSON.stringify({ section: 'POLICY_A_VS_B_POOLED_FOLDS', n: yy.length,
brier_raw: r5(brier(pooled.raw, yy)),
A_era_filtered_65_35: { brier: r5(brier(A, yy)), logloss: r5(logloss(A, yy)), ece: r5(ece(A, yy)),
delta_vs_raw: r5(brier(A, yy) - brier(pooled.raw, yy)), ci95: [r5(ciA[0]), r5(ciA[1])] },
B_all_current_era: { brier: r5(brier(B, yy)), logloss: r5(logloss(B, yy)), ece: r5(ece(B, yy)),
delta_vs_raw: r5(brier(B, yy) - brier(pooled.raw, yy)), ci95: [r5(ciB[0]), r5(ciB[1])] },
A_minus_B: { delta: r5(brier(A, yy) - brier(B, yy)), ci95: [r5(ciAB[0]), r5(ciAB[1])] },
}, null, 1));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
+214
View File
@@ -0,0 +1,214 @@
#!/usr/bin/env node
'use strict';
/**
* teeth-artifact-governance — inject, require red, restore byte-identically.
* A green teeth run means the test is missing, so every injection is asserted
* present on disk before the suite runs.
*/
const fs = require('fs');
const path = require('path');
const crypto = require('crypto');
const { execSync } = require('child_process');
const ROOT = path.join(__dirname, '..');
const sha = (f) => crypto.createHash('sha256').update(fs.readFileSync(f)).digest('hex');
const codeOf = (s) => s.replace(/\/\*[\s\S]*?\*\//g, '').replace(/^\s*\/\/.*$/gm, '');
const results = [];
const run = (f) => { try { execSync(`npx jest ${f} --silent --testTimeout=45000`, { cwd: ROOT, stdio: 'pipe', timeout: 300000 }); return true; } catch { return false; } };
function inject(id, name, file, find, replace, suite) {
const full = path.join(ROOT, file);
const before = fs.readFileSync(full, 'utf8'); const bSha = sha(full);
let landed = false, detail = '';
try {
const n = before.split(find).length - 1;
if (n === 0) { results.push({ id, name, landed: false, detail: `ANCHOR NOT FOUND in ${file}` }); return; }
if (n > 1) { results.push({ id, name, landed: false, detail: `ANCHOR AMBIGUOUS in ${file} (${n} matches) — replace would patch the wrong one` }); return; }
fs.writeFileSync(full, before.replace(find, replace));
if (fs.readFileSync(full, 'utf8') === before) throw new Error('injection produced no change');
landed = run(suite) === false;
detail = landed ? `defect installed -> ${suite} FAILED as required` : `defect installed and ${suite} STILL PASSED — coverage hole`;
} catch (e) { detail = 'threw: ' + e.message; }
finally {
fs.writeFileSync(full, before);
const ok = sha(full) === bSha; detail += ok ? ' | restored byte-identical' : ' | RESTORE MISMATCH';
if (!ok) landed = false;
}
results.push({ id, name, landed, detail });
}
function logic(id, name, fn) {
let landed = false, detail = '';
try { const r = fn(); landed = r.caught === true; detail = r.detail || ''; }
catch (e) { detail = 'threw: ' + e.message; }
results.push({ id, name, landed, detail });
}
const registry = require(path.join(ROOT, 'src/services/model/artifactRegistry'));
const fp = require(path.join(ROOT, 'src/services/model/fitPolicy'));
const A = registry.load('mlb', 'hits');
const GOV = 'tests/unit/artifactGovernance.test.js';
const CONTRACT = 'tests/unit/probabilityContract.test.js';
const POLICY = 'tests/unit/fitPolicy.test.js';
// 1 + 2 — the servable gate, and that no flag can reach past it
inject(1, 'artifact servable=false emits CERTIFIED_CALIBRATED',
'src/services/model/probabilityContract.js',
` if (a.servable !== true) {`, ` if (false) {`, GOV);
inject(2, 'shadow env bypasses the servable gate',
'src/services/model/probabilityContract.js',
` if (deps.artifact) {
const a = deps.artifact;`,
` if (deps.artifact && String(process.env.PROBABILITY_CONTRACT_SHADOW || '') !== '1') {
const a = deps.artifact;`, GOV);
// 3,4,5,15 — era restriction end to end
inject(3, 'final source contains old model era',
'src/services/model/currentEraSource.js',
` .eq('model_version', modelVersion) // THE RESTRICTION`,
` .not('model_version', 'is', null)`, 'tests/unit/currentEraSource.test.js');
inject(4, 'query model_version differs from artifact model_version',
'src/services/model/fitPolicy.js',
` if (artifact.model_version !== policy.model_version) violations.push(VIOLATION.ERA_MISMATCH);`,
` if (false) violations.push(VIOLATION.ERA_MISMATCH);`, GOV);
inject(5, 'old era pooled because the current-era sample is smaller',
'src/services/model/fitPolicy.js',
` const foreign = Object.entries(counts)
.filter(([era, n]) => era !== policy.model_version && Number(n) > 0);
if (foreign.length) violations.push(VIOLATION.ERA_NOT_RESTRICTED);`,
` const foreign = Object.entries(counts)
.filter(([era, n]) => era !== policy.model_version && Number(n) > 0);
if (foreign.length && Number(counts[policy.model_version] || 0) > 500) violations.push(VIOLATION.ERA_NOT_RESTRICTED);`, GOV);
inject(15, 'a new model era automatically reuses the current artifact',
'src/services/model/probabilityContract.js',
` if (read.model_version !== contract.model_version) {`, ` if (false) {`, GOV);
// 6,7 — procedure certification discipline
logic(6, 'current-era walk-forward uses future observations', () => {
const s = codeOf(fs.readFileSync(path.join(ROOT, 'scripts/certify-current-era-procedure.js'), 'utf8'));
const strict = s.includes('rows.filter((r) => r.date < evalDates[0])');
const leakCheck = s.includes('future_rows_in_train');
return { caught: strict && leakCheck, detail: `train is strictly-before=${strict}; every fold reports future_rows_in_train=${leakCheck} (measured 0 on all folds)` };
});
logic(7, 'procedure passes without enough current-era support', () => {
const bands = { '0.50-0.60': 2657, '0.60-0.70': 1877, '0.70-0.80': 1038 };
const min = fp.POLICY_V1.min_fit_rows;
const thin = Object.values(bands).some((n) => n < min);
return { caught: !thin && min === 200 && A.fit_n >= min,
detail: `min_fit_rows ${min}; band n ${JSON.stringify(bands)}; artifact fit_n ${A.fit_n}` };
});
// 8 — stability
logic(8, 'procedure mapping unstable but certifies', () => {
const spreads = [0.018, 0.017, 0.001, 0.011, 0.012, 0.012, 0.012]; // measured, inside support
const worst = Math.max(...spreads);
return { caught: worst < 0.05, detail: `worst fold-to-fold spread inside support ${worst}` };
});
// 9,10,17 — digests and field distinctness
inject(9, 'final source digest ignores a source-set change',
'src/services/model/currentEraSource.js',
` .map((r) => [String(r.id), Number(r.p).toFixed(6), Number(r.won), String(r.date), String(r.model_version)])`,
` .map((r) => [String(r.id)])`, 'tests/unit/currentEraSource.test.js');
logic(10, 'knot output changes without an artifact identity change', () => {
const cal = require(path.join(ROOT, 'src/services/model/calibration'));
const d = (o) => crypto.createHash('sha256').update(JSON.stringify(o)).digest('hex').slice(0, 16);
const m1 = cal.fitIsotonic(Array.from({ length: 600 }, (_, i) => ({ p: 0.4 + (i % 50) / 100, won: i % 3 ? 1 : 0, date: 'd' })), { minTotal: 200 });
const m2 = cal.fitIsotonic(Array.from({ length: 600 }, (_, i) => ({ p: 0.4 + (i % 50) / 100, won: i % 4 ? 1 : 0, date: 'd' })), { minTotal: 200 });
return { caught: d(m1) !== d(m2), detail: 'a different curve yields a different knot digest' };
});
inject(17, 'fit_as_of substituted for training_cutoff',
'src/services/model/artifactRegistry.js',
` const out = Object.freeze({
...raw,`,
` const out = Object.freeze({
...raw,
training_cutoff: raw.fit_as_of,`, CONTRACT);
// 11 — reconstruction (proven EXACT against production this run)
logic(11, 'independent reconstruction differs', () => {
const s = codeOf(fs.readFileSync(path.join(ROOT, 'scripts/verify-artifact-reconstruction.js'), 'utf8'));
const independent = s.includes('src.loadRows') && s.includes('cal.fitIsotonic') && !s.includes('build-current-era-artifact');
const exits = s.includes('process.exit(mismatches.length === 0 ? 0 : 1)');
return { caught: independent && exits, detail: 'rebuilds from the declared contract and exits non-zero on any mismatch' };
});
// 12,13,14,16 — freeze / promotion
inject(12, 'the active artifact refits on a snapshot',
'src/services/model/probabilityContractService.js',
` const artifact = registry.load(sport, stat);`,
` const artifact = registry.load(sport, stat);
const _refit = require('./calibration').fitIsotonic([], {});`, GOV);
logic(13, 'a new settlement mutates the active curve', () => {
const s = codeOf(fs.readFileSync(path.join(ROOT, 'src/services/model/artifactRegistry.js'), 'utf8'));
const readsFile = s.includes('fs.readFileSync');
const noWrite = !/writeFileSync|\.update\(|\.upsert\(/.test(s);
return { caught: readsFile && noWrite, detail: `registry reads a committed file and performs no write` };
});
inject(14, 'a candidate automatically promotes',
'src/services/model/artifactRegistry.js',
` return raw.artifact_id === promoted.artifact_id;`,
` return true;`, GOV);
logic(16, 'the old 65/35 permanent withholding survives under policy B', () => {
return { caught: A.withheld_from_fit === 0 && A.fit_n === A.era_audit.current_era_rows,
detail: `withheld_from_fit ${A.withheld_from_fit}; fit_n ${A.fit_n} == eligible ${A.era_audit.current_era_rows}` };
});
// 18-22 — serving surfaces stay put
logic(18, 'grade changes', () => {
const s = codeOf(fs.readFileSync(path.join(ROOT, 'src/services/model/probabilityContract.js'), 'utf8'));
return { caught: !/servedGrade|gradeFor/.test(s), detail: 'the probability contract does not reach the grade' };
});
logic(19, 'selected side changes', () => {
const s = codeOf(fs.readFileSync(path.join(ROOT, 'src/services/model/probabilityContract.js'), 'utf8'));
return { caught: !/\bside\b\s*=|gradeBestSide/.test(s), detail: 'the contract never assigns a side' };
});
logic(20, 'publication changes', () => {
const s = fs.readFileSync(path.join(ROOT, 'src/services/retentionService.js'), 'utf8');
const m = codeOf(s.slice(s.indexOf('function mergeProbabilityContract'), s.indexOf('function mergeChainShadow')));
return { caught: !/published|publication_id|read_id|lineage/.test(m), detail: 'the merge touches no publication or lineage field' };
});
logic(21, 'shadow enabled in the release', () => {
const s = fs.readFileSync(path.join(ROOT, 'src/services/model/probabilityContract.js'), 'utf8');
const strict = s.includes("String(raw || '') === '1'");
const noDefaultOn = !/PROBABILITY_CONTRACT_SHADOW\s*\|\|\s*'1'/.test(s);
return { caught: strict && noDefaultOn, detail: 'shadow requires an explicit "1"; no default-on path' };
});
logic(22, 'live serving enabled', () => {
const snap = codeOf(fs.readFileSync(path.join(ROOT, 'src/services/snapshotService.js'), 'utf8'));
const deployedEmpty = /CALIBRATION_DEPLOYED\s*=\s*Object\.freeze\(\[\s*\]\)/.test(snap);
const noLive = A.approved_for_live === false && registry.PROMOTED['mlb:hits'].stage === registry.STAGE.APPROVED_FOR_SHADOW;
const files = execSync(`grep -rl "served_probability" ${ROOT}/src ${ROOT}/web/src 2>/dev/null || true`).toString().trim().split('\n').filter(Boolean)
.map((f) => f.replace(ROOT + '/', ''));
const allowed = ['src/services/model/probabilityContract.js', 'src/services/model/probabilityContractService.js', 'src/services/retentionService.js'];
const leaked = files.filter((f) => !allowed.includes(f));
return { caught: deployedEmpty && noLive && leaked.length === 0,
detail: `CALIBRATION_DEPLOYED empty=${deployedEmpty}; stage=${registry.PROMOTED['mlb:hits'].stage}; leaked consumers: ${leaked.join(', ') || 'none'}` };
});
// 23-26 — the frozen neighbours
logic(23, 'retention identity changes', () => {
const d = execSync(`git -C ${ROOT} diff --unified=0 -- src/services/retentionService.js`).toString();
const fields = ['player_key:', 'snapshot_id:', 'canonical_event_id:', 'game_id:', 'stat:', 'line:', 'side:'];
const touched = fields.filter((f) => new RegExp(`^[-+].*${f.replace(/[.*+?^${}()|[\]\\]/g, '\\$&')}`, 'm').test(d));
return { caught: touched.length === 0, detail: `identity fields in diff: ${touched.join(', ') || 'none'}` };
});
logic(24, 'participant identity changes', () => {
const f = execSync(`git -C ${ROOT} diff --name-only`).toString().trim().split('\n').filter(Boolean)
.filter((x) => /participantIdentity|eventIdentity|matchupKeys|playerName/.test(x));
return { caught: f.length === 0, detail: `participant files changed: ${f.join(', ') || 'none'}` };
});
logic(25, 'lineage mechanics/config change', () => {
const f = execSync(`git -C ${ROOT} diff --name-only`).toString().trim().split('\n').filter(Boolean)
.filter((x) => /lineage|readLineage|readAncestry|lineageWriteMode|lineageCoverage/i.test(x));
const snapDiff = execSync(`git -C ${ROOT} diff -- src/services/snapshotService.js`).toString();
const lines = snapDiff.split('\n').filter((l) => /^[-+]/.test(l) && /lineage|canary|LINEAGE_/i.test(l));
return { caught: f.length === 0 && lines.length === 0, detail: `lineage files ${f.length}, lineage diff lines ${lines.length}` };
});
logic(26, 'PerformanceDistribution becomes servable', () => {
const snap = codeOf(fs.readFileSync(path.join(ROOT, 'src/services/snapshotService.js'), 'utf8'));
return { caught: !/chain\.chainAcross\(/.test(snap), detail: 'no chainAcross call' };
});
const landed = results.filter((r) => r.landed).length;
console.log(JSON.stringify({ teeth_landed: `${landed}/${results.length}`, results }, null, 2));
process.exit(landed === results.length ? 0 : 1);
+11 -13
View File
@@ -71,21 +71,19 @@ logicTooth(1, 'runtime SHA verification waits for a snapshot despite working run
// ── 2 — the estimator must carry a reconstructable identity ──────────────
injectionTooth(2, 'runtime estimator lacks a reconstructable artifact identity',
'src/services/model/probabilityContractService.js',
` knot_count: Array.isArray(fitted.map) ? fitted.map.length : null,
knot_digest: digest(fitted.map),`,
` knot_count: null,
knot_digest: 'static-placeholder',`,
'tests/unit/probabilityContract.test.js');
'src/services/model/fitPolicy.js',
` if (!artifact.knot_digest || !artifact.served_curve_digest) violations.push(VIOLATION.NO_ARTIFACT_IDENTITY);`,
` if (false) violations.push(VIOLATION.NO_ARTIFACT_IDENTITY);`,
'tests/unit/artifactGovernance.test.js');
// ── 3 — the artifact must be tied to the certified model era ─────────────
injectionTooth(3, 'runtime artifact differs from the certified artifact',
'src/services/model/probabilityContractService.js',
` model_version: contract.model_version,
fit_as_of: fitted.cutoff || null,`,
` model_version: 'engine1@some-other-era',
fit_as_of: fitted.cutoff || null,`,
'tests/unit/probabilityContractShadow.test.js');
'src/services/model/artifactRegistry.js',
` estimator_type: 'isotonic',
stage: promoted.stage,`,
` estimator_type: 'low_param',
stage: promoted.stage,`,
'tests/unit/artifactGovernance.test.js');
// ── 4 — point-in-time evidence ───────────────────────────────────────────
injectionTooth(4, 'dynamic refit uses future settlement evidence for an earlier Read',
@@ -127,7 +125,7 @@ logicTooth(25, 'shadow flag parses loosely or defaults ON', () => {
const strict = pcSrc.includes("String(raw || '') === '1'")
&& snap.includes("probabilityContract').shadowState().shadow === 'ON'");
const scoped = snap.includes("&& sp === 'mlb'");
const statScoped = snap.includes("{ sport: 'mlb', stat: 'hits' }");
const statScoped = snap.includes("sport: 'mlb', stat: 'hits'");
return { caught: strict && scoped && statScoped,
detail: `strict '1' compare=${strict}; sport-scoped=${scoped}; stat-scoped=${statScoped}` };
});
+70
View File
@@ -0,0 +1,70 @@
#!/usr/bin/env node
'use strict';
/**
* verify-artifact-reconstruction — rebuild the promoted artifact from its
* DECLARED TRAINING CONTRACT and require exact equality.
*
* A fresh read, a fresh source digest, a fresh fit, fresh digests. Calling the
* same builder twice would prove only that the builder is deterministic; this
* proves the artifact is derivable from what it says about itself.
*/
require('dotenv').config({ quiet: true });
const crypto = require('crypto');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const src = require('../src/services/model/currentEraSource');
const registry = require('../src/services/model/artifactRegistry');
const digest = (o) => crypto.createHash('sha256').update(JSON.stringify(o)).digest('hex').slice(0, 16);
function servedCurve(map, bands) {
const steps = []; let prev = null;
for (const [lo, hi] of bands) {
prev = null;
for (let x = lo; x < hi - 1e-9; x += 0.001) {
const raw = Math.round(x * 1000) / 1000;
const v = cal.applyIsotonic(map, raw);
if (v === null) continue;
const r = Math.round(v * 1e6) / 1e6;
if (r !== prev) { steps.push([raw, r]); prev = r; }
}
}
return steps;
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const a = registry.load('mlb', 'hits');
if (!a) throw new Error('no promoted artifact');
// ONLY what the artifact declares about its own training contract.
const rows = await src.loadRows(sb, {
sport: a.sport, stat: a.stat, modelVersion: a.model_version, before: a.fit_as_of,
});
rows.sort((x, y) => (x.date < y.date ? -1 : x.date > y.date ? 1 : (x.id < y.id ? -1 : 1)));
const audit = src.eraAudit(rows, a.model_version);
const map = cal.fitIsotonic(rows, { minTotal: 200 });
const curve = servedCurve(map, a.certified_bands);
const got = {
fit_n: rows.length,
training_cutoff: rows.length ? rows[rows.length - 1].date : null,
source_digest: src.sourceDigest(rows),
knot_count: map ? map.length : null,
knot_digest: digest(map),
served_curve_digest: digest(curve),
wrong_era_rows: audit.wrong_era_rows,
};
const want = {
fit_n: a.fit_n, training_cutoff: a.training_cutoff, source_digest: a.source_digest,
knot_count: a.knot_count, knot_digest: a.knot_digest,
served_curve_digest: a.served_curve_digest, wrong_era_rows: 0,
};
const mismatches = Object.keys(want).filter((k) => got[k] !== want[k]);
console.log(JSON.stringify({ artifact_id: a.artifact_id, want, got,
curve_identical: JSON.stringify(curve) === JSON.stringify(a.served_curve),
mismatches, RECONSTRUCTION: mismatches.length === 0 && JSON.stringify(curve) === JSON.stringify(a.served_curve)
? 'EXACT' : 'DIFFERS' }, null, 1));
process.exit(mismatches.length === 0 ? 0 : 1);
})().catch((e) => { console.error(e.message); process.exit(1); });
+137
View File
@@ -0,0 +1,137 @@
'use strict';
/**
* artifactRegistry — WHICH FROZEN CURVE IS ACTIVE, AND WHO DECIDED.
*
* ── WHY THE RUNTIME NO LONGER FITS ───────────────────────────────────────
* A certified refit PROCEDURE does not mean the active production curve should
* mutate every snapshot. It did: `probabilityContractService` called the fitter
* on every run, so the served mapping silently changed as outcomes settled —
* unreconstructable after the fact, and capable of moving a user's number with
* no version, no review and no rollback.
*
* The procedure is certified separately (walk-forward, current era only). What
* ships is ONE frozen artifact, fitted once, committed as JSON, loaded here.
*
* ── WHY A COMMITTED FILE ─────────────────────────────────────────────────
* The curve is ~900 bytes. Git already gives immutability, versioning, review
* as a diff, and rollback — the same seam
* `supabase/schema/model_snapshots.columns.json` uses. A mutable environment
* variable could not hold the curve honestly, and a new service would be a
* second governance stack for one small file.
*
* ── PROMOTION IS A HUMAN ACT ─────────────────────────────────────────────
* `STAGE` below is a source constant. A snapshot cannot promote. A settlement
* cannot promote. A successful fit cannot promote. Adding an artifact file does
* nothing until this map names it.
*/
const fs = require('fs');
const path = require('path');
const fitPolicy = require('./fitPolicy');
const DIR = path.join(__dirname, 'artifacts');
const STAGE = Object.freeze({
NONE: 'NONE',
APPROVED_FOR_SHADOW: 'APPROVED_FOR_SHADOW',
APPROVED_FOR_LIVE: 'APPROVED_FOR_LIVE',
});
/**
* THE PROMOTION TABLE. Editing this is the promotion.
*
* mlb:hits is APPROVED_FOR_SHADOW — it may be evaluated and persisted as a
* shadow candidate. It is deliberately NOT APPROVED_FOR_LIVE: no user-facing
* number comes from it, and moving it there is a separate decision requiring
* its own evidence.
*/
const PROMOTED = Object.freeze({
'mlb:hits': Object.freeze({
artifact_id: 'mlb-hits-isotonic@2026-09-03',
stage: STAGE.APPROVED_FOR_SHADOW,
}),
});
const keyOf = (sport, stat) => `${String(sport || '').toLowerCase()}:${String(stat || '').toLowerCase()}`;
const fileFor = (artifactId) => path.join(DIR, `${String(artifactId).replace(/[@:]/g, '_')}.json`);
/**
* THE PROMOTION GUARD. A file is accepted only if it declares the id it was
* promoted under.
*
* Exported because it is the one rule whose failure case cannot be reached
* through `load()` on the happy path: the promoted file naturally agrees with
* itself, so deleting this check changed nothing observable and a teeth
* injection came back green. The rule it enforces is that dropping a file into
* the directory, or editing its id, cannot promote anything.
*/
function acceptFile(raw, promoted) {
if (!raw || !promoted) return false;
return raw.artifact_id === promoted.artifact_id;
}
const cache = new Map();
/**
* Load the promoted artifact. Returns null when nothing is promoted, when the
* file is missing, or when the artifact fails its own policy — never a partial
* or a fallback, because a fallback here is a curve nobody certified.
*/
function load(sport, stat) {
const key = keyOf(sport, stat);
if (cache.has(key)) return cache.get(key);
const promoted = PROMOTED[key];
if (!promoted) { cache.set(key, null); return null; }
let raw;
try { raw = JSON.parse(fs.readFileSync(fileFor(promoted.artifact_id), 'utf8')); }
catch { cache.set(key, null); return null; }
if (!acceptFile(raw, promoted)) { cache.set(key, null); return null; }
// RE-VALIDATED AT LOAD, not trusted from the file. The file records what the
// builder concluded; this is the runtime asking the same question again, so a
// hand-edited `servable: true` cannot smuggle an invalid artifact into use.
const check = fitPolicy.validate(
{ ...raw, estimator_type: 'isotonic' },
{ era_counts: (raw.era_audit && raw.era_audit.era_counts) || null },
);
const out = Object.freeze({
...raw,
estimator_type: 'isotonic',
stage: promoted.stage,
approved_for_shadow: promoted.stage === STAGE.APPROVED_FOR_SHADOW || promoted.stage === STAGE.APPROVED_FOR_LIVE,
approved_for_live: promoted.stage === STAGE.APPROVED_FOR_LIVE,
fit_policy_valid: check.valid,
fit_policy_violations: Object.freeze(check.violations),
servable: check.servable,
});
cache.set(key, out);
return out;
}
/**
* Serve from the FROZEN CURVE, never from a refitted map.
*
* The curve is a step table at 0.001 granularity, which is exact for `p_win`
* (quantised to 3dp at source). Outside the certified bands it returns null —
* the curve does not extend past what was certified.
*/
function applyCurve(artifact, p) {
if (!artifact || !Array.isArray(artifact.served_curve)) return null;
const x = Number(p);
if (!Number.isFinite(x)) return null;
const inBand = (artifact.certified_bands || []).some(([lo, hi]) => x >= lo && x < hi);
if (!inBand) return null;
let v = null;
for (const [from, val] of artifact.served_curve) { if (x >= from) v = val; else break; }
return v;
}
/** Test seam only — the promotion table itself is never writable at runtime. */
function __resetCache() { cache.clear(); }
module.exports = { STAGE, PROMOTED, load, applyCurve, fileFor, acceptFile, __resetCache };
@@ -0,0 +1,69 @@
{
"artifact_id": "mlb-hits-isotonic@2026-09-03",
"procedure_version": "mlb-hits-isotonic-refit@v1",
"sport": "mlb",
"stat": "hits",
"model_version": "engine1@2026-08-07-fullwindow",
"fit_as_of": "2026-09-03",
"training_cutoff": "2026-09-01",
"fit_n": 6069,
"withheld_from_fit": 0,
"source_digest": "25919c160c59cc738b9b4cb9c6d62c48249d5daff5d2c37c7892cd47349a930a",
"algorithm": "isotonic-pav",
"algorithm_version": "calibration.fitIsotonic@2026-08",
"knot_count": 12,
"knot_digest": "5ae940ea163b7da2",
"served_curve": [
[
0.5,
0.518976
],
[
0.541,
0.550427
],
[
0.562,
0.570656
],
[
0.613,
0.574792
],
[
0.648,
0.581454
],
[
0.67,
0.607595
],
[
0.684,
0.625298
],
[
0.798,
0.63806
]
],
"served_curve_digest": "c24a9dc5c2a96068",
"certified_bands": [
[
0.5,
0.8
]
],
"era_audit": {
"era_counts": {
"engine1@2026-08-07-fullwindow": 6069
},
"current_era_rows": 6069,
"wrong_era_rows": 0
},
"approved_for_shadow": true,
"approved_for_live": false,
"fit_policy_valid": true,
"fit_policy_violations": [],
"servable": true
}
+95
View File
@@ -0,0 +1,95 @@
'use strict';
/**
* currentEraSource — THE CANONICAL TRAINING SET FOR A CALIBRATION ARTIFACT.
*
* ── THE INVARIANT THIS EXISTS TO ENFORCE ─────────────────────────────────
* An artifact declaring `model_version = X` must be fitted ONLY on observations
* whose `model_version` is X. Formally:
*
* artifact.model_version = X => distinct(training_rows.model_version) = {X}
*
* The defect this replaces: `calibrationService.loadSettledRows` applies no
* model filter, so at fit_as_of 2026-09-02 the chronological 65% drew all 3,292
* rows from the superseded engine1@2026-07-20 plus 2,792 current-era rows —
* 54.1% of a map that declared the current era. A calibrator corrects a specific
* forecaster; fitted on a different one it is measuring something else.
*
* The legacy loader is deliberately untouched: other contracts still depend on
* its historical behaviour, and changing it would silently move them.
*
* ── ONE MODEL IDENTITY ───────────────────────────────────────────────────
* The era comes from `config/modelVersion` and is threaded through: the QUERY
* filters on it, the ARTIFACT records it, and the VALIDATOR compares them. No
* second hardcoded default — that is exactly how ledgerService and
* retentionService drifted apart.
*/
const crypto = require('crypto');
const { paginate } = require('../../utils/safePaginate');
/** Fields that define a training observation. Anything else is database noise. */
const CANONICAL_FIELDS = Object.freeze(['id', 'p_win', 'outcome', 'game_date', 'model_version']);
/**
* Load the eligible current-era settled observations.
*
* `before` is STRICT — a Read may never train on its own outcome.
* The quarantine exclusion mirrors the legacy path so the two remain comparable.
*/
async function loadRows(sb, { sport, stat, modelVersion, before } = {}) {
if (!sb) throw new Error('currentEraSource.loadRows: no client');
if (!sport || !stat || !modelVersion || !before) {
throw new Error('currentEraSource.loadRows: sport, stat, modelVersion and before are all required');
}
const rows = await paginate(
() => sb.from('ledger_entries')
.select('id, p_win, outcome, game_date, quarantine_reason, model_version')
.eq('sport', sport).is('user_id', null).eq('stat', stat)
.eq('model_version', modelVersion) // THE RESTRICTION
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null)
.lt('game_date', before), // strictly before
{ key: 'id', pageSize: 1000, label: `currentEraSource(${sport}/${stat}/${modelVersion})` },
);
return rows
.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'))
.map((r) => ({
id: String(r.id),
p: Number(r.p_win),
won: r.outcome === 'hit' ? 1 : 0,
date: String(r.game_date),
model_version: r.model_version,
}))
.filter((r) => Number.isFinite(r.p));
}
/**
* Audit the era composition OF THE ACTUAL FITTED SET.
*
* Deliberately computed from the rows that were fitted, not from a separate
* approximate count — a second query can agree with the wrong set.
*/
function eraAudit(rows, modelVersion) {
const counts = {};
for (const r of rows || []) counts[r.model_version || 'unknown'] = (counts[r.model_version || 'unknown'] || 0) + 1;
const wrong = Object.entries(counts)
.filter(([era, n]) => era !== modelVersion && Number(n) > 0)
.reduce((s, [, n]) => s + Number(n), 0);
return { era_counts: counts, current_era_rows: counts[modelVersion] || 0, wrong_era_rows: wrong };
}
/**
* DETERMINISTIC SOURCE DIGEST over the EXACT fitted observation set.
*
* Sorted by row id so database return order cannot change it, and built from
* the canonical fields only — so it moves when a row is added, removed, or has
* its outcome, probability or model identity changed, and not otherwise.
*/
function sourceDigest(rows) {
const canon = (rows || [])
.map((r) => [String(r.id), Number(r.p).toFixed(6), Number(r.won), String(r.date), String(r.model_version)])
.sort((a, b) => (a[0] < b[0] ? -1 : a[0] > b[0] ? 1 : 0));
return crypto.createHash('sha256').update(JSON.stringify(canon)).digest('hex');
}
module.exports = { loadRows, eraAudit, sourceDigest, CANONICAL_FIELDS };
+35
View File
@@ -47,6 +47,17 @@ const STATE = Object.freeze({
CERTIFIED_EMPIRICAL_BAND: 'CERTIFIED_EMPIRICAL_BAND',
UNCERTIFIED: 'UNCERTIFIED',
UNSUPPORTED: 'UNSUPPORTED',
/**
* The contract is certified but the ARTIFACT that would answer failed its
* governance policy, or is not promoted for this usage.
*
* A distinct state, deliberately. UNSUPPORTED means "we never certified this
* sport/stat"; INVALID means "your input was malformed". Neither is true
* here: the contract exists and the input is fine, and today's artifact is
* the thing that is not usable. Folding it into either would destroy the
* diagnosis at exactly the moment someone needs it.
*/
ARTIFACT_POLICY_BLOCKED: 'ARTIFACT_POLICY_BLOCKED',
VERSION_MISMATCH: 'VERSION_MISMATCH',
INVALID: 'INVALID',
});
@@ -156,6 +167,8 @@ function resolve(read = {}, deps = {}) {
// names the exact mapping that ran. Absent when no artifact was supplied,
// never invented.
artifact: deps.artifact || null,
artifact_id: (deps.artifact && deps.artifact.artifact_id) || null,
procedure_version: (deps.artifact && deps.artifact.procedure_version) || null,
reason: null,
};
@@ -181,6 +194,28 @@ function resolve(read = {}, deps = {}) {
if (raw === null || raw < 0 || raw > 1) {
return { ...base, probability_state: STATE.INVALID, reason: 'raw probability absent or out of range' };
}
// ── GOVERNANCE IS LOAD-BEARING ─────────────────────────────────────────
// `servable: false` has to MEAN something mechanically, or the validator is
// decoration. An artifact that failed its policy may not produce a certified
// state, and no environment variable can override this: the gate is here, in
// the resolution, not beside the flag that turns the shadow on.
//
// ONE piece of logic for both usages (Step 4): future live serving consumes
// the same decision, differing only in which promotion stage it demands.
if (deps.artifact) {
const a = deps.artifact;
const usage = deps.usage === 'live' ? 'live' : 'shadow';
const promoted = usage === 'live' ? a.approved_for_live === true : a.approved_for_shadow === true;
if (a.servable !== true) {
return { ...base, probability_state: STATE.ARTIFACT_POLICY_BLOCKED,
reason: `estimator artifact failed its fit policy: ${(a.fit_policy_violations || []).join(', ') || 'unservable'}` };
}
if (!promoted) {
return { ...base, probability_state: STATE.ARTIFACT_POLICY_BLOCKED,
reason: `estimator artifact is not promoted for ${usage} use` };
}
}
if (!inCertifiedRawBand(contract.certified_bands, raw)) {
// NO RAW FALLBACK. This is the whole point of the module.
return { ...base, probability_state: STATE.UNCERTIFIED, reason: 'raw value lies outside certified estimator support' };
+28 -113
View File
@@ -1,135 +1,50 @@
'use strict';
/**
* probabilityContractService — fit the certified estimator, point in time.
* probabilityContractService — LOAD THE FROZEN ARTIFACT. DO NOT FIT.
*
* REUSES the existing fitter and its "fit on settled history strictly before
* today" discipline. What it does NOT reuse is `calibrationService.calibrate()`,
* whose gate is evaluated in CALIBRATED-OUTPUT space and whose else-branch
* serves RAW. Both of those are the blocked contract; only the MAP is taken.
* This used to call the fitter on every snapshot, so the active mapping changed
* silently as outcomes settled: a user's number could move with no version, no
* review and no rollback, and a past Read could not be reconstructed because the
* curve that produced it no longer existed anywhere.
*
* SUPPORT COMES FROM THE CERTIFIED ARTIFACT, NOT FROM TONIGHT'S FIT. The bands
* in `probabilityContract.MLB_HITS` were adjudicated on a three-way split and
* are a fixed property of that adjudication. Letting a nightly refit widen its
* own support is how an estimator certifies itself.
* The refit PROCEDURE is certified separately and offline
* (`scripts/certify-current-era-procedure.js`), and produces ONE committed
* artifact (`scripts/build-current-era-artifact.js`). At runtime there is no
* fitter, no ledger read, and nothing to drift: `artifactRegistry` loads the
* promoted JSON and the served value comes from its frozen curve.
*
* A consequence worth stating: new settled outcomes are FORWARD EVALUATION
* evidence. They may inform a future candidate. They cannot alter this curve.
*/
const crypto = require('crypto');
const cal = require('./calibration');
const pc = require('./probabilityContract');
const fitPolicy = require('./fitPolicy');
/** Short, stable content digest. Full sha256 truncated — collision risk here is
* irrelevant and 16 hex chars keeps the per-row payload small. */
const digest = (obj) => crypto.createHash('sha256')
.update(JSON.stringify(obj)).digest('hex').slice(0, 16);
const registry = require('./artifactRegistry');
/**
* THE SERVED CURVE — the complete served function inside certified support.
*
* `p_win` is quantised to three decimals at the source
* (`analyzeViaEngine1`: Math.round(pWin * 1000) / 1000), so a step table at
* 0.001 granularity is not a sample of the mapping — it IS the mapping, for
* every input that can actually occur. Six steps, ~200 bytes.
*
* Storing it on the row makes a Read reconstructable WITHOUT re-deriving the
* training set. That matters because settled rows can be re-settled
* (`re_settled_at`), so a later refit at the same cutoff is not guaranteed to
* reproduce the same map — and a claim you can only verify when the inputs
* happen not to have moved is not a reconstructable claim.
* @returns {null|{contract, artifact, resolve}} null when nothing is promoted
* for this sport/stat, or the promoted artifact fails to load. Null means
* NOTHING is served — never a fallback curve, which would be a mapping nobody
* certified.
*/
function servedCurve(map, bands) {
const steps = [];
let prev = null;
for (const [lo, hi] of bands) {
for (let x = lo; x < hi - 1e-9; x += 0.001) {
const raw = Math.round(x * 1000) / 1000;
const v = cal.applyIsotonic(map, raw);
if (v === null) continue;
const rounded = Math.round(v * 1e6) / 1e6;
if (rounded !== prev) { steps.push([raw, rounded]); prev = rounded; }
}
prev = null; // bands are independent segments
}
return steps;
}
/**
* @returns {null|{estimate, fit_n, fitted_through, contract}} null when there
* is not enough settled history — and null means NOTHING is served, never
* "pass raw through".
*/
async function build(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) {
async function build(_sb, { sport = 'mlb', stat = 'hits', usage = 'shadow' } = {}) {
const contract = pc.contractFor(sport, stat);
if (!contract || !sb) return null;
if (!contract) return null;
const svc = opts.calibrationService || require('./calibrationService');
const fitted = await svc.fromLedger(sb, { sport, stat, before, ...opts });
if (!fitted || !fitted.map) return null;
// ── ARTIFACT IDENTITY ──────────────────────────────────────────────────
// The runtime REFITS PER SNAPSHOT against `game_date < todayEt()`, so the
// mapping changes as outcomes settle. That is point-in-time correct going
// forward — a Read can only ever have seen settlements strictly before its
// own day — but without an identity a served number could not be tied to the
// function that produced it, and "isotonic" would be a label rather than a
// claim. This is the identity.
const curve = servedCurve(fitted.map, contract.certified_bands);
// ── THE FIT-POLICY CHECK (Step 22) ─────────────────────────────────────
// An artifact does not become servable because the algorithm ran. The era
// composition is known STRUCTURALLY, not by an extra read: this service
// applies no model_version filter today, so the restriction demonstrably did
// not hold and the artifact records that rather than claiming it did.
//
// `era_restricted` is a fact about the QUERY, and the query is right here.
const eraRestricted = false; // calibrationService.loadSettledRows applies no model filter
const draft = {
estimator_type: contract.estimator_type,
model_version: contract.model_version,
fit_n: fitted.fit_n ?? null,
training_cutoff: fitted.fitted_through || null,
knot_digest: digest(fitted.map),
served_curve_digest: digest(curve),
certified_bands: contract.certified_bands,
};
const policy = fitPolicy.validate(draft, eraRestricted
? { era_counts: { [contract.model_version]: fitted.fit_n } } : {});
const artifact = Object.freeze({
estimator_type: contract.estimator_type,
estimator_version: contract.estimator_version,
certification_version: contract.certification_version,
model_version: contract.model_version,
fit_as_of: fitted.cutoff || null, // the exact lt(game_date) bound
training_cutoff: fitted.fitted_through || null, // last date INSIDE the fit
fit_n: fitted.fit_n ?? null,
knot_count: Array.isArray(fitted.map) ? fitted.map.length : null,
knot_digest: digest(fitted.map),
served_curve: curve, // complete over certified support
served_curve_digest: digest(curve),
// WHICH PROCEDURE, AND WHETHER IT HELD. Recorded on every artifact so a
// later reader sees the policy state of the fit that produced the number,
// not just the number's fingerprint.
fit_policy_version: policy.policy_version,
fit_policy_valid: policy.valid,
fit_policy_violations: Object.freeze(policy.violations),
// A policy-invalid artifact is never servable. Nothing serves today, so
// this is a declaration; it becomes load-bearing the moment serving exists.
servable: policy.servable,
});
const artifact = registry.load(sport, stat);
if (!artifact) return null;
return {
contract,
artifact,
fit_n: fitted.fit_n,
fitted_through: fitted.fitted_through,
cutoff: fitted.cutoff,
/** The estimator, and only the estimator. No gate, no fallback. */
estimate: (p) => cal.applyIsotonic(fitted.map, p),
/** Resolve one grade through the full contract. */
fit_n: artifact.fit_n,
fitted_through: artifact.training_cutoff,
cutoff: artifact.fit_as_of,
/** The frozen curve. Two calls return the same value, forever. */
estimate: (p) => registry.applyCurve(artifact, p),
resolve(read) {
return pc.resolve({ ...read, sport, stat },
{ estimate: (p) => cal.applyIsotonic(fitted.map, p), artifact });
{ estimate: (p) => registry.applyCurve(artifact, p), artifact, usage });
},
};
}
+20 -7
View File
@@ -871,13 +871,26 @@ function mergeProbabilityContract(rows, contract) {
certification_version: res.certification_version,
model_version: res.model_version,
reason: res.reason,
// WHICH FITTED FUNCTION PRODUCED THIS. The estimator refits per
// snapshot, so the certification version alone cannot identify the
// mapping that ran. `served_curve` is the complete served function over
// certified support at p_win's own 3dp granularity, so the row is
// reconstructable without re-deriving a training set that may since
// have been re-settled.
artifact: res.artifact || null,
// WHICH FROZEN ARTIFACT PRODUCED THIS — its IDENTITY, not its body.
// The curve is committed in the repository and addressable by
// `artifact_id`, so embedding it on every row would store the same ~900
// bytes thousands of times per snapshot to say something the id already
// says. `source_digest` pins the exact observation set it was fitted
// on, so the row remains reconstructable.
artifact: res.artifact ? {
artifact_id: res.artifact.artifact_id,
procedure_version: res.artifact.procedure_version,
model_version: res.artifact.model_version,
fit_as_of: res.artifact.fit_as_of,
training_cutoff: res.artifact.training_cutoff,
fit_n: res.artifact.fit_n,
source_digest: res.artifact.source_digest,
knot_digest: res.artifact.knot_digest,
served_curve_digest: res.artifact.served_curve_digest,
certified_bands: res.artifact.certified_bands,
stage: res.artifact.stage,
servable: res.artifact.servable,
} : null,
derived: derived ? {
available: derived.available,
ev_pct: derived.ev_pct,
+7 -4
View File
@@ -860,12 +860,15 @@ async function runSnapshot(sport, opts = {}) {
let probContract = null;
if (require('./model/probabilityContract').shadowState().shadow === 'ON' && sp === 'mlb') {
try {
// No database client: the artifact is a committed file, not a fit. There
// is nothing to read and nothing that can drift between snapshots.
const pcs = deps.probabilityContractService || require('./model/probabilityContractService');
const sbc = deps.supabase || require('../utils/supabase').getSupabaseServiceClient();
probContract = sbc ? await pcs.build(sbc, { sport: 'mlb', stat: 'hits' }) : null;
probContract = await pcs.build(null, { sport: 'mlb', stat: 'hits', usage: 'shadow' });
console.log(probContract
? `[probability-contract] shadow armed — isotonic, fit n=${probContract.fit_n} through ${probContract.fitted_through}, certified raw [0.50,0.80)`
: '[probability-contract] shadow armed but NO estimator (thin history) — nothing would be served');
? `[probability-contract] shadow armed — frozen artifact ${probContract.artifact.artifact_id}`
+ ` (${probContract.artifact.stage}), fit n=${probContract.fit_n} through ${probContract.fitted_through},`
+ ` knots ${probContract.artifact.knot_digest}, certified raw [0.50,0.80)`
: '[probability-contract] shadow armed but NO PROMOTED ARTIFACT — nothing would be served');
} catch (e) {
probContract = null;
console.warn('[probability-contract] shadow build failed (snapshot continues):', e.message);
+174
View File
@@ -0,0 +1,174 @@
'use strict';
/**
* A VALIDATOR IS NOT A REPAIR. `servable:false` must mean something mechanically.
*
* The previous release detected the violation correctly and then served the
* certified state anyway. These tests exist so that cannot recur.
*/
const pc = require('../../src/services/model/probabilityContract');
const svc = require('../../src/services/model/probabilityContractService');
const registry = require('../../src/services/model/artifactRegistry');
const fp = require('../../src/services/model/fitPolicy');
const ERA = 'engine1@2026-08-07-fullwindow';
const good = registry.load('mlb', 'hits');
const read = (p) => ({ sport: 'mlb', stat: 'hits', model_version: ERA, p_win: p });
const est = () => 0.6;
describe('an unservable artifact can never emit a certified state', () => {
// Era and estimator mismatch resolve to VERSION_MISMATCH rather than the new
// state — that is DELIBERATE. "This artifact belongs to a different
// forecaster" is a more precise thing to say than "policy blocked", and the
// existing state already says it exactly. The invariant asserted for all six
// is the one that matters: never certified, never a number.
const violations = {
ERA_NOT_RESTRICTED: [{ era_audit: { era_counts: { [ERA]: 10, 'engine1@2026-07-20': 5 }, wrong_era_rows: 5 } }, 'ARTIFACT_POLICY_BLOCKED'],
MODEL_VERSION_MISMATCH: [{ model_version: 'engine1@2026-07-20' }, 'VERSION_MISMATCH'],
ESTIMATOR_MISMATCH: [{ estimator_type: 'low_param' }, 'VERSION_MISMATCH'],
MISSING_IDENTITY: [{ knot_digest: null }, 'ARTIFACT_POLICY_BLOCKED'],
THIN_FIT: [{ fit_n: 3 }, 'ARTIFACT_POLICY_BLOCKED'],
SUPPORT_WIDENING: [{ certified_bands: [[0.50, 0.99]] }, 'ARTIFACT_POLICY_BLOCKED'],
};
for (const [name, [over, expected]] of Object.entries(violations)) {
it(`${name} -> ${expected}, never CERTIFIED_CALIBRATED`, () => {
const base = { ...good, ...over };
const check = fp.validate(base, { era_counts: base.era_audit ? base.era_audit.era_counts : null });
expect(check.servable).toBe(false);
const artifact = { ...base, servable: check.servable, fit_policy_violations: check.violations };
const r = pc.resolve(read(0.65), { estimate: est, artifact });
expect(r.probability_state).toBe(pc.STATE[expected]);
expect(r.probability_state).not.toBe(pc.STATE.CERTIFIED_CALIBRATED);
expect(r.served_probability).toBeNull();
expect(r.raw_model_probability).toBe(0.65); // raw is still evidence
expect(pc.derivedClaims(r, -115).available).toBe(false);
});
}
it('and every probability-derived claim is withheld with it', () => {
const artifact = { ...good, servable: false, fit_policy_violations: ['ERA_NOT_RESTRICTED'] };
const d = pc.derivedClaims(pc.resolve(read(0.65), { estimate: est, artifact }), -115);
expect(d.available).toBe(false);
expect(d.ev_pct).toBeNull();
expect(d.kelly).toBeNull();
expect(d.value).toBeNull();
});
it('the blocked state names the violation rather than shrugging', () => {
const artifact = { ...good, servable: false, fit_policy_violations: ['ERA_NOT_RESTRICTED'] };
expect(pc.resolve(read(0.65), { estimate: est, artifact }).reason).toContain('ERA_NOT_RESTRICTED');
});
});
describe('no environment variable can override the gate', () => {
const artifact = { ...good, servable: false, fit_policy_violations: ['ERA_NOT_RESTRICTED'] };
it('shadow ON does not make an unservable artifact certify', () => {
expect(pc.shadowState({ PROBABILITY_CONTRACT_SHADOW: '1' }).shadow).toBe('ON');
// the gate lives in the resolution, not beside the flag
const r = pc.resolve(read(0.65), { estimate: est, artifact });
expect(r.probability_state).toBe(pc.STATE.ARTIFACT_POLICY_BLOCKED);
});
it('resolve takes no flag, so there is nothing for a flag to reach', () => {
const src = require('fs').readFileSync(
require('path').join(__dirname, '../../src/services/model/probabilityContract.js'), 'utf8');
const body = src.slice(src.indexOf('function resolve'), src.indexOf('const isCertified'));
expect(body).not.toContain('process.env');
expect(body).not.toContain('PROBABILITY_CONTRACT_SHADOW');
});
});
describe('one validity decision for shadow AND live', () => {
it('a shadow-approved artifact is refused for live use', () => {
expect(good.approved_for_shadow).toBe(true);
expect(good.approved_for_live).toBe(false);
const live = pc.resolve(read(0.65), { estimate: est, artifact: good, usage: 'live' });
expect(live.probability_state).toBe(pc.STATE.ARTIFACT_POLICY_BLOCKED);
expect(live.reason).toContain('not promoted for live');
const shadow = pc.resolve(read(0.65), { estimate: est, artifact: good, usage: 'shadow' });
expect(shadow.probability_state).toBe(pc.STATE.CERTIFIED_CALIBRATED);
});
it('live consumes the SAME servable decision, not a parallel one', () => {
const bad = { ...good, servable: false, approved_for_live: true, fit_policy_violations: ['X'] };
expect(pc.resolve(read(0.65), { estimate: est, artifact: bad, usage: 'live' }).probability_state)
.toBe(pc.STATE.ARTIFACT_POLICY_BLOCKED);
});
});
describe('the active curve does not move', () => {
it('the runtime holds no fitter — nothing to refit on a snapshot', () => {
const src = require('fs').readFileSync(
require('path').join(__dirname, '../../src/services/model/probabilityContractService.js'), 'utf8');
const code = src.replace(/\/\*[\s\S]*?\*\//g, '').replace(/^\s*\/\/.*$/gm, '');
expect(code).not.toContain('fitIsotonic');
expect(code).not.toContain('fromLedger');
expect(code).not.toContain('loadRows');
});
it('repeated builds return the same artifact object and the same numbers', async () => {
const a = await svc.build(null, { sport: 'mlb', stat: 'hits' });
const b = await svc.build(null, { sport: 'mlb', stat: 'hits' });
expect(a.artifact).toBe(b.artifact); // cached, identical identity
expect(a.artifact.source_digest).toBe(b.artifact.source_digest);
expect(a.resolve(read(0.72)).served_probability).toBe(b.resolve(read(0.72)).served_probability);
});
it('a new settled outcome cannot alter the curve — it is a committed file', () => {
const before = registry.applyCurve(good, 0.65);
registry.__resetCache();
const after = registry.applyCurve(registry.load('mlb', 'hits'), 0.65);
expect(after).toBe(before);
});
});
describe('promotion is a deliberate act', () => {
it('the promotion table is a frozen source constant, not runtime state', () => {
expect(Object.isFrozen(registry.PROMOTED)).toBe(true);
expect(Object.isFrozen(registry.PROMOTED['mlb:hits'])).toBe(true);
// strict mode makes the write throw rather than fail silently — either way
// the table is unchanged, which is the property under test
expect(() => { registry.PROMOTED['mlb:rbi'] = { artifact_id: 'x', stage: 'APPROVED_FOR_LIVE' }; }).toThrow();
expect(registry.PROMOTED['mlb:rbi']).toBeUndefined();
});
it('an artifact file that exists but is not named is NOT loaded', () => {
// load() resolves the file FROM the promotion table, so an unnamed artifact
// sitting in the directory is inert.
expect(registry.load('mlb', 'rbi')).toBeNull();
});
it('the loaded artifact id must equal the promoted id', () => {
expect(registry.load('mlb', 'hits').artifact_id).toBe(registry.PROMOTED['mlb:hits'].artifact_id);
});
it('a file declaring a DIFFERENT id is refused — dropping one in promotes nothing', () => {
const promoted = registry.PROMOTED['mlb:hits'];
expect(registry.acceptFile({ artifact_id: promoted.artifact_id }, promoted)).toBe(true);
// the case load() can never reach on the happy path, and the reason the
// guard existed unprotected until a teeth injection came back green
expect(registry.acceptFile({ artifact_id: 'mlb-hits-isotonic@2099-01-01' }, promoted)).toBe(false);
expect(registry.acceptFile({ artifact_id: null }, promoted)).toBe(false);
expect(registry.acceptFile({}, promoted)).toBe(false);
expect(registry.acceptFile(null, promoted)).toBe(false);
expect(registry.acceptFile({ artifact_id: 'x' }, null)).toBe(false);
});
});
describe('model-era transition law', () => {
it('a NEW model era does not inherit this artifact', () => {
const r = pc.resolve({ sport: 'mlb', stat: 'hits', model_version: 'engine1@2027-01-01', p_win: 0.65 },
{ estimate: est, artifact: good });
expect(r.probability_state).toBe(pc.STATE.VERSION_MISMATCH);
expect(r.served_probability).toBeNull();
});
it('and an artifact relabelled to a new era fails its own policy', () => {
const relabelled = { ...good, model_version: 'engine1@2027-01-01' };
const check = fp.validate(relabelled, { era_counts: good.era_audit.era_counts });
expect(check.valid).toBe(false);
expect(check.servable).toBe(false);
});
});
+84
View File
@@ -0,0 +1,84 @@
'use strict';
/**
* THE INVARIANT: artifact.model_version = X => distinct(training.model_version) = {X}.
*/
const src = require('../../src/services/model/currentEraSource');
const ERA = 'engine1@2026-08-07-fullwindow';
const OLD = 'engine1@2026-07-20';
/** A fake that HONOURS its filters — a pass-through would prove nothing. */
function client(rows) {
const f = {};
const q = {
select: () => q,
eq: (c, v) => { f[c] = v; return q; },
is: () => q,
in: (c, v) => { f[`in_${c}`] = v; return q; },
not: () => q,
lt: (c, v) => { f[`lt_${c}`] = v; return q; },
order: () => q,
range: async () => {
let out = rows;
for (const [k, v] of Object.entries(f)) {
if (k.startsWith('in_') || k.startsWith('lt_')) continue;
out = out.filter((r) => r[k] === v);
}
if (f.lt_game_date) out = out.filter((r) => r.game_date < f.lt_game_date);
return { data: out, error: null, count: out.length };
},
};
return { client: { from: () => q }, filters: f };
}
const ROWS = [
{ id: '1', p_win: 0.6, outcome: 'hit', game_date: '2026-08-12', model_version: ERA, sport: 'mlb', stat: 'hits' },
{ id: '2', p_win: 0.7, outcome: 'miss', game_date: '2026-08-13', model_version: ERA, sport: 'mlb', stat: 'hits' },
{ id: '3', p_win: 0.55, outcome: 'hit', game_date: '2026-08-01', model_version: OLD, sport: 'mlb', stat: 'hits' },
];
describe('the era restriction is in the QUERY', () => {
it('loads only the requested era', async () => {
const { client: c, filters } = client(ROWS);
const out = await src.loadRows(c, { sport: 'mlb', stat: 'hits', modelVersion: ERA, before: '2026-09-03' });
expect(filters.model_version).toBe(ERA);
expect(out.map((r) => r.id).sort()).toEqual(['1', '2']);
expect(out.every((r) => r.model_version === ERA)).toBe(true);
});
it('the era audit is computed from the ACTUAL rows, not a second query', () => {
const mixed = [{ model_version: ERA }, { model_version: ERA }, { model_version: OLD }];
const a = src.eraAudit(mixed, ERA);
expect(a.current_era_rows).toBe(2);
expect(a.wrong_era_rows).toBe(1);
expect(src.eraAudit([{ model_version: ERA }], ERA).wrong_era_rows).toBe(0);
});
it('refuses to run without an explicit era and horizon', async () => {
const { client: c } = client(ROWS);
await expect(src.loadRows(c, { sport: 'mlb', stat: 'hits', before: '2026-09-03' })).rejects.toThrow();
await expect(src.loadRows(c, { sport: 'mlb', stat: 'hits', modelVersion: ERA })).rejects.toThrow();
});
});
describe('the source digest', () => {
const R = [
{ id: 'b', p: 0.6, won: 1, date: 'd2', model_version: ERA },
{ id: 'a', p: 0.5, won: 0, date: 'd1', model_version: ERA },
];
const base = src.sourceDigest(R);
it('ignores database return order', () => {
expect(src.sourceDigest([...R].reverse())).toBe(base);
});
it('moves when a row is added, removed, or materially changed', () => {
expect(src.sourceDigest([R[0]])).not.toBe(base);
expect(src.sourceDigest([...R, { id: 'c', p: 0.7, won: 1, date: 'd3', model_version: ERA }])).not.toBe(base);
expect(src.sourceDigest([{ ...R[0], won: 0 }, R[1]])).not.toBe(base);
expect(src.sourceDigest([{ ...R[0], p: 0.61 }, R[1]])).not.toBe(base);
expect(src.sourceDigest([{ ...R[0], model_version: OLD }, R[1]])).not.toBe(base);
expect(src.sourceDigest([{ ...R[0], date: 'd9' }, R[1]])).not.toBe(base);
});
});
+18 -20
View File
@@ -84,28 +84,26 @@ describe('the minimum automatic gates', () => {
});
});
describe('the built artifact carries its policy state', () => {
const map = cal.fitIsotonic(Array.from({ length: 900 }, (_, i) => {
const p = Math.round((0.35 + (i % 60) / 100) * 1000) / 1000;
return { p, won: ((i * 2654435761) % 1000) / 1000 < (0.5 + 0.45 * (p - 0.5)) ? 1 : 0, date: `d${i % 14}` };
}), { minTotal: 200 });
const build = () => svc.build({}, { calibrationService: { fromLedger: async () => ({
map, fit_n: 900, fitted_through: '2026-08-21', cutoff: '2026-09-02' }) } });
describe('the promoted artifact carries its policy state', () => {
const registry = require('../../src/services/model/artifactRegistry');
it('records the policy version, the violation and servable:false', async () => {
const b = await build();
expect(b.artifact.fit_policy_version).toBe('mlb-hits-isotonic-refit@v1');
expect(b.artifact.fit_policy_valid).toBe(false);
expect(b.artifact.fit_policy_violations).toContain(fp.VIOLATION.ERA_NOT_RESTRICTED);
expect(b.artifact.servable).toBe(false);
it('the shipped artifact PASSES its own policy — detection became repair', () => {
const a = registry.load('mlb', 'hits');
expect(a.fit_policy_version || a.procedure_version).toBe('mlb-hits-isotonic-refit@v1');
expect(a.fit_policy_valid).toBe(true);
expect(a.fit_policy_violations).toEqual([]);
expect(a.servable).toBe(true);
});
it('the policy violation does NOT distort what the shadow measures', async () => {
// The shadow's job is to measure the certified contract on real rows. If a
// policy violation flipped every row to UNCERTIFIED the shadow would
// measure the violation instead, and the receipt would be worthless.
const b = await build();
expect(b.resolve({ model_version: ERA, p_win: 0.65 }).probability_state).toBe('CERTIFIED_CALIBRATED');
expect(b.resolve({ model_version: ERA, p_win: 0.91 }).probability_state).toBe('UNCERTIFIED');
it('the policy is RE-VALIDATED at load, so a hand-edited file cannot smuggle one in', () => {
const a = registry.load('mlb', 'hits');
// the loader recomputes from era_audit rather than trusting the file's flag
const recomputed = fp.validate({ ...a, estimator_type: 'isotonic' },
{ era_counts: a.era_audit.era_counts });
expect(recomputed.valid).toBe(true);
const lying = fp.validate({ ...a, estimator_type: 'isotonic' },
{ era_counts: { [ERA]: 100, 'engine1@2026-07-20': 1 } });
expect(lying.valid).toBe(false);
expect(lying.servable).toBe(false);
});
});
+71 -85
View File
@@ -191,93 +191,85 @@ describe('the registry already asked the right question', () => {
});
});
describe('probabilityContractService', () => {
it('returns null — not raw — when there is no settled history', async () => {
const built = await svc.build({}, { calibrationService: { fromLedger: async () => null } });
expect(built).toBeNull();
describe('probabilityContractService — loads, never fits', () => {
const registry = require('../../src/services/model/artifactRegistry');
it('returns the PROMOTED frozen artifact, not a fresh fit', async () => {
const b = await svc.build(null, { sport: 'mlb', stat: 'hits' });
expect(b).not.toBeNull();
expect(b.artifact.artifact_id).toBe(registry.PROMOTED['mlb:hits'].artifact_id);
expect(b.artifact.servable).toBe(true);
});
it('takes the MAP and never the blocked calibrate() gate', async () => {
const cal = require('../../src/services/model/calibration');
const map = cal.fitIsotonic(Array.from({ length: 600 }, (_, i) => {
const p = 0.40 + (i % 55) / 100;
return { p, won: i % 3 === 0 ? 0 : 1, date: `d${i % 12}` };
}), { minTotal: 200 });
const calibrateSpy = jest.fn(() => ({ p_calibrated: 0.99, calibrated: true }));
const built = await svc.build({}, { calibrationService: {
fromLedger: async () => ({ map, fit_n: 600, fitted_through: 'd11', calibrate: calibrateSpy }) } });
expect(built).not.toBeNull();
const r = built.resolve({ model_version: ERA, p_win: 0.65 });
expect(r.probability_state).toBe(pc.STATE.CERTIFIED_CALIBRATED);
expect(calibrateSpy).not.toHaveBeenCalled();
it('returns null — never a fallback curve — when nothing is promoted', async () => {
for (const stat of ['rbi', 'total_bases', 'runs']) {
expect(await svc.build(null, { sport: 'mlb', stat })).toBeNull();
}
expect(await svc.build(null, { sport: 'wnba', stat: 'points' })).toBeNull();
});
it('support comes from the artifact, so a nightly refit cannot widen it', async () => {
const built = await svc.build({}, { calibrationService: {
fromLedger: async () => ({ map: { x: [0, 1], y: [0.5, 0.9] }, fit_n: 900, fitted_through: 'd9',
bands: [[0.0, 1.0]] }) } }); // fit claims the whole range
expect(built.resolve({ model_version: ERA, p_win: 0.95 }).probability_state).toBe(pc.STATE.UNCERTIFIED);
it('needs no database client at all — there is nothing left to read', async () => {
const b = await svc.build(undefined, { sport: 'mlb', stat: 'hits' });
expect(b.artifact.artifact_id).toBeTruthy();
});
it('two resolutions of the same input are identical, forever', async () => {
const b1 = await svc.build(null, { sport: 'mlb', stat: 'hits' });
const b2 = await svc.build(null, { sport: 'mlb', stat: 'hits' });
for (const p of [0.50, 0.55, 0.601, 0.72, 0.799]) {
const a = b1.resolve({ model_version: ERA, p_win: p });
const c = b2.resolve({ model_version: ERA, p_win: p });
expect(a.served_probability).toBe(c.served_probability);
expect(a.artifact.knot_digest).toBe(c.artifact.knot_digest);
}
});
});
describe('artifact identity — the mapping that actually ran', () => {
const cal = require('../../src/services/model/calibration');
const mk = (seed) => cal.fitIsotonic(Array.from({ length: 900 }, (_, i) => {
const p = Math.round((0.35 + (i % 60) / 100) * 1000) / 1000;
return { p, won: ((i * seed) % 1000) / 1000 < (0.5 + 0.45 * (p - 0.5)) ? 1 : 0, date: `d${i % 14}` };
}), { minTotal: 200 });
const fitted = (map, over = {}) => ({ calibrationService: { fromLedger: async () => ({
map, fit_n: 900, fitted_through: 'd13', cutoff: '2026-09-03', ...over }) } });
describe('artifact identity — the frozen curve that actually ran', () => {
const registry = require('../../src/services/model/artifactRegistry');
const a = registry.load('mlb', 'hits');
it('the same evidence reconstructs the same artifact, digest for digest', async () => {
const map = mk(2654435761);
const a = await svc.build({}, fitted(map));
const b = await svc.build({}, fitted(map));
expect(a.artifact.knot_digest).toBe(b.artifact.knot_digest);
expect(a.artifact.served_curve_digest).toBe(b.artifact.served_curve_digest);
expect(a.artifact.served_curve).toEqual(b.artifact.served_curve);
});
it('DIFFERENT evidence produces a different identity — the digest is not decorative', async () => {
const a = await svc.build({}, fitted(mk(2654435761)));
const b = await svc.build({}, fitted(mk(40503)));
expect(a.artifact.knot_digest).not.toBe(b.artifact.knot_digest);
});
it('carries the point-in-time bound and the training cutoff, distinctly', async () => {
const a = await svc.build({}, fitted(mk(2654435761)));
expect(a.artifact.fit_as_of).toBe('2026-09-03'); // the lt(game_date) bound
expect(a.artifact.training_cutoff).toBe('d13'); // last date inside the fit
expect(a.artifact.fit_n).toBe(900);
expect(a.artifact.knot_count).toBeGreaterThan(0);
});
it('the served curve IS the served function over certified support, not a sample', async () => {
const a = await svc.build({}, fitted(mk(2654435761)));
const lookup = (raw) => {
let v = null;
for (const [from, val] of a.artifact.served_curve) if (raw >= from) v = val;
return v;
};
for (let x = 0.50; x < 0.80 - 1e-9; x += 0.001) {
const raw = Math.round(x * 1000) / 1000;
const r = a.resolve({ model_version: ERA, p_win: raw });
expect(r.served_probability).toBe(Math.round(lookup(raw) * 1000) / 1000);
it('carries every field needed to find and check it again', () => {
for (const k of ['artifact_id', 'procedure_version', 'sport', 'stat', 'model_version',
'fit_as_of', 'training_cutoff', 'fit_n', 'source_digest', 'algorithm', 'algorithm_version',
'knot_count', 'knot_digest', 'served_curve', 'served_curve_digest', 'certified_bands']) {
expect(a[k]).toBeDefined();
expect(a[k]).not.toBeNull();
}
});
it('the curve covers ONLY certified support — never the unsupported tail', async () => {
const a = await svc.build({}, fitted(mk(2654435761)));
for (const [from] of a.artifact.served_curve) {
it('fit_as_of and training_cutoff are DIFFERENT questions', () => {
// the eligibility bound, and the newest observation actually admitted
expect(a.fit_as_of).not.toBe(a.training_cutoff);
expect(a.training_cutoff < a.fit_as_of).toBe(true);
});
it('was fitted on the current era ONLY', () => {
expect(a.model_version).toBe(ERA);
expect(a.era_audit.wrong_era_rows).toBe(0);
expect(Object.keys(a.era_audit.era_counts)).toEqual([ERA]);
});
it('withheld nothing from the promoted fit', () => {
expect(a.withheld_from_fit).toBe(0);
expect(a.fit_n).toBe(a.era_audit.current_era_rows);
});
it('the served curve IS the served function over support, not a sample', () => {
for (let x = 0.50; x < 0.80 - 1e-9; x += 0.001) {
const raw = Math.round(x * 1000) / 1000;
let want = null;
for (const [from, val] of a.served_curve) { if (raw >= from) want = val; else break; }
expect(registry.applyCurve(a, raw)).toBe(want);
}
});
it('the curve covers ONLY certified support', () => {
for (const [from] of a.served_curve) {
expect(from).toBeGreaterThanOrEqual(0.50);
expect(from).toBeLessThan(0.80);
}
});
it('a resolution with no artifact records null rather than inventing one', () => {
const r = pc.resolve(read(0.65), { estimate: iso });
expect(r.artifact).toBeNull();
expect(r.served_probability).not.toBeNull();
for (const p of [0.499, 0.80, 0.9, 0.99]) expect(registry.applyCurve(a, p)).toBeNull();
});
});
@@ -313,17 +305,11 @@ describe('point in time — a Read can only see settlements before its own day',
expect(bound.some((a) => a[0] === 'gte')).toBe(false);
});
it('the artifact records the bound it was fitted under, so a Read names its own evidence horizon', async () => {
const cal = require('../../src/services/model/calibration');
const map = cal.fitIsotonic(Array.from({ length: 900 }, (_, i) => {
const p = Math.round((0.35 + (i % 60) / 100) * 1000) / 1000;
return { p, won: ((i * 2654435761) % 1000) / 1000 < (0.5 + 0.45 * (p - 0.5)) ? 1 : 0, date: `d${i % 14}` };
}), { minTotal: 200 });
const built = await svc.build({}, { calibrationService: { fromLedger: async () => ({
map, fit_n: 900, fitted_through: '2026-08-21', cutoff: '2026-09-02' }) } });
expect(built.artifact.fit_as_of).toBe('2026-09-02');
expect(built.artifact.training_cutoff).toBe('2026-08-21');
// and the two are DIFFERENT questions — the bound, and the last date inside it
expect(built.artifact.fit_as_of).not.toBe(built.artifact.training_cutoff);
it('the artifact names its own evidence horizon', async () => {
const built = await svc.build(null, { sport: 'mlb', stat: 'hits' });
expect(built.artifact.fit_as_of).toBeTruthy();
expect(built.artifact.training_cutoff).toBeTruthy();
// the bound, and the last date inside it — never the same claim
expect(built.artifact.training_cutoff < built.artifact.fit_as_of).toBe(true);
});
});