9809626c99
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.
So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.
TERMINAL RETENTION STATES (retentionService.classifyPersist):
NOTHING_TO_PERSIST attempted 0 — a refusal-only slate is still a cycle
SKIPPED_NO_DATABASE no database configured; not a failure
COMPLETE attempted > 0, written === attempted, no error
FAILED_ZERO_WRITE written === 0 — first chunk failed
FAILED_PARTIAL 0 < written < attempted — a later chunk failed
FAILED_UNRESOLVED_ERROR counts look complete but an error is unresolved;
unreachable through today's loop, and kept because
the alternative is reporting COMPLETE holding an error
The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.
classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.
`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.
VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.
OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.
RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.
Six teeth, each with the injection verified present, against a green baseline:
1 written>0 as generic success -> 6 fail
2 later-chunk failure reports COMPLETE -> 5 fail
3 FAILED_PARTIAL does not alert -> 3 fail
4 status reports a recalculated count -> 1 fail
5 row presence treated as completion -> 1 fail
6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).
Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.
Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.
383 suites / 5,118 tests pass. web tsc exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
226 lines
9.9 KiB
JavaScript
226 lines
9.9 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* RETENTION SCHEMA CONTRACT — the guard that would have caught the outage.
|
|
*
|
|
* WHAT HAPPENED. `createCollector.onPublished` set `published_side` alongside
|
|
* `published`. `published_side` is not a `model_snapshots` column, supabase-js
|
|
* declares the UNION of row keys in the `columns=` parameter, and PostgREST
|
|
* rejected the entire batch with a 400 — for every sport, silently, because
|
|
* retention is best-effort. Verified in production edge logs.
|
|
*
|
|
* WHY 381 GREEN SUITES MISSED IT. Every retention test injects a permissive
|
|
* fake client — `{ from: () => ({ upsert: async () => ({ error: null }) }) }` —
|
|
* which accepts any column set. The suite proved the logic and never once
|
|
* compared a row against the database contract.
|
|
*
|
|
* WHY THE FIRST MANUAL CHECK ALSO MISSED IT. It sampled the collector after
|
|
* `onGraded` only and never called `onPublished`, so the offending key was not
|
|
* yet on the row. It inspected a PRE-PUBLICATION shape and reported the FINAL
|
|
* outbound shape as clean.
|
|
*
|
|
* So this test does two things the old ones could not:
|
|
* 1. it validates the payload ACTUALLY HANDED TO `.upsert()`, captured by a
|
|
* spy, after the full production call order including `onPublished`;
|
|
* 2. it validates against a contract DERIVED from the migration chain, not a
|
|
* hand-maintained list.
|
|
*/
|
|
|
|
const path = require('path');
|
|
const fs = require('fs');
|
|
const retention = require('../../src/services/retentionService');
|
|
|
|
const ROOT = path.resolve(__dirname, '..', '..');
|
|
const CONTRACT = JSON.parse(fs.readFileSync(
|
|
path.join(ROOT, 'supabase/schema/model_snapshots.columns.json'), 'utf8',
|
|
));
|
|
|
|
/** Captures the exact array supabase-js would send. */
|
|
function spyClient() {
|
|
const seen = [];
|
|
return {
|
|
seen,
|
|
client: {
|
|
from: (table) => ({
|
|
upsert: async (rows, opts) => { seen.push({ table, rows, opts }); return { error: null }; },
|
|
}),
|
|
},
|
|
};
|
|
}
|
|
|
|
const CTX = {
|
|
snapshotId: '00000000-0000-0000-0000-000000000000',
|
|
capturedAt: '2026-08-27T22:00:00Z',
|
|
gameDate: '2026-08-27',
|
|
gameIdFor: () => 'mlb:2026-08-27:BostonRedSox@MiamiMarlins',
|
|
};
|
|
const BASE = {
|
|
player: 'Aaron Judge', stat_type: 'hits', line: 0.5, sport: 'mlb', book: 'draftkings',
|
|
canonical_event_id: 'mlb:gamepk:823825', event_identity_source: 'MLB_STATSAPI_GAMEPK',
|
|
event_identity_method: 'CANONICAL', event_identity_version: 'evid@1', event_occurrence: 1,
|
|
};
|
|
const OVER = { ...BASE, direction: 'over', grade: 'B', confidence: 61, p_win: 0.61 };
|
|
const UNDER = { ...BASE, direction: 'under', grade: 'C', confidence: 39, p_win: 0.39 };
|
|
|
|
/** The FULL production call order — onGraded for both sides, then onPublished. */
|
|
function finalOutboundRows() {
|
|
const c = retention.createCollector(CTX);
|
|
c.onGraded(BASE, [OVER, UNDER]);
|
|
c.onPublished(BASE, OVER);
|
|
return c.rows;
|
|
}
|
|
|
|
describe('the contract itself is derived, not hand-written', () => {
|
|
test('it declares its generator and its derivation', () => {
|
|
expect(CONTRACT.table).toBe('model_snapshots');
|
|
expect(CONTRACT.generated_by).toBe('scripts/generate-schema-contract.js');
|
|
expect(CONTRACT.derived_from).toMatch(/supabase\/migrations/);
|
|
expect(fs.existsSync(path.join(ROOT, 'scripts/generate-schema-contract.js'))).toBe(true);
|
|
expect(fs.existsSync(path.join(ROOT, 'scripts/verify-schema-contract.js'))).toBe(true);
|
|
});
|
|
|
|
test('it is non-trivial and self-consistent', () => {
|
|
expect(CONTRACT.columns.length).toBe(CONTRACT.column_count);
|
|
expect(CONTRACT.columns.length).toBeGreaterThan(50);
|
|
expect(CONTRACT.columns).toContain('published');
|
|
expect(CONTRACT.columns).toContain('canonical_event_id');
|
|
// The offending field must NOT be in the contract — that is the fact.
|
|
expect(CONTRACT.columns).not.toContain('published_side');
|
|
});
|
|
|
|
test('known production drift is recorded rather than hidden', () => {
|
|
const d = CONTRACT.known_production_drift;
|
|
expect(d.columns_in_production_not_in_migrations).toEqual(
|
|
expect.arrayContaining(['quarantine_reason', 're_settled_at', 'settlement_source']),
|
|
);
|
|
// The migration-derived set is the stricter of the two, so a writer inside
|
|
// it is valid against production as well.
|
|
expect(d.explanation).toMatch(/stricter/);
|
|
});
|
|
});
|
|
|
|
describe('FINAL OUTBOUND payload is within the schema contract', () => {
|
|
test('the captured upsert payload uses only real columns', async () => {
|
|
const spy = spyClient();
|
|
await retention.persist(finalOutboundRows(), { getClient: () => spy.client });
|
|
|
|
expect(spy.seen.length).toBeGreaterThan(0);
|
|
const call = spy.seen[0];
|
|
expect(call.table).toBe('model_snapshots');
|
|
|
|
// supabase-js sends the UNION of keys across the batch — validate the union.
|
|
const union = new Set();
|
|
for (const r of call.rows) for (const k of Object.keys(r)) union.add(k);
|
|
const invalid = [...union].filter((k) => !CONTRACT.columns.includes(k)).sort();
|
|
expect(invalid).toEqual([]);
|
|
});
|
|
|
|
test('the payload is captured AFTER onPublished, not before', async () => {
|
|
// The pre-publication shape is what the earlier manual check inspected.
|
|
const pre = retention.createCollector(CTX);
|
|
pre.onGraded(BASE, [OVER, UNDER]);
|
|
const preKeys = new Set(pre.rows.flatMap((r) => Object.keys(r)));
|
|
|
|
const post = finalOutboundRows();
|
|
const postKeys = new Set(post.flatMap((r) => Object.keys(r)));
|
|
|
|
// Both must be valid; the point is that this test exercises the later one.
|
|
for (const set of [preKeys, postKeys]) {
|
|
expect([...set].filter((k) => !CONTRACT.columns.includes(k))).toEqual([]);
|
|
}
|
|
// And onPublished must actually have marked a row, or the test is vacuous.
|
|
expect(post.filter((r) => r.published === true)).toHaveLength(1);
|
|
expect(post.filter((r) => r.published === false)).toHaveLength(1);
|
|
});
|
|
|
|
test('every declared key is present on EVERY row', async () => {
|
|
// A key on only some rows is dropped for the whole batch by PostgREST, so
|
|
// a ragged batch is its own defect class.
|
|
const rows = finalOutboundRows();
|
|
const union = new Set(rows.flatMap((r) => Object.keys(r)));
|
|
for (const r of rows) {
|
|
expect(new Set(Object.keys(r))).toEqual(union);
|
|
}
|
|
});
|
|
|
|
test('published_side is gone from the source entirely', () => {
|
|
const src = fs.readFileSync(path.join(ROOT, 'src/services/retentionService.js'), 'utf8')
|
|
.replace(/\/\*[\s\S]*?\*\//g, '')
|
|
.replace(/(^|[^:])\/\/.*$/gm, '$1');
|
|
expect(src).not.toMatch(/published_side/);
|
|
});
|
|
});
|
|
|
|
describe('RETENTION FAILURE IS NOT SILENT', () => {
|
|
test('a PostgREST 400 is reported, never swallowed as success', async () => {
|
|
const failing = {
|
|
from: () => ({
|
|
upsert: async () => ({
|
|
error: { message: "Could not find the 'published_side' column of 'model_snapshots'", code: 'PGRST204' },
|
|
}),
|
|
}),
|
|
};
|
|
const out = await retention.persist(finalOutboundRows(), { getClient: () => failing });
|
|
expect(out.written).toBe(0);
|
|
expect(out.error).toMatch(/published_side/);
|
|
// Never reported as a success.
|
|
expect(out.skipped).toBe(false);
|
|
});
|
|
|
|
test('the pipeline emits a high-severity structured event on that failure', () => {
|
|
// MECHANISM CHANGED after 35da190: the alert condition was
|
|
// `r.error || (!r.skipped && r.attempted > 0 && r.written === 0)`, which
|
|
// could not distinguish a PARTIAL cohort from a complete one. It is now
|
|
// driven by the terminal status classified from the exact persist result,
|
|
// so FAILED_PARTIAL alerts as loudly as FAILED_ZERO_WRITE.
|
|
const src = fs.readFileSync(path.join(ROOT, 'src/services/snapshotService.js'), 'utf8');
|
|
const i = src.indexOf('const terminal = retention.recordTerminal');
|
|
expect(i).toBeGreaterThan(-1);
|
|
const block = src.slice(i, i + 2200);
|
|
for (const field of ['stage=model_snapshots', 'snapshot_id=', 'code_sha=', 'at=', 'error=']) {
|
|
expect(block).toContain(field);
|
|
}
|
|
expect(block).toMatch(/priority: 'high'/);
|
|
expect(block).toMatch(/product is unaffected/);
|
|
});
|
|
|
|
test('a skipped write (no database configured) is NOT alerted', async () => {
|
|
const out = await retention.persist(finalOutboundRows(), { getClient: () => null });
|
|
expect(out.skipped).toBe(true);
|
|
expect(out.error).toBeNull();
|
|
// The skip exemption now lives in the classifier, not in an inline
|
|
// `!r.skipped` test at the call site.
|
|
expect(retention.classifyPersist(out)).toBe(retention.TERMINAL.SKIPPED_NO_DATABASE);
|
|
expect(retention.isRetentionFailure(retention.classifyPersist(out))).toBe(false);
|
|
});
|
|
|
|
test('zero rows written with candidates present also alerts', () => {
|
|
const st = retention.classifyPersist({ attempted: 600, written: 0, skipped: false, error: null });
|
|
expect(st).toBe(retention.TERMINAL.FAILED_ZERO_WRITE);
|
|
expect(retention.isRetentionFailure(st)).toBe(true);
|
|
});
|
|
});
|
|
|
|
describe('CYCLE PERSISTENCE MECHANISM', () => {
|
|
test('a cycle is written in chunks, and a failed chunk aborts the rest', async () => {
|
|
// Documented for the dark-cycle gate: persistence is CHUNKED (250), so a
|
|
// partial cycle is possible in principle — the loop breaks on the first
|
|
// error rather than continuing. There is no terminal completion marker.
|
|
const src = fs.readFileSync(path.join(ROOT, 'src/services/retentionService.js'), 'utf8');
|
|
const body = src.slice(src.indexOf('async function persist('));
|
|
expect(body).toMatch(/const CHUNK = 250/);
|
|
expect(body).toMatch(/if \(error\) \{ out\.error = error\.message; break; \}/);
|
|
// `written` is the exact count of rows in successfully committed chunks, so
|
|
// written === attempted is the completion signal available today.
|
|
expect(body).toMatch(/out\.written \+= chunk\.length/);
|
|
});
|
|
|
|
test('written vs attempted distinguishes a complete cycle from a partial one', async () => {
|
|
const spy = spyClient();
|
|
const rows = finalOutboundRows();
|
|
const out = await retention.persist(rows, { getClient: () => spy.client });
|
|
expect(out.attempted).toBe(rows.length);
|
|
expect(out.written).toBe(rows.length);
|
|
});
|
|
});
|