9809626c99
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.
So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.
TERMINAL RETENTION STATES (retentionService.classifyPersist):
NOTHING_TO_PERSIST attempted 0 — a refusal-only slate is still a cycle
SKIPPED_NO_DATABASE no database configured; not a failure
COMPLETE attempted > 0, written === attempted, no error
FAILED_ZERO_WRITE written === 0 — first chunk failed
FAILED_PARTIAL 0 < written < attempted — a later chunk failed
FAILED_UNRESOLVED_ERROR counts look complete but an error is unresolved;
unreachable through today's loop, and kept because
the alternative is reporting COMPLETE holding an error
The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.
classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.
`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.
VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.
OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.
RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.
Six teeth, each with the injection verified present, against a green baseline:
1 written>0 as generic success -> 6 fail
2 later-chunk failure reports COMPLETE -> 5 fail
3 FAILED_PARTIAL does not alert -> 3 fail
4 status reports a recalculated count -> 1 fail
5 row presence treated as completion -> 1 fail
6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).
Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.
Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.
383 suites / 5,118 tests pass. web tsc exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
72 lines
3.4 KiB
JavaScript
72 lines
3.4 KiB
JavaScript
'use strict';
|
|
|
|
/**
|
|
* RELEASE-AUTHORIZED model_snapshots INSERT CONTRACT — drift verifier.
|
|
*
|
|
* The committed contract is derived from the MIGRATION CHAIN, and it is the
|
|
* RELEASE AUTHORITY: it says which columns this release is permitted to write.
|
|
* Production is NOT the authority. A column that exists in production but in no
|
|
* migration is drift, and drift does not silently become permission — that is
|
|
* how an accidental artifact turns into schema.
|
|
*
|
|
* Classification:
|
|
* RELEASE COLUMN MISSING IN PROD -> HARD FAILURE (exit 1). Retention will
|
|
* 400 on the whole batch.
|
|
* PROD-ONLY COLUMN -> DRIFT WARNING / RECORDED DEBT (exit 0).
|
|
* OUTBOUND INSERT KEY OUTSIDE
|
|
* THE RELEASE CONTRACT -> CONTRACT FAILURE. Enforced by
|
|
* tests/unit/retentionSchemaContract.test.js
|
|
* against the real .upsert() payload.
|
|
*
|
|
* This script is READ-ONLY. It never rewrites the contract from live schema.
|
|
* Regenerating is a deliberate act: scripts/generate-schema-contract.js against
|
|
* a disposable database with the migration chain applied.
|
|
*
|
|
* SUPABASE_URL=... SUPABASE_SERVICE_KEY=... node scripts/verify-schema-contract.js
|
|
*/
|
|
|
|
const fs = require('fs');
|
|
const path = require('path');
|
|
const { createClient } = require('@supabase/supabase-js');
|
|
|
|
const TABLE = process.env.SCHEMA_TABLE || 'model_snapshots';
|
|
|
|
async function main() {
|
|
const contract = JSON.parse(fs.readFileSync(
|
|
path.join(__dirname, '..', 'supabase', 'schema', `${TABLE}.columns.json`), 'utf8'));
|
|
const url = process.env.SUPABASE_URL;
|
|
const key = process.env.SUPABASE_SERVICE_KEY || process.env.SUPABASE_SERVICE_ROLE_KEY;
|
|
if (!url || !key) throw new Error('SUPABASE_URL + service key required');
|
|
const sb = createClient(url, key, { auth: { persistSession: false } });
|
|
|
|
// One row is enough to learn the live column set from the response shape.
|
|
const { data, error } = await sb.from(TABLE).select('*').limit(1);
|
|
if (error) throw new Error(`live read failed: ${error.message}`);
|
|
if (!data || data.length === 0) throw new Error(`${TABLE} is empty — cannot infer live columns`);
|
|
|
|
const live = new Set(Object.keys(data[0]));
|
|
const missingInProd = contract.columns.filter((c) => !live.has(c));
|
|
const prodOnly = [...live].filter((c) => !contract.columns.includes(c)).sort();
|
|
|
|
console.log(`RELEASE-AUTHORIZED INSERT CONTRACT (${TABLE})`);
|
|
console.log(` release columns (migration-derived) : ${contract.columns.length}`);
|
|
console.log(` live production columns : ${live.size}`);
|
|
console.log(` PROD-ONLY COLUMN (drift/debt) : ${prodOnly.length}${prodOnly.length ? ` -> ${prodOnly.join(', ')}` : ''}`);
|
|
console.log(` RELEASE COLUMN MISSING IN PROD : ${missingInProd.length}${missingInProd.length ? ` -> ${missingInProd.join(', ')}` : ''}`);
|
|
|
|
if (prodOnly.length) {
|
|
console.log(' NOTE: prod-only columns are RECORDED DEBT. They are NOT release-authorized');
|
|
console.log(' and must not be added to the contract from live schema.');
|
|
}
|
|
if (missingInProd.length) {
|
|
console.error(' HARD FAILURE: a release-authorized column does not exist in production.');
|
|
process.exit(1);
|
|
}
|
|
console.log(' RESULT: PASS (no release column missing in production)');
|
|
process.exit(0);
|
|
}
|
|
|
|
if (require.main === module) {
|
|
main().catch((e) => { console.error('FAILED:', e.message); process.exit(1); });
|
|
}
|