Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.
So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.
TERMINAL RETENTION STATES (retentionService.classifyPersist):
NOTHING_TO_PERSIST attempted 0 — a refusal-only slate is still a cycle
SKIPPED_NO_DATABASE no database configured; not a failure
COMPLETE attempted > 0, written === attempted, no error
FAILED_ZERO_WRITE written === 0 — first chunk failed
FAILED_PARTIAL 0 < written < attempted — a later chunk failed
FAILED_UNRESOLVED_ERROR counts look complete but an error is unresolved;
unreachable through today's loop, and kept because
the alternative is reporting COMPLETE holding an error
The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.
classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.
`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.
VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.
OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.
RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.
Six teeth, each with the injection verified present, against a green baseline:
1 written>0 as generic success -> 6 fail
2 later-chunk failure reports COMPLETE -> 5 fail
3 FAILED_PARTIAL does not alert -> 3 fail
4 status reports a recalculated count -> 1 fail
5 row presence treated as completion -> 1 fail
6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).
Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.
Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.
383 suites / 5,118 tests pass. web tsc exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
@@ -189,7 +189,7 @@ router.get('/snapshot/status', async (req, res) => {
|
||||
const ticker = await cacheGet('ticker:items');
|
||||
redis_keys['ticker:items'] = !!ticker;
|
||||
// Same helpers the pipeline itself uses — one build identity, one canary parser.
|
||||
const { codeSha } = require('../services/retentionService');
|
||||
const { codeSha, lastRetention } = require('../services/retentionService');
|
||||
const lineageCanary = require('../services/lineageCanaryConfig');
|
||||
// Session 56 — surface the missed-cron signal in the health probe.
|
||||
const mlbTs = last_snapshot.mlb && last_snapshot.mlb.updated_at;
|
||||
@@ -204,6 +204,12 @@ router.get('/snapshot/status', async (req, res) => {
|
||||
// change without a new process, so `runtime.started_at` is a defensible
|
||||
// lower bound for how long this state has held.
|
||||
lineage_canary: lineageCanary.state(),
|
||||
// Latest TERMINAL retention result per sport, exactly as the persistence
|
||||
// function reported it. Chunked writes stop at the first failed chunk, so
|
||||
// rows existing under a snapshot_id does not mean the cohort is complete —
|
||||
// this is the only place that distinction is observable in production.
|
||||
// Counts and status only: no row payloads, no credentials.
|
||||
last_retention: lastRetention(),
|
||||
cron_armed: process.env.SNAPSHOT_CRON === '1',
|
||||
cron_hours_utc: HOURS_UTC,
|
||||
last_snapshot,
|
||||
|
||||
@@ -727,8 +727,107 @@ function newSnapshotId() {
|
||||
return crypto.randomUUID();
|
||||
}
|
||||
|
||||
/* ------------------------------------------------------------------ *
|
||||
* TERMINAL RETENTION STATE
|
||||
*
|
||||
* `persist()` writes in CHUNKS and STOPS ON THE FIRST FAILED CHUNK. Chunks
|
||||
* committed before the failure are already durable, so a failed cycle can
|
||||
* leave real, valid-looking rows behind. Row presence under a snapshot_id is
|
||||
* therefore NOT completion evidence, and neither is a uniform captured_at.
|
||||
*
|
||||
* The invariant: written < attempted with attempted > 0 is a FAILED cycle.
|
||||
* A partial cycle is never degraded success.
|
||||
*
|
||||
* `written` counts rows in successfully COMMITTED CHUNKS — not rows inserted.
|
||||
* The upsert uses ignoreDuplicates, so a re-run legitimately inserts far fewer
|
||||
* database rows than it writes. Comparing `written` to count(*) for a
|
||||
* snapshot_id will disagree by design; that is not a partial write.
|
||||
* ------------------------------------------------------------------ */
|
||||
|
||||
const TERMINAL = Object.freeze({
|
||||
/** attempted === 0 — a refusal-only slate is still a legitimate cycle. */
|
||||
NOTHING_TO_PERSIST: 'NOTHING_TO_PERSIST',
|
||||
/** No database configured (dev/test). Not a failure. */
|
||||
SKIPPED_NO_DATABASE: 'SKIPPED_NO_DATABASE',
|
||||
/** attempted > 0, written === attempted, no unresolved error. */
|
||||
COMPLETE: 'COMPLETE',
|
||||
/** attempted > 0, written === 0 — the first chunk failed. */
|
||||
FAILED_ZERO_WRITE: 'FAILED_ZERO_WRITE',
|
||||
/** attempted > 0, 0 < written < attempted — a later chunk failed. */
|
||||
FAILED_PARTIAL: 'FAILED_PARTIAL',
|
||||
/**
|
||||
* Counts look complete but an error is unresolved. Unreachable through
|
||||
* today's loop (it breaks before crediting a failed chunk), and kept
|
||||
* because the alternative is reporting COMPLETE with an error in hand.
|
||||
*/
|
||||
FAILED_UNRESOLVED_ERROR: 'FAILED_UNRESOLVED_ERROR',
|
||||
});
|
||||
|
||||
const FAILURE_STATUSES = Object.freeze([
|
||||
TERMINAL.FAILED_ZERO_WRITE, TERMINAL.FAILED_PARTIAL, TERMINAL.FAILED_UNRESOLVED_ERROR,
|
||||
]);
|
||||
|
||||
function isRetentionFailure(status) { return FAILURE_STATUSES.includes(status); }
|
||||
|
||||
/**
|
||||
* Classify the EXACT object `persist()` returned. It never recomputes
|
||||
* attempted or written — a second calculation could disagree with the writer,
|
||||
* and then the status would describe something that did not happen.
|
||||
*/
|
||||
function classifyPersist(result) {
|
||||
if (!result || typeof result.attempted !== 'number' || typeof result.written !== 'number') {
|
||||
throw new TypeError('classifyPersist requires the exact persist() result');
|
||||
}
|
||||
const { attempted, written, skipped, error } = result;
|
||||
if (attempted === 0) return TERMINAL.NOTHING_TO_PERSIST;
|
||||
if (skipped) return TERMINAL.SKIPPED_NO_DATABASE;
|
||||
if (written === 0) return TERMINAL.FAILED_ZERO_WRITE;
|
||||
if (written < attempted) return TERMINAL.FAILED_PARTIAL;
|
||||
// written >= attempted from here.
|
||||
if (error) return TERMINAL.FAILED_UNRESOLVED_ERROR;
|
||||
return written === attempted ? TERMINAL.COMPLETE : TERMINAL.FAILED_UNRESOLVED_ERROR;
|
||||
}
|
||||
|
||||
/**
|
||||
* Latest terminal retention result per sport, for the internal status probe.
|
||||
* In-memory and per-process on purpose: this is an observability surface, not
|
||||
* a record. The record is model_snapshots.
|
||||
*/
|
||||
const lastTerminal = new Map();
|
||||
|
||||
function recordTerminal({ sport, snapshotId, result, completedAt }) {
|
||||
const status = classifyPersist(result);
|
||||
const entry = Object.freeze({
|
||||
sport: sport || null,
|
||||
snapshot_id: snapshotId || null,
|
||||
attempted: result.attempted,
|
||||
written: result.written,
|
||||
status,
|
||||
completed_at: completedAt || new Date().toISOString(),
|
||||
code_sha: codeSha(),
|
||||
// Message only — never row payloads.
|
||||
error_summary: result.error ? String(result.error).slice(0, 300) : null,
|
||||
});
|
||||
if (sport) lastTerminal.set(sport, entry);
|
||||
return entry;
|
||||
}
|
||||
|
||||
function lastRetention() {
|
||||
const out = {};
|
||||
for (const [sport, entry] of lastTerminal) out[sport] = entry;
|
||||
return out;
|
||||
}
|
||||
|
||||
function resetTerminal() { lastTerminal.clear(); }
|
||||
|
||||
module.exports = {
|
||||
MODEL_VERSION,
|
||||
TERMINAL,
|
||||
classifyPersist,
|
||||
isRetentionFailure,
|
||||
recordTerminal,
|
||||
lastRetention,
|
||||
resetTerminal,
|
||||
REPAIRED_CHAMPION_VERSION,
|
||||
codeSha,
|
||||
rowsFromSides,
|
||||
|
||||
@@ -658,23 +658,35 @@ async function runSnapshot(sport, opts = {}) {
|
||||
// commitPublication AFTER the slate write, so at this point no row has a
|
||||
// read_id yet — building an index here would produce an empty one and
|
||||
// quietly leave every ledger row unlinked.
|
||||
console.log(`[snapshot] retention ${sp}: ${r.written}/${r.attempted} rows${r.skipped ? ' (skipped — no supabase env)' : ''}${r.error ? ` ERROR: ${r.error}` : ''}`);
|
||||
// RETENTION FAILURE MUST NOT BE SILENT.
|
||||
// TERMINAL RETENTION STATE — classified from the EXACT persist() result,
|
||||
// never recomputed. Chunked writes stop at the first failed chunk, so a
|
||||
// failed cycle can leave durable rows behind: presence of rows under a
|
||||
// snapshot_id proves nothing about whether the cohort is complete.
|
||||
const terminal = retention.recordTerminal({
|
||||
sport: sp,
|
||||
snapshotId: retentionCtx.snapshotId,
|
||||
result: r,
|
||||
completedAt: deps.now(),
|
||||
});
|
||||
console.log(`[snapshot] retention ${sp}: ${terminal.status} ${r.written}/${r.attempted} rows snapshot_id=${terminal.snapshot_id}${r.error ? ` ERROR: ${r.error}` : ''}`);
|
||||
// RETENTION FAILURE MUST NOT BE SILENT — INCLUDING A PARTIAL ONE.
|
||||
//
|
||||
// Retention stays best-effort — the product keeps publishing to Redis and
|
||||
// that is deliberate. But a failed batch previously reached only this log
|
||||
// Retention stays best-effort: the product keeps publishing to Redis and
|
||||
// that is deliberate. But a failed batch previously reached only a log
|
||||
// line, and a PostgREST 400 (`published_side` is not a column) killed
|
||||
// every batch for every sport with nothing surfacing anywhere. The
|
||||
// measurement record died quietly while the product looked healthy.
|
||||
// every batch for every sport with nothing surfacing anywhere.
|
||||
//
|
||||
// `skipped` is NOT a failure: it means no database is configured, which is
|
||||
// the normal state in dev and test.
|
||||
if (r.error || (!r.skipped && r.attempted > 0 && r.written === 0)) {
|
||||
// FAILED_PARTIAL is the dangerous successor to that bug: earlier chunks
|
||||
// committed, the cohort is short, and the surviving rows make it look
|
||||
// healthy. It alerts exactly as loudly as a total failure.
|
||||
//
|
||||
// NOTHING_TO_PERSIST and SKIPPED_NO_DATABASE are not failures.
|
||||
if (retention.isRetentionFailure(terminal.status)) {
|
||||
await deps.notify(
|
||||
`Retention write FAILED for ${sp.toUpperCase()} — ${r.written}/${r.attempted} rows persisted. `
|
||||
+ `stage=model_snapshots snapshot_id=${retentionCtx.snapshotId} code_sha=${retention.codeSha() || 'unknown'} `
|
||||
+ `at=${deps.now()} error=${r.error || 'zero rows written with candidates present'}. `
|
||||
+ 'The product is unaffected; the historical record for this cycle is missing.',
|
||||
`Retention ${terminal.status} for ${sp.toUpperCase()} — ${terminal.written}/${terminal.attempted} rows persisted. `
|
||||
+ `stage=model_snapshots snapshot_id=${terminal.snapshot_id} code_sha=${terminal.code_sha || 'unknown'} `
|
||||
+ `at=${terminal.completed_at} error=${terminal.error_summary || 'none reported'}. `
|
||||
+ 'The product is unaffected; this retention cohort is INCOMPLETE and must not be used as evidence.',
|
||||
{ title: 'VYNDR pipeline', priority: 'high', tags: ['rotating_light'] },
|
||||
);
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user