Retention completion: a cohort is complete only when the writer says N of N

The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.

So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.

TERMINAL RETENTION STATES (retentionService.classifyPersist):
  NOTHING_TO_PERSIST       attempted 0 — a refusal-only slate is still a cycle
  SKIPPED_NO_DATABASE      no database configured; not a failure
  COMPLETE                 attempted > 0, written === attempted, no error
  FAILED_ZERO_WRITE        written === 0 — first chunk failed
  FAILED_PARTIAL           0 < written < attempted — a later chunk failed
  FAILED_UNRESOLVED_ERROR  counts look complete but an error is unresolved;
                           unreachable through today's loop, and kept because
                           the alternative is reporting COMPLETE holding an error

The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.

classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.

`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.

VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.

OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.

RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.

Six teeth, each with the injection verified present, against a green baseline:
  1 written>0 as generic success        -> 6 fail
  2 later-chunk failure reports COMPLETE -> 5 fail
  3 FAILED_PARTIAL does not alert        -> 3 fail
  4 status reports a recalculated count  -> 1 fail
  5 row presence treated as completion   -> 1 fail
  6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).

Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.

Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.

383 suites / 5,118 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
Kev
2026-08-27 19:12:59 -04:00
parent 35da190f2c
commit 9809626c99
8 changed files with 455 additions and 39 deletions
+7 -1
View File
@@ -189,7 +189,7 @@ router.get('/snapshot/status', async (req, res) => {
const ticker = await cacheGet('ticker:items');
redis_keys['ticker:items'] = !!ticker;
// Same helpers the pipeline itself uses — one build identity, one canary parser.
const { codeSha } = require('../services/retentionService');
const { codeSha, lastRetention } = require('../services/retentionService');
const lineageCanary = require('../services/lineageCanaryConfig');
// Session 56 — surface the missed-cron signal in the health probe.
const mlbTs = last_snapshot.mlb && last_snapshot.mlb.updated_at;
@@ -204,6 +204,12 @@ router.get('/snapshot/status', async (req, res) => {
// change without a new process, so `runtime.started_at` is a defensible
// lower bound for how long this state has held.
lineage_canary: lineageCanary.state(),
// Latest TERMINAL retention result per sport, exactly as the persistence
// function reported it. Chunked writes stop at the first failed chunk, so
// rows existing under a snapshot_id does not mean the cohort is complete —
// this is the only place that distinction is observable in production.
// Counts and status only: no row payloads, no credentials.
last_retention: lastRetention(),
cron_armed: process.env.SNAPSHOT_CRON === '1',
cron_hours_utc: HOURS_UTC,
last_snapshot,
+99
View File
@@ -727,8 +727,107 @@ function newSnapshotId() {
return crypto.randomUUID();
}
/* ------------------------------------------------------------------ *
* TERMINAL RETENTION STATE
*
* `persist()` writes in CHUNKS and STOPS ON THE FIRST FAILED CHUNK. Chunks
* committed before the failure are already durable, so a failed cycle can
* leave real, valid-looking rows behind. Row presence under a snapshot_id is
* therefore NOT completion evidence, and neither is a uniform captured_at.
*
* The invariant: written < attempted with attempted > 0 is a FAILED cycle.
* A partial cycle is never degraded success.
*
* `written` counts rows in successfully COMMITTED CHUNKS — not rows inserted.
* The upsert uses ignoreDuplicates, so a re-run legitimately inserts far fewer
* database rows than it writes. Comparing `written` to count(*) for a
* snapshot_id will disagree by design; that is not a partial write.
* ------------------------------------------------------------------ */
const TERMINAL = Object.freeze({
/** attempted === 0 — a refusal-only slate is still a legitimate cycle. */
NOTHING_TO_PERSIST: 'NOTHING_TO_PERSIST',
/** No database configured (dev/test). Not a failure. */
SKIPPED_NO_DATABASE: 'SKIPPED_NO_DATABASE',
/** attempted > 0, written === attempted, no unresolved error. */
COMPLETE: 'COMPLETE',
/** attempted > 0, written === 0 — the first chunk failed. */
FAILED_ZERO_WRITE: 'FAILED_ZERO_WRITE',
/** attempted > 0, 0 < written < attempted — a later chunk failed. */
FAILED_PARTIAL: 'FAILED_PARTIAL',
/**
* Counts look complete but an error is unresolved. Unreachable through
* today's loop (it breaks before crediting a failed chunk), and kept
* because the alternative is reporting COMPLETE with an error in hand.
*/
FAILED_UNRESOLVED_ERROR: 'FAILED_UNRESOLVED_ERROR',
});
const FAILURE_STATUSES = Object.freeze([
TERMINAL.FAILED_ZERO_WRITE, TERMINAL.FAILED_PARTIAL, TERMINAL.FAILED_UNRESOLVED_ERROR,
]);
function isRetentionFailure(status) { return FAILURE_STATUSES.includes(status); }
/**
* Classify the EXACT object `persist()` returned. It never recomputes
* attempted or written — a second calculation could disagree with the writer,
* and then the status would describe something that did not happen.
*/
function classifyPersist(result) {
if (!result || typeof result.attempted !== 'number' || typeof result.written !== 'number') {
throw new TypeError('classifyPersist requires the exact persist() result');
}
const { attempted, written, skipped, error } = result;
if (attempted === 0) return TERMINAL.NOTHING_TO_PERSIST;
if (skipped) return TERMINAL.SKIPPED_NO_DATABASE;
if (written === 0) return TERMINAL.FAILED_ZERO_WRITE;
if (written < attempted) return TERMINAL.FAILED_PARTIAL;
// written >= attempted from here.
if (error) return TERMINAL.FAILED_UNRESOLVED_ERROR;
return written === attempted ? TERMINAL.COMPLETE : TERMINAL.FAILED_UNRESOLVED_ERROR;
}
/**
* Latest terminal retention result per sport, for the internal status probe.
* In-memory and per-process on purpose: this is an observability surface, not
* a record. The record is model_snapshots.
*/
const lastTerminal = new Map();
function recordTerminal({ sport, snapshotId, result, completedAt }) {
const status = classifyPersist(result);
const entry = Object.freeze({
sport: sport || null,
snapshot_id: snapshotId || null,
attempted: result.attempted,
written: result.written,
status,
completed_at: completedAt || new Date().toISOString(),
code_sha: codeSha(),
// Message only — never row payloads.
error_summary: result.error ? String(result.error).slice(0, 300) : null,
});
if (sport) lastTerminal.set(sport, entry);
return entry;
}
function lastRetention() {
const out = {};
for (const [sport, entry] of lastTerminal) out[sport] = entry;
return out;
}
function resetTerminal() { lastTerminal.clear(); }
module.exports = {
MODEL_VERSION,
TERMINAL,
classifyPersist,
isRetentionFailure,
recordTerminal,
lastRetention,
resetTerminal,
REPAIRED_CHAMPION_VERSION,
codeSha,
rowsFromSides,
+25 -13
View File
@@ -658,23 +658,35 @@ async function runSnapshot(sport, opts = {}) {
// commitPublication AFTER the slate write, so at this point no row has a
// read_id yet — building an index here would produce an empty one and
// quietly leave every ledger row unlinked.
console.log(`[snapshot] retention ${sp}: ${r.written}/${r.attempted} rows${r.skipped ? ' (skipped — no supabase env)' : ''}${r.error ? ` ERROR: ${r.error}` : ''}`);
// RETENTION FAILURE MUST NOT BE SILENT.
// TERMINAL RETENTION STATE — classified from the EXACT persist() result,
// never recomputed. Chunked writes stop at the first failed chunk, so a
// failed cycle can leave durable rows behind: presence of rows under a
// snapshot_id proves nothing about whether the cohort is complete.
const terminal = retention.recordTerminal({
sport: sp,
snapshotId: retentionCtx.snapshotId,
result: r,
completedAt: deps.now(),
});
console.log(`[snapshot] retention ${sp}: ${terminal.status} ${r.written}/${r.attempted} rows snapshot_id=${terminal.snapshot_id}${r.error ? ` ERROR: ${r.error}` : ''}`);
// RETENTION FAILURE MUST NOT BE SILENT — INCLUDING A PARTIAL ONE.
//
// Retention stays best-effort — the product keeps publishing to Redis and
// that is deliberate. But a failed batch previously reached only this log
// Retention stays best-effort: the product keeps publishing to Redis and
// that is deliberate. But a failed batch previously reached only a log
// line, and a PostgREST 400 (`published_side` is not a column) killed
// every batch for every sport with nothing surfacing anywhere. The
// measurement record died quietly while the product looked healthy.
// every batch for every sport with nothing surfacing anywhere.
//
// `skipped` is NOT a failure: it means no database is configured, which is
// the normal state in dev and test.
if (r.error || (!r.skipped && r.attempted > 0 && r.written === 0)) {
// FAILED_PARTIAL is the dangerous successor to that bug: earlier chunks
// committed, the cohort is short, and the surviving rows make it look
// healthy. It alerts exactly as loudly as a total failure.
//
// NOTHING_TO_PERSIST and SKIPPED_NO_DATABASE are not failures.
if (retention.isRetentionFailure(terminal.status)) {
await deps.notify(
`Retention write FAILED for ${sp.toUpperCase()} — ${r.written}/${r.attempted} rows persisted. `
+ `stage=model_snapshots snapshot_id=${retentionCtx.snapshotId} code_sha=${retention.codeSha() || 'unknown'} `
+ `at=${deps.now()} error=${r.error || 'zero rows written with candidates present'}. `
+ 'The product is unaffected; the historical record for this cycle is missing.',
`Retention ${terminal.status} for ${sp.toUpperCase()} — ${terminal.written}/${terminal.attempted} rows persisted. `
+ `stage=model_snapshots snapshot_id=${terminal.snapshot_id} code_sha=${terminal.code_sha || 'unknown'} `
+ `at=${terminal.completed_at} error=${terminal.error_summary || 'none reported'}. `
+ 'The product is unaffected; this retention cohort is INCOMPLETE and must not be used as evidence.',
{ title: 'VYNDR pipeline', priority: 'high', tags: ['rotating_light'] },
);
}