Retention completion: a cohort is complete only when the writer says N of N

The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.

So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.

TERMINAL RETENTION STATES (retentionService.classifyPersist):
  NOTHING_TO_PERSIST       attempted 0 — a refusal-only slate is still a cycle
  SKIPPED_NO_DATABASE      no database configured; not a failure
  COMPLETE                 attempted > 0, written === attempted, no error
  FAILED_ZERO_WRITE        written === 0 — first chunk failed
  FAILED_PARTIAL           0 < written < attempted — a later chunk failed
  FAILED_UNRESOLVED_ERROR  counts look complete but an error is unresolved;
                           unreachable through today's loop, and kept because
                           the alternative is reporting COMPLETE holding an error

The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.

classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.

`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.

VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.

OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.

RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.

Six teeth, each with the injection verified present, against a green baseline:
  1 written>0 as generic success        -> 6 fail
  2 later-chunk failure reports COMPLETE -> 5 fail
  3 FAILED_PARTIAL does not alert        -> 3 fail
  4 status reports a recalculated count  -> 1 fail
  5 row presence treated as completion   -> 1 fail
  6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).

Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.

Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.

383 suites / 5,118 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
Kev
2026-08-27 19:12:59 -04:00
parent 35da190f2c
commit 9809626c99
8 changed files with 455 additions and 39 deletions
+99
View File
@@ -727,8 +727,107 @@ function newSnapshotId() {
return crypto.randomUUID();
}
/* ------------------------------------------------------------------ *
* TERMINAL RETENTION STATE
*
* `persist()` writes in CHUNKS and STOPS ON THE FIRST FAILED CHUNK. Chunks
* committed before the failure are already durable, so a failed cycle can
* leave real, valid-looking rows behind. Row presence under a snapshot_id is
* therefore NOT completion evidence, and neither is a uniform captured_at.
*
* The invariant: written < attempted with attempted > 0 is a FAILED cycle.
* A partial cycle is never degraded success.
*
* `written` counts rows in successfully COMMITTED CHUNKS — not rows inserted.
* The upsert uses ignoreDuplicates, so a re-run legitimately inserts far fewer
* database rows than it writes. Comparing `written` to count(*) for a
* snapshot_id will disagree by design; that is not a partial write.
* ------------------------------------------------------------------ */
const TERMINAL = Object.freeze({
/** attempted === 0 — a refusal-only slate is still a legitimate cycle. */
NOTHING_TO_PERSIST: 'NOTHING_TO_PERSIST',
/** No database configured (dev/test). Not a failure. */
SKIPPED_NO_DATABASE: 'SKIPPED_NO_DATABASE',
/** attempted > 0, written === attempted, no unresolved error. */
COMPLETE: 'COMPLETE',
/** attempted > 0, written === 0 — the first chunk failed. */
FAILED_ZERO_WRITE: 'FAILED_ZERO_WRITE',
/** attempted > 0, 0 < written < attempted — a later chunk failed. */
FAILED_PARTIAL: 'FAILED_PARTIAL',
/**
* Counts look complete but an error is unresolved. Unreachable through
* today's loop (it breaks before crediting a failed chunk), and kept
* because the alternative is reporting COMPLETE with an error in hand.
*/
FAILED_UNRESOLVED_ERROR: 'FAILED_UNRESOLVED_ERROR',
});
const FAILURE_STATUSES = Object.freeze([
TERMINAL.FAILED_ZERO_WRITE, TERMINAL.FAILED_PARTIAL, TERMINAL.FAILED_UNRESOLVED_ERROR,
]);
function isRetentionFailure(status) { return FAILURE_STATUSES.includes(status); }
/**
* Classify the EXACT object `persist()` returned. It never recomputes
* attempted or written — a second calculation could disagree with the writer,
* and then the status would describe something that did not happen.
*/
function classifyPersist(result) {
if (!result || typeof result.attempted !== 'number' || typeof result.written !== 'number') {
throw new TypeError('classifyPersist requires the exact persist() result');
}
const { attempted, written, skipped, error } = result;
if (attempted === 0) return TERMINAL.NOTHING_TO_PERSIST;
if (skipped) return TERMINAL.SKIPPED_NO_DATABASE;
if (written === 0) return TERMINAL.FAILED_ZERO_WRITE;
if (written < attempted) return TERMINAL.FAILED_PARTIAL;
// written >= attempted from here.
if (error) return TERMINAL.FAILED_UNRESOLVED_ERROR;
return written === attempted ? TERMINAL.COMPLETE : TERMINAL.FAILED_UNRESOLVED_ERROR;
}
/**
* Latest terminal retention result per sport, for the internal status probe.
* In-memory and per-process on purpose: this is an observability surface, not
* a record. The record is model_snapshots.
*/
const lastTerminal = new Map();
function recordTerminal({ sport, snapshotId, result, completedAt }) {
const status = classifyPersist(result);
const entry = Object.freeze({
sport: sport || null,
snapshot_id: snapshotId || null,
attempted: result.attempted,
written: result.written,
status,
completed_at: completedAt || new Date().toISOString(),
code_sha: codeSha(),
// Message only — never row payloads.
error_summary: result.error ? String(result.error).slice(0, 300) : null,
});
if (sport) lastTerminal.set(sport, entry);
return entry;
}
function lastRetention() {
const out = {};
for (const [sport, entry] of lastTerminal) out[sport] = entry;
return out;
}
function resetTerminal() { lastTerminal.clear(); }
module.exports = {
MODEL_VERSION,
TERMINAL,
classifyPersist,
isRetentionFailure,
recordTerminal,
lastRetention,
resetTerminal,
REPAIRED_CHAMPION_VERSION,
codeSha,
rowsFromSides,