Retention completion: a cohort is complete only when the writer says N of N

The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.

So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.

TERMINAL RETENTION STATES (retentionService.classifyPersist):
  NOTHING_TO_PERSIST       attempted 0 — a refusal-only slate is still a cycle
  SKIPPED_NO_DATABASE      no database configured; not a failure
  COMPLETE                 attempted > 0, written === attempted, no error
  FAILED_ZERO_WRITE        written === 0 — first chunk failed
  FAILED_PARTIAL           0 < written < attempted — a later chunk failed
  FAILED_UNRESOLVED_ERROR  counts look complete but an error is unresolved;
                           unreachable through today's loop, and kept because
                           the alternative is reporting COMPLETE holding an error

The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.

classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.

`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.

VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.

OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.

RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.

Six teeth, each with the injection verified present, against a green baseline:
  1 written>0 as generic success        -> 6 fail
  2 later-chunk failure reports COMPLETE -> 5 fail
  3 FAILED_PARTIAL does not alert        -> 3 fail
  4 status reports a recalculated count  -> 1 fail
  5 row presence treated as completion   -> 1 fail
  6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).

Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.

Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.

383 suites / 5,118 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
Kev
2026-08-27 19:12:59 -04:00
parent 35da190f2c
commit 9809626c99
8 changed files with 455 additions and 39 deletions
+15 -7
View File
@@ -168,14 +168,19 @@ describe('RETENTION FAILURE IS NOT SILENT', () => {
});
test('the pipeline emits a high-severity structured event on that failure', () => {
// MECHANISM CHANGED after 35da190: the alert condition was
// `r.error || (!r.skipped && r.attempted > 0 && r.written === 0)`, which
// could not distinguish a PARTIAL cohort from a complete one. It is now
// driven by the terminal status classified from the exact persist result,
// so FAILED_PARTIAL alerts as loudly as FAILED_ZERO_WRITE.
const src = fs.readFileSync(path.join(ROOT, 'src/services/snapshotService.js'), 'utf8');
const block = src.slice(src.indexOf('Retention write FAILED') - 400,
src.indexOf('Retention write FAILED') + 700);
const i = src.indexOf('const terminal = retention.recordTerminal');
expect(i).toBeGreaterThan(-1);
const block = src.slice(i, i + 2200);
for (const field of ['stage=model_snapshots', 'snapshot_id=', 'code_sha=', 'at=', 'error=']) {
expect(block).toContain(field);
}
expect(block).toMatch(/priority: 'high'/);
// Best-effort semantics preserved: the product is explicitly unaffected.
expect(block).toMatch(/product is unaffected/);
});
@@ -183,13 +188,16 @@ describe('RETENTION FAILURE IS NOT SILENT', () => {
const out = await retention.persist(finalOutboundRows(), { getClient: () => null });
expect(out.skipped).toBe(true);
expect(out.error).toBeNull();
const src = fs.readFileSync(path.join(ROOT, 'src/services/snapshotService.js'), 'utf8');
expect(src).toMatch(/!r\.skipped/);
// The skip exemption now lives in the classifier, not in an inline
// `!r.skipped` test at the call site.
expect(retention.classifyPersist(out)).toBe(retention.TERMINAL.SKIPPED_NO_DATABASE);
expect(retention.isRetentionFailure(retention.classifyPersist(out))).toBe(false);
});
test('zero rows written with candidates present also alerts', () => {
const src = fs.readFileSync(path.join(ROOT, 'src/services/snapshotService.js'), 'utf8');
expect(src).toMatch(/r\.attempted > 0 && r\.written === 0/);
const st = retention.classifyPersist({ attempted: 600, written: 0, skipped: false, error: null });
expect(st).toBe(retention.TERMINAL.FAILED_ZERO_WRITE);
expect(retention.isRetentionFailure(st)).toBe(true);
});
});