3 Commits

Author SHA1 Message Date
builtbykev cc5bdf5797 A canary is temporary: bound activation with a fail-closed lease
`LINEAGE_CANARY_SPORTS=mlb` was parsed once at module load and frozen, so an
enabled canary stayed writable for the whole process lifetime. On 2026-08-29 it
was left set and the 22:00Z, 01:00Z and 03:00Z scheduled snapshots each wrote
lineage overnight with nobody watching. That history happened to be correct.
Lineage is append-only evidence, so a defective canary would have written
irreversible WRONG evidence exactly as quietly. Correct history was luck, not a
safety property.

NEW CONTRACT:  LINEAGE_CANARY_SPORTS=mlb@2026-08-30T18:00:00Z

Absolute UTC instant only -- no duration, no local timezone. A relative "4h"
would silently restart on every redeploy, which is the defect being removed.

LEGACY `mlb` NO LONGER ACTIVATES ANYTHING. It is INVALID_MISSING_EXPIRY. Leaving
it working would have left the defect in place behind a nicer-looking
alternative, so this is the load-bearing half of the repair.

EXPIRY IS EVALUATED AT THE WRITE GATE, not at startup. `isEnabled(sport, now)`
re-reads the clock on every call and `lineageCanaryEnabled` threads it through,
so a lease turns itself off with no operator, no restart, no Redis and no
network. A design where an expired canary keeps writing until someone restarts
is the same failure in a different hat.

MAX_LEASE is FOUR HOURS, and the bound is proven rather than chosen. MLB ticks
are [14,19,22,1,3] UTC; exhaustively over a minute grid across UTC date
boundaries, the shortest span enclosing THREE consecutive ticks is
22:00 -> 01:00 -> 03:00 = five hours. Four hours encloses at most two, with an
hour of margin. (An earlier note claimed six hours admitted two. It admits
three; that claim was false and the test now pins the arithmetic.) The bound is
a function of the scheduler, so a test reads the hours from sportCadence and
fails if a cadence change invalidates the proof.

Everything fails closed: absent, empty, bare sport, comma list, two leases,
duplicate @, non-leasable sport, missing timezone, numeric offset, malformed,
over-long, already-expired, expiry-equals-now, and an unreadable clock.
Configured-but-EXPIRED stays distinguishable from never-configured so automatic
containment is auditable.

No Redis, no Supabase, no counter, no network in the lease path -- an
unreachable dependency must never decide whether lineage may write.

Lineage algorithms, cache-date, participant and retention identity untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 14:58:11 -04:00
builtbykev 9809626c99 Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.

So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.

TERMINAL RETENTION STATES (retentionService.classifyPersist):
  NOTHING_TO_PERSIST       attempted 0 — a refusal-only slate is still a cycle
  SKIPPED_NO_DATABASE      no database configured; not a failure
  COMPLETE                 attempted > 0, written === attempted, no error
  FAILED_ZERO_WRITE        written === 0 — first chunk failed
  FAILED_PARTIAL           0 < written < attempted — a later chunk failed
  FAILED_UNRESOLVED_ERROR  counts look complete but an error is unresolved;
                           unreachable through today's loop, and kept because
                           the alternative is reporting COMPLETE holding an error

The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.

classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.

`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.

VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.

OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.

RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.

Six teeth, each with the injection verified present, against a green baseline:
  1 written>0 as generic success        -> 6 fail
  2 later-chunk failure reports COMPLETE -> 5 fail
  3 FAILED_PARTIAL does not alert        -> 3 fail
  4 status reports a recalculated count  -> 1 fail
  5 row presence treated as completion   -> 1 fail
  6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).

Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.

Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.

383 suites / 5,118 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 19:12:59 -04:00
builtbykev ceaa896f77 Runtime observability: report the build and canary state the system acts on
The rollout stalled at RUNTIME_UNVERIFIED because two facts were answerable only
as a side effect of a scheduled snapshot writing a row: which build is running,
and whether MLB lineage is effectively enabled. Every state transition therefore
waited on cron rather than on asking the service.

- src/services/lineageCanaryConfig.js — THE canary resolver. Parsed once at
  module load (unchanged semantics), normalised sorted/deduped/trimmed, frozen.
  snapshotService's write gate now delegates to it, and the status probe reads
  the SAME state. A route that parsed the environment itself would be a second
  version of the truth, free to drift from the gate it claims to report.
- GET /api/internal/snapshot/status gains runtime.code_sha (the production
  codeSha resolver — never git, never gitea/main; null when unavailable),
  runtime.started_at (computed ONCE at module load, so it marks a boundary
  rather than reading as now; deliberately not called deployed_at), and
  lineage_canary {enabled, sports, configuration_source}.
- No raw environment value is returned; sports is the normalised set and
  configuration_source says only ENVIRONMENT vs DEFAULT. Router-wide
  requireInternalAuth is unchanged: 200 with key, 401 without.
- Effective lineage config is fixed for the process lifetime, so
  runtime.started_at is a defensible lower bound for how long that state held.

Strictly observational — the handler still only reads Redis.

Model and decision code byte-identical to 8c6aef1: analyzeViaEngine1,
probabilityEstimator, gradeRanking, eventIdentity, gradeSlateService,
retentionService, ledgerService, mlbStatsAdapter.

Suite 381/5,081/0 from the release worktree; web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 17:16:05 -04:00