84f1fc075c2ebb07cdfd41783a2e23b93ab3a27e
3 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cc5bdf5797 |
A canary is temporary: bound activation with a fail-closed lease
`LINEAGE_CANARY_SPORTS=mlb` was parsed once at module load and frozen, so an enabled canary stayed writable for the whole process lifetime. On 2026-08-29 it was left set and the 22:00Z, 01:00Z and 03:00Z scheduled snapshots each wrote lineage overnight with nobody watching. That history happened to be correct. Lineage is append-only evidence, so a defective canary would have written irreversible WRONG evidence exactly as quietly. Correct history was luck, not a safety property. NEW CONTRACT: LINEAGE_CANARY_SPORTS=mlb@2026-08-30T18:00:00Z Absolute UTC instant only -- no duration, no local timezone. A relative "4h" would silently restart on every redeploy, which is the defect being removed. LEGACY `mlb` NO LONGER ACTIVATES ANYTHING. It is INVALID_MISSING_EXPIRY. Leaving it working would have left the defect in place behind a nicer-looking alternative, so this is the load-bearing half of the repair. EXPIRY IS EVALUATED AT THE WRITE GATE, not at startup. `isEnabled(sport, now)` re-reads the clock on every call and `lineageCanaryEnabled` threads it through, so a lease turns itself off with no operator, no restart, no Redis and no network. A design where an expired canary keeps writing until someone restarts is the same failure in a different hat. MAX_LEASE is FOUR HOURS, and the bound is proven rather than chosen. MLB ticks are [14,19,22,1,3] UTC; exhaustively over a minute grid across UTC date boundaries, the shortest span enclosing THREE consecutive ticks is 22:00 -> 01:00 -> 03:00 = five hours. Four hours encloses at most two, with an hour of margin. (An earlier note claimed six hours admitted two. It admits three; that claim was false and the test now pins the arithmetic.) The bound is a function of the scheduler, so a test reads the hours from sportCadence and fails if a cadence change invalidates the proof. Everything fails closed: absent, empty, bare sport, comma list, two leases, duplicate @, non-leasable sport, missing timezone, numeric offset, malformed, over-long, already-expired, expiry-equals-now, and an unreadable clock. Configured-but-EXPIRED stays distinguishable from never-configured so automatic containment is auditable. No Redis, no Supabase, no counter, no network in the lease path -- an unreachable dependency must never decide whether lineage may write. Lineage algorithms, cache-date, participant and retention identity untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
9809626c99 |
Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.
So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.
TERMINAL RETENTION STATES (retentionService.classifyPersist):
NOTHING_TO_PERSIST attempted 0 — a refusal-only slate is still a cycle
SKIPPED_NO_DATABASE no database configured; not a failure
COMPLETE attempted > 0, written === attempted, no error
FAILED_ZERO_WRITE written === 0 — first chunk failed
FAILED_PARTIAL 0 < written < attempted — a later chunk failed
FAILED_UNRESOLVED_ERROR counts look complete but an error is unresolved;
unreachable through today's loop, and kept because
the alternative is reporting COMPLETE holding an error
The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.
classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to
|
||
|
|
ceaa896f77 |
Runtime observability: report the build and canary state the system acts on
The rollout stalled at RUNTIME_UNVERIFIED because two facts were answerable only
as a side effect of a scheduled snapshot writing a row: which build is running,
and whether MLB lineage is effectively enabled. Every state transition therefore
waited on cron rather than on asking the service.
- src/services/lineageCanaryConfig.js — THE canary resolver. Parsed once at
module load (unchanged semantics), normalised sorted/deduped/trimmed, frozen.
snapshotService's write gate now delegates to it, and the status probe reads
the SAME state. A route that parsed the environment itself would be a second
version of the truth, free to drift from the gate it claims to report.
- GET /api/internal/snapshot/status gains runtime.code_sha (the production
codeSha resolver — never git, never gitea/main; null when unavailable),
runtime.started_at (computed ONCE at module load, so it marks a boundary
rather than reading as now; deliberately not called deployed_at), and
lineage_canary {enabled, sports, configuration_source}.
- No raw environment value is returned; sports is the normalised set and
configuration_source says only ENVIRONMENT vs DEFAULT. Router-wide
requireInternalAuth is unchanged: 200 with key, 401 without.
- Effective lineage config is fixed for the process lifetime, so
runtime.started_at is a defensible lower bound for how long that state held.
Strictly observational — the handler still only reads Redis.
Model and decision code byte-identical to
|