MLB canonical event identity, impossible-binding refusal, event-aware dedupe, publication commit

Release-isolated slice built from 41ba38e. Ships ONLY the event-integrity +
publication + lineage-canary closure; the 90-path development tree stays
undeployed.

- canonical MLB event identity from statsapi gamePk (mlb:gamepk:<pk>), with
  event_identity_source/method/version recorded. The id is canonical; the
  binding is derived and says so.
- IMPOSSIBLE-BINDING REFUSAL. Verified in prod 2026-08-26: Joe Mack (Marlins)
  was bound to Dodgers@Braves and Yandy Diaz (Rays) to Rangers@WhiteSox, both
  from one book in the 01:00/03:01 UTC cycles after their own games began. Root
  cause is source market data, not the binder. A prop whose player's team is not
  an event participant now refuses; unknown team preserves uncertainty.
- event-aware dedupe: books still collapse, events no longer do. An unresolved
  MLB event fractures rather than falling back to the collision-prone
  date+teams key.
- publication commit moved AFTER the authoritative Redis slate write, with
  exact parity-gap identity when the product publishes and the record does not.
- lineage dual-write behind LINEAGE_CANARY_SPORTS, DISABLED for this deploy.

Excluded deliberately: WNBA feed/chain, market ontology, PerformanceDistribution,
calibration certification, truth diagnostics, applyRevision Phase-1, and the
analyzeViaEngine1 confidence-rounding change (a served field).

Suite 380/5,040/0 from this worktree; web tsc exit 0; champion output identical
to production.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
Kev
2026-08-27 16:27:36 -04:00
parent 41ba38ed4e
commit 352016790a
16 changed files with 4217 additions and 6 deletions
+89
View File
@@ -0,0 +1,89 @@
-- Migration 045: append-only Read lineage (ADDITIVE ONLY).
--
-- Every statement is ADD COLUMN IF NOT EXISTS or CREATE INDEX. No column is
-- dropped, renamed, retyped or backfilled. No historical row is rewritten.
--
-- ── WHY THIS IS COLUMNS AND NOT A NEW CLAIM TABLE ──────────────────────────
-- `model_snapshots` is ALREADY an append-only store of published claims: it is
-- upserted with ignoreDuplicates on (snapshot_id, player_key, stat, line, side),
-- so a row is written once per snapshot cycle and never overwritten, and
-- `captured_at` is NOT NULL. Its only post-insert write is settlement
-- (outcome / actual_value / settled_at), which is not part of the claim.
--
-- MEASURED (mlb, 2026-08-26): 5,376 props, 4,286 with more than one row, and
-- 183 whose GRADE CHANGED across those rows. The immutable claims exist. What
-- was missing is the lineage over them, which is what these columns are.
--
-- Building a separate revision table would have duplicated a claim that is
-- already immutable, and given VYNDR two stores that can disagree.
--
-- ── THE ROW'S OWN id IS THE REVISION IDENTITY ─────────────────────────────
-- `model_snapshots.id` is a bigserial primary key on an append-only row. It is
-- already immutable and already unique, so a parallel text `revision_id` would
-- be a second spelling of the same fact. `supersedes_id` and `recaptures_id`
-- point at it directly.
--
-- ── CAPTURE IS NOT REVISION ───────────────────────────────────────────────
-- Most cycles re-observe an unchanged claim. `claim_digest` (sha256 over the
-- market terms plus gradeFreeze.SERVED_FIELDS) is what separates a genuine new
-- published state from a re-observation, so the chronology reports the ~183
-- revisions that happened rather than the ~13,000 captures that were written.
-- The logical Read: stable across cycles, and deliberately NOT derived from the
-- game_id string. Measured over 30,746 identity groups, 1,328 carried more than
-- one game_id spelling for the same real game while only 18 carried genuinely
-- different team pairs. Identity that included the raw spelling would have
-- fractured 1,310 Reads that are one Read.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS read_id text;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS read_natural_key text;
-- The published state.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS claim_digest text;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS revision_ordinal integer;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS supersedes_id bigint;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS recaptures_id bigint;
-- ORIGIN | REVISION | RECAPTURE. A row says which it is rather than leaving a
-- reader to infer it from ordinals.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS lineage_action text;
-- LIVE | LEGACY_UNVERIFIED. Rows written before this migration keep NULL, which
-- is the honest value: their position in a chronology is unknown and is not
-- reconstructed.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS lineage_state text;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS lineage_version text;
-- Walking one Read's chronology in order.
CREATE INDEX IF NOT EXISTS model_snapshots_lineage_idx
ON model_snapshots (read_id, revision_ordinal) WHERE read_id IS NOT NULL;
-- Resolving a candidate's family on write.
CREATE INDEX IF NOT EXISTS model_snapshots_natural_key_idx
ON model_snapshots (read_natural_key) WHERE read_natural_key IS NOT NULL;
-- FORK PREVENTION, not merely fork detection.
--
-- Only ONE row may supersede a given parent. Two workers that each read the
-- same standing revision before either wrote cannot see one another, so no
-- in-memory check can catch that race — the second insert must fail, and this
-- is what makes it fail. Without it the chain could silently become
-- R1 -> {R2a, R2b} and no reader could say which claim VYNDR actually published.
CREATE UNIQUE INDEX IF NOT EXISTS model_snapshots_supersedes_unique
ON model_snapshots (supersedes_id) WHERE supersedes_id IS NOT NULL;
-- ── LEDGER LINKAGE (SHADOW) ───────────────────────────────────────────────
-- A receipt should be able to name the EXACT revision it settles. These are
-- written in shadow alongside the existing settlement path and are read by
-- nothing; current settlement outcomes are unchanged. NULL on every existing
-- row, and no historical receipt is re-pointed.
ALTER TABLE ledger_entries ADD COLUMN IF NOT EXISTS read_id text;
ALTER TABLE ledger_entries ADD COLUMN IF NOT EXISTS read_revision_id bigint;
CREATE INDEX IF NOT EXISTS ledger_entries_read_id_idx
ON ledger_entries (read_id) WHERE read_id IS NOT NULL;
COMMENT ON COLUMN model_snapshots.read_id IS
'The stable identity of a logical Read across every published revision. Deliberately NOT derived from game_id: measured, 1,328 of 30,746 identity groups carry more than one game_id spelling for the same real game.';
COMMENT ON COLUMN model_snapshots.claim_digest IS
'sha256 over the market terms (line, side, book, locked_odds) plus gradeFreeze.SERVED_FIELDS. Separates a genuine new published claim from a re-observation of the standing one.';
COMMENT ON COLUMN model_snapshots.lineage_action IS
'ORIGIN | REVISION | RECAPTURE. Most cycles re-observe an unchanged claim; only a changed claim_digest advances the chain.';
COMMENT ON COLUMN ledger_entries.read_revision_id IS
'SHADOW: the exact model_snapshots row this receipt settles. Read by nothing; settlement outcomes are unchanged by migration 045.';
@@ -0,0 +1,52 @@
-- Migration 046: publication semantics (ADDITIVE ONLY).
--
-- Every statement is ADD COLUMN IF NOT EXISTS or CREATE INDEX. Nothing is
-- dropped, renamed, retyped or backfilled. Historical rows keep NULL.
--
-- ── WHY 045 WAS NOT ENOUGH ────────────────────────────────────────────────
-- Migration 045 gave `model_snapshots` a lineage. Tracing the publication path
-- afterwards showed the lineage would have described the wrong population:
-- `gradeSlateService` fires its retention hook with BOTH sides, graded AND
-- refused, BEFORE any filtering, and only the higher-confidence graded side
-- becomes the served Read.
--
-- MEASURED, mlb 2026-08-26: of 13,012 captured rows, 3,859 were refusals and
-- only 4,569 matched a prop that reached the ledger. **8,443 rows (64.9%)
-- describe a state no user was ever shown**, and nothing on the row said which
-- was which. Building chronology over all of them would have produced
-- "publication history" for model activity that was never published.
--
-- `published` is the smallest missing link. It is set by the collector's
-- publication signal at the one place that knows the answer — the moment the
-- winner is selected — and defaults FALSE, because a capture is not a
-- publication until something says it is.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS published boolean;
-- ── WHY THE PUBLISHED STATE ADVANCED ──────────────────────────────────────
-- The revision ordinal says WHEN. It cannot say WHY, and the difference is the
-- whole product question: VYNDR's belief can hold still while the market moves.
-- A replay of one real date found 4,507 digest changes of which only 301 moved
-- the grade — 4,206 were price alone. One undifferentiated "revision" count
-- conflates a changed opinion with a changed price.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS change_type text;
-- ── CLAIM SEMANTICS ARE VERSIONED ─────────────────────────────────────────
-- A digest is only interpretable against the claim definition that produced it.
-- The first digest listed only `locked_odds` while this table stores
-- `over_odds`/`under_odds`, so it was blind to price on exactly the rows it
-- digested — caught by a replay crashing on a missing column, i.e. by luck.
-- These stamps make the next such change visible instead of silent. A
-- historical digest is never recomputed under a newer definition.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS claim_schema_version text;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS digest_algorithm_version text;
-- Publication chronology is read published-only; the partial index matches.
CREATE INDEX IF NOT EXISTS model_snapshots_published_lineage_idx
ON model_snapshots (read_id, revision_ordinal) WHERE published IS TRUE;
COMMENT ON COLUMN model_snapshots.published IS
'TRUE only for the side that actually became the served Read. Measured: 64.9% of captured rows describe a state no user was shown, and before this column nothing on the row distinguished them. NULL on rows written before migration 046 - unknown, not reconstructed.';
COMMENT ON COLUMN model_snapshots.change_type IS
'INITIAL_PUBLICATION | BELIEF_CHANGE | MARKET_REPRICE | MARKET_LINE_CHANGE | COMPARISON_CHANGE | PROVENANCE_CHANGE | PRESENTATION_CHANGE | MIXED_CHANGE | NO_MATERIAL_PUBLISHED_CHANGE. Ordinal says when; this says why.';
COMMENT ON COLUMN model_snapshots.claim_schema_version IS
'The claim definition in force when the digest was computed. Never recompute a historical digest under a newer definition and present it as original.';
@@ -0,0 +1,58 @@
-- Migration 047: canonical event identity + publication commit (ADDITIVE ONLY).
--
-- Every statement is ADD COLUMN IF NOT EXISTS or CREATE INDEX. Nothing is
-- dropped, renamed, retyped or backfilled. Historical rows keep NULL.
--
-- ── WHY: EVENT LABELS ARE NOT EVENT IDENTITY ─────────────────────────────
-- `ledgerService.gameIdFor` derives `sport:date:away@home`. That has no
-- occurrence component, so both halves of a doubleheader produce a
-- BYTE-IDENTICAL string.
--
-- VERIFIED against statsapi, MLB 2026 through 2026-08-27: 19 doubleheaders /
-- 38 real games collapse into 19 derived labels. Real case 2026-08-17,
-- St. Louis @ Cincinnati: gamePk 824514 (game 1, 17:40Z) and 824478 (game 2,
-- 22:40Z). VYNDR recorded ONE id holding 171 ledger rows and FOUR distinct
-- starting pitchers -- both games merged into one event.
--
-- `canonical_event_id` is namespaced (`mlb:gamepk:824514`) so two sports'
-- numeric ids can never collide. The legacy `game_id` is KEPT for diagnostics
-- and compatibility and is explicitly not authoritative.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS canonical_event_id text;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS event_identity_source text;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS event_identity_method text;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS event_identity_version text;
-- The source's own occurrence number (statsapi `gameNumber`). Never inferred.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS event_occurrence integer;
-- ── WHY: PUBLISHED MUST MEAN COMMITTED ───────────────────────────────────
-- The authoritative served slate is ONE atomic Redis write,
-- `cacheSet('snapshot:{sport}:latest')`. Retention persists BEFORE it
-- (snapshotService:1137 vs :1152) and the winner is chosen earlier still, so
-- neither capture nor selection means published. `published_at` records the
-- commit instant and `publication_id` the slate it committed in -- the write is
-- slate-atomic, so publication is slate-level and every served row shares one.
--
-- `captured_at` answers when the model state was captured. `published_at`
-- answers when the composed Read became authoritative. They are different
-- questions and are stored separately.
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS published_at timestamptz;
ALTER TABLE model_snapshots ADD COLUMN IF NOT EXISTS publication_id text;
-- Chronology is walked per canonical event.
CREATE INDEX IF NOT EXISTS model_snapshots_canonical_event_idx
ON model_snapshots (canonical_event_id) WHERE canonical_event_id IS NOT NULL;
-- Publication parity is read per slate.
CREATE INDEX IF NOT EXISTS model_snapshots_publication_idx
ON model_snapshots (publication_id) WHERE publication_id IS NOT NULL;
-- Ledger linkage to the canonical event, additive and read by nothing.
ALTER TABLE ledger_entries ADD COLUMN IF NOT EXISTS canonical_event_id text;
CREATE INDEX IF NOT EXISTS ledger_entries_canonical_event_idx
ON ledger_entries (canonical_event_id) WHERE canonical_event_id IS NOT NULL;
COMMENT ON COLUMN model_snapshots.canonical_event_id IS
'Namespaced canonical event identity, e.g. mlb:gamepk:824514. MLB only; other sports record UNSUPPORTED_SPORT rather than a guessed id. The legacy game_id is a derived LABEL and is not authoritative -- it cannot separate a doubleheader.';
COMMENT ON COLUMN model_snapshots.published_at IS
'When the composed Read became authoritative -- i.e. when the atomic slate write to snapshot:{sport}:latest succeeded. NOT the capture time, and never set by winner selection alone.';
COMMENT ON COLUMN model_snapshots.publication_id IS
'The slate whose atomic Redis write committed this row. Publication is slate-level because the write is.';