Event-aware retention identity: two games, two receipts

One player prop in Game 1 and the same-looking prop in Game 2 are two different
historical claims. The retention conflict identity did not know that.

SEMANTIC IDENTITY FIRST. Two outbound rows are the same retention proposition
within one cycle when they share the cycle, the EVENT, the participant, the
stat, the line and the side. Book is deliberately absent — collapsing books is
dedupeProps's actual job and the price anchor is chosen later. The database
index is enforcement of that answer, never the definition of it.

THE EVENT COMPONENT NEVER FABRICATES. canonical_event_id where a sport has a
resolver — MLB's admission gate rejects unresolved/ambiguous/contradicted props
BEFORE grading, so every row that can reach retention has one — and game_id
otherwise, which is NOT NULL in the schema and is the only event label sports
without a resolver possess. Both are in the identity, so the weaker label still
discriminates where the stronger is absent.

NULLS NOT DISTINCT IS LOAD-BEARING, NOT STYLISTIC. canonical_event_id is NULL
for every non-MLB row. Measured on a disposable PG17: under PostgreSQL's default
semantics the same NBA proposition inserted twice produced TWO rows — every
retry duplicating for ever. With NULLS NOT DISTINCT the same test yields one.
That measurement is what rejected the plain composite option.

MIXED-FLEET BRIDGE. A rollout serves both builds at once (measured 11/12 new,
1 old). Old and new writers need different indexes and NO schema state satisfies
both: with the legacy index present a new writer fails 23505 on a doubleheader;
with it gone an old writer fails 42P10. A bare ON CONFLICT DO NOTHING would have
bridged this, and PostgREST does not emit one — `ignoreDuplicates` WITHOUT
`onConflict` was measured raising a real duplicate-key error, so that bridge does
not exist through this client.

So the writer bridges it. It targets the event-aware identity and, on exactly
the two errors meaning "the schema is not in the state I expect" (42P10, or
23505 NAMING the legacy index), retries the SAME chunk on the legacy target. A
failed chunk rolls back atomically — measured 0 rows — so the retry cannot
double-write. Correct in every schema state: legacy-only and both-present
degrade to legacy semantics with no outage; new-only keeps both games.

The bridge is deliberately narrow. A supersedes conflict is ALSO a 23505, and
swallowing it would destroy the forked-history guard, so the legacy index must
be named. All three model_snapshots writers (persist, commitPublication,
recoverFromFork) go through it; no hardcoded legacy target survives.

MEASURED, through the real supabase-js -> PostgREST -> Postgres path on
production-shaped PG17:
  * 1,000 REAL propositions from the verified 2026-08-17 STL@CIN doubleheader
    (1,738 retained rows under ONE game_id), replayed across both real gamePks:
    OLD index materialized 1,000 of 2,000 — 1,000 LOST. NEW index materialized
    2,000 of 2,000 — 0 lost.
  * retry idempotency, over/under, line, stat, player, non-MLB same-game and
    non-MLB different-game all behave correctly under the new index.
  * ORDINARY-SLATE PARITY over ALL 434 real cohorts / 328,262 retained rows:
    old identities 328,262, new identities 328,262, delta 0, cohorts changed 0.
    The index is therefore guaranteed creatable and nothing historical splits.

CONFLICT_IDENTITY is now DERIVED from RETENTION_CONFLICT rather than restated —
a test caught them silently disagreeing, which is exactly how the materialization
check could have expected an identity the database no longer enforced.

EXPAND/CONTRACT are separate files on purpose. 048 is additive and retires
nothing; 049 drops the legacy index and must not be applied until fleet
convergence is proven by sampling, never assumed from a fast rollout.

NO BACKFILL. Legacy rows keep NULL canonical_event_id and remain LEGACY
EVENT-AGNOSTIC RETENTION, which is what that NULL truthfully says.

The materialization defence is untouched and now reports the bridge honestly:
while the legacy index still collapses a doubleheader, expected 4 vs actual 2
yields MATERIALIZATION_MISSING and the cohort is refused.

Nine teeth, injections verified present, against a green baseline of 97:
1 event distinction removed (10) · 2 phases collapsed (2) · 3 bridge swallows
everything (6) · 4 NULLS NOT DISTINCT removed (1) · 5 old-container error as
success (3) · 6 semantic/DB identity disagree (8) · 7 collision detector removed
(2) · 8 partial transport usable (3) · 9 collision unannounced (1).
Restored byte-identically.

Model and product untouched: gradeSlateService (event-aware dedupe), event
identity, ledger, calibration, chain, lineage config and the status route all
UNCHANGED. Zero cacheSet changes, zero web paths, schema contract unchanged (no
new columns). Lineage stays OFF.

385 suites / 5,178 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
This commit is contained in:
Kev
2026-08-27 20:30:45 -04:00
parent 11277a1b99
commit 048e4eaa3f
5 changed files with 481 additions and 37 deletions
@@ -0,0 +1,47 @@
-- 048 — EXPAND: event-aware retention proposition identity.
--
-- WHY. The retention conflict identity is
-- model_snapshots_cycle_prop_uniq (snapshot_id, player_key, stat, line, side)
-- which carries no event component. An MLB doubleheader — the same hitter, the
-- same stat, the same line, in two genuinely different games — is therefore ONE
-- identity, and `ignoreDuplicates` discards the second row with no error while
-- transport reports COMPLETE. Verified on production data: 2026-08-17
-- St. Louis @ Cincinnati holds 1,738 retained rows under a single game_id.
--
-- WHAT. An additive unique index that adds the EVENT and nothing else.
--
-- game_id AND canonical_event_id are both included, on purpose:
-- * canonical_event_id splits a doubleheader (game_id is byte-identical for
-- both halves), and MLB's admission gate makes it non-null for every row
-- that can reach retention;
-- * game_id is NOT NULL in the schema and is the only event label sports
-- without a canonical resolver possess, so the weaker label still
-- discriminates where the stronger one is absent. Nothing is fabricated.
--
-- NULLS NOT DISTINCT is required, not stylistic. canonical_event_id is NULL for
-- every non-MLB row, and under PostgreSQL's default NULL semantics two NULLs
-- are DISTINCT — measured on a disposable PG17: the same NBA proposition
-- inserted twice produced TWO rows, i.e. every retry would duplicate for ever.
-- With NULLS NOT DISTINCT the same test yields one row.
--
-- SAFE TO CREATE. Measured over all 434 real cohorts / 328,262 retained rows,
-- the new identity produces exactly 328,262 identities — identical to the old
-- count, 0 cohorts changed. Adding columns can only split, and legacy rows all
-- carry canonical_event_id IS NULL with one game_id per proposition, so no
-- existing row can violate it.
--
-- THIS MIGRATION DOES NOT RETIRE THE OLD INDEX. Both coexist deliberately so a
-- rolling deploy keeps working; the old index is dropped in 049 only after the
-- fleet is proven converged onto a writer that no longer needs it.
--
-- NO BACKFILL. No historical value is written or rewritten. Legacy rows remain
-- LEGACY EVENT-AGNOSTIC RETENTION, and their NULL canonical_event_id continues
-- to say exactly that.
create unique index if not exists model_snapshots_cycle_event_prop_uniq
on public.model_snapshots
(snapshot_id, game_id, canonical_event_id, player_key, stat, line, side)
nulls not distinct;
comment on index public.model_snapshots_cycle_event_prop_uniq is
'Event-aware retention proposition identity. Supersedes model_snapshots_cycle_prop_uniq, which merged doubleheaders. NULLS NOT DISTINCT so a NULL canonical_event_id (every non-MLB row) still deduplicates on retry.';