Commit Graph

16 Commits

Author SHA1 Message Date
builtbykev be8e16aca9 Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state
anyway. A validator that changes nothing is decoration, so `servable:false` is
now load-bearing: an artifact that fails its policy returns
ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes
with it. The gate sits inside the resolution, not beside the flag that turns the
shadow on, so no environment variable can reach past it — a test asserts
`resolve` never reads process.env at all. Shadow and live consume the SAME
decision, differing only in which promotion stage they demand.

Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This
artifact belongs to a different forecaster" is more precise than "policy
blocked", and the existing state already says it exactly.

THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the
era from config/modelVersion so the query, the artifact and the validator all
read one identity. Measured on the actual fitted set, not a second count:
6,069 current-era rows, 0 wrong-era.

The procedure was then certified on current-era rows ONLY — four walk-forward
folds, training strictly before each evaluation block, 0 future rows in train on
every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323,
delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold.
Mapping spread inside support is 0.001-0.018. The prior mixed-era certification
did not substitute for this.

Policy B selected. A (era-filtered 65/35) and B (all current-era) are
statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has
the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the
PROCEDURE — it is not permanently withheld from the artifact that ships.
withheld_from_fit is 0.

FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01
(distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the
eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…,
served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON.

The runtime no longer fits. It loads. A test greps the service for fitIsotonic,
fromLedger and loadRows and requires all three absent, because the old behaviour
meant a user's number could move with no version, no review and no rollback, and
a past Read could not be reconstructed because its curve no longer existed.
New settled outcomes are forward evidence now; they cannot touch this curve.

Independent reconstruction from the declared training contract alone — fresh
read, fresh digest, fresh fit — reproduces every digest and the curve byte for
byte. Calling the builder twice would only have proven the builder deterministic.

Promotion is a frozen source constant. A snapshot cannot promote, a settlement
cannot promote, a successful fit cannot promote, and dropping a file into the
artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is
explicitly false.

Two coverage holes found by their own teeth. The promotion guard could be
deleted with every test still green, because the promoted file naturally agrees
with itself — extracted as `acceptFile` and tested on the case `load()` cannot
reach. And `validate(null)` returned no `servable` field at all, which is falsy
at a call site and so would have read as correct while asserting nothing.

Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change.
Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:06:18 -04:00
builtbykev 22cf51c4b0 The mapping that ran now has a name, and the row carries it
Runtime probes say the fleet is on a8de676, one process generation. But
"isotonic" was a label, not a claim: probabilityContractService refits per
snapshot against `game_date < todayEt()`, so the mapping changes as outcomes
settle, and nothing on a row could say WHICH mapping produced its number.

The artifact now has an identity:

  estimator_type / estimator_version / certification_version / model_version
  fit_as_of         the exact lt(game_date) bound   2026-09-02
  training_cutoff   last date INSIDE the fit        2026-08-21
  fit_n / knot_count                                6,084 / 28
  knot_digest       d9d571d728ba76de
  served_curve      the COMPLETE served function over [0.50,0.80)
  served_curve_digest

The served curve is not a sample. p_win is quantised to three decimals at the
source, so a step table at 0.001 granularity is the mapping itself for every
input that can occur — six steps, ~200 bytes. Storing it makes a Read
reconstructable WITHOUT re-deriving a training set that may since have been
re-settled, and a claim you can only verify when the inputs happen not to have
moved is not a reconstructable claim.

Proven, not asserted: the production construction path run twice gives an
identical digest, and an INDEPENDENT reconstruction — re-walk 9,361 settled
ledger rows at the declared bound, refit from scratch — reproduces
d9d571d728ba76de exactly, 28 knots for 28.

A teeth injection found a real defect behind a coverage hole. `resolve` checked
the CONTRACT's model era and never the ARTIFACT's, so a mapping fitted for a
different era could be recorded beside a served number with every test green.
Both the era and the estimator type are now checked, and a mismatch serves
nothing rather than serving quietly.

OBSERVED AND NOT CHANGED: calibrationService splits 65/35 to certify its own
bands, a step this contract does not consume because support comes from the
frozen artifact. So the served map is fitted through 2026-08-21 while 3,277
more recent settled rows sit unused, and that lag grows with history. Changing
it would change the fitted function, which this tranche froze.

Shadow still defaults OFF. CALIBRATION_DEPLOYED still []. served_probability is
referenced by nothing outside the contract layer — asserted by a tooth.

Suite 401/401, 5,593 passed, 4 skipped. Teeth 23/23 (prior) + 7/7 (new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:44:52 -04:00
builtbykev a8de676756 A probability is served because evidence supports it, not because nothing else answered
The band gate was blocked for its `else` branch. It read:

    candidate = F(raw)
    served    = inCertifiedBand(candidate) ? candidate : RAW

and above raw 0.60 the model is measured overconfident — holdout raw 0.80-0.90
predicts 0.843 and realizes 0.639. So "the calibrator is not supported here" was
being answered with a number already proven wrong. Unsupported calibration does
not make raw true.

Four candidates were adjudicated on ONE split — fit on the earliest 60% of
train, decide support on the last 40%, evaluate on a holdout that saw neither:

  A low-param      80.2% coverage  0.24374  REFUTED — its extra region
                   (raw 0.80-0.90) certified on cert (err +0.040, n=55) and
                   refuted on holdout (served 0.754 vs observed 0.639), and it
                   leaves a hole at 0.70-0.80 while serving the island above it
  B isotonic       91.3% coverage  0.24337  CERTIFIED, contiguous raw [0.50,0.80)
  C empirical band 91.3% coverage  0.24335  REFUTED — refitted point-in-time on
                   current-model hits the realized rates INVERT in grade order
                   (B+ 0.593 < B 0.614 < C+ 0.623), so the served function steps
                   down at raw 0.78. Its shipped constants come from 3,417 props
                   pooled across four batter stats and do not reproduce here
  D raw identity   43.1% coverage  0.24866  certifies raw 0.50-0.60 and only there

Raw is candidate D, not a fallback. It earns exactly one region (holdout error
+0.010 on n=1,316), which is why the law is "raw must earn its region" rather
than "raw is never true". B already covers that region, so no hybrid is built.

Above raw 0.80 nothing is certified and nothing is served. That is the region
where raw is most wrong, isotonic over-corrects (cert err -0.093) and its LODO
mapping at 0.95 has spread 0.180. 8.7% of holdout rows land there.

The registry did not need changing. `serves(stat, p)` already tested certified
bands against the RAW p_win — support in the input domain, the correct question —
and returned {serve:false, reason}. It never said "serve raw". The output-space
gate and the raw fallback were both invented downstream in calibrationService.

ACTIVATION IS OFF. PROBABILITY_CONTRACT_SHADOW defaults to 0, CALIBRATION_DEPLOYED
stays frozen empty, and every served field is byte-identical. This releases the
support first, which is the required order. The shadow records raw belief, the
candidate served value, the state, the estimator identity, and what EV/Kelly/VALUE
would be under the actionability law — into its own column, read by nothing.

Migration 051 was applied to production BEFORE retentionService named the column.
PostgREST builds a bulk insert from the first row's shape, so a key whose column
does not exist 400s the whole batch silently — that is how migration 038 took
retention down for three days.

The user-facing contradiction is NOT fixed here. A B+ still says "realized about
66%" beside a confidence of 84. Fixing that is activation, and activation costs
32% of VALUE flags and 46% of Kelly recommendations on the holdout.

Suite 401/401, 5,580 passed, 4 skipped, deterministic across three runs.
Teeth 23/23, each independently injected and restored byte-identically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:03:16 -04:00
builtbykev 4aca33deb6 One human, one semantic identity — MLB participant convergence
The collision autopsy left two unrepaired defects, running in OPPOSITE
directions, and `outbound_collision_count` can only ever see one of them.

UNDER-COLLAPSE. Dedupe keys on `mlb:<personId>` when the participant is
proven and on the RAW PROVIDER SPELLING when it is not. Mickey Gasper
(681508) is on Boston's 40-man and not on its active roster, so an
active-only index could not identify him and every book's spelling of him
survived dedupe as its own proposition — retention was the first layer to
notice, far too late, and could only discard the loser.

SPLIT. The mirror image, and invisible to the collision metric because it
makes MORE identities, not fewer: Leo Jiménez (677870) is published as both
"Leo Jiménez" and "Leonardo Jimenez", so one human became two semantic
players in one game. Measured across the 15 MLB cohort slices since the
canonical-participant repair, this is a recurring class, not one case:
cam/cameron smith (5 slices), mitch/mitchell bratt, zac/zachary thornton,
leo/leonardo jimenez.

THE REPAIR READS MLB'S OWN RECORD. `hydrate=person` on the roster call the
pipeline already makes returns firstName / useName / useLastName, so the
legitimate name forms for a human come from the league rather than from an
alias table. An alias table is a list of the mistakes we happened to notice.
`nickName` is DELIBERATELY EXCLUDED: over 821 people it produced 14
ambiguous keys, because MLB's nickname field carries bare surnames and
shared clubhouse names — `nameKey('Smitty Smith')` is one string for both
Burch Smith and Will Smith. The four forms kept produce ZERO ambiguity.

Canonical participant reach widens to the 40-man; TEAM EVIDENCE still reads
the ACTIVE roster alone, so event admission and the impossible-binding
refusal are unchanged. Identity still fails closed: a name matching more
than one person in the event resolves to nobody.

CONTINUITY, MEASURED BEFORE WRITING ANY CODE. Over the real 19:00 cohort,
208 of 209 player_keys are unchanged and the one that moves is the defect —
`leonardo jimenez` converging onto `leo jimenez`, a key that already exists.
No new lineage family. The natural key contains game_date, so chains never
span dates and a forward change cannot fork a closed one.

DETERMINISTIC REPRESENTATIVE. Which book's payload survives was decided by
position. It is now decided by the existing MODEL_BOOKS declaration order —
reused, not authored; inventing a sportsbook ranking to settle a tiebreak
would be a market judgement smuggled in as a bug fix — with book name and a
content tiebreak. Stable under every input permutation.

TWO GUARDS, BOTH DIRECTIONS. split (one person, many identities) and merge
(one identity, many people). A merge is refused at the same single admission
seam event identity already uses; a split is counted and alerted but does not
cut the board, because it duplicates an identity rather than asserting a
falsehood.

RETENTION REMAINS AN INDEPENDENT CHECK. The old assertion grepped the source
for `player_key: nameKey(player)`. That expression stood in for a PROPERTY,
and a grep verifies a spelling. Replaced with the property itself, asserted
in both modes: when the producer emits two rows for one human, retention
still files them under one identity and still reports the collision.

Replay of the real cohort through the repair: 3,129 offerings, 100%
participants resolved, every one of 207 participants on exactly ONE semantic
key, collision 0, split 0, merge 0.

Suite 396/5,455/0 · tsc 0 · 15/15 teeth. Tooth 12 came back green first
time and that was a coverage hole, not a safe defect: nothing asserted
retention's append-only upsert. It does now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 20:37:01 -04:00
builtbykev 7efb04e280 Lineage family lookup: bound it to a slate, and say what an action is
TWO DEFECTS, one lookup.

SCALE. The family lookup sent 100 natural keys as a PostgREST IN-list.
`read_natural_key` has NO pg_stats row at all -- the table's last autoanalyze
(2026-08-26) predates the column ever being populated -- so the planner used a
default per-value selectivity, estimated 172,409 rows and chose a sequential
scan of 344,818: 8.5s, then 57014. At 50 keys the same shape returned in
~357ms. The cliff is a statistics artifact, not a volume one, which is why the
repair does not depend on the estimate improving and is not CH=50.

`readNaturalKey` builds `sport|game_date|player_key|stat|side|line[|#event]`,
so SPORT AND GAME_DATE ARE COMPONENTS OF THE KEY. Two rows sharing a key
necessarily share both, and scoping the lookup to the (sport, game_date) pairs
present in the requested keys is LOSSLESS BY CONSTRUCTION. One index-backed
range per date, walked with safePaginate; cost is bounded by ONE SLATE however
long the chronology gets. Measured: 5,000 keys -> 1 scope, and the plan is
`Index Scan using model_snapshots_lineage_family_idx, cost 0.28..1.92`.

VALIDITY. A row carrying `read_natural_key` is not history: the key is stamped
on every candidate BEFORE the lookup, so a failure leaves it on a row that
never became an action. Proven this was not cosmetic -- fed the raw rows the
old lookup returned, the resolver produced a REVISION with a NULL read_id (an
orphaned chain node) and labelled a brand-new Read LEGACY_UNVERIFIED.
`isValidLineageAction` states what a completed action IS: all nine fields, in
the query and again in code.

ATOMICITY. A failed attempt now leaves NO lineage-specific state.
`publication_id`/`published_at` are untouched -- the slate really was
published, and erasing a true fact to tidy a false one is the wrong repair.

Replayed the exact failed 19:00Z cohort through the real resolver, side-effect
free: 119 NEW / 379 CHANGED / 621 UNCHANGED -> ORIGIN 119 / REVISION 379 /
RECAPTURE 621, 0 wrong parent, 0 wrong ordinal, 0 null read_id, 0 forks --
byte-identical with all 1,119 failed partial rows present. Clean-head parity
1,024/1,024.

Migration 050 is CONCURRENTLY + IF NOT EXISTS, drops nothing, rewrites nothing.
Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 19:56:56 -04:00
builtbykev 048e4eaa3f Event-aware retention identity: two games, two receipts
One player prop in Game 1 and the same-looking prop in Game 2 are two different
historical claims. The retention conflict identity did not know that.

SEMANTIC IDENTITY FIRST. Two outbound rows are the same retention proposition
within one cycle when they share the cycle, the EVENT, the participant, the
stat, the line and the side. Book is deliberately absent — collapsing books is
dedupeProps's actual job and the price anchor is chosen later. The database
index is enforcement of that answer, never the definition of it.

THE EVENT COMPONENT NEVER FABRICATES. canonical_event_id where a sport has a
resolver — MLB's admission gate rejects unresolved/ambiguous/contradicted props
BEFORE grading, so every row that can reach retention has one — and game_id
otherwise, which is NOT NULL in the schema and is the only event label sports
without a resolver possess. Both are in the identity, so the weaker label still
discriminates where the stronger is absent.

NULLS NOT DISTINCT IS LOAD-BEARING, NOT STYLISTIC. canonical_event_id is NULL
for every non-MLB row. Measured on a disposable PG17: under PostgreSQL's default
semantics the same NBA proposition inserted twice produced TWO rows — every
retry duplicating for ever. With NULLS NOT DISTINCT the same test yields one.
That measurement is what rejected the plain composite option.

MIXED-FLEET BRIDGE. A rollout serves both builds at once (measured 11/12 new,
1 old). Old and new writers need different indexes and NO schema state satisfies
both: with the legacy index present a new writer fails 23505 on a doubleheader;
with it gone an old writer fails 42P10. A bare ON CONFLICT DO NOTHING would have
bridged this, and PostgREST does not emit one — `ignoreDuplicates` WITHOUT
`onConflict` was measured raising a real duplicate-key error, so that bridge does
not exist through this client.

So the writer bridges it. It targets the event-aware identity and, on exactly
the two errors meaning "the schema is not in the state I expect" (42P10, or
23505 NAMING the legacy index), retries the SAME chunk on the legacy target. A
failed chunk rolls back atomically — measured 0 rows — so the retry cannot
double-write. Correct in every schema state: legacy-only and both-present
degrade to legacy semantics with no outage; new-only keeps both games.

The bridge is deliberately narrow. A supersedes conflict is ALSO a 23505, and
swallowing it would destroy the forked-history guard, so the legacy index must
be named. All three model_snapshots writers (persist, commitPublication,
recoverFromFork) go through it; no hardcoded legacy target survives.

MEASURED, through the real supabase-js -> PostgREST -> Postgres path on
production-shaped PG17:
  * 1,000 REAL propositions from the verified 2026-08-17 STL@CIN doubleheader
    (1,738 retained rows under ONE game_id), replayed across both real gamePks:
    OLD index materialized 1,000 of 2,000 — 1,000 LOST. NEW index materialized
    2,000 of 2,000 — 0 lost.
  * retry idempotency, over/under, line, stat, player, non-MLB same-game and
    non-MLB different-game all behave correctly under the new index.
  * ORDINARY-SLATE PARITY over ALL 434 real cohorts / 328,262 retained rows:
    old identities 328,262, new identities 328,262, delta 0, cohorts changed 0.
    The index is therefore guaranteed creatable and nothing historical splits.

CONFLICT_IDENTITY is now DERIVED from RETENTION_CONFLICT rather than restated —
a test caught them silently disagreeing, which is exactly how the materialization
check could have expected an identity the database no longer enforced.

EXPAND/CONTRACT are separate files on purpose. 048 is additive and retires
nothing; 049 drops the legacy index and must not be applied until fleet
convergence is proven by sampling, never assumed from a fast rollout.

NO BACKFILL. Legacy rows keep NULL canonical_event_id and remain LEGACY
EVENT-AGNOSTIC RETENTION, which is what that NULL truthfully says.

The materialization defence is untouched and now reports the bridge honestly:
while the legacy index still collapses a doubleheader, expected 4 vs actual 2
yields MATERIALIZATION_MISSING and the cohort is refused.

Nine teeth, injections verified present, against a green baseline of 97:
1 event distinction removed (10) · 2 phases collapsed (2) · 3 bridge swallows
everything (6) · 4 NULLS NOT DISTINCT removed (1) · 5 old-container error as
success (3) · 6 semantic/DB identity disagree (8) · 7 collision detector removed
(2) · 8 partial transport usable (3) · 9 collision unannounced (1).
Restored byte-identically.

Model and product untouched: gradeSlateService (event-aware dedupe), event
identity, ledger, calibration, chain, lineage config and the status route all
UNCHANGED. Zero cacheSet changes, zero web paths, schema contract unchanged (no
new columns). Lineage stays OFF.

385 suites / 5,178 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 20:30:45 -04:00
builtbykev 11277a1b99 Materialization truth: a completed write is not a complete cohort
Transport truth says every intended write succeeded. Materialization truth says
every identity that should exist actually exists. The retention writer could
only report the first, and the gap is not theoretical.

THE CONFLICT IDENTITY, traced to the real index:
  model_snapshots_cycle_prop_uniq UNIQUE (snapshot_id, player_key, stat, line, side)
written by `upsert(..., { onConflict: same five columns, ignoreDuplicates: true })`.
Measured in production: no key column is ever NULL (0 of 328,262 rows), so
NULLS DISTINCT never applies and the identity is plain column equality. `line`
is an unconstrained numeric, so identity normalises it — 0.5 in and "0.50" out
must not read as two identities for one stored row.

PRIOR-CYCLE COLLISION IS IMPOSSIBLE. `snapshot_id` is in the identity and is a
fresh UUID per cycle, so no row can be suppressed by an earlier cycle.
Append-only chronology across cycles is safe, and `captured_at` is not in the
identity, so a cohort cannot be split by timing.

INTRA-CYCLE COLLISION IS REAL, AND WE CAUSED IT. `canonical_event_id` is NOT in
the identity. A doubleheader — same hitter, same stat, same line, two genuinely
different games — is ONE identity. Demonstrated through the real collector: 4
outbound rows, 2 distinct identities, 2 rows discarded by ignoreDuplicates with
no error, `written` counting all 4 and the terminal status reading COMPLETE.
Before event-aware dedupe the second game was dropped before grading, so the
collision could not arise; that fix moved the loss downstream into retention.

The conflict identity is NOT changed here — that is a separate decision with its
own before/after. This makes the loss visible instead of silent.

EXPECTED vs ACTUAL. `expectedMaterialization(rows)` derives the identity set
from the FINAL outbound payload using the exact database identity — never from
`attempted`, which counts rows sent, not identities that can exist.
`reconcileMaterialization` compares SETS, not counts: two sets of equal size can
still differ, and a cohort that swapped one identity for another passes every
count test ever written. A collision passes set equality by construction (the
discarded row was never in the expected set) while real rows were lost, so
collision_count > 0 fails the cohort on its own.

A cohort is evidence-complete only when transport is COMPLETE, missing = 0,
extra = 0, and collisions = 0.

OBSERVABILITY stayed minimal. `last_retention` was already PER SPORT (a Map
keyed by sport), so no fix was needed there and the route is UNCHANGED — the new
fields ride the existing entry: outbound_rows, expected_materialized_count,
outbound_collision_count, expected_identity_digest. Counts and a digest only,
never the identities, which carry player names. The expected set is the one
materialization fact unrecoverable from the database afterwards, which is why it
is the only thing recorded at runtime.

A collision leaves transport COMPLETE, so the existing failure alert could never
see it. It now has its own high-severity alert naming the counts, the cycle and
the build, and says the cohort is not evidence-complete.

Seven teeth, each injection verified present, against a GREEN baseline of 63:
  1 attempted===written as evidence completeness   -> 1 fail
  2 COUNT(*) equality instead of set equality       -> 1 fail
  3 snapshot_id dropped from expected identity      -> 4 fail
  4 unexpected collision allowed to qualify         -> 1 fail
  5 single global last_retention slot               -> 2 fail
  6 partial chunk failure treated as usable         -> 3 fail
  7 collision loses its announcement                -> 1 fail
Restored byte-identically (retention 742f116473d97f49, snapshot 81129facbabeb280).

Three brittle assertions repaired, with the reason recorded: two windowed on a
byte count that a neighbouring block outgrew — a test failing because of its
neighbour, not its subject — now windowed to syntactic landmarks; and one
counted TERMINAL.COMPLETE occurrences, which a legitimate comparison
incremented. It now asserts one DECISION and one READ.

persist() and createCollector are BYTE-IDENTICAL. onConflict and
ignoreDuplicates appear in the diff only as prose. Model, event, ledger,
calibration, chain, lineage config, and the status route: UNCHANGED. Zero
lineage/publication files, zero cacheSet changes, zero web paths. Lineage OFF.

Schema contract unchanged: release 64, prod 67, prod-only 3 (debt, not
authorized), missing in prod 0.

384 suites / 5,144 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 19:44:56 -04:00
builtbykev 9809626c99 Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.

So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.

TERMINAL RETENTION STATES (retentionService.classifyPersist):
  NOTHING_TO_PERSIST       attempted 0 — a refusal-only slate is still a cycle
  SKIPPED_NO_DATABASE      no database configured; not a failure
  COMPLETE                 attempted > 0, written === attempted, no error
  FAILED_ZERO_WRITE        written === 0 — first chunk failed
  FAILED_PARTIAL           0 < written < attempted — a later chunk failed
  FAILED_UNRESOLVED_ERROR  counts look complete but an error is unresolved;
                           unreachable through today's loop, and kept because
                           the alternative is reporting COMPLETE holding an error

The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.

classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.

`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.

VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.

OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.

RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.

Six teeth, each with the injection verified present, against a green baseline:
  1 written>0 as generic success        -> 6 fail
  2 later-chunk failure reports COMPLETE -> 5 fail
  3 FAILED_PARTIAL does not alert        -> 3 fail
  4 status reports a recalculated count  -> 1 fail
  5 row presence treated as completion   -> 1 fail
  6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).

Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.

Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.

383 suites / 5,118 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 19:12:59 -04:00
builtbykev 35da190f2c Retention hotfix: drop published_side, derive the schema contract, break the silence
`createCollector.onPublished` set `published_side` beside `published`.
`published_side` is not a model_snapshots column. supabase-js declares the
UNION of row keys in the `columns=` parameter, so one invalid key made
PostgREST reject the ENTIRE batch with a 400 — every sport, every cycle.
Retention is best-effort, so nothing surfaced. Confirmed in edge logs.

The field was redundant as well as invalid: `side` is already on the row.
Deleted rather than added to the schema — a column would preserve an
accidental artifact.

Three things missed it, and each is now closed:

1. WRONG SHAPE INSPECTED. The manual check sampled the collector after
   onGraded only and never called onPublished, so the offending key was
   not yet on the row. It read a pre-publication shape and reported the
   final outbound shape as clean. The new test captures the array actually
   handed to .upsert(), after the full production call order.

2. NO CONTRACT. Every retention test injects a permissive fake client that
   accepts any column set, so 381 suites proved the logic and never once
   compared a row against the database. The contract is now DERIVED — the
   migration chain applied to a disposable postgres, read out of
   information_schema (scripts/generate-schema-contract.js). A
   hand-maintained list would be a second opinion about the schema, and a
   second opinion is what let this through. scripts/verify-schema-contract.js
   checks the contract still describes a live database.

3. SILENT FAILURE. A failed batch reached one console.log. It now emits a
   high-severity structured event carrying sport, snapshot id, stage,
   error, code_sha and timestamp. Best-effort semantics are unchanged —
   the product continues and says so — but the failure is observable.
   `skipped` (no database configured) is not a failure and does not alert.

Teeth, each with the injection verified present before the run:
  - published_side back into the final payload -> 4 tests fail; restored
    byte-identically (sha 6a0ced7c52134135 both sides)
  - settled_at (a REAL contract column) -> accepted, so the guard
    discriminates by contract membership, not by novelty
  - alert block deleted -> 3 tests fail; restored byte-identically

Model and product behaviour untouched: analyzeViaEngine1,
probabilityEstimator, gradeSlateService, lineageCanaryConfig all unchanged.
Lineage stays OFF. Net source change is one behavioural line plus the alert.

382 suites / 5,094 tests pass. web tsc exit 0 (zero web paths touched).

Measurement blackout recorded, NOT backfilled: last good retention write
2026-08-27T19:08:32Z; ceaa896 started 21:16:41Z; the 22:00 UTC cycle ran
(ledger wrote 22:05:02) and persisted zero snapshot rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 18:53:37 -04:00
builtbykev 352016790a MLB canonical event identity, impossible-binding refusal, event-aware dedupe, publication commit
Release-isolated slice built from 41ba38e. Ships ONLY the event-integrity +
publication + lineage-canary closure; the 90-path development tree stays
undeployed.

- canonical MLB event identity from statsapi gamePk (mlb:gamepk:<pk>), with
  event_identity_source/method/version recorded. The id is canonical; the
  binding is derived and says so.
- IMPOSSIBLE-BINDING REFUSAL. Verified in prod 2026-08-26: Joe Mack (Marlins)
  was bound to Dodgers@Braves and Yandy Diaz (Rays) to Rangers@WhiteSox, both
  from one book in the 01:00/03:01 UTC cycles after their own games began. Root
  cause is source market data, not the binder. A prop whose player's team is not
  an event participant now refuses; unknown team preserves uncertainty.
- event-aware dedupe: books still collapse, events no longer do. An unresolved
  MLB event fractures rather than falling back to the collision-prone
  date+teams key.
- publication commit moved AFTER the authoritative Redis slate write, with
  exact parity-gap identity when the product publishes and the record does not.
- lineage dual-write behind LINEAGE_CANARY_SPORTS, DISABLED for this deploy.

Excluded deliberately: WNBA feed/chain, market ontology, PerformanceDistribution,
calibration certification, truth diagnostics, applyRevision Phase-1, and the
analyzeViaEngine1 confidence-rounding change (a served field).

Suite 380/5,040/0 from this worktree; web tsc exit 0; champion output identical
to production.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 16:29:38 -04:00
builtbykev 6c34af3414 checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion
recon (Tony resurrection, STEP 0). This work existed only on the
laptop disk.

- chain shadow accrual + probe script (038_chain_shadow.sql)
- WNBA possession feed: ESPN adapter, usage service, verify script
  (039_wnba_player_game.sql)
- baseball chain
- retention/snapshot service updates, tableKeys, matchupKeys
- specs: chain-v1, wnba-possession-feed, wnba-source-survey
- unit tests for the above

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
2026-08-14 16:53:37 -04:00
builtbykev f61ec6b391 Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 22:49:56 -04:00
builtbykev 494c83cf76 Hunt the window-bug class: three more paths, and the forward re-audit rule
in code

PHASE 0 — getStatRows is the single base-rate path, so every branch is
audited, plus the feature builders since l20_avg is the season reference
projectionFor reads:

  getStatRows MLB -> estimator base    fullLog            CORRECT (929fd81)
  mlbGameLogFeatures l5/l10/l20        last10 = 10        DEFECTIVE
  espnStatsAdapter.parseGameLog        slice(0,20)        DEFECTIVE
  getStatRows NBA/WNBA ESPN branch     inherits 20-cap    DEFECTIVE via source
  getStatRows NBA/WNBA python branch   getGameLogs(...,20) dormant (offline)
  pitcherEngine / skillProjection      statcast profiles  N/A
  pitcher props via getStatRows MLB    fullLog            CORRECT
  settleSource                         full log (S64)     CORRECT

THE PITCHER ANSWER IS GOOD NEWS: pitcher props run through the same
getStatRows MLB branch, so 929fd81 repaired them too. There is no separate
defective pitcher base-rate path.

THE ONE HIDING IN PLAIN SIGHT: mlbGameLogFeatures carries the comment
"l20 = all available (the season per-game reference projectionFor needs)"
while building from last10 -- so l20_avg was a TEN-GAME AVERAGE WEARING A
SEASON LABEL, feeding both the consistency pull inside the estimator and
projectionFor, which decides refusals. It survived the previous repair
because that fix touched only getStatRows.

PHASE 1 — mlbGameLogFeatures now reads fullLog; espnStatsAdapter drops its
slice(0,20) cap. ZERO new API calls on both: each widens data already
fetched and then discarded, the same shape as the original repair. The
python branch is left alone -- the service is offline in prod and fixing it
would be speculative.

Their before/after resolution is NOT measured, deliberately: the only way
to measure today is to reconstruct the repaired forecast over old rows,
which is the reconstruction-vs-served trap this order refuses. Code fix
now, measurement at accrual.

PHASE 2 — MODEL_VERSION bumped to engine1@2026-08-07-fullwindow, so every
forward snapshot is self-identifying (retentionService already stamps it;
no new plumbing). model/reAuditEligibility.js encodes the rule: isEligible
accepts only the repaired marker, assess counts eligible DATES not rows,
and ACCRUAL is frozen at calibration 10 / hits-lift 10 / verdict-reaudit
14 / rbi-gate 14. A test locks the invisible case -- a MIXED table of 330
rows with 30 repaired returns eligible_dates 3, not 330 rows of false
confidence. Once both generations share a table a naive count would fit a
map on a blend of two forecasters.

PHASE 3 — the board, each consequence labelled: calibration WITHDRAWN
(refits at 10 dates, never on reconstructions); factor verdicts SUSPECT
(all measured against a champion worse than a frequency table, direction
UNKNOWN, not pre-priced, 14 dates); hits factor lift UN-REMEASURABLE (10
dates, factors still wired and transmitting); rbi lineup-slot RE-QUEUED
(14 dates). Pre-registered order: calibration, hits lift, verdict
re-audit, rbi gate.

Then STOP and accrue. Nothing further can be honestly measured until the
board fills with rows the repaired champion produced.

Serving-path changes by design for the MLB feature path and NBA/WNBA logs;
eleven frozen model modules verified unchanged. p_win never mutated. No
Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:40:20 -04:00
builtbykev 1bcdd8b305 Fix the ROOT: bind props to their real game, never date by the grade clock
Order 1.5 Phase 1. PropLine emits NO commence_time (grep-verified: zero
hits in proplineAdapter), so ledgerService's
`dateET(prop.game_time) || dateET(gradedTs)` always fell through to the
GRADE timestamp — and a 01:00/03:00 UTC snapshot is 21:00/23:00 ET the
PREVIOUS day. Tonight's props were filed under yesterday, settlement
correctly found no game there, and Order 1's void logic turned that into
64 destroyed results.

gameBinder.attachGameTimes() now matches every prop to a scheduled game by
TEAMS across the plausible ET window (grade date, +1, -1) and attaches the
GAME'S OWN time/date/id. It runs in snapshotService before grading and
before the ledger write, so ledger, retention and settlement all inherit
the correct date from one place.

HARD CONTRACT: an unbindable prop returns NOTHING. ledgerService no longer
has a grade-clock fallback — a row with no real game time is SKIPPED and
counted, because a mis-dated row is fabricated data and the ledger holds
real values or nothing. A slate that binds nothing pages.

DOUBLEHEADERS are reported, never guessed: two games with the same teams
on one date mark the binding `ambiguous` so settlement can decline rather
than attribute a prop to the wrong game. (Real example already in the
data: mlb:2026-07-11:MilwaukeeBrewers@PittsburghPirates(Game1).)

Also fixes retention, which had the SAME bug from last night — I had
dated model_snapshots rows with the snapshot clock. Rows now take the ET
date of the bound game_time.

Suite 282/3381 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 03:25:01 -04:00
builtbykev 5a5e37e32e Retention: fill enrichment fields + page on a zero-write slot
PHASE 1 — cron capture needed NO wiring. Verified in code: the scheduler
tick calls runAll = snapshotService.runAllSnapshots, which loops
runSnapshot per sport, which already carries the onGraded -> retention
hook. The scheduled path and the manual path are the SAME function. The
reason no cron cycle had been captured is simply that no slot has fired
since retention deployed (slots are 14/19/22/1/3 UTC; retention landed
~02:55). Induced proof follows the deploy.

PHASE 2 — archetype/team/opponent were permanently null because retention
persisted at GRADE time, before enrichment attaches them. Retention still
COLLECTS at grade time (the only moment the feature vector exists) but now
PERSISTS after enrichment, merging those three fields via
retentionService.mergeEnrichment. The merge is pure and fills ONLY those
three fields — features and every model output are grade-time values and
must never be rewritten by enrichment; a test asserts that. Unmatched rows
(refusals not in the enriched slate) keep nulls rather than guesses. The
empty-slate early return now persists too: a refusal-only slate is still
history worth keeping.

PHASE 3 — ZERO-WRITE ALARM. opsWatch.retentionZeroWriteAlarm pages at
missed-snapshot severity when a slot GRADED props but retention wrote
fewer rows than the slate (or nothing). runSnapshot now returns
retentionRows so the scheduler can evaluate it. Retention is best-effort
by design so it can never break a snapshot — which means a broken write is
silent by construction. This is the counterweight. A slot that graded
nothing never false-pages; an absent count reads as NOTHING and still
pages, distinct from a reported 0.

Suite 280/3349 green, build exit 0. Outcome stamping deliberately NOT
implemented (depends on the settlement fix).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 02:00:40 -04:00
builtbykev d3ffa1b8c2 Retention: model_snapshots live + base64 SSH key support
RETENTION (Phase 2, priority zero). History starts compounding tonight.

migration 025 model_snapshots — APPLIED to prod. Append-only, one row per
graded prop PER SIDE PER CYCLE, with a unique index on
(snapshot_id, player_key, stat, line, side) so a retried cycle cannot
duplicate. RLS on, service-role writes only.

What it captures that the ledger never did:
- features jsonb — the model's INPUTS. Without these a backtest can only
  grade our own homework; with them any future model can be replayed
  against the exact conditions this one faced.
- REFUSALS (refused + refusal_reason). The ledger drops them, so a gate
  refusing props that would have WON is invisible — unmeasurable lost
  edge. Captured via a new onGraded hook in gradeSlateService that fires
  with BOTH sides before any filtering.
- grade_11, the pre-collapse grade. The 4-letter map throws away the
  entire live C-/C/C+/B- range.
- model_version + code_sha on every row. ledger_entries mixes pre/post-fix
  grades with no marker and cannot be separated retroactively.
- p_win / ev_pct / fair_odds / takeable / value — none of which any
  permanent store held.

Wiring: analyzeViaEngine1 attaches _features/_grade_11 (underscore =
internal); gradeSlateService fires onGraded then STRIPS them so they never
reach a cache or API payload; snapshotService builds rows and persists
best-effort. Retention reuses the LEDGER's dateET/gameIdFor helpers so
rows share the ledger's natural key exactly — otherwise the settle pass
could never join outcomes onto them. Rows are written BEFORE the empty-
slate early return: a slate that refused everything is exactly the case
worth recording.

CONTRACT HELD: retention is injectable and every path is caught. persist()
returns errors, never throws; a missing Supabase client is SKIPPED, not an
error. A retention failure can never break a snapshot.

BACKUP: backup-db.sh now accepts BACKUP_SSH_KEY as base64 (recommended —
survives env-var newline mangling, which is how injected SSH keys usually
break silently) OR raw PEM, detected by decoding and looking for the PEM
header. Verified both forms detect correctly against a real generated key.

Suite 279/3325 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 23:01:12 -04:00