23 Commits

Author SHA1 Message Date
builtbykev a8de676756 A probability is served because evidence supports it, not because nothing else answered
The band gate was blocked for its `else` branch. It read:

    candidate = F(raw)
    served    = inCertifiedBand(candidate) ? candidate : RAW

and above raw 0.60 the model is measured overconfident — holdout raw 0.80-0.90
predicts 0.843 and realizes 0.639. So "the calibrator is not supported here" was
being answered with a number already proven wrong. Unsupported calibration does
not make raw true.

Four candidates were adjudicated on ONE split — fit on the earliest 60% of
train, decide support on the last 40%, evaluate on a holdout that saw neither:

  A low-param      80.2% coverage  0.24374  REFUTED — its extra region
                   (raw 0.80-0.90) certified on cert (err +0.040, n=55) and
                   refuted on holdout (served 0.754 vs observed 0.639), and it
                   leaves a hole at 0.70-0.80 while serving the island above it
  B isotonic       91.3% coverage  0.24337  CERTIFIED, contiguous raw [0.50,0.80)
  C empirical band 91.3% coverage  0.24335  REFUTED — refitted point-in-time on
                   current-model hits the realized rates INVERT in grade order
                   (B+ 0.593 < B 0.614 < C+ 0.623), so the served function steps
                   down at raw 0.78. Its shipped constants come from 3,417 props
                   pooled across four batter stats and do not reproduce here
  D raw identity   43.1% coverage  0.24866  certifies raw 0.50-0.60 and only there

Raw is candidate D, not a fallback. It earns exactly one region (holdout error
+0.010 on n=1,316), which is why the law is "raw must earn its region" rather
than "raw is never true". B already covers that region, so no hybrid is built.

Above raw 0.80 nothing is certified and nothing is served. That is the region
where raw is most wrong, isotonic over-corrects (cert err -0.093) and its LODO
mapping at 0.95 has spread 0.180. 8.7% of holdout rows land there.

The registry did not need changing. `serves(stat, p)` already tested certified
bands against the RAW p_win — support in the input domain, the correct question —
and returned {serve:false, reason}. It never said "serve raw". The output-space
gate and the raw fallback were both invented downstream in calibrationService.

ACTIVATION IS OFF. PROBABILITY_CONTRACT_SHADOW defaults to 0, CALIBRATION_DEPLOYED
stays frozen empty, and every served field is byte-identical. This releases the
support first, which is the required order. The shadow records raw belief, the
candidate served value, the state, the estimator identity, and what EV/Kelly/VALUE
would be under the actionability law — into its own column, read by nothing.

Migration 051 was applied to production BEFORE retentionService named the column.
PostgREST builds a bulk insert from the first row's shape, so a key whose column
does not exist 400s the whole batch silently — that is how migration 038 took
retention down for three days.

The user-facing contradiction is NOT fixed here. A B+ still says "realized about
66%" beside a confidence of 84. Fixing that is activation, and activation costs
32% of VALUE flags and 46% of Kelly recommendations on the holdout.

Suite 401/401, 5,580 passed, 4 skipped, deterministic across three runs.
Teeth 23/23, each independently injected and restored byte-identically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:03:16 -04:00
builtbykev 7efb04e280 Lineage family lookup: bound it to a slate, and say what an action is
TWO DEFECTS, one lookup.

SCALE. The family lookup sent 100 natural keys as a PostgREST IN-list.
`read_natural_key` has NO pg_stats row at all -- the table's last autoanalyze
(2026-08-26) predates the column ever being populated -- so the planner used a
default per-value selectivity, estimated 172,409 rows and chose a sequential
scan of 344,818: 8.5s, then 57014. At 50 keys the same shape returned in
~357ms. The cliff is a statistics artifact, not a volume one, which is why the
repair does not depend on the estimate improving and is not CH=50.

`readNaturalKey` builds `sport|game_date|player_key|stat|side|line[|#event]`,
so SPORT AND GAME_DATE ARE COMPONENTS OF THE KEY. Two rows sharing a key
necessarily share both, and scoping the lookup to the (sport, game_date) pairs
present in the requested keys is LOSSLESS BY CONSTRUCTION. One index-backed
range per date, walked with safePaginate; cost is bounded by ONE SLATE however
long the chronology gets. Measured: 5,000 keys -> 1 scope, and the plan is
`Index Scan using model_snapshots_lineage_family_idx, cost 0.28..1.92`.

VALIDITY. A row carrying `read_natural_key` is not history: the key is stamped
on every candidate BEFORE the lookup, so a failure leaves it on a row that
never became an action. Proven this was not cosmetic -- fed the raw rows the
old lookup returned, the resolver produced a REVISION with a NULL read_id (an
orphaned chain node) and labelled a brand-new Read LEGACY_UNVERIFIED.
`isValidLineageAction` states what a completed action IS: all nine fields, in
the query and again in code.

ATOMICITY. A failed attempt now leaves NO lineage-specific state.
`publication_id`/`published_at` are untouched -- the slate really was
published, and erasing a true fact to tidy a false one is the wrong repair.

Replayed the exact failed 19:00Z cohort through the real resolver, side-effect
free: 119 NEW / 379 CHANGED / 621 UNCHANGED -> ORIGIN 119 / REVISION 379 /
RECAPTURE 621, 0 wrong parent, 0 wrong ordinal, 0 null read_id, 0 forks --
byte-identical with all 1,119 failed partial rows present. Clean-head parity
1,024/1,024.

Migration 050 is CONCURRENTLY + IF NOT EXISTS, drops nothing, rewrites nothing.
Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 19:56:56 -04:00
builtbykev 048e4eaa3f Event-aware retention identity: two games, two receipts
One player prop in Game 1 and the same-looking prop in Game 2 are two different
historical claims. The retention conflict identity did not know that.

SEMANTIC IDENTITY FIRST. Two outbound rows are the same retention proposition
within one cycle when they share the cycle, the EVENT, the participant, the
stat, the line and the side. Book is deliberately absent — collapsing books is
dedupeProps's actual job and the price anchor is chosen later. The database
index is enforcement of that answer, never the definition of it.

THE EVENT COMPONENT NEVER FABRICATES. canonical_event_id where a sport has a
resolver — MLB's admission gate rejects unresolved/ambiguous/contradicted props
BEFORE grading, so every row that can reach retention has one — and game_id
otherwise, which is NOT NULL in the schema and is the only event label sports
without a resolver possess. Both are in the identity, so the weaker label still
discriminates where the stronger is absent.

NULLS NOT DISTINCT IS LOAD-BEARING, NOT STYLISTIC. canonical_event_id is NULL
for every non-MLB row. Measured on a disposable PG17: under PostgreSQL's default
semantics the same NBA proposition inserted twice produced TWO rows — every
retry duplicating for ever. With NULLS NOT DISTINCT the same test yields one.
That measurement is what rejected the plain composite option.

MIXED-FLEET BRIDGE. A rollout serves both builds at once (measured 11/12 new,
1 old). Old and new writers need different indexes and NO schema state satisfies
both: with the legacy index present a new writer fails 23505 on a doubleheader;
with it gone an old writer fails 42P10. A bare ON CONFLICT DO NOTHING would have
bridged this, and PostgREST does not emit one — `ignoreDuplicates` WITHOUT
`onConflict` was measured raising a real duplicate-key error, so that bridge does
not exist through this client.

So the writer bridges it. It targets the event-aware identity and, on exactly
the two errors meaning "the schema is not in the state I expect" (42P10, or
23505 NAMING the legacy index), retries the SAME chunk on the legacy target. A
failed chunk rolls back atomically — measured 0 rows — so the retry cannot
double-write. Correct in every schema state: legacy-only and both-present
degrade to legacy semantics with no outage; new-only keeps both games.

The bridge is deliberately narrow. A supersedes conflict is ALSO a 23505, and
swallowing it would destroy the forked-history guard, so the legacy index must
be named. All three model_snapshots writers (persist, commitPublication,
recoverFromFork) go through it; no hardcoded legacy target survives.

MEASURED, through the real supabase-js -> PostgREST -> Postgres path on
production-shaped PG17:
  * 1,000 REAL propositions from the verified 2026-08-17 STL@CIN doubleheader
    (1,738 retained rows under ONE game_id), replayed across both real gamePks:
    OLD index materialized 1,000 of 2,000 — 1,000 LOST. NEW index materialized
    2,000 of 2,000 — 0 lost.
  * retry idempotency, over/under, line, stat, player, non-MLB same-game and
    non-MLB different-game all behave correctly under the new index.
  * ORDINARY-SLATE PARITY over ALL 434 real cohorts / 328,262 retained rows:
    old identities 328,262, new identities 328,262, delta 0, cohorts changed 0.
    The index is therefore guaranteed creatable and nothing historical splits.

CONFLICT_IDENTITY is now DERIVED from RETENTION_CONFLICT rather than restated —
a test caught them silently disagreeing, which is exactly how the materialization
check could have expected an identity the database no longer enforced.

EXPAND/CONTRACT are separate files on purpose. 048 is additive and retires
nothing; 049 drops the legacy index and must not be applied until fleet
convergence is proven by sampling, never assumed from a fast rollout.

NO BACKFILL. Legacy rows keep NULL canonical_event_id and remain LEGACY
EVENT-AGNOSTIC RETENTION, which is what that NULL truthfully says.

The materialization defence is untouched and now reports the bridge honestly:
while the legacy index still collapses a doubleheader, expected 4 vs actual 2
yields MATERIALIZATION_MISSING and the cohort is refused.

Nine teeth, injections verified present, against a green baseline of 97:
1 event distinction removed (10) · 2 phases collapsed (2) · 3 bridge swallows
everything (6) · 4 NULLS NOT DISTINCT removed (1) · 5 old-container error as
success (3) · 6 semantic/DB identity disagree (8) · 7 collision detector removed
(2) · 8 partial transport usable (3) · 9 collision unannounced (1).
Restored byte-identically.

Model and product untouched: gradeSlateService (event-aware dedupe), event
identity, ledger, calibration, chain, lineage config and the status route all
UNCHANGED. Zero cacheSet changes, zero web paths, schema contract unchanged (no
new columns). Lineage stays OFF.

385 suites / 5,178 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 20:30:45 -04:00
builtbykev 9809626c99 Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.

So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.

TERMINAL RETENTION STATES (retentionService.classifyPersist):
  NOTHING_TO_PERSIST       attempted 0 — a refusal-only slate is still a cycle
  SKIPPED_NO_DATABASE      no database configured; not a failure
  COMPLETE                 attempted > 0, written === attempted, no error
  FAILED_ZERO_WRITE        written === 0 — first chunk failed
  FAILED_PARTIAL           0 < written < attempted — a later chunk failed
  FAILED_UNRESOLVED_ERROR  counts look complete but an error is unresolved;
                           unreachable through today's loop, and kept because
                           the alternative is reporting COMPLETE holding an error

The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.

classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.

`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.

VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.

OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.

RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.

Six teeth, each with the injection verified present, against a green baseline:
  1 written>0 as generic success        -> 6 fail
  2 later-chunk failure reports COMPLETE -> 5 fail
  3 FAILED_PARTIAL does not alert        -> 3 fail
  4 status reports a recalculated count  -> 1 fail
  5 row presence treated as completion   -> 1 fail
  6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).

Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.

Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.

383 suites / 5,118 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 19:12:59 -04:00
builtbykev 35da190f2c Retention hotfix: drop published_side, derive the schema contract, break the silence
`createCollector.onPublished` set `published_side` beside `published`.
`published_side` is not a model_snapshots column. supabase-js declares the
UNION of row keys in the `columns=` parameter, so one invalid key made
PostgREST reject the ENTIRE batch with a 400 — every sport, every cycle.
Retention is best-effort, so nothing surfaced. Confirmed in edge logs.

The field was redundant as well as invalid: `side` is already on the row.
Deleted rather than added to the schema — a column would preserve an
accidental artifact.

Three things missed it, and each is now closed:

1. WRONG SHAPE INSPECTED. The manual check sampled the collector after
   onGraded only and never called onPublished, so the offending key was
   not yet on the row. It read a pre-publication shape and reported the
   final outbound shape as clean. The new test captures the array actually
   handed to .upsert(), after the full production call order.

2. NO CONTRACT. Every retention test injects a permissive fake client that
   accepts any column set, so 381 suites proved the logic and never once
   compared a row against the database. The contract is now DERIVED — the
   migration chain applied to a disposable postgres, read out of
   information_schema (scripts/generate-schema-contract.js). A
   hand-maintained list would be a second opinion about the schema, and a
   second opinion is what let this through. scripts/verify-schema-contract.js
   checks the contract still describes a live database.

3. SILENT FAILURE. A failed batch reached one console.log. It now emits a
   high-severity structured event carrying sport, snapshot id, stage,
   error, code_sha and timestamp. Best-effort semantics are unchanged —
   the product continues and says so — but the failure is observable.
   `skipped` (no database configured) is not a failure and does not alert.

Teeth, each with the injection verified present before the run:
  - published_side back into the final payload -> 4 tests fail; restored
    byte-identically (sha 6a0ced7c52134135 both sides)
  - settled_at (a REAL contract column) -> accepted, so the guard
    discriminates by contract membership, not by novelty
  - alert block deleted -> 3 tests fail; restored byte-identically

Model and product behaviour untouched: analyzeViaEngine1,
probabilityEstimator, gradeSlateService, lineageCanaryConfig all unchanged.
Lineage stays OFF. Net source change is one behavioural line plus the alert.

382 suites / 5,094 tests pass. web tsc exit 0 (zero web paths touched).

Measurement blackout recorded, NOT backfilled: last good retention write
2026-08-27T19:08:32Z; ceaa896 started 21:16:41Z; the 22:00 UTC cycle ran
(ledger wrote 22:05:02) and persisted zero snapshot rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 18:53:37 -04:00
builtbykev 352016790a MLB canonical event identity, impossible-binding refusal, event-aware dedupe, publication commit
Release-isolated slice built from 41ba38e. Ships ONLY the event-integrity +
publication + lineage-canary closure; the 90-path development tree stays
undeployed.

- canonical MLB event identity from statsapi gamePk (mlb:gamepk:<pk>), with
  event_identity_source/method/version recorded. The id is canonical; the
  binding is derived and says so.
- IMPOSSIBLE-BINDING REFUSAL. Verified in prod 2026-08-26: Joe Mack (Marlins)
  was bound to Dodgers@Braves and Yandy Diaz (Rays) to Rangers@WhiteSox, both
  from one book in the 01:00/03:01 UTC cycles after their own games began. Root
  cause is source market data, not the binder. A prop whose player's team is not
  an event participant now refuses; unknown team preserves uncertainty.
- event-aware dedupe: books still collapse, events no longer do. An unresolved
  MLB event fractures rather than falling back to the collision-prone
  date+teams key.
- publication commit moved AFTER the authoritative Redis slate write, with
  exact parity-gap identity when the product publishes and the record does not.
- lineage dual-write behind LINEAGE_CANARY_SPORTS, DISABLED for this deploy.

Excluded deliberately: WNBA feed/chain, market ontology, PerformanceDistribution,
calibration certification, truth diagnostics, applyRevision Phase-1, and the
analyzeViaEngine1 confidence-rounding change (a served field).

Suite 380/5,040/0 from this worktree; web tsc exit 0; champion output identical
to production.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 16:29:38 -04:00
builtbykev 6c34af3414 checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion
recon (Tony resurrection, STEP 0). This work existed only on the
laptop disk.

- chain shadow accrual + probe script (038_chain_shadow.sql)
- WNBA possession feed: ESPN adapter, usage service, verify script
  (039_wnba_player_game.sql)
- baseball chain
- retention/snapshot service updates, tableKeys, matchupKeys
- specs: chain-v1, wnba-possession-feed, wnba-source-survey
- unit tests for the above

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
2026-08-14 16:53:37 -04:00
builtbykev f61ec6b391 Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 22:49:56 -04:00
builtbykev 7c6fd95e68 Build 2 Phase A: real founder cap — atomic claim, race PROVEN, flags collapsed
DB only. No Stripe call, no checkout/webhook rewire (Phase B). Migrations 035,
036, 037 applied to prod and tracked; repo files added.

035 SCHEMA TRUTH — user_profiles gains stripe_customer_id and
stripe_subscription_id (G3 proved the webhook stores neither today, yet
finalize and grandfather reconciliation both key off the subscription id), plus
a partial unique index so a subscription id resolves to exactly one profile.

036 THE MECHANISM — founder_slots is a real TABLE replacing the decorative view.
The claim is a single UPDATE whose target row is chosen FOR UPDATE SKIP LOCKED;
no count is read in the decision path. UNIQUE(slot_number) plus a PARTIAL
UNIQUE(user_id) WHERE status <> 'free' (one live slot per user). Seeded 100 free.
Q1 global pool: the slot travels with the user, so analyst->desk keeps founder
with no second claim. Q2: release_expired_slots handles TTL abandonment ONLY —
cancelled slots retire, so the counter only rises. A6 redirects
founder_pricing_seats to count claimed slots, capped 100.

PRICE IDS ARE NOT IN SQL. claim_founder_slot returns a price KEY
(analyst_founder / analyst_standing / desk_founder / desk_standing) and the Node
layer maps it to STRIPE_PRICE_* env with a boot assertion — adopted over
hardcoding so a typo fails at boot instead of becoming a permanent mis-charge.

A7 FLAG COLLAPSE — finalize_founder_slot is now the SINGLE writer of both
founder flags in ONE transaction: user_profiles.founder_pricing is canonical and
users.founder_status mirrors it. founder_status is NOT dropped (G5 proved it
live: written at stripeService:163, served at routes/stripe:95, loaded in
middleware/auth:24 PROFILE_COLUMNS). Only the independent write is retired —
the two flags had already drifted in prod (1 vs 0).

A9 RACE TEST, run in Supabase before any Stripe:
  - pool squeezed to ONE free slot; three distinct users claimed concurrently
    -> EXACTLY ONE is_founder=true on slot 100, two returned analyst_standing,
    zero double-allocation.
  - idempotency: the winner claiming again returned the SAME slot 100 and still
    held exactly 1 live slot (two tabs cannot take two seats).
  - constraint layer proven directly: a raw UPDATE granting that user a SECOND
    live slot was REJECTED by the partial unique index, and verify-after-write
    confirmed state unchanged (1 live slot, target row untouched).
  HONEST LIMIT: the three claims contend within one transaction via LATERAL, so
  this proves the claim logic, the SKIP LOCKED path and the constraint that makes
  parallel safe — but it is not N genuinely parallel backend sessions. True
  multi-session concurrency is not drivable through this SQL interface and should
  be exercised once in Phase B against the test key.

037 NEXAPAY DROP — own migration, evidence-led (G4: zero code refs, column
empty). VYNDR is Stripe-only.

CLEAN BASELINE (Q3) verified after the test: 100 free slots, 0 non-free, counter
0/100, and BOTH founder flags cleared to 0 across user_profiles and users — the
inconsistent test record is no longer enshrined as a founder.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 17:19:39 -04:00
builtbykev 2bfaeff572 Ledger takeable tagging (deferred C2); efficiency challenger BLOCKED
Champion grade UNCHANGED. Push scoring untouched. Additive tags only — nothing
deleted, nothing re-settled.

PART A — THE EFFICIENCY CHALLENGER: BLOCKED, NOT BUILT.
Review Zero came back ABSENT on all three inputs:
  0.1 efficiency scores DO NOT EXIST (zero occurrences of market_efficiency /
      marketEfficiency / efficiency_score in src/ or web/src/).
  0.2 base thresholds DO NOT EXIST (engine1.js has zero `edge` references — the
      grade is not an edge-vs-threshold comparison; grade_thresholds.json holds
      PROBABILITY bands).
  0.3 the +/-0.05 additive efficiency nudge DOES NOT EXIST. The only 0.05s on
      the grade path are featureCache.teammate_absence_bump, a bvp_advantage
      cutoff, and p*0.9+0.05 inside probabilityEstimator (the 0.5*0.1 term of
      the shrink-toward-0.5). There is no additive scaling to replace.

So a challenger differing from the champion in EXACTLY ONE thing cannot be
constructed: there is no additive scaling to swap, no base threshold to
multiply, and engine1.js has zero `sport` references so market cannot reach the
grade. A threshold must exist first — that is R1 of
specs/full-output-grade-mapping.md, an explicitly held separate order. Shipping
R1+R4 together would make the Phase-3 delta report misleading: the re-letter
would be driven mostly by switching to probability grading while being
presented as the efficiency fix.

0.4 coverage: the spec names 5 scores; the live ledger has 11 markets and only
MLB total_bases maps to one. 9 of 11 have no score, so "all scored markets"
cannot be satisfied without inventing 9 numbers.

PART B — LEDGER TAKEABLE TAGGING: BUILT (the deferred C2).
New src/config/takeableStandard.js: floor on the minus side, UNCAPPED plus.
Deliberately NOT valueEngine.isTakeable (the -160..+200 PROMOTION band) — a
+400 prop is not promotable but IS takeable; a test asserts the two diverge on
the plus side and agree at the floor so they can never quietly merge. Absent
price returns null, never false (Number(null) === 0 would tag a missing price
takeable). The floor is POLICY not derived (C1 could not derive one) and is
labelled so; each row records takeable_floor so a re-derivation can re-tag.

Migration 034 (applied + tracked): ledger_entries.takeable boolean +
takeable_floor numeric, nullable, partial index. Forward tagging in
ledgerService at row build; backfill in one statement.
Result: 1254 rows, 1246 tagged (781 takeable / 465 below floor), 8 NULL with
null_despite_price = 0 (the NULLs are genuinely priceless rows). Settled 1163
and graded 1254 unchanged.

PART C — the model-version boundary tag is DELIBERATELY NOT APPLIED: no scaling
change shipped, so no boundary exists, and stamping one would mark a model
transition that never happened. modelEras.js is its home when a real one lands.

Floor: 312 suites / 3890 tests green (8 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:19:13 -04:00
builtbykev c7067c80c4 Persist lock-time multi-book lines to lock_lines (unblocks the staleness audit)
The over-side skew audit's confirming check — was our locked line stale-high vs
consensus AT LOCK — was BLOCKED because multi-book lines at lock were never
persisted (bookprices is Redis current-only). This persists them.

- migration 033: lock_lines table (tracked + applied to prod). One row per
  (graded prop × book) with both odds + a lock timestamp. RLS enabled, NO
  policies -> service-role only (fence). UNIQUE key -> idempotent re-runs.
- lockLineCapture.js: buildLockRows (pure, graded-props only, honest-absent
  single-book) + idempotent upsert persist. Built from the in-memory props at
  the LOCK moment (ts) -> no Redis re-read, no TTL race.
- snapshotService: persist right after `enriched` (the lock moment; gradedAt
  uses the same ts). Best-effort + fenced.

FENCE (measurement-only): lock_lines is read by NOTHING on the grade path
(gradeSlateService, snapshot dedup/indexOdds, challengers, selector, ledger) —
a grep test asserts it, and RLS locks it to the service role. Grade byte-
identical proven: runSnapshot grades are identical with persist on/off (test).

Volume ~1.5-3k rows/day (graded props x books x 5 snapshots); weeks retained,
no pruning needed short-term. Does NOT retroactively fix the existing 62 rows —
future accrual only; confirmation still needs weeks of settled rows. Full suite
3842 green, web build exit 0. No grade/locked_odds/outcome/served surface changed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 02:47:47 -04:00
builtbykev 6386e737b9 proj-v1: absolute matchup projection challenger (distribution + full ladder)
A THIRD challenger (after arch-v1, contact-v1), MLB batting v1. Champion is
market-relative P(stat>LINE); proj-v1 is ABSOLUTE — what the hitter will DO —
emitted as a full distribution from which the WHOLE LADDER (P≥1,P≥2,P≥3) derives.
Champion untouched; nothing claimed; the ledger decides per rung, per stat.

- projection/distribution.js — Bayesian Gamma-Poisson → negative-binomial
  predictive. Admits over-dispersion; under-dispersion → Poisson approx
  (conservative, documented). Uncertainty scales with sample by construction
  (r=α): thin → WIDE (real mass on P≥1, honestly thin P≥3), thick → tight.
  NEVER abstains — width carries the honesty.
- projection/matchupRead.js — the input the book doesn't use. HONEST FIDELITY:
  pitcher repertoire is rich (97% pitch-mix) but hitters have NO pitch-type
  performance, so TRUE repertoire-vs-profile is impossible today. This is the
  COARSE version (arsenal buckets fastball/sinker/breaking + whiff/hard-hit
  tendency × hitter whiff/chase/gb-fb/hard-hit) — beats generic L/R, derived +
  documented + TESTED two-sided. A hitter pitch-type feed unlocks the true form.
- projectionChallenger.js — park RELATIVE to the player's own log exposure
  (isHome→own park, away→opp park; Phase B's raw-multiply bug solved), recency-
  weighted fit, per-factor breakdown (form/park/weather/platoon/matchup — show
  your work), full rung set + book-implied per rung. Combined non-form
  multiplier bounded.
- Wired after contact-v1, own try, flag PROJ_V1_ENABLED, reusing arch-v1's
  already-computed park/weather/platoon (no duplicate env I/O). Own ledger
  columns (migration 032, applied to prod): distribution, ladder, point, line,
  our-P, book-implied, factor breakdown — measurable per rung/stat after settle.

Phase 0 (prod-verified): venue join via isHome; NB family; uncertainty-as-width;
coarse matchup honest fidelity; no lineup-slot (per-game rate, volume implicit).
Sanity: thin-hot → wide (credible low rung, thin high rung); .300 hitter ≠ 3.0;
matchup two-sided; champion byte-identical. proj-v1 suites 23/23; snapshot/
ledger/siblings 74 green. Forward-only, version-stamped, PROJ_V1_ENABLED kill.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-23 03:59:36 -04:00
builtbykev b6f12daa98 Contact-quality challenger (contact-v1) — nominate, don't swap
Phase A #2: the champion grade (l5/l20 result-based form) is a HYPOTHESIS that
contact quality predicts better — unmeasured on our props, with zero settled
p_win yet. Swapping l5/l20 (the champion's two heaviest ±1.0 factors) blind
could degrade the core grade undetectably for weeks. So this NOMINATES contact
quality as a second challenger, records what it WOULD project per prop, and lets
the settled ledger decide. Nothing users see changes; the champion is untouched.

- src/services/contactChallenger.js — pure, mirrors challengerProjection. Log-
  odds lean (capped, never a re-forecast) from SEASON contact quality vs league
  percentiles. Metric→prop mapping is the whole game: barrel_pct→HR,
  hard_hit_pct→TB/doubles, k_pct-INVERSE→hits (singles resolve on contact
  frequency, not barrels), k_pct→batter K. rbi/runs/walks ABSTAIN (opportunity/
  discipline — no clean contact predictor). Honest-absent: thin (<50 PA)/absent/
  unmapped/non-batter → p_win_contact NULL (no projection), never a fallback;
  "measured but unremarkable" is distinct (equals champion, delta 0).
- Wired in snapshotService AFTER arch-v1, reusing the already-loaded statcast
  rows; its own try so a second challenger can't break the pipeline. Reads
  g.p_win, never writes it.
- Retained SEPARATELY on the ledger (p_win_contact/contact_delta/
  contact_adjustments/contact_version='contact-v1') so each challenger's marginal
  contribution is measured independently; ledger_entries.stat gives per-prop-type
  segmentation. Migration 031 (applied to prod).

Phase 0 (prod-verified): statcast_aggregates is SEASON cumulative (not rolling),
48h stale now but season-scoped so ~8 PA/600 is negligible; 100% of graded
hitters covered, 92% at ≥50 PA; no xBA/xwOBA in the feed. Forward-only,
version-stamped (contact_version null on pre-nomination rows). Promotion is a
LATER decision on settled evidence, per prop type — never asserted here.

contactChallenger 14/14; snapshot/ledger/arch-v1 suites 80 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 22:42:55 -04:00
builtbykev a011ae79fe Statcast: role belongs in the key (two-way players)
Found by inducing the real job on the server, not by review: the first chunk
wrote, the second failed with 'ON CONFLICT DO UPDATE command cannot affect row
a second time'. A player can legitimately appear in BOTH the batter and the
pitcher feeds — two-way players, position players who pitch, pitchers who bat —
so (sport, season, source_id) collapsed two real profiles into one key and a
single batch hit the same row twice.

Ohtani has a real batter profile and a real pitcher profile. Merging them would
invent one player out of two genuinely different sets of measurements, so role
goes in the primary key rather than one profile winning. Migration 031 applied;
conflict target updated; a two-way case is now a test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 21:49:53 -04:00
builtbykev 528cb1a6d0 Layer 1: Statcast mechanism-data ingestion (backfill + nightly refresh)
The data foundation for the archetype and projection layers, built as the
pattern every sport inherits. Layers 2 and 3 are not touched.

PHASE 0 GATE — both match rates measured live, both 100%. Batters 40/40;
PITCHERS 66/66 across five real rosters (CLE, DET, MIN, NYY, LAD) joined by
MLBAM id against the 713-pitcher Savant feed. Zero honest-absent on identity,
because the join is an integer both systems use natively — and the snapshot
pipeline already stores it per graded row.

SOURCE — five Baseball Savant CSV leaderboards, free and public, pulled with
axios and the CSV parser savantAdapter already runs in prod. pybaseball is
deliberately NOT used: it is an MIT wrapper over these same URLs, and adding it
would reintroduce a Python runtime in a stack where the existing Python service
is already offline. min=1 on every feed, not Savant's default min=q, so the
long tail arrives and OUR minimum-sample gate decides what is thin — explicit
and testable rather than silently dropped upstream.

Measured: 1,354 rows per season (604 batters, 750 pitchers), all five feeds in
about five seconds. Pitcher mechanism includes arm angle, GB/FB/LD, chase and
whiff; batters get exit velo, launch angle, barrel and hard-hit, chase and
z-swing. Handedness rides in free on the movement feed (677 pitchers); batter
handedness stays absent pending a roster join rather than being guessed.

BACKFILL AND REFRESH ARE THE SAME CALL — a full re-pull upserted on
(sport, season, source_id). Idempotent and self-healing: a missed night
self-corrects on the next run, with no incremental who-played bookkeeping to
drift out of sync. At 1,354 rows the simple thing is also the robust one.

HONESTY RULES, each with a test: a metric the feed did not carry is null and
never 0; a thin sample is STORED and flagged rather than dropped or inflated,
because thin and missing are different claims; an unjoined player is stored
with a null player_key and joins later; and if every feed comes back empty the
job REFUSES to write, so a bad night can never blank a good table.

Freshness is treated as a truth property. updated_at on every row, and the
scheduler pages on a failed run AND on silent staleness — a job that stops
being scheduled never produces a failure, so staleness has to alarm on its own.
Never-built is deliberately not stale: different condition, different fix, and
paging on a fresh install teaches the operator to ignore the alarm.

Nightly at STATCAST_HOUR_UTC (default 11 UTC, after every game is final), kill
switch STATCAST=0, and induce-able at POST /api/internal/statcast/refresh with
a freshness probe at /statcast/status — we verify a refresh by running it, not
by waiting for the slot.

Migration 030 applied. Promoted columns for the classification-critical metrics
plus a metrics JSONB carrying every raw field, so Layer 2 can reach something we
did not promote without a re-ingest. Raw per-pitch stays out of Postgres on
purpose: one season is ~0.85 GB against a 500 MB plan ceiling, and it is
re-pullable from the free source if Layer 3 ever needs it.

Pattern documented in docs/MECHANISM-DATA.md for NBA tracking and NFL Next Gen.

Tests 3581 passed / 292 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 21:37:30 -04:00
builtbykev d3ffa1b8c2 Retention: model_snapshots live + base64 SSH key support
RETENTION (Phase 2, priority zero). History starts compounding tonight.

migration 025 model_snapshots — APPLIED to prod. Append-only, one row per
graded prop PER SIDE PER CYCLE, with a unique index on
(snapshot_id, player_key, stat, line, side) so a retried cycle cannot
duplicate. RLS on, service-role writes only.

What it captures that the ledger never did:
- features jsonb — the model's INPUTS. Without these a backtest can only
  grade our own homework; with them any future model can be replayed
  against the exact conditions this one faced.
- REFUSALS (refused + refusal_reason). The ledger drops them, so a gate
  refusing props that would have WON is invisible — unmeasurable lost
  edge. Captured via a new onGraded hook in gradeSlateService that fires
  with BOTH sides before any filtering.
- grade_11, the pre-collapse grade. The 4-letter map throws away the
  entire live C-/C/C+/B- range.
- model_version + code_sha on every row. ledger_entries mixes pre/post-fix
  grades with no marker and cannot be separated retroactively.
- p_win / ev_pct / fair_odds / takeable / value — none of which any
  permanent store held.

Wiring: analyzeViaEngine1 attaches _features/_grade_11 (underscore =
internal); gradeSlateService fires onGraded then STRIPS them so they never
reach a cache or API payload; snapshotService builds rows and persists
best-effort. Retention reuses the LEDGER's dateET/gameIdFor helpers so
rows share the ledger's natural key exactly — otherwise the settle pass
could never join outcomes onto them. Rows are written BEFORE the empty-
slate early return: a slate that refused everything is exactly the case
worth recording.

CONTRACT HELD: retention is injectable and every path is caught. persist()
returns errors, never throws; a missing Supabase client is SKIPPED, not an
error. A retention failure can never break a snapshot.

BACKUP: backup-db.sh now accepts BACKUP_SSH_KEY as base64 (recommended —
survives env-var newline mangling, which is how injected SSH keys usually
break silently) OR raw PEM, detected by decoding and looking for the PEM
header. Verified both forms detect correctly against a real generated key.

Suite 279/3325 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 23:01:12 -04:00
builtbykev 219167eebf A1: migration 021 — partner attribution (pre-apply commit, per docs/PARTNERS.md)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 19:54:16 -04:00
builtbykev b20145c215 S10 (a1): public ledger profiles v1
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-11 19:32:22 -04:00
builtbykev c96e74c54b Session 59: migration 020 — ledger team/opponent (pre-apply commit)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 22:02:48 -04:00
builtbykev 2c79373a3b Session 58: Phase 1 spec + ledger_entries migration (pre-apply commit)
Migration 019: the truth-infrastructure table. Committed BEFORE it runs,
per the Phase 1 GO instructions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 20:49:28 -04:00
builtbykev 1fa04dc776 Sessions 5-7a: 955 tests, deployment ready 2026-06-08 18:35:13 -04:00
builtbykev 2366660f5e feat: Feature 2.2 — Line Movement + Cascade Detection
Line movement system:
- Baseline capture on first odds fetch of the day
- Movement detection >= 0.5 points with direction (up/down)
- Sharp money heuristic (sharp_action/public_action/unknown)
- GET /api/movements with player, stat_type, min_movement filters
- Movements included in GET /api/odds/nba live responses

Cascade detection system:
- Scratch detection: player props disappear from 2+ books
- Affected user lookup via scan_sessions + picks
- Parlay re-grade without scratched legs
- cascade_alerts created for affected users
- GET /api/alerts (Analyst/Desk only), PATCH /api/alerts/:id/read

Zero extra Odds API credits — all detection piggybacks on existing fetches.
Migration 002: line_baselines, line_movements, cascade_alerts tables.

30 new tests, 188 total (161 Node.js + 27 Python), all passing.
Phase 2 Core Product COMPLETE.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 14:21:34 -04:00
builtbykev 3da1b4242c feat: Feature 1.2 (NBA stats FastAPI service) + Feature 1.4 (database schema)
Feature 1.2: Python FastAPI microservice wrapping nba_api
- GET /stats/season-avg, /stats/last-n, /stats/splits, /players/search
- Redis caching (24hr/1hr/6hr/7day), 0.6s rate limiting, PRA derived stat
- 27 Python tests passing

Feature 1.4: Complete Supabase database schema
- 6 tables: users, picks, scan_sessions, bets, outcomes, performance
- RLS enabled on all tables with auth.uid() policies
- 3 triggers: auto-create user, updated_at, scan count reset
- 37 schema validation tests passing
- Migration SQL ready, pending manual apply (WSL2 DNS blocker)

Total: 92 tests (65 Node.js + 27 Python), all passing

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 10:58:58 -04:00