91911cfb1c
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2806 lines
206 KiB
Markdown
2806 lines
206 KiB
Markdown
# VYNDR — STATE OF THE WORLD
|
||
### As of `4f3f433` (main, DEPLOYED + fingerprinted live), 2026-07-22. This file opens every future session. **Start with the CURRENT STATUS + OPEN ITEMS block below.**
|
||
|
||
> ## 📊 CANONICAL BOARD → `outputs/VYNDR-COMPLETION-MATRIX.md`
|
||
> The scannable five-part (DESIGNED·BUILT·WIRED·LIVE·HONEST) matrix of every surface +
|
||
> model component **supersedes this file's narrative board** for "what's done" (built
|
||
> 2026-07-27 from repo `b0a51c8` + prod, not from this file's claims). 15/26 surfaces
|
||
> fully done. Read the matrix first; this narrative is the history behind it.
|
||
>
|
||
> **HONESTY PASS shipped 2026-07-27 (6bc18d8, deployed):** every KNOWN live
|
||
> fabrication removed/hidden — /compare (was hardcoded Jokić/Wembanyama grades)
|
||
> → honest in-dev + pulled from nav; founder pricing corrected (Desk $44.99,
|
||
> Analyst $14.99, struck/unwired numbers removed); FAQ NexaPay→Stripe (NexaPay
|
||
> code fully PURGED 2026-07-27 — VYNDR is Stripe-only; env keys + orphaned
|
||
> `user_profiles.nexapay_customer_id` column flagged for Kev); FAQ+Features
|
||
> "Brier/CLV published" claim removed; MobileEdgeBoard placeholder-edge% removed;
|
||
> price-triplet null model/EV → honest NO_MODEL (no false "poisoned" copy). No
|
||
> feature cut (media/newsletter/WIRE kept). Honest state = "no KNOWN live
|
||
> fabrications," not "provably none." See the matrix's HONESTY PASS section.
|
||
|
||
---
|
||
# 🎯 CLV REDIRECT — investigated 2026-07-26. FINDING: it was ALREADY BUILT.
|
||
*The premise "CLV is dead, redirect it to closing_captures" is STALE. Nothing needed building.*
|
||
|
||
> ## 🔧 CLV INSTRUMENT REPAIRED 2026-07-28 (commit 6552281, deployed)
|
||
> The starvation was NOT capture (per-prop coverage is 94.9% MLB / 95.8% WNBA) and NOT
|
||
> the join (key `player_key|stat|side|game_date` is PERFECT, 0 mismatches in-range). It
|
||
> was `attachClosingProb`: `.limit(50000)` with no ORDER BY on a 730k-row table (86%
|
||
> refusal rows) saw ~7% for MLB, and `market_unavailable_reason` was write-once/terminal
|
||
> so wrongly-declared rows never recovered. FIX: read only PRICED captures scoped to the
|
||
> candidate dates (complete, small); make market-unavailable a re-checkable absence that
|
||
> a real close UPGRADES; closing_prob stays write-once. **`closing_prob` 59 → 406**
|
||
> (MLB 30→248, WNBA 29→158; all with provenance; 0 fabricated; grade/locked_odds/outcome
|
||
> untouched). **CLV finding, reported straight: MLB UNDERS LAG THE CLOSE — mean −9.1
|
||
> prob-pts, 74% lose. MLB overs +2.0. WNBA flat both sides.** So the +4.57% MLB-C and the
|
||
> over/under asymmetry ARE substantially stale-line artifacts on the unders (line 38-39
|
||
> below is now RESOLVED → confirmed, not "cannot determine"). Held: proj-v1.1 promotion +
|
||
> p_win/ev on served grades = the PROOF order, gated on this.
|
||
|
||
> ## 🧪 proj-v1.1 PROOF verdict 2026-07-29 (report-only): **NOT PROVEN**
|
||
> N-gate PASSED (overlap n=45 takeable MLB overs; the CLV repair made n≥30 reachable).
|
||
> Edge = `proj_p_over_line − proj_book_implied` (de-vigged, VERIFIED). Phase 2 partial
|
||
> correlation controlling for price: raw r(edge,CLV)=0.455 → **partial 0.245 (n.s.)** —
|
||
> ~half the raw signal is the shared −fair_prob_lock term (mechanical). The CHAMPION
|
||
> out-predicts proj on the SAME 45 rows (champ partial-CLV 0.380 SIG p≈0.01; champ-edge→hit
|
||
> 0.25 vs proj 0.12). MLB unders (n=20) contaminated (CLV −9.3, r=−0.03); WNBA — proj-v1.1
|
||
> doesn't run (MLB-only). **Verdict NOT PROVEN — proj-v1.1 loses to the champion before the
|
||
> under-audit even applies, so the audit is moot for its promotion.** The one defensible
|
||
> edge signal is the CHAMPION's p_win→CLV (0.380), itself still audit-gated. Promotion HELD;
|
||
> re-test needs ~20-25 more settled takeable MLB-over rows AND proj must beat the champion.
|
||
|
||
> ## 🔬 OVER-SIDE SKEW AUDIT 2026-07-29 (report-only): champion signal **SURVIVES BASELINE**
|
||
> Gates the champion's over-CLV (partial r=0.375 SIG p≈0.003, n=62 takeable MLB overs).
|
||
> Three numbers: mechanical baseline CLV (no-edge, champ-edge≤0) **+1.51pt (n=20)**;
|
||
> high-edge CLV **+8.64pt (n=37)**; **DIFFERENCE +7.14pt** = real edge (~5× the floor).
|
||
> Skew is ONE-SIDED (unders lag −7; over baseline only +1.5, not a symmetric +7). De-vig
|
||
> CLEAN (`analyzeViaEngine1.js:539` pairs over+under same-book/same-fetch — no fresh/stale
|
||
> pairing). Close well-defined (DK vs MGM over-prob r=0.921). **VERDICT: SURVIVES → greenlights
|
||
> building the takeable-edge grade ON THE CHAMPION (engine1 p_win), NOT proj-v1.1** — the
|
||
> project's first edge signal to survive an adversarial audit. FLAGGED promising-NOT-confirmed:
|
||
> thin n; Phase-2 lock-staleness check BLOCKED (no multi-book retained AT LOCK — `bookprices`
|
||
> is Redis current-only); pinnacle sharp ref only n=8 (+3.17pt, hint). Fix that would
|
||
> strengthen it (NOT built): retain multi-book at lock + accrue sample. Promotion HELD.
|
||
|
||
> ## 🔓 LOCK-LINE PERSISTENCE shipped 2026-07-29 (commit c7067c8): unblocks the staleness audit
|
||
> The audit's BLOCKED check (was our locked line stale-high vs consensus AT LOCK?) is now
|
||
> being unblocked. New `lock_lines` table (migration 033, applied+tracked) — the lock-time
|
||
> analog of `closing_captures`. `lockLineCapture.buildLockRows` persists each graded prop's
|
||
> per-book lines (both odds + a lock timestamp) from the in-memory `props` at the lock
|
||
> moment (`ts`, no Redis re-read → no TTL race), in `snapshotService` right after
|
||
> `enriched`. FENCED: RLS-no-policies (service-role only) + a grep test — read by nothing on
|
||
> the grade path; grade byte-identical proven (runSnapshot grades identical with persist
|
||
> on/off). **Accrues FUTURE rows only — does NOT retroactively fix the existing 62; the
|
||
> confirming audit still needs weeks of settled lock+close+outcome rows.** Volume ~1.5-3k
|
||
> rows/day. Populates on the next snapshot tick (14:00 UTC); no promotion, no served edge,
|
||
> no capture-of-close change. Next: let it accrue, then re-run the staleness audit.
|
||
|
||
> ## 📚 BOOK COMPARISON WIRED 2026-07-29 (commit 2ab2eea): matrix row 18 → DONE (16/26)
|
||
> The dead `BookComparison.tsx` is routed to the GradeResultCard via a new self-fetching
|
||
> `BookComparisonPanel` (reads live `/api/books` = fenced `bookprices`). Contract fix: books
|
||
> frequently sit at DIFFERENT lines (WNBA DK 21.5/FD 18.5), so it renders EACH book's own
|
||
> line per-row (no false shared-line comparison). Honest states: single-book → one book +
|
||
> "One book posting this prop.", no crown; multi-book → all books, none crowned
|
||
> (`BOOK_CROWN_ENABLED=false`); no books → NULL (self-hide). Display-only: only the
|
||
> always-empty inline `d.books` section was replaced — grade/projection/line/PriceTriplet
|
||
> untouched (web-only, grade byte-identical trivially). Freshness equal (same snapshot as
|
||
> the lock; intraday touches neither). HELD: push-to-book, movement strip, crown enable.
|
||
> VERIFY: floor green (3851 tests, web build 0). Live-card SCREENSHOT is BLOCKED — the
|
||
> GradeResultCard is only on the auth-gated `/scan` (redirects to `/signup`), and creating
|
||
> an account is a prohibited action. Verified instead via the LIVE data feed the deployed
|
||
> panel renders: `/api/books` multi-book different-lines (WNBA DK o21.5 / FD o18.5), MLB
|
||
> multi-book, honest-absent (`books:null`→NULL), crown OFF (`bestLines:0`) — plus the
|
||
> component/panel/wiring unit tests + the type-checked build. Fingerprint: no unauthenticated
|
||
> card surface (frontend-behind-auth); the consumed feed is live + unchanged.
|
||
|
||
> ## 🎯 HERO RANKING FIX 2026-07-29 (commit 41b86e3): p_win among takeable, no arbitrary pick
|
||
> Review Zero found the hero's REAL behavior was worse than "unknown": it ranked on ev_pct
|
||
> (v2), but ev_pct is NULL on served grades and **`Number(null) === 0`** made every prop tie
|
||
> at EV 0 → the "top read" was the FIRST takeable A/B prop in cache order (ARBITRARY, dressed
|
||
> as ranked; prod was serving Kelsey Mitchell A p_win 0.755 = the #6 read). 0.2/0.3 PASS:
|
||
> champion p_win is on 100% of recent ledger rows (same graded source as the snapshot cache;
|
||
> `/api/snapshot` null is just the `stripModelPrice` tier strip), book_odds present. v3:
|
||
> rank by champion **p_win** among A/B TAKEABLE reads (`isTakeable` −160..+200, same as the
|
||
> proof/audit), strict-null guard kills `Number(null)=0`, takeable filter excludes chalk, NO
|
||
> backfill → honest empty state (`available:false, reason:'no_qualifying_read'`). p_win is
|
||
> RANKING-ONLY (never exposed; route strips it). Display-only: reads caches, writes nothing.
|
||
> Framing = model-vs-book number + grade (no proven-edge/+EV/best-bet/CLV/ROI claim).
|
||
> FINGERPRINT (VERIFIED): the current MLB snapshot has 1 takeable A/B grade (Schanuel −135)
|
||
> which OLD code (mlb-first iteration + `Number(null)=0` first-wins) would pick — but the live
|
||
> hero is Kelsey Mitchell (WNBA A), the max-p_win read ACROSS sports, and untakeable chalk
|
||
> (Yainer Diaz −200, Altuve C) is excluded → new code confirmed serving. Visual is auth-gated
|
||
> (landing hero public, dashboard not) — data fingerprint used, `cf-cache-status: DYNAMIC`.
|
||
> ## ✅ PUSH-SCORING PREMISE VERIFY 2026-07-29 (report-only): **RULING HOLDS — row closed**
|
||
> Tested whether the standing "push scoring is correct — do not touch" ruling still rests on a
|
||
> true premise, or whether whole-number lines had entered the feed since it was made (which
|
||
> would make 0 pushes a real mis-scoring bug the ruling was shielding). **The premise HOLDS.**
|
||
> 0.1 basis confirmed verbatim (`VYNDR-CANONICAL-STATE.md:69`): "*no `push`* (pushes structurally
|
||
> impossible: half-number lines)". **Phase 1 — the feed is STILL 100% half-numbers, on 4
|
||
> independent populations, per sport AND per market (no blend): `closing_captures` **117,970
|
||
> priced market lines → 0 whole** (5 books incl. `sharp`/pinnacle, 16 sport/stat groups, Jul 20–29
|
||
> continuous); `model_snapshots` 6,050 → 0 (incl. grader refusals, so not survivorship);
|
||
> `ledger_entries` 1,141 public → 0 (11 markets: MLB hits/doubles/TB/SB/ER/HR/outs, WNBA pts/reb/
|
||
> ast/3s — every min AND max ends in `.5`); `lock_lines` 173 → 0 (TODAY's lock, freshest feed).
|
||
> **Phase 2** — push branch is `outcomeService.js:151` `if (a === l) return 'push'` **after**
|
||
> `Number()` + `Number.isFinite` guards on both operands → sound numeric compare, NOT the
|
||
> `Number(null)===0` string-vs-number class. It is the **single scoring chokepoint** (`ledgerService.js:31`
|
||
> imports `settleResult`; no parallel scorer derives hit/miss anywhere in `src/`), it is **unit-tested
|
||
> live** (`outcomeService.test.js:38` `settleResult('over',2,2)==='push'`, `nbaSettlement.test.js:104`),
|
||
> and both `ledger_entries.outcome` + `outcomes.result` CHECK constraints **include `'push'`** → a real
|
||
> push would score, write, and persist. **Phase 2.6 / the decisive number: 0 exact ties in 992 settled
|
||
> rows, 0 fractional actuals, and the SMALLEST actual-vs-line gap across all 992 rows is 0.5** — the
|
||
> arithmetic minimum between an integer result and a half-number line. **VERDICT: RULING HOLDS.**
|
||
> Expected push rate is **exactly 0 (P=0), not "low" — 0/992 is FORCED, not chance**, so the
|
||
> "implausible" flag is retired: it mistook an arithmetic impossibility for a suspicious absence.
|
||
> The open-items row is CLOSED honestly. Nothing was modified (no scoring, settlement, re-settle,
|
||
> or backfill). NOTE the stale n: the flag said 470 settled; it is now **1,097 settled** (593 hit /
|
||
> 399 miss / 105 void / 44 unsettled-today). **No latent bug either** — the branch is correct and
|
||
> covered, so if a whole-number market ever DOES enter (NFL/NHL are code-wired but out of season;
|
||
> whole-number K props exist at some books), it scores as a push automatically. Residual risk is
|
||
> a monitoring gap, not a scoring gap: nothing ALERTS on the first whole-number line.
|
||
|
||
> ## 🔬 EDGE_PCT SCALE DIAGNOSIS 2026-07-29 (report-only): **not a units bug — a mislabelled metric that DRIVES A LIVE SORT**
|
||
> **0.1 What it is.** `analyzeViaEngine1.js:265-270 edgePctFor()`:
|
||
> `signed = over ? (projection − line) : (line − projection)`; `Math.round((signed/line)*1000)/10`
|
||
> → **edge_pct = ((projection − line) / line) × 100**, signed by direction, 1dp. `projection` =
|
||
> `projectionFor` = `l5_avg ?? l20_avg ?? {stat}_per_90 ?? xg_per_90` (`:252-257`). **It is
|
||
> INDEPENDENT of `p_win`** — so it is NOT tainted by the overconfidence that damns `ev_pct`. But it
|
||
> takes **NO price input at all**, so it cannot express edge in the betting sense. **Verdict:
|
||
> arithmetically correct, MISLABELLED. Honest as "% the model's projection differs from the line";
|
||
> a lie at ANY scale as "EDGE"** — which is exactly how every surface labels it. **TWO independent
|
||
> implementations exist** — backend `edgePctFor` AND `web/src/lib/gradeAdapter.js:25-31 computeEdge`
|
||
> (same formula); the grade card renders the WEB one, so a backend-only fix would not reach it.
|
||
> **0.2 The cap.** `SANE_EDGE_MAX = 40`, `deskShowcaseService.js:31`, comment: *"beyond this the
|
||
> (model-line)/line value isn't a market edge"*. Mirrored `slateAdapter.js:613 EDGE_BOARD_SANE_MAX=40`
|
||
> + `MobileEdgeBoard.tsx:45`. It is a self-declared plausibility bound from an earlier order, **not a
|
||
> derived statistical bound**.
|
||
> **0.3 Mechanism = SMALL-DENOMINATOR EXPLOSION** (not units, not inverted, not missing ×100 — the
|
||
> ×100 is present and correct). `line` is the denominator and **562 of 655 MLB rows (86%) sit at line
|
||
> 0.5**, where every 0.1 of projection is ±20 points. Max 620 = projection ≈3.6 on a 0.5 line. PROVEN
|
||
> per-line: MLB 0.5 → 73.0% over cap, max 620 · MLB 1.5 → 43.8%, max 153 · WNBA 12.5 → 6.7% · **WNBA
|
||
> 26.5 → max 1.9**. Monotone decay with line size.
|
||
> **PHASE 1 — matrix figures RE-VERIFIED, partly stale.** Blended over cap-40 = **502/1141 = 44.0%**
|
||
> (matrix said 51.5% — direction right, number stale). **Worst value 620 = EXACT match.** Per sport:
|
||
> **MLB n=655** — 13.6% >100, 65.8% >50, median **60** (already 1.5× the cap), p95 180, p99 238,
|
||
> min −86.7, max 620. **WNBA n=486 — 0% >100**, median 12, p95 49, max 77.8. **Shape: structurally
|
||
> broken for MLB, essentially SANE for WNBA.** Not a mild calibration — and note what that means:
|
||
> the scale is a FUNCTION OF LINE SIZE, so the metric is incomparable across markets **by
|
||
> construction**. No rescaling fixes that; only changing what the metric IS would.
|
||
> **PHASE 2 — the ~13-surface count is NOT confirmed. Only 3 surfaces RENDER it:**
|
||
> (a) **`GradeResultCard.tsx:182,216,325`** — LIVE (3 importers), **auth-gated `/scan`** →
|
||
> **TAGGED FOR THE CHROME AUDIT, no visual faked** — renders **"+620% edge" RAW and GREEN**, no cap;
|
||
> (b) **`DeskShowcase.tsx:40`** — LIVE, **PUBLIC** `/pricing` — **already honest**, shows "—"
|
||
> (service nulls >40); (c) **`SoccerGradeResult.tsx:229`** — uncapped, on the ORPHAN `/soccer`
|
||
> (0 nav links, public by URL). **Dead:** `MobileEdgeBoard` (0 importers, pulled in the honesty pass),
|
||
> `DemoScan` (0 importers). **TYPE-ONLY, never rendered:** `PropRow.tsx:45`, `GradeCard.tsx:32`,
|
||
> `ledger/page.tsx:40`. **API-only:** `contentTemplateService.js:164` emits an uncapped "+620% edge"
|
||
> string on public `/api/content` — **no page fetches it** (verified).
|
||
> **🔴 THE SORT ANSWER — YES, IT DRIVES TWO LIVE SORTS.** (1) `slateAdapter.selectTopGrades:469-471`
|
||
> sorts `grade → confidence → |edge| desc`, consumed by `dashboard/page.tsx:419` for the dashboard
|
||
> **TOP GRADES top-10**. **MEASURED: 97.3% of rows (1110/1141) sit in a (date,sport,grade,confidence)
|
||
> tie group of ≥2 (biggest 56)** → the |edge| key is operative for essentially the whole slate, so it
|
||
> is the **de facto ordering** of that leaderboard. (2) `analyzeViaEngine1.js:506` sorts the Desk
|
||
> alt-line ladder by `edge_pct` desc. **TWO SCALE-INDEPENDENT DEFECTS FOUND INSIDE THAT SORT:**
|
||
> `edge: Math.abs(numOr(g.edge, -Infinity))` — (i) **abs()** on an already-direction-signed value ranks
|
||
> the model's strongest DISAGREEMENTS equal to its strongest agreements (**177 ledger rows carry a
|
||
> negative edge**: 58 B / 118 C / 1 F, most negative −86.7); (ii) `Math.abs(-Infinity) = Infinity`, so
|
||
> a **MISSING edge sorts FIRST** — the `Number(null)` fabrication class again, in a new costume.
|
||
> **PHASE 4 — the matrix's "nothing renders `ev_pct`" is WRONG.** `PriceTriplet.tsx:60,67,76` renders
|
||
> `${pct(ev)} EV` and is LIVE + wired (scan → `gradeAdapter:143` → PriceTriplet). It shows nothing only
|
||
> because `ev_pct` is NULL on served grades, so `valueState.js:121` correctly falls to **NO_MODEL
|
||
> honest-absent**. So: **the right metric already has a live, honest render site starved of data** —
|
||
> and the card's "EDGE" row sits exactly where a price-aware number belongs.
|
||
> **PHASE 3 — THE FORK (reported, NOT chosen).** **FIX is not honest**: no rescaling turns a
|
||
> price-free projection gap into an edge — you would be renaming, not fixing; it silently re-ranks the
|
||
> dashboard top-10 (the hero-class bug just fixed); and it must land in TWO implementations.
|
||
> **HIDE is cheap**: only 3 render sites, every one already has a null branch (so no layout breaks),
|
||
> and **DeskShowcase already proves the honest "—" pattern in-product**. **RECOMMENDED: HIDE the
|
||
> number, and re-point the sort at `p_win`** (the hero order established p_win is on 100% of recent
|
||
> ledger rows and is the one signal that survived an adversarial audit) rather than repair a key that
|
||
> is a 0.5-line artifact. Coherent end state: **p_win ranks · ev_pct displays · edge_pct retires.**
|
||
> The abs()/null-first sort defects are worth their own small order regardless of the fork.
|
||
|
||
> ## 🔧 GRADE-BOARD SORT FIXED 2026-07-29 (spec `specs/grade-board-sort.md`): signed signal, missing sorts LAST
|
||
> Display ORDERING only — no grade, ledger, lock_line, scoring, or edge_pct scale/display change.
|
||
> **THREE PREMISE BREAKS found REPORT-FIRST, before code:** (1) **`/api/props/top-graded` returns 404
|
||
> in prod** — it does not exist in `src/` (only 3 axios *callers*), so the Next proxy catches →
|
||
> `{props:[]}` → the dashboard board renders `proofMode`/empty and **the edge sort orders nothing on
|
||
> that surface today**. CORRECTION to my prior order: its "97.3% of rows tie → the edge key decides the
|
||
> board" was a LEDGER measurement I extrapolated to this board — wrong; the board has no rows. The fix
|
||
> is still correct-in-itself and lands the moment the feed is restored. (2) **p_win CANNOT be a
|
||
> client-side sort key for all tiers** — `utils/snapshotGating.stripModelPrice` (Session 67) strips
|
||
> `p_win`/`ev_pct`/`model_odds`/`value`/`takeable` for unentitled tiers because *"shipping p_win is
|
||
> shipping the model price in a different base"*, and `selectTopGrades` runs in the BROWSER. Ranking
|
||
> there by p_win for everyone would REVERSE that gate. **VERIFIED LIVE: prod `/api/snapshot` returns
|
||
> p_win on 0/8 MLB and 0/25 WNBA grades** (stripped, as designed). (3) **The ladder cannot take the
|
||
> takeable gate** — rungs carry NO per-rung price (books price each line differently; we don't fetch
|
||
> them) and `isTakeable` is a property of price alone.
|
||
> **WHAT SHIPPED.** `selectTopGrades`: `grade → confidence → takeable-gated p_win (nulls LAST) →
|
||
> SIGNED edge (nulls LAST) → input order`. `Math.abs()` GONE — `edge` is signed by direction upstream
|
||
> so positive = the model AGREES; `|edge|` had been ranking the model's strongest DISAGREEMENTS level
|
||
> with its agreements (177 public ledger rows carry a negative edge). `Math.abs(-Infinity)=Infinity`
|
||
> GONE — a missing signal sorted FIRST (absent data as the top pick, the `Number(null)` class); it now
|
||
> sorts LAST and rows are never dropped. **Scales are never mixed** (p_win 0..1 vs edge % — 0.62 vs 62
|
||
> is not a comparison). Takeable band = `web/src/lib/valueState.isTakeable`, asserted by test to be
|
||
> byte-equal to the hero's `config/valueEngine.isTakeable` (−160..+200) incl. strict-null.
|
||
> **Alt-line ladder** (`analyzeViaEngine1:506`): was `edge_pct desc` (and `Number(x)||0` collapsed
|
||
> absent edges to mid-pack); now ordered **highest-p_win-first derived ANALYTICALLY at zero added
|
||
> compute** — `P(stat ≥ k)` is monotone non-increasing in k, so p_win-desc is exactly line-ASC for an
|
||
> over and line-DESC for an under. `base` stays marked, so order never implies a recommendation; no
|
||
> consumer depends on `alt_lines[0]` (grepped), and `deskShowcaseService.rungsOf:40` already re-sorted
|
||
> by line anyway.
|
||
> **FINGERPRINTED ON REAL DATA (both sports, both paths).** Unentitled path, live prod snapshots:
|
||
> **WNBA (n=25) ordering CHANGED**, MLB (n=8) unchanged (its top edges were already positive);
|
||
> **within every (grade,confidence) tie group the signed edge is non-increasing — 20 adjacent pairs,
|
||
> 0 violations**. Entitled path, 40 real ledger rows carrying p_win+locked_odds: p_win-descending
|
||
> within tie groups, **untakeable chalk correctly NOT promoted** (Trea Turner p_win .757 @−275 does
|
||
> not beat Rhyne Howard .745 @−120) — **36 pairs, 0 violations**.
|
||
> **HERO CONSISTENCY — stated honestly, they are NOT identical and should not be.** Same signal, same
|
||
> gate, DIFFERENT precedence by contract: the board is `top GRADES` (grade-tier first), the hero is
|
||
> `top read` (p_win first). On the real rows the board leads with Angel Reese (A, p_win .555) while
|
||
> the hero picks Brionna Jones (B, p_win .90). **Within the leading tier they agree exactly (verified
|
||
> true).** That is a difference of question, not a contradiction — do not "fix" it by making the board
|
||
> ignore grade.
|
||
> **FLOOR: 310 suites / 3864 tests green, web build exit 0.** New `tests/unit/gradeBoardSort.test.js`
|
||
> locks: disagreement never outranks agreement, missing sorts LAST and stays present, null/'' edge is
|
||
> absent-not-zero, takeable p_win beats untakeable higher-p_win chalk, scales never mixed, the two
|
||
> takeable bands match, grade tier still dominates, and the ladder order for BOTH directions.
|
||
> ### ✅ POST-DEPLOY FINGERPRINT (commit `b85b351`, both halves VERIFIED)
|
||
> **Client half — before/after captured, not inferred.** Polled the deployed `/dashboard` chunks every
|
||
> 15s: attempts 1-3 served the OLD key (`Math.abs(...-1/0)` present, new comparator absent); **attempt
|
||
> 4 flipped — new comparator present, OLD abs-key GONE.** Both states observed across the deploy
|
||
> boundary. API health after deploy: `/api/snapshot/mlb` 200, `/api/snapshot/wnba` 200, `/api/accuracy` 200.
|
||
> **Backend half — REAL LADDER PATH INDUCED ON DEMAND** (not waiting on a cron slot), `analyzeViaEngine1`
|
||
> with live MLB game logs: **OVER** Junior Caminero hits @0.5 → grade C, p_win 0.644, ladder
|
||
> `0.5•(C) 1(C) 1.5(F)` = **line-ASC PASS**; **UNDER** @1.5 → grade C, p_win 0.662, ladder
|
||
> `2.5(C) 2(C) 1.5•(C) 1(C) 0.5(F)` = **line-DESC PASS**. Both directions correct on the real engine,
|
||
> `base` rung marked in both. (The Supabase-backed side-signals logged lookup failures locally — the
|
||
> known WSL2 `*.supabase.co` DNS blocker, BLOCKER-003 — the grade + ladder still computed from real
|
||
> stats.) `model_snapshots.code_sha` will independently stamp `b85b351` on the next tick (latest is
|
||
> `9b5235c`); the on-demand induction above is the verification, not that stamp.
|
||
|
||
> **HELD:** edge_pct rescale/display retirement (Order B) · building the missing
|
||
> `/api/props/top-graded` server selector (that is what would make the board render at all) ·
|
||
> exposing p_win to unentitled tiers. **Dashboard + Desk visuals are auth/feed-gated → TAGGED FOR
|
||
> THE CHROME AUDIT, no visual faked.**
|
||
|
||
> ## 🟢 `/api/props/top-graded` SERVER SELECTOR BUILT 2026-07-29 (spec `specs/top-graded-selector.md`)
|
||
> The dashboard TOP GRADES board finally has a feed. New READ endpoint; no grade/ledger/
|
||
> lock_line/scoring write. **0.1 CORRECTS the premise: the handler NEVER EXISTED** — searched
|
||
> every commit (`git rev-list --all`) for a `/top-graded` definition in `src/`, **zero hits**. Not
|
||
> "removed": the three axios callers (`cheatsheetGenerator`, `gradeOfTheDay`, `widget`) and the Next
|
||
> proxy were written against a phantom endpoint, so **those three content generators have silently
|
||
> received `[]` for their entire life** — a second, previously-unnoticed casualty now fixed.
|
||
> **CONTRACT recovered from consumers, not guessed:** envelope `{props:[...]}`; params `sport`
|
||
> (UPPERCASE NBA|MLB|WNBA, absent = all sports, which `gradeOfTheDay` relies on) + `limit`; rows carry
|
||
> `player/stat/line/direction/sport/grade/confidence?` plus the `player_name`/`stat_type` aliases and
|
||
> `game_id` the other callers read.
|
||
> **🔴 0.5 POPULATED-PATH RISK FOUND (the board's populated branch had never run in prod):**
|
||
> `dashboard/page.tsx:463` calls **`g.stat.replace(/_/g,' ')` UNGUARDED**, and `g.player` feeds the row
|
||
> key + `/scan` URL + heading, and `sport` must be UPPERCASE for `SportPill` (`type Sport =
|
||
> 'NBA'|'MLB'|'WNBA'`). A row missing any of those would have CRASHED the board on first populated
|
||
> render. `toRow` therefore requires non-empty string `player`+`stat` and a finite `line`, uppercases
|
||
> `sport`, and **DROPS** an unrenderable row — a shorter board beats a broken one.
|
||
> **THE ORDER OF OPERATIONS (the leak surface):** read cache → **RANK with `p_win` for EVERY tier
|
||
> server-side** → map rows *including* model fields → **`stripModelPrice(rows, tier)`** → serialize. So
|
||
> a free caller receives the paid RANKING without the paid VALUES. Tier via
|
||
> `resolveTierFromRequest` which **FAILS CLOSED to 'free'** (a resolution failure can only withhold,
|
||
> never leak). `Cache-Control` = `private` with a bearer token, `public` otherwise (the `/api/snapshot`
|
||
> precedent — a CDN must never hand a paid payload to an anonymous viewer).
|
||
> **ONE SHARED DEFINITION, no drift:** extracted `src/utils/gradeRanking.js`
|
||
> (`takeablePWin`/`descNullsLast`/`rankGrades`). **`heroPropService` now imports `takeablePWin`
|
||
> instead of its inline copy** (behaviour unchanged — it was that logic verbatim); the selector imports
|
||
> `rankGrades`; `web/src/lib/slateAdapter` keeps its mirror (browser can't import `src/`, S25) and a
|
||
> test **cross-checks the two on identical fixtures** — the `playerName.js` precedent.
|
||
> **0.3 VERIFIED LIVE that the server HAS p_win:** `/api/hero-prop` returns `available:true` (Brionna
|
||
> Jones, B, wnba) and the hero rule REQUIRES non-null p_win + a takeable price. Public
|
||
> `/api/snapshot` shows 0/8 MLB + 0/25 WNBA only because it is stripped on the way out.
|
||
> **⚠️ HONEST LIMIT — via the dashboard, EVERY viewer gets the free-tier payload.** The Next proxy
|
||
> (`cachedBackendJson`) sends only `{Accept}` — **no Authorization header** — and caches under a
|
||
> SHARED key (`todayKey(sport,'top_graded')`). That is the SAFE behaviour: forwarding auth into a shared
|
||
> cache is exactly how a paid payload leaks to anonymous viewers. So the board shows the correct ORDER
|
||
> with no paid values for everyone; entitled payloads are served on a direct authenticated API call
|
||
> (proven by route test). Wiring per-tier delivery through the proxy would need a tier-keyed cache — NOT
|
||
> this order.
|
||
> **VERIFIED ON REAL PROD SNAPSHOT DATA** (anonymous path, what the board will actually render):
|
||
> MLB 8 props (B c57 Chandler Simpson edge 140 → …), WNBA 10 (B c69 Rhyne Howard u17.5 points → …),
|
||
> **0 paid-field leaks, render-contract safe on every row, sport uppercase**.
|
||
> **FLOOR: 311 suites / 3882 tests green (18 new), web build exit 0.** The leak test uses POPULATED
|
||
> p_win fixtures (not today's nulls): entitled → p_win present and it drove the order; unentitled →
|
||
> **byte-identical order, all five MODEL_FIELDS absent, `JSON.stringify` carries no trace**, while
|
||
> market facts (book/fair) SURVIVE — the fair leg is never the paywall. Also locked: chalk (p_win .95
|
||
> @−300) never tops the board, nulls last but PRESENT, refusals excluded, unrenderable rows dropped,
|
||
> thin slate → `200 {props:[]}` never a 404, and board-vs-hero differ by design yet agree within tier.
|
||
> ### ✅ POST-DEPLOY FINGERPRINT (commit `72a14dc`) — the boundary PROVEN IN PRODUCTION
|
||
> **404 → 200 captured across the deploy boundary** (polled every 15s: attempts 1-4 `HTTP 404`,
|
||
> attempt 5 `HTTP 200`). `Cache-Control: public, max-age=60` for the anonymous caller.
|
||
> **🔑 THE PROOF THAT THE STRIP-AFTER-RANK BOUNDARY WORKS, on live data, not a fixture.** The
|
||
> anonymous live order is **1. Brionna Jones (edge 29.4) · 2. Rhyne Howard (edge 42.9)** — both B/c69.
|
||
> An edge-only sort would put Rhyne Howard FIRST (42.9 > 29.4), and that is exactly what the local
|
||
> induction over the *stripped* public snapshot produced. Production ranks Brionna first because
|
||
> **server-side p_win ordered it** (Brionna 0.90 @−106 > Howard 0.745 @−120, both takeable — matching
|
||
> the ledger rows exactly). So the free caller received the **paid RANKING while the payload carries
|
||
> `PAID FIELDS: NONE`**. That is the acceptance criterion demonstrated end-to-end in prod.
|
||
> Also verified live: **bogus bearer token → still no paid fields** (`resolveTierFromRequest` fails
|
||
> closed, proven in prod, not just in test); render-contract safe on every row + `sport` uppercase;
|
||
> market facts (book/fair) survive; MLB 3 rows (Chandler Simpson B first); the no-sport merged call
|
||
> (gradeOfTheDay's shape) returns rows; and **the board's own proxy path `vyndr.app/api/props/top-graded`
|
||
> now returns 10 props → the TOP GRADES board RENDERS instead of falling back to empty.**
|
||
> The rendered board itself is a client-side render → TAGGED FOR THE CHROME AUDIT, not faked.
|
||
|
||
> **HELD:** edge_pct rescale/retirement (Order B) · board columns/contract unchanged · tier-keyed proxy
|
||
> caching. Dashboard visual is auth-gated → TAGGED FOR THE CHROME AUDIT, not faked.
|
||
|
||
> ## 🔬 C1 — TAKEABLE FLOOR DERIVATION 2026-07-30 (report-only): **CANNOT DERIVE — the data does not support ANY floor number**
|
||
> Read-only. Nothing tagged, built, or changed. **The SHAPE (floor on minus, uncapped plus) is Kev's
|
||
> ratified call and is not in question — this order could not supply the NUMBER, and says so rather
|
||
> than dressing a guess as a derivation.**
|
||
> **0.1/0.2 POPULATION — VERIFIED.** MLB decided (hit/miss) **OVERS = 296**, Jul 11-30, and **100%
|
||
> carry `locked_odds`** (0 missing — the floor is about the price you could actually bet). **METHOD
|
||
> NOTE:** derived over ALL prices, NOT pre-filtered to the current −160..+200 band — pre-filtering
|
||
> would be CIRCULAR and would empty the very deep buckets the floor must judge.
|
||
> **PHASE 1 — ROI BY LOCKED PRICE (flat 1u, 95% CI on the mean return):**
|
||
> | bucket | n | hit% | breakeven% | **ROI%** | 95% CI |
|
||
> |---|---|---|---|---|---|
|
||
> | +200 and up | 39 | 20.5 | 22.0 | **−13.15** | [−68.6, +42.3] |
|
||
> | +100..+199 | 51 | 52.9 | 44.5 | **+19.31** | [−12.0, +50.7] |
|
||
> | −110..+99 | 8 | 62.5 | 51.3 | +21.60 | [−48.2, +91.4] **TOO THIN** |
|
||
> | −111..−135 | 32 | 50.0 | 55.8 | **−10.55** | [−42.1, +21.0] |
|
||
> | −136..−160 | 43 | 53.5 | 60.1 | **−11.68** | [−36.6, +13.2] |
|
||
> | −161..−200 | 71 | 69.0 | 64.6 | **+6.46** | [−10.3, +23.2] |
|
||
> | worse than −200 | 52 | 73.1 | 69.7 | **+4.75** | [−12.7, +22.2] |
|
||
> **🔴 THREE FINDINGS THAT BLOCK A DERIVED FLOOR:**
|
||
> **(1) EVERY 95% CI CONTAINS ZERO.** Not one bucket is statistically distinguishable from zero ROI.
|
||
> Phase 2 asked for "the deepest price where ROI is still positive WITH MARGIN" — **no such bucket
|
||
> exists.** Any floor drawn here is drawn through noise, which this order explicitly forbids.
|
||
> **(2) THE CURVE IS NON-MONOTONE AND RUNS OPPOSITE TO THE PREMISE'S MODEL.** The premise reasons that
|
||
> deep-negative prices are structurally −EV ("the price ate the edge"). The data shows the **deepest
|
||
> buckets are the POSITIVE ones** (−161..−200 → +6.5%; worse-than−200 → +4.8%, hitting 73.1% against a
|
||
> 69.7% breakeven) while the **most negative ROI sits INSIDE the current takeable band** (−111..−160 →
|
||
> ≈ −11 to −13%). Read literally this data says "avoid −111..−160," which is almost certainly noise —
|
||
> and that is exactly why it must not be turned into a floor.
|
||
> **(3) A REAL CONFOUND: price bucket is entangled with MARKET.** +200-and-up is `doubles/home_runs`
|
||
> (rare-event longshots, avg line 0.53); the deep-negative buckets are `hits/total_bases` ("will he get
|
||
> a hit"). Holding the market constant (hits + total_bases only) the non-monotone shape PERSISTS
|
||
> (+100up +22.6 · −111..−135 **−13.4** · −136..−160 **−13.5** · −161..−200 +6.5 · worse-than−200 +4.8),
|
||
> so the confound is not the whole story — but it means a price-only floor would partly be encoding
|
||
> "avoid HR/doubles props," which is a market rule, not a price rule.
|
||
> **POWER — how far the data runs out.** Rows needed PER BUCKET to resolve a 5-point ROI edge at 95%:
|
||
> **661-2,285.** Actual bucket sizes: **8-71.** We are 10x-100x short. At the observed accrual (18.5
|
||
> decided MLB overs/day; 4.44/day into the −161..−200 bucket) reaching 828 rows in that ONE bucket
|
||
> takes **~187 days** — and the MLB season ends well before that, so it will not accrue continuously.
|
||
> **THE INHERITED −160 IS ALSO UNVALIDATED (neither confirmed nor refuted).** In-band (−160..+200) ROI
|
||
> **+3.84%** [−13.0, +20.7] vs out-of-band **−0.04%** [−16.1, +16.0] — a ~3.9-point gap with massively
|
||
> overlapping intervals. Post-fix-only (≥2026-07-19, the model-version cutoff): in-band +7.91%
|
||
> [−11.7, +27.5] (n=102) vs out-of-band −2.07% (n=112). Directionally friendly to the current band,
|
||
> statistically silent.
|
||
> **PHASE 2.6 — THE NO-CEILING CALL IS NOT SUPPORTED BY THIS DATA EITHER (stated straight).** The
|
||
> +200-and-up bucket is the WORST performer (ROI −13.15%, hitting 20.5% against a 22.0% breakeven).
|
||
> It is confounded (doubles/HR) and n=39 with a [−68.6, +42.3] interval, so it does not REFUTE the
|
||
> no-cap call — but it certainly does not support it. **Uncapped plus-money remains a defensible
|
||
> design/risk choice; it should not be described as data-backed.**
|
||
> **RECOMMENDATION FOR C2 (so it is not blocked on a number that cannot be derived):** ship the floor
|
||
> as an explicitly-labelled **POLICY** floor — keep **−160** (inherited, already the band everywhere,
|
||
> and directionally the better half of the only comparison available) — and **label it in code and copy
|
||
> as a policy choice pending derivation, NOT as derived**. **RE-DERIVATION TRIGGER (so "provisional"
|
||
> cannot silently become permanent): re-run this order when ANY negative bucket reaches n ≥ 300, or at
|
||
> the end of the MLB regular season, whichever comes first — and only adopt a data-derived floor when a
|
||
> bucket's 95% CI EXCLUDES zero.**
|
||
> **0.3 — EVERY `takeable` DEFINITION SITE C2 MUST UNIFY (9):** **DEFINITIONS (2):**
|
||
> `src/config/valueEngine.js:21-22,30-35` (env-tunable `TAKEABLE_ODDS_CEILING`/`TAKEABLE_ODDS_MAX`,
|
||
> defaults −160/200) · `web/src/lib/valueState.js:42-43,72-77` (**HARDCODED −160/200**).
|
||
> **🔴 LIVE DRIFT HAZARD FOR C2:** the backend is env-tunable and the frontend is hardcoded, so
|
||
> changing `TAKEABLE_ODDS_CEILING` in prod TODAY would silently desync the two. **CONSUMERS (7):**
|
||
> `src/utils/gradeRanking.js:25,56` (hero + server board) · `src/services/heroPropService.js` (via
|
||
> gradeRanking) · **`src/services/intelligence/analyzeViaEngine1.js:567`** (stamps `legacy.takeable`
|
||
> onto every graded prop — this is the one that reaches the snapshot/ledger) ·
|
||
> `src/config/valueEngine.js:40` (`isValue`) · `web/src/lib/slateAdapter.js:474,478` (client board) ·
|
||
> `web/src/lib/valueState.js:82` (`isValue` mirror) · **`web/src/components/vyndr/PriceTriplet.tsx:9-10,68`
|
||
> (renders the band IN USER-VISIBLE COPY — "Takeable band −160 to +200", so a floor change is a copy
|
||
> change too).**
|
||
|
||
> ## 🗺️ MODEL ARCHITECTURE RECOVERY MAP 2026-07-30 (report-only) → **`specs/model-architecture-recovery-map.md`**
|
||
> Archaeology of the specced three-layer engine vs the live grade path. Nothing built,
|
||
> reconnected, promoted, or changed. **HEADLINE: the live grade uses 0 of the 3 specced layers —
|
||
> every metric this project has audited (calibration, CLV, the skew audit, the takeable floor, the
|
||
> champion's p_win→CLV r=0.375) is measured on the SHADOW model, not the specced engine.** Those
|
||
> findings stand (the shadow model served every real grade) but none of them are evidence about the
|
||
> specced architecture, which has never been measured.
|
||
> **0.1 THE CHAMPION = two DISCONNECTED estimates.** The letter is `engine1.gradeProp`:
|
||
> `idx = NEUTRAL_INDEX(3); for (f of factors) idx += f.delta` — and **`engine1.js` has ZERO
|
||
> references to `p_win` or any probability** (grep count 0). The probability is
|
||
> `probabilityEstimator` (`frequencyOver` + 5 heuristic layers, clamp [.10,.95]) computed AFTER and
|
||
> merely attached. **p_win never feeds the letter; the letter never feeds p_win.** The live grade path
|
||
> makes **no call to the Python service at all**.
|
||
> **🔴 THE PYTHON ENGINE IS NOT DEPLOYED.** `Dockerfile` has no python/pip/requirements — the image is
|
||
> Node-only; `app.js:92-95` only health-CHECKS it and reports `down`. Layers 1-2 are not "bypassed at
|
||
> runtime", they have **never shipped**. `grade_thresholds.json` carries `"shadow_mode": true`.
|
||
> **LAYER STATE:** L1 Similarity `python/utils/similarity.py` (101 ln, MIN_SIMILARITY .7, NBA+MLB
|
||
> factor weights) **BUILT·NOT WIRED·NOT DEPLOYED** — the live grade skips straight to season/recent
|
||
> averages. L2 Bayesian `python/utils/bayesian.py` (320 ln) **BUILT·NOT WIRED·NOT DEPLOYED** — and the
|
||
> spec's "sport-agnostic math, per-sport parameters" claim is **TRUE of the built code**
|
||
> (`BAYESIAN_WEIGHTS` prior/recent/context for 11 stat types across both sports, skewness,
|
||
> data-sufficiency degradation, ±0.15 calibration clamp, a calibration disclaimer). L3
|
||
> **BUILT·WIRED BACKWARDS** — `grade_thresholds.json` maps **PROBABILITY→GRADE** (A+ = p .85-1.00);
|
||
> the live JS reads it in REVERSE, picking a letter from the factor index then looking up that
|
||
> letter's band MIDPOINT to manufacture `confidence`. **That one artifact IS the drift, provable in a
|
||
> single file.** Per-sport **market-efficiency scaling = SPECCED-BUT-ABSENT** (one global scale; the
|
||
> spec's MLB .55 / NBA-stars .80 exists nowhere; the only `efficiency` in code is *usage*-efficiency).
|
||
> **⚠️ The spec's "min 15 instances" is NOT in the code — CANNOT DETERMINE its origin** (code's
|
||
> abstention is `similar_games_below: 3`; the only 15 is NBA `min_minutes_per_game`; the founding
|
||
> three-layer spec text is not in the repo). **NOTE: STATE.md's own "Layer 1/2/3" references are the
|
||
> Session 73-77 ARCHETYPE build train — a different numbering. Do not conflate them.**
|
||
> **🔴 SPORT BOUNDARY — TESTED, NOT ASSERTED: it is NOT clean on the live path.** Adding a sport to the
|
||
> Python spine costs nothing (the spine is dead); adding one to the LIVE path is a **CORE EDIT across
|
||
> ~10+ files** — proplineAdapter.MARKETS · oddsNormalizer.MARKET_MAP (miss → silent ZERO) ·
|
||
> oddsService.SPORT_KEYS/SPORT_MARKETS · **three stat_type whitelists** (analyze.js, scan.js,
|
||
> validation.py — miss one → silent 400s) · featureCache log-field · outcomeService MLB_LOG_FIELD +
|
||
> SPORTS · liveTrackingService LIVE_BOX_FIELD · a projection-emitting extractor (**without it every
|
||
> prop REFUSES**) · emptyState/statFilters/teams · and a settled feed or the sport never settles.
|
||
> Making "a sport" a real module is itself a prerequisite build.
|
||
> **PER-SPORT RECORDS EXIST (2.7 satisfied):** `getModelAggregate` filters on `opts.sport`
|
||
> (`ledgerService:809,819`) and `/api/accuracy` carries a real breakdown — live **`sports.mlb` n=526
|
||
> 62%** alongside **`overall` n=937 58%**, each with its own `min_sample:20`. Two caveats: the pooled
|
||
> `overall` would absorb a new sport's rows, and **`/api/accuracy` IGNORES the `?sport=` query param**
|
||
> (identical payload for mlb vs wnba — consumers must read `sports.{sport}`).
|
||
> **PARK×WEATHER CONFIRMED CHALLENGER, NOT CHAMPION** — attached in `snapshotService` (S77) onto
|
||
> `env_*`/`challenger_*` columns; `ledgerService.js:253` says it outright ("measured, never served"),
|
||
> and proj-v1.1, which consumes park/weather/platoon/arsenal, came back **NOT PROVEN** on 2026-07-29.
|
||
> MLB champion inputs are ONLY l5/l20 + opp_rank + rest + usage. **xwOBA = 0 files (absent); leash = 0
|
||
> (absent);** barrel/hard-hit/launch-angle/exit-velo/statcast/platoon all present but challenger-only.
|
||
> **RECOVERY ORDER (dependency-sequenced, full detail in the map):** 0 decide the grading BASIS
|
||
> (probability-vs-factor-index — a founder call; note that moving to probability grading legitimately
|
||
> changes the letter distribution because it is NEW INFORMATION, categorically unlike rescaling
|
||
> thresholds to mint A's) → 1 pick a runtime (**recommend PORTING `bayesian.py` to Node** rather than
|
||
> resurrecting a service every live path already routes around) → 2 wire Layer 2 FIRST (its
|
||
> prior/recent/context inputs are already available from `featureCache.getStatRows`, no new I/O) →
|
||
> 3 reconnect Layer 3 FORWARD (grade FROM probability; kills the `confidence` tautology) → 4 Layer 1
|
||
> Similarity (now feasible — S73-77 produce the context store it needs) → 5 per-sport efficiency
|
||
> scaling → 6 challenger promotion on ledger evidence only → 7 make "a sport" a real module →
|
||
> 8 NFL/CFB. **MLB = the REFERENCE MODULE** (deepest stack + a working settled feed + its own 526-row
|
||
> record); the map lists the exact contract a new sport must declare.
|
||
> **HONEST PER-SPORT LABELING: mechanism EXISTS, rule must be enforced** — every future sport renders
|
||
> `sports.{sport}` and reads ACCRUING until its own n≥20 clears; never `overall`, which would silently
|
||
> borrow MLB/WNBA credibility.
|
||
|
||
> ## 🧪 FULL-OUTPUT GRADE MAPPING + COLLAPSE COST 2026-07-30 (report-only) → **`specs/full-output-grade-mapping.md`**
|
||
> Track-B 1 of 3. Nothing built, reconnected, or promoted. **PREMISE CORRECTED AGAIN: the three-layer
|
||
> engine is BUILT but NOT WIRED and NOT DEPLOYED** (re-verified independently — 0 python refs in every
|
||
> grade-path file, 0 python lines in `Dockerfile`; there is no `engine1Adapter`, the real files are
|
||
> `utils/gradeAdapter.js` + `analyzeViaEngine1.js`). So no posterior/CI/similarity prior is produced
|
||
> today; Phase 1 could not inventory a live three-layer output and Phase 3 could not diff against one.
|
||
> Phase 3 was run instead against the collapse that ACTUALLY exists — which is **worse** than the
|
||
> premise describes, and measurable now.
|
||
> **THREE COLLAPSES (VERIFIED), not one.** (A) `estimateProbability` returns
|
||
> `{p_over,p_under,components{base,recency,weighted,opp_adjustment,home_adjustment,consistency_adjustment,cv}}`
|
||
> and `analyzeViaEngine1:521-524` keeps ONLY the scalar — `est.components` is attached to nothing.
|
||
> **(B) THE SEVERE ONE: `p_win` never reaches the grade at all** — the letter is `engine1`'s factor
|
||
> index and `engine1.js` has ZERO probability references, so the probability isn't collapsed INTO the
|
||
> grade, it's excluded FROM it. (C) `grade_thresholds.json` (PROBABILITY→GRADE) is read BACKWARDS to
|
||
> manufacture `confidence`. **0.2: market-efficiency scaling is NEVER COMPUTED** — a gap, not a
|
||
> second collapse.
|
||
> **🔴 PHASE 3 — THE COLLAPSE IS COSTLY ON MLB, AND UN-COLLAPSING DOES NOT HELP WNBA.** 354 settled
|
||
> rows carrying both the served letter and the locked pre-game `p_win` (forward test, NOT lookahead).
|
||
> Grade→outcome point-biserial r: **champion letter 0.0050 (p≈0.93, NULL)** vs **probability letter
|
||
> 0.1313 (p≈0.013)**. Per sport: **MLB champ 0.0686 n.s. vs prob 0.2356 (p≈0.0004, n=224)** —
|
||
> **WNBA champ −0.0986 vs prob −0.1258 (n=130, BOTH INVERSE)**. The pooled number is MLB's signal
|
||
> diluted by WNBA's inversion; this independently corroborates the 07-26 calibration finding that the
|
||
> champion does not discriminate on WNBA. **→ the full-output challenger must be MLB-FIRST; shipping
|
||
> it for WNBA on the pooled number would ship an anti-predictive grade.**
|
||
> **THE SERVED LETTER IS INVERTED BETWEEN ITS ONLY TWO POPULATED TIERS: B hits 52.4% (n=168), C hits
|
||
> 56.9% (n=174).** A user reading B as better than C is reading noise. Probability letters spread
|
||
> 30.8% (D) → 70.0% (B+), use 10-11 of 11 letters vs the champion's 3-4, and split ROI **−1.42%
|
||
> (A-family, n=78) vs −26.62% (C-/D/F, n=51) — a 25-point spread**. Banding is NOT the lossy part
|
||
> (0.1313 banded vs 0.1349 raw).
|
||
> **PHASE 2 — five EXPLICIT, FALSIFIABLE rules specced** (none assumed to be an improvement):
|
||
> R1 posterior→letter via the table read FORWARD · **R2 uncertainty grades DOWN, stated not smuggled:
|
||
> `p_adj = 0.5 + (p−0.5)·(1 − k·min(1, 1.96·SE/W0))`, k=1, W0=0.15 — falsifiable: rows R2 moves down
|
||
> must hit closer to their NEW band or R2 is WRONG and gets dropped** · R3 archetype adjusts the
|
||
> PROJECTION never the letter (drop if MAE doesn't improve) · R4 efficiency scales the THRESHOLD per
|
||
> sport, `E_sport` FIT from each sport's own record — success = equal hit rate per letter ACROSS
|
||
> sports · R5 abstain below min instances, never a default C.
|
||
> **🔴 HARD REQUIREMENT ON THE NEXT ORDER: R2/R3/R4 are UNMEASURABLE today — per-row instance count
|
||
> `n` is NOT stored on the ledger.** The challenger build MUST emit and persist `n`, `SE`, and the
|
||
> pre-adjustment `p`, or the CI-width and efficiency rules can never be adjudicated.
|
||
> **3.7 CONSUMERS:** nothing consumes a distribution, so a distribution-based grade is safe IF it
|
||
> still emits a letter; what changes is the letter DISTRIBUTION — `tierGating`, byGrade buckets
|
||
> (AccuracyBadge/ModelRecord/TierRecord), `outcomeService.gradeBucket`, `isAB` in hero +
|
||
> deskShowcase (the hero pool grows), `GRADE_RANK`/`selectTopGrades`, `capper_minimum_grade:'A-'`,
|
||
> newsletter templates. **A-GRADES: probability grading emits 53 A-family rows where the champion
|
||
> emitted 2. This is NOT the forbidden "rescale to mint A's"** — it is a measurably more informative
|
||
> basis (new information), and the A's earn it (A 65.5%, A+ 62.5%, B+ 70.0% vs D 30.8%) — **but A-
|
||
> hits 44.0%, breaking top-tier monotonicity, so the A-RATED marketing hold STAYS until the
|
||
> challenger's own forward record shows a monotone top tier.**
|
||
> **PHASE 4 RE-ADJUDICATION LIST (flagged, not re-run):** champion p_win→CLV 0.375 · the over-side
|
||
> skew audit · **proj-v1.1's "NOT PROVEN" — it was judged against the COLLAPSED champion, so its
|
||
> death is NOT final** · the C1 takeable-floor derivation · the 07-26 calibration curves (its WNBA
|
||
> finding is corroborated and looks robust; MLB needs re-running) · **ROI-by-grade — with B and C
|
||
> inverted, "MLB-C is the +4.57% profitable segment" is very likely an artifact of a meaningless
|
||
> letter, not a real segment** · `confidence` on every historical row (never use it as a weight).
|
||
|
||
> ## ⚖️ MARKET-EFFICIENCY SCALING CHECK 2026-07-30 (report-only): **VERDICT — FLAT. The specced mechanism is STRUCTURALLY ABSENT.**
|
||
> **⚠️ FIRST, A PREMISE CORRECTION (measured, not asserted).** This order opens with "the collapse
|
||
> concern was disproven — full-output and collapsed grades agree 100%." **That is not what the data
|
||
> shows.** On the 512 public ledger rows carrying both, collapsed (champion letter) and full-output
|
||
> (probability letter) agree on **17.8%**, and **33.8% differ by THREE OR MORE tiers**. The prior
|
||
> order's discrimination result stands unchanged: champion letter→outcome **r=0.0050 (p≈0.93, null)**
|
||
> vs probability letter **r=0.1313 (p≈0.013)**; MLB **0.0686 n.s. vs 0.2356 (p≈0.0004)**. Nothing in
|
||
> the repo changed between the two orders (HEAD is still `708f0fd`, clean tree), so this is not a
|
||
> stale reading — **the collapse was NOT a phantom, and the Phase-4 re-adjudication list stays open.**
|
||
> Reporting this straight because Phase 3.7 ("the model-soundness thread closes") is gated on an input
|
||
> that does not hold.
|
||
> **0.1 `marketEfficiency.js` DOES NOT EXIST — VERIFIED.** Zero occurrences of `market_efficiency`,
|
||
> `marketEfficiency`, or `efficiency_score` in `src/` or `web/src/`. **TRAP FLAGGED:** the spec's
|
||
> values 0.85 / 0.60 / 0.55 DO appear in `grade_thresholds.json` — but those are the **probability
|
||
> bands** (`A+ low 0.85`, `B low 0.60`, `B- low 0.55`), a coincidental numeric overlap, **not**
|
||
> efficiency scores. Do not mistake them for a wired implementation.
|
||
> **0.2 THE BASE THRESHOLDS DO NOT EXIST EITHER — VERIFIED.** There is no "MLB A: 5% edge / NBA A: 7%
|
||
> edge" table anywhere. **`engine1.js` contains ZERO references to `edge`** — the live grade is not an
|
||
> edge-%-vs-threshold comparison at all; it is the additive factor index
|
||
> (`idx = NEUTRAL_INDEX(3); idx += f.delta`). The specced rule
|
||
> `threshold = base_threshold × market_efficiency_score` therefore has **no host**: no base threshold,
|
||
> no efficiency score, and no threshold comparison to scale.
|
||
> **PHASE 1 — DISPOSITIVE: `engine1.js` contains ZERO references to `sport`.** `computeFactors({
|
||
> features, trap, consistency, prop })` receives no sport and no market; `prop` supplies only
|
||
> line/direction. **Sport cannot influence the grade, so per-market OR per-sport scaling is
|
||
> structurally impossible in the live grader** — this is not "specced but unwired", it is
|
||
> "the input isn't even in the function." It applies to the CHAMPION (what users see).
|
||
> **PHASE 2 — the matched-edge test is CONFOUNDED BY CONSTRUCTION (stated, not hidden).** Same-edge
|
||
> pairs return a mix of same/different grades, but that cannot discriminate flat-vs-scaled because
|
||
> **edge is not the grading input** — the SAME market emits both B and C at the same edge (mlb
|
||
> total_bases @ edge 20 → B n=5 AND C n=4). The aggregate that CAN discriminate: mean grade index by
|
||
> market — **wnba points 4.71 (mean edge 10.2) · mlb hits 4.58 (69.5) · mlb total_bases 4.32 (84.9)**.
|
||
> **WNBA points earns the HIGHEST mean grade on ONE-EIGHTH the edge.** Under the spec, the efficient
|
||
> market (NBA/WNBA points, 0.80) should grade TIGHTER than the lazy one (MLB total bases, 0.55) — the
|
||
> observed pattern is the OPPOSITE, and confirms the grade is untethered from both edge and market.
|
||
> **VERDICT: FLAT — one grading rule applied to every sport and every market.** An A does NOT mean the
|
||
> same thing across markets; it does not currently mean anything market-relative at all.
|
||
> **PHASE 3.6 — SCOPE, with an honest correction to "bounded".** The guardrail's framing is CORRECT
|
||
> and worth stating plainly: **a flat threshold is a grade-CALIBRATION gap. It does NOT touch the
|
||
> projection, and it does NOT touch the CLV edge** (the CLV/skew work measured `p_win`, never the
|
||
> letter), so this is **not** a third shadow-model alarm. **BUT the fix is NOT an independent bounded
|
||
> change in the current architecture:** since there is no edge-vs-threshold step to multiply, wiring
|
||
> efficiency scaling presupposes probability-based grading first. **It is exactly rule R4 of
|
||
> `specs/full-output-grade-mapping.md` and should ship as part of that MLB-first challenger, not as a
|
||
> standalone order** — with `E_sport` FIT from each sport's own accrued record, never hand-set, and
|
||
> success defined as equal realized hit-rate per letter across sports.
|
||
> **GATING NOTE FOR THE MULTI-SPORT VISION:** the premise is right that a new sport would grade wrong
|
||
> from day one — but the cause is broader than missing efficiency scaling. The grader takes no sport
|
||
> input at all, so NFL passing yards and MLB total bases would receive the identical rule. Sequence:
|
||
> probability grading (R1) → per-sport efficiency (R4) → then a new sport module.
|
||
|
||
> ## 🏷️ TAKEABLE TAGGING BUILT · EFFICIENCY CHALLENGER **BLOCKED** 2026-07-31 → `specs/takeable-tagging.md`
|
||
> **Champion grade UNCHANGED — verified byte-identical by diff** (`src/services/intelligence/`,
|
||
> `gradeSlateService`, `gradeAdapter`, `grade_thresholds.json` = zero diff). Additive tags only.
|
||
> **🔴 PART A — THE EFFICIENCY CHALLENGER COULD NOT BE BUILT. Review Zero came back ABSENT on all
|
||
> THREE inputs.** (0.1) efficiency scores **DO NOT EXIST** — zero occurrences of
|
||
> `market_efficiency`/`marketEfficiency`/`efficiency_score` in `src/` or `web/src/`; `src/config/` has
|
||
> no efficiency file (its one `0.80` is `providers.js WARN_PCT`, a quota warning). (0.2) base
|
||
> thresholds **DO NOT EXIST** — no "MLB A:5% / NBA A:7%" table; **`engine1.js` has zero `edge`
|
||
> references**, so the grade is not an edge-vs-threshold comparison at all. (0.3) **the ±0.05 additive
|
||
> efficiency nudge DOES NOT EXIST** — searched the whole grade path; the only `0.05`s are
|
||
> `featureCache.teammate_absence_bump` (a feature), a `bvp_advantage` matchup cutoff, and
|
||
> `p*0.9 + 0.05` in `probabilityEstimator` (the `0.5×0.1` term of the shrink-toward-0.5). **There is
|
||
> no additive efficiency application to replace.**
|
||
> **WHY IT STOPPED RATHER THAN GUESSED:** the order requires a challenger differing from the champion
|
||
> in EXACTLY ONE thing (multiplicative vs additive). There is no additive scaling to swap, no base
|
||
> threshold to multiply, and **`engine1.js` has zero `sport` references** so market cannot reach the
|
||
> grade. A threshold must exist FIRST — that is **R1 of `specs/full-output-grade-mapping.md`**, an
|
||
> explicitly held separate order. Shipping R1+R4 together would make the Phase-3 delta report
|
||
> **actively misleading**: the re-letter would be driven mostly by the switch to probability grading
|
||
> while being presented to Kev as the efficiency fix. **(0.4) COVERAGE would sink it anyway** — the
|
||
> spec names 5 scores; the live ledger has **11 markets** (MLB hits 376 · doubles 247 · total_bases 44
|
||
> · stolen_bases 17 · earned_runs 6 · home_runs 4 · outs 2; WNBA points 246 · rebounds 156 · assists
|
||
> 142 · threes 10) and **only MLB total_bases maps to a specced score**. MLB K's and NFL passing don't
|
||
> appear at all. **9 of 11 markets have NO score**, so "all scored markets" can't be met without
|
||
> inventing 9 numbers — fabricated calibration.
|
||
> **✅ PART B — LEDGER TAKEABLE TAGGING SHIPPED (the deferred C2).** New
|
||
> `src/config/takeableStandard.js`: `isLedgerTakeable(p) = p >= floor` — **FLOOR on the minus side,
|
||
> UNCAPPED plus**. **🔴 Deliberately NOT `valueEngine.isTakeable`**: that is the −160..+200 PROMOTION
|
||
> band ("may we hero it?"); this is the RECORD standard ("could a bettor have taken it?"). **A +400
|
||
> prop is NOT promotable but IS takeable — both true at once**, and a test asserts the two diverge on
|
||
> the plus side and agree at the minus floor so they can never quietly merge. **Absent price → `null`,
|
||
> never `false`** (without the strict guard `Number(null) === 0` is ≥ −160 and a MISSING price would be
|
||
> tagged takeable — fabricated data in the record). **The floor is POLICY, not derived, and is labelled
|
||
> so everywhere** (C1 could not derive one: every bucket's 95% ROI interval contained zero, the curve
|
||
> was non-monotone, and 661-2,285 rows/bucket are needed vs 8-71 actual). Each row stores
|
||
> **`takeable_floor`** so a re-derivation can re-tag safely instead of guessing the standard.
|
||
> **Migration 034 applied + tracked** (`takeable boolean`, `takeable_floor numeric`, partial index).
|
||
> Forward tagging in `ledgerService` at row build; backfill in one statement.
|
||
> **BACKFILL VERIFIED IN PROD: 1,254 rows → 1,246 tagged (781 takeable / 465 below floor), 8 NULL, and
|
||
> `null_despite_price = 0`** — every priced row tagged, the 8 NULLs genuinely priceless. One distinct
|
||
> floor (−160). **Settled 1,163 and graded 1,254 both unchanged** — no outcome, grade, or price touched.
|
||
> **🚫 PART C — THE MODEL-VERSION BOUNDARY TAG WAS DELIBERATELY NOT APPLIED.** It is meant to separate
|
||
> pre-fix (additive) from post-fix (multiplicative) grades — **no scaling change shipped, so no such
|
||
> boundary exists**, and stamping one would mark a model transition that never happened: a fabricated
|
||
> era in the permanent record. `src/config/modelEras.js` (S64) is its home the day a real one promotes.
|
||
> **FLOOR: 312 suites / 3,890 tests green (8 new), web build exit 0.** Grade visuals auth-gated →
|
||
> Chrome audit; this order changed nothing visual.
|
||
|
||
> ## 🧪 EDGE-SHADING CHALLENGER BUILT + MEASURED 2026-07-31 → `specs/edge-shading-challenger.md`
|
||
> **Verdict: the mechanic is sound and built — but it does NOT fix the flooding, and the measurement
|
||
> proves the bug was never the multiply's placement.** Champion byte-identical (diff empty across
|
||
> `intelligence/`, `gradeSlateService`, `gradeAdapter`, `snapshotService`, `grade_thresholds.json`).
|
||
> Nothing promoted, nothing re-lettered.
|
||
> **0.1 THIRD RE-CHECK: the flat additive ±0.05 efficiency nudge STILL DOES NOT EXIST** — no
|
||
> `efficiency` reference anywhere in `src/services/intelligence/` or `gradeAdapter`. There was nothing
|
||
> to remove. **0.2 no live base threshold exists**, so the challenger DEFINES one that never moves
|
||
> (A+≥10, A≥5, B≥3, C≥1, D≥0, F<0), identical for every market — the spec's per-sport 5%/7% variant is
|
||
> deliberately unused because a bar differing by sport IS a moving bar.
|
||
> **BUILT** (`src/services/challengers/efficiencyShading.js`, measured-never-served, fence-tested):
|
||
> `adjusted = raw_edge × f(e)`, `f(e) = E_SOFTEST/e` **bounded to (0,1]** — soft intact (never
|
||
> amplified), sharp shaded toward but not past zero (`f(.60)=.917`, `f(.80)=.688`, `f(.85)=.647`),
|
||
> unscored → `f=1` FLAGGED, absent edge → null. **Cross-market behaviour is unit-proven: the same raw
|
||
> 6% edge grades A in soft `mlb:total_bases` and B in sharp `nba:points`.**
|
||
> **🔴 PHASE 2.5 ANSWER: THE FLOODING IS NOT GONE.** On 1,250 live rows the challenger grades **79.0%
|
||
> A / 80.9% A-or-B** (MLB **93.4% A**) against the champion's 0.2% A / 58.1% A-B. The naive formula's
|
||
> 86% A/B is not meaningfully improved on.
|
||
> **TWO FINDINGS EXPLAIN IT — and they are the point of the order. (1) THE SHADING IS A NO-OP ON THE
|
||
> LIVE BOARD: `rows_actually_shaded = 0` of 1,250.** 96.5% (1,206) of rows are **UNSCORED** → `f=1`,
|
||
> and the one scored market present (`mlb:total_bases`) is the ANCHOR so its `f` is 1.0 by
|
||
> construction. `mlb:strikeouts` and `nba:points` **do not appear in the ledger at all** — our
|
||
> basketball is **wnba**, not `nba`. **Challenger vs baseline: 0 rows changed.**
|
||
> **(2) PLACEMENT WAS NEVER THE BUG — THE INPUT SCALE IS.** Against a fixed 5% bar the **RAW** edge
|
||
> already clears A on **100% of MLB doubles, 100% stolen_bases, 89.6% hits, 88.6% total_bases —
|
||
> before any shading**. MLB's median raw edge is **60%, twelve times the bar**. **DECISIVE TEST: apply
|
||
> the SHARPEST score in the whole spec (f=0.647) to EVERY row — the maximum the design permits — and
|
||
> 75.8% STILL clear A (MLB 91.7%).** Since `f` is bounded ≤1, **no achievable shading can close a 12×
|
||
> overshoot**; moving the multiply from threshold to edge cannot change that. This is `edge_pct`
|
||
> behaving exactly as the 07-29 diagnosis described — a price-free `(proj−line)/line` gap whose scale
|
||
> is a FUNCTION OF LINE SIZE. **It is not a betting edge, so no fixed betting-edge bar is meaningful
|
||
> against it.**
|
||
> **2.6 efficient-market over-suppression: CANNOT DETERMINE** — zero live rows are shaded, so there is
|
||
> no efficient market in the data to over-suppress.
|
||
> **PHASE 3:** takeable tagging was completed in the PREVIOUS order (migration 034, 1,246/1,254 rows,
|
||
> floor −160 recorded per row) — **not repeated, it would be a no-op**. The model-version boundary is
|
||
> **again NOT applied**: nothing promoted, so no boundary exists; stamping one would record a
|
||
> transition that never happened.
|
||
> **UNBLOCKING NEEDS THE INPUT REPLACED, NOT THE MULTIPLY MOVED:** a bar-based grade needs a real edge
|
||
> — model probability vs the de-vigged market price (**`p_win` vs `fair_prob`, both already computed
|
||
> and stored**) — plus scores **FIT from our own record** for the markets we actually grade (9 of 11
|
||
> live markets have none; inventing them would be fabricated calibration). The mechanic itself is
|
||
> sound and already built, so that re-run is cheap.
|
||
> **FLOOR: 313 suites / 3,899 tests green (9 new), web build exit 0.**
|
||
|
||
> ## 🛑 PROMOTION GATE **NOT PASSED — NO FLIP** · ORDER B SHIPPED 2026-07-31 → `specs/edge-pct-display-retirement.md`
|
||
> **The champion grade was NOT flipped.** Grade path byte-identical (diff empty across
|
||
> `intelligence/`, `gradeSlateService`, `snapshotService`). Projection, `p_win`, CLV untouched.
|
||
> **THREE OF FOUR REVIEW-ZERO PREREQUISITES FAIL — and the order's own rule for 0.2 is "STOP".**
|
||
> **0.1** scores are **ESTIMATED priors** from the founding spec, not measured — and the premise's
|
||
> cited values **are not in the code**: the module holds `nba:points .80` / `mlb:total_bases .55`;
|
||
> there is **no NBA 0.72 and no WNBA score at all**. **0.2 ❌ THE VERSION-BOUNDARY TAG NEVER LANDED**
|
||
> — `config/modelEras.js` has ZERO shading references; it was deliberately not applied twice
|
||
> (nothing had been promoted) and reported both times. **0.3 ❌ NO ROLLBACK FLAG EXISTS** — zero
|
||
> occurrences of `SHADING_ENABLED`/`EDGE_SHADING`/`shadingEnabled` in `src/`. **0.4 ✅** takeable tags
|
||
> landed (migration 034, 1,246/1,254).
|
||
> **🔴 AND THE APPROVED DELTA DOES NOT MATCH THE MEASURED ONE.** Approved: "43.6% re-letter, efficient
|
||
> tighten / soft hold." Measured on all 1,250 live rows: **97.4% change (1,217), 79.8% move UP**, 17.6%
|
||
> down → **79.0% A-family (MLB 93.4%)** vs the champion's 0.2%. **And `rows_actually_shaded = 0` of
|
||
> 1,250** — 96.5% of markets are unscored (f=1) and the one scored market present is the ANCHOR
|
||
> (f=1.0 by construction). **The entire re-letter comes from switching to edge-vs-fixed-bar grading,
|
||
> NOT from efficiency shading, which is inert on this board** — an unapproved grading-basis change
|
||
> riding along, which the order's own "no new scaling changes riding along" guardrail forbids.
|
||
> **WHY THIS IS A REFUSAL, NOT A DELAY:** flipping would re-letter **97.4% of an append-only public
|
||
> record**, move 79.8% of grades UP and mint A's on 79% of the board — on a letter whose measured
|
||
> correlation with outcomes is **r ≈ 0.005**. That is precisely what the permanent founder ruling
|
||
> forbids ("a relabelled B sold as an A corrupts an append-only ledger"). **To unblock:** land the
|
||
> version tag, add a rollback flag, **replace the INPUT** (`p_win` vs `fair_prob` — a real price-aware
|
||
> edge, both already stored) so a fixed bar is meaningful, and fit scores for markets we actually
|
||
> grade. Then re-measure and re-approve against the true delta.
|
||
> **✅ ORDER B SHIPPED (independent, and a live falsehood).** `edge_pct` display retired from
|
||
> **GradeResultCard** (confidence-strip figure, `EDGE` cell → honest-absent `—`, alt-ladder rung — the
|
||
> rung's LINE and GRADE still render) and **SoccerGradeResult**. **DeskShowcase KEPT** (already honest).
|
||
> **COMPUTATION + THE SIGNED-EDGE SORT FALLBACK SURVIVE** (`edgePctFor`, `computeEdge`,
|
||
> `descNullsLast(a.edge,b.edge)`) — deleting them would re-break the sort fixed 2026-07-29; a test
|
||
> asserts all three survive and that the sort still orders `agrees → disagrees → absent`.
|
||
> **Two build-breakers the retirement caused were found and fixed** (orphaned `edgeColor` import,
|
||
> orphaned `edge_pct` destructure — `edge_pct` stays on the props CONTRACT). **Two pre-existing tests
|
||
> SUPERSEDED, not deleted** (`ds4Billboards`, `colorContract`): they asserted the edge figure is
|
||
> sign-coloured; they now assert the stronger property that **no edge percentage renders at all**.
|
||
> `edgeColor` itself is untouched and still enforced for `MarketBreadth`.
|
||
> **FLOOR: 314 suites / 3,908 tests green (9 new), web build exit 0.**
|
||
|
||
> ## 📋 PRE-AUDIT STATUS PULL 2026-07-31 (report-only) → **`specs/pre-audit-status-pull.md`**
|
||
> Matrix re-derived from repo `10aaaeb` + live prod probes; nothing inherited, nothing built.
|
||
> **🔴 0.1 — THE 92.9% RE-LETTER DID NOT HAPPEN. NOTHING WAS PROMOTED.** There is no live
|
||
> re-lettering to explain. Verified three ways: **HEAD is `10aaaeb`** ("promotion gate not passed —
|
||
> no flip"), clean tree, nothing landed since · **`efficiencyShading` is imported by ZERO production
|
||
> files** (still measured-never-served) · **live grades carry 0.0% A-family** (MLB n=5 `{B:1,C:4}`,
|
||
> WNBA n=25 `{B:15,C:10}`; a promoted shading model would show ~79% A). **Neither 92.9% nor 43.6% is
|
||
> a figure this project measured** — the challenger's real numbers were 97.4% would-change / 79.8% UP
|
||
> / 79.0% A-family, with **0 of 1,250 rows actually shaded**. The live board is the unchanged additive
|
||
> factor-index champion.
|
||
> **0.2 TAGS — takeable ✅ LANDED** (1,245/1,250 tagged, `takeable_floor` on all 1,245, one floor
|
||
> −160; the 5 untagged have no locked price = honest NULL). **Model-version boundary ⚪ ABSENT AND
|
||
> CORRECTLY SO** — `model_version` holds only the S64 era labels; **there was no recalibration, so
|
||
> there is no boundary to mark.** NOT a pre-audit gap; it becomes a hard prerequisite the day a
|
||
> promotion actually happens.
|
||
> **PHASE 1 — 15 of 26 FULLY DONE.** **Book comparison RESOLVED** (row 18 — `BookComparisonPanel` has
|
||
> a real importer; the headline dead component is now routed). **Grade card HONEST improved** — the
|
||
> edge_pct display was retired 07-31, `EDGE` renders `—`.
|
||
> **DESIGNED-but-not-BUILT: only Offseason** (`/offseason`→404 with a full artboard).
|
||
> **BUILT-but-DEAD:** `ShareCard` (**0 real importers**), `MobileEdgeBoard` (0 — correctly hidden),
|
||
> `DemoScan` (0). **LIVE-but-not-HONEST: NONE FOUND** — every known fabrication removed or hidden.
|
||
> **ORPHANS UNRESOLVED (live, 0 nav links):** `/intelligence`, `/soccer`, `/marketplace`,
|
||
> `/notifications`, `/slip`, `/compare`, and `/parlay` (drawer-only). `/system` + `/offseason` are
|
||
> **404**; `/terminal`→307 by design.
|
||
> **PHASE 2 — DESIGN IS NOT COMPLETE. Three surfaces are shipped but UNDESIGNED** (design is the gap,
|
||
> not build): **Live tracking · Slip reader · Newsletter** — all spec-only, no artboard in the
|
||
> S-series bundle. Partial: Parlay lab (only "PARLAY BUILDER" inside `System.dc.html`), Compare,
|
||
> Alt-line ladder. **The core funnel IS fully designed** (landing → dashboard → scan → grade card →
|
||
> ledger → pricing).
|
||
> **PHASE 3 — CHROME AUDIT MANIFEST assembled: 11 items, each with the session state to check in.**
|
||
> **🔑 FOUR REQUIRE AN ENTITLED (Desk) SESSION — only Kev can drive them:** the grade card's entitled
|
||
> half, the **alt-line ladder** (Desk-gated), the **entitled top-graded board**, and the **ledger**.
|
||
> **#4 is the highest-value item**: the strip-after-rank boundary was proven in the payload
|
||
> (logged-out gets correct ORDER with no paid values) but its **entitled rendering has never been
|
||
> seen** — that is the exact path unverifiable from the backend.
|
||
|
||
> ## 🛑 RECALIBRATION BOUNDARY **NOT WRITTEN** (no promotion exists) + BUILD TRIAGE 2026-07-31 → `specs/incomplete-surface-triage.md`
|
||
> The order asked to tag grades `pre/post edge-shading` at "the true promotion timestamp." **There is
|
||
> no such timestamp**, and writing the marker would insert a **fabricated model transition into an
|
||
> append-only public record** — the exact corruption the order exists to prevent. **THREE INDEPENDENT
|
||
> PRODUCTION PROOFS:** (A) **`model_snapshots.code_sha`** — every sha that ran the pipeline in the last
|
||
> 5 days is a documented commit from this session (**`f3bf300` currently live**, then `f310608`,
|
||
> `9b5235c`, `b8ee216`, `afb56b1`, `3592aba`, `914a057`), **all stamped `engine1@2026-07-20`** — no
|
||
> promotion commit exists. (B) **daily A-family share 07-24→07-30 = 0.0 · 0.0 · 0.0 · 0.0 · 1.0 · 1.6 ·
|
||
> 0.0** — **flat at zero, no step change on any date**; a 92.9% re-letter would have driven the board
|
||
> to ~79-93% A overnight. (C) `efficiencyShading` has **zero production importers**; HEAD `d54eca0`,
|
||
> clean tree.
|
||
> **THE DELTA, WITH ROW COUNTS: neither 92.9% nor 43.6% is attested in any measurement here, and no
|
||
> re-lettering occurred at any scale** — so "partial slate vs full board" cannot explain a gap that
|
||
> does not exist. The only measured numbers, on **all 1,250** rows (the FULL board): **97.4% would
|
||
> change · 79.8% UP · 79.0% A-family · and 0 of 1,250 rows actually shaded.** The hypothetical
|
||
> re-letter would have come **entirely from an unapproved grading-basis switch**, which is why the flip
|
||
> was refused. **Measurement integrity needs nothing new today** — `model_version` already separates the
|
||
> S64 eras and takeable tags are complete (1,245/1,250). The boundary becomes a hard prerequisite **the
|
||
> day a promotion actually ships**; `config/modelEras.js` is its home.
|
||
> **BUILD TRIAGE — 14 surfaces classified FROM CODE, five dependency-ordered waves, nothing cut:**
|
||
> **W1 pure wiring** (`/intelligence` 160ln real 0-placeholder, `/slip` 526ln, `/parlay` direct link,
|
||
> `/marketplace` after a copy-honesty pass — it advertises unbuilt capability). **W2 design-only gaps**
|
||
> — artboards for **Live tracking · Slip reader · Newsletter** (all built+live+honest; DESIGN is the
|
||
> only missing column). **W3 self-contained** — `/compare` real two-player fetch, `ShareCard` host,
|
||
> `/notifications` (`RouteStub`, needs a store + delivery). **W4 MODEL-GATED** — price-triplet MODEL leg
|
||
> + calibration/edge board **both need `p_win` vs `fair_prob`; building either on `edge_pct` would
|
||
> re-ship the retired 620% lie.** **W5 sport/quota-gated** — **`/soccer` blocked on odds-api quota
|
||
> (0/500), not code**, plus a settled feed; `/system` (decide the `/intelligence` overlap first);
|
||
> Offseason.
|
||
> **MATRIX CORRECTIONS:** Live tracking / Newsletter / Slip reader are built-live-honest and incomplete
|
||
> ONLY on DESIGN · `/parlay` is drawer-reachable, not unreachable · `/compare` is already honest.
|
||
> **No live surface is showing fabricated data today.**
|
||
|
||
> ## 🔌 WAVE 1 WIRED 2026-07-31 → `specs/wave1-wiring.md` — three surfaces proven, then linked
|
||
> No grade/ledger/model/scoring change (diff empty across `intelligence/`, `ledgerService`,
|
||
> `outcomeService`, `gradeSlateService`).
|
||
> **REVIEW ZERO PROVED EACH BEFORE WIRING.** **0.1 `/intelligence` vs `/system` are NOT duplicates** —
|
||
> `System.dc.html` is a MULTI-surface artboard (TERMINAL + INTELLIGENCE + WIRE sections), not a design
|
||
> for a distinct `/system` route; its INTELLIGENCE section is **already realised** as the live
|
||
> `/intelligence` page, so **no `/system` should be built as a second copy — the prod 404 is correct**.
|
||
> **0.2 `/intelligence` gates SERVER-side, not by blur** — the proxy requires auth and limits by tier
|
||
> (**desk 50 / non-desk 8**), returning **401 to anonymous, verified live** → wiring it leaks nothing.
|
||
> **0.3 `/slip` parses a REAL slip end to end** — DraftKings fixture → **3/3 legs, `needs_review:
|
||
> false`**, `Aaron Judge · total_bases · over 1.5 @ -115`. **Honest limit: parsers are layout-rigid; an
|
||
> unsupported layout yields ZERO legs, never wrong ones** (never-guess) — so real-world OCR hit-rate
|
||
> across layouts is **CANNOT DETERMINE** until user slips arrive. **0.4 `/parlay` direct route hits the
|
||
> real correlation builder** on the same `ParlayContext` as the drawer. **0.5 `/marketplace` made NO
|
||
> performance/profit claim** (the worse category — clean) and its capture was **already real**
|
||
> (`/api/waitlist` → upsert to a `waitlist` table); the gap was TENSE, not fabrication.
|
||
> **WIRED:** Nav MORE gains **Intelligence · Slip Reader · Marketplace**; **Parlay Lab re-pointed from
|
||
> the `#parlay` drawer hash to `/parlay`** (drawer unaffected — `ParlayPanel` stays mounted with its
|
||
> floating badge). **GATING:** `/intelligence` added to `GATED_ROUTES` (its feed 401s signed-out, so an
|
||
> ungated link would land visitors on a permanently empty page). **`/parlay` stays OPEN deliberately** —
|
||
> it is the free parlay funnel; gating it would be a monetization regression.
|
||
> **`/marketplace` HONESTY PASS:** every item body now opens **"Not built yet." / "Not written yet." /
|
||
> "Not produced yet."** + what is planned; the subhead states plainly it is **"not a purchase, not a
|
||
> pre-order, and not a promise of a ship date"**; the playbook item carries **"No profit claim, no
|
||
> promised return."** Capture stays REAL — no fake button. Unit-locked (every body must carry a
|
||
> not-yet marker; no guaranteed/win-rate/ROI/beat-the-book language).
|
||
> **FLOOR: 315 suites / 3,920 tests green (12 new), web build exit 0.** **FINGERPRINT:** deploy
|
||
> CONFIRMED (new marketplace copy live; all four routes 200; intelligence feed still 401s anon).
|
||
> **⚠️ The nav ENTRIES could NOT be verified by curl and are not claimed** — a control test shows the
|
||
> PRE-EXISTING `"Parlay Lab"`/`"The Report"` labels are **equally absent** from the 18 sampled chunks
|
||
> (the dropdown is client-rendered from a lazily-fetched chunk), so this is a probe limitation, not a
|
||
> failure → **added to the Chrome manifest** (confirm the dropdown shows the four entries, and that a
|
||
> signed-out Intelligence click lands on `/login`).
|
||
|
||
> ## 🔀 WAVE 3 — /compare BUILT · RESOLUTION TAIL **SCOPED, NOT SHIPPED** 2026-07-31 → `specs/wave3-compare-and-resolution-tail.md`
|
||
> No grade/ledger/scoring change (diff empty across `intelligence/`, `ledgerService`, `outcomeService`).
|
||
> **🔴 0.3/0.4 — THE RESOLUTION TAIL DOES NOT FIRE.** The resolver is `POST /api/grading/resolve`
|
||
> (`routes/grading.js:208`, "called by the ESPN poller at FINAL"); its fanout at `:356-371` holds
|
||
> webPush + telegram + discord. State of all five outputs: **share-card generation = SPEC'D-NOT-BUILT**
|
||
> (`grep shareCard grading.js` = **0** — not in the fanout at all; `shareCards/renderer.js` has **zero
|
||
> callers**) · **push = BUILT-NOT-FIRING** (gated on `webPush.configured()`; **`push_subscriptions` 0
|
||
> rows, `user_notifications` 0 rows** — nothing ever subscribed or delivered) · **Telegram = BUILT-NOT-
|
||
> FIRING** · **Discord = BUILT-NOT-FIRING** · **recap = SPEC'D-NOT-BUILT** (no recap file in `src/`).
|
||
> **AND THE WHOLE TAIL IS UNREACHABLE — nothing calls `/api/grading/resolve`; there is NO ESPN poller
|
||
> in the repo.** The live settle path is the scheduler's `settleAllOutcomes`/`settleAllLedgers`, which
|
||
> fans out to **`opsNotify` only** (ops alerts), with **zero user-facing output**. Per the order's own
|
||
> rule, **ShareCard · /notifications · result posts · recap are ALL SCOPED, none shipped** — no dead
|
||
> shells over a silent pipeline.
|
||
> **✅ /compare BUILT.** Semantics (0.2): a **same-market head-to-head** — two players, every row a
|
||
> measure BOTH sides are scored on, aligned by `alignRows`; deliberately NOT two disconnected graded
|
||
> props. Reads the live `/api/stats/player/:name?sport=` aggregate. **Honest-absent three ways:** an
|
||
> unresolved side reads **NO DATA** while the other still renders · a measure only one side has renders
|
||
> **`—`, never 0** · if NEITHER resolves the page **refuses to compare**. **NO VERDICT** — it shows the
|
||
> measures and says the reader draws the call.
|
||
> **FINGERPRINT:** deploy boundary captured (attempts 1-4 `IN DEVELOPMENT`, attempt 5 gone). Live
|
||
> head-to-head through the deployed proxy — **Judge vs Ohtani: AVG .248/.289 · HR 17/23 · RBI 38/65 ·
|
||
> OPS .908/.931 · GP 59/103, 5 of 5 shared measures**; bogus name → `found:false` → NO DATA.
|
||
> **⚠️ HARNESS ERROR RECORDED:** my first live check reported `found=false` for both real players —
|
||
> that was **my script's bare `except` returning a fabricated negative**, not the service (curl via
|
||
> both proxy and Express: `found:true`, 3/3). Same class as `Number(null)===0`: a verification harness
|
||
> must fail loudly, never return a plausible-looking false.
|
||
> **FLOOR: 316 suites / 3,930 tests green (10 new), web build exit 0.** Two pre-existing tests
|
||
> (`vyndrPhaseE`, `vyndrParityQA`) asserted the in-dev placeholder — **superseded, not deleted**; they
|
||
> now assert the stronger properties on the real page (live fetch, no sample players, NO VERDICT,
|
||
> NO DATA, "not a zero").
|
||
|
||
- **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows`→`closing_captures` (append-only,
|
||
provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221`
|
||
+ internal endpoint; `ledgerService.attachClosingProb`→`closing_prob` (de-vigs both raw sides,
|
||
write-once, honest-absent `market_unavailable_reason`) via `snapshotScheduler:310`.
|
||
- **Honest-absent is already implemented + is the invalidation marker:** 651,624 capture rows
|
||
(86% `missed`, no fabricated price); ledger has **59 genuine `closing_prob`, 870
|
||
`market_unavailable_reason`**. Of 841 old-`clv` rows, **791 are already marked unavailable**;
|
||
**745 (89%) of old `clv` = exactly 0/'flat'** (the fake "never captured" signal). Old broken
|
||
`clv`/`closing_line` left in place, unread by the new instrument (but STILL consumed by ledger
|
||
ROW_COLUMNS/UI — repointing display is a FUTURE order, not done here).
|
||
- **0.4 GATE:** captures **differ from locked** (10/10 comparable rows moved) → NOT a
|
||
captures==locked timing defect. BUT the capture **86%-misses** → CLV is STARVED: only **~10-50
|
||
usable rows**, and only **10 have both `closing_prob` + `fair_prob_lock` (all WNBA, 0 MLB)**.
|
||
- **LINE-LAG HYPOTHESIS → CANNOT DETERMINE.** The 10 WNBA comparable rows move ~1pp, roughly
|
||
symmetric (mean CLV gap −0.008 over / −0.011 under). **The +4.57% MLB-C and the over/under
|
||
asymmetry are NEITHER confirmed nor refuted by CLV** — there is essentially no MLB CLV data.
|
||
- **The real open defect (separate order):** WHY 86% of captures miss (props leave the feed before
|
||
the 45-min window? matching failure?). Until that's fixed the instrument stays starved.
|
||
- Docs corrected: CLAUDE.md (WNBA settles, not MLB-only), STATE.md grade_11 (model_snapshots only).
|
||
Orphan `next start -p 3111` (6-day-old, the "1 shell running") killed.
|
||
|
||
---
|
||
# 📐 CALIBRATION DIAGNOSIS — measured 2026-07-26 (read-only; nothing changed)
|
||
*Run on BOTH populations: unselected both-sides (model_snapshots, p_win per side) vs published selected (ledger). Never pooled across sport. Full detail in chat log of that order.*
|
||
|
||
**Feasibility caveats that govern every number below:**
|
||
- p_win window ≈ **last 10 games** (`probabilityEstimator`, last-5 double-weighted) → **SE ≈ 0.16 at p=0.5** (noisy estimator).
|
||
- Continuous p_win coverage is thin+recent (revived ~07-20): **~18 MLB games, ~5 WNBA games** of joinable calibration data. WNBA calibration is **effectively anecdotal (~5 games)**; MLB is suggestive not confident (~18 games).
|
||
- Effective sample generally: MLB 31 games / 190 players, **WNBA 12 games / 77 players** — prop outcomes cluster within a game; do not treat rows as independent.
|
||
|
||
**Findings (VERIFIED where measured; sample caveats attached):**
|
||
1. **Selection bias, measured directly** = both-sides→selected Brier: **MLB 0.212→0.233, WNBA 0.258→0.278 (~+0.02 each)**. Published grade is inflated ~1 sub-tier by the side-pick: mean idx **MLB C→C+**, **WNBA C+→B−**. Honest A-or-better ≈ 0 → the "caps at ~B+/A-" story survives (optimistic if anything).
|
||
2. **MLB calibration** (both-sides): Reliability 0.010 (good), Resolution 0.047 (moderate), Uncertainty 0.249 → **discriminates + roughly calibrated**; middle deciles near-perfect. SELECTED side is **overconfident, growing with p** (+0.02 at p<.5 → **+0.19 at p≥.8**).
|
||
3. **WNBA calibration** (both-sides): Brier 0.258 **> uncertainty 0.25 = worse than always-predicting-0.5**; Resolution 0.019 (near-noise), Reliability 0.027 (poor). Realized ~flat vs predicted. **Champion does not discriminate on WNBA** — but ~5 games, so unproven.
|
||
4. **Null baselines (real prices):** **unders lose ~−14% flat both sports; overs near break-even** (MLB −2.8%, WNBA −1.4%). **Ledger strongly supports "no unders by default."** MLB model beats always-over (MLB-C +4.6% vs −2.8%); **WNBA model LOSES to always-over** (−5% vs −1.4%). Both within noise at 31/12 games.
|
||
5. **ev_pct = p_win×decimal−1 uses the overconfident p_win** → inflated; hero ranks on it → picks the most-overconfident reads. No current metric reliably predicts ROI; price-aware EV requires recalibrated p_win first.
|
||
|
||
**BUILD FORK (declared before data; sample-tempered):** MLB → *resolution good / reliability poor-when-selected* → recalibration mapping + EV ranking — **but ~18 games can't validate a holdout yet (test half 15 games)**. WNBA → *resolution poor* → champion doesn't discriminate / price-aware replacement — **but ~5 games = accrue volume before concluding.** **CLV redirect ships next regardless.** Most robust actionable finding now: **kill default unders.**
|
||
|
||
|
||
---
|
||
|
||
# 🧭 CURRENT STATUS + OPEN ITEMS — orientation block (2026-07-20, ~03:20 UTC)
|
||
*Written to orient a fresh session. Ground truth, not optimism. Everything below
|
||
was verified live unless explicitly marked UNVERIFIED.*
|
||
|
||
## What shipped tonight (all deployed + fingerprinted)
|
||
|
||
| # | Shipped | Proof |
|
||
|---|---|---|
|
||
| 1 | **Probability layer revived** — it was DEAD in prod | `p_win`/`ev_pct`/`model_odds`/`value` on **32/32** live grades (was 0/8) |
|
||
| 2 | **Value engine Arc 1** — de-vig, EV, value triplet, hero v2 | live: book −140 · fair −125 · model −109, `value:false` correctly |
|
||
| 3 | **Grade-range work** — L20 symmetry, `refreshTeamStats` wired, consistency CV guard, `confidence_basis`, dead `mlbGrader.js` removed | suite 279/3325; D emitted on real WNBA data locally |
|
||
| 4 | **Backup durable ON-BOX** | dump read back: **643 rows in dump == 643 live**, 18 MB |
|
||
| 5 | **`model_snapshots` retention LIVE** (migration 025) | **100 rows** captured over 2 cycles, incl. **36 refusals** |
|
||
| 6 | ESPN team-stats parser fixed (`buckets is not iterable`, 0/15 → 15/15) | live refresh |
|
||
|
||
### ⚠️ Two honest qualifiers on the above
|
||
- **A STILL DOES NOT EMIT IN PRODUCTION.** The 11-step grade is stored in
|
||
`model_snapshots.grade_11` ONLY (CORRECTED 2026-07-26: it is NOT in
|
||
`ledger_entries` — `_grade_11` is deleted at `gradeSlateService.js:97` before
|
||
the ledger write; sub-tier resolution on a SETTLED ledger row is recoverable
|
||
only by joining model_snapshots). A/D are *arithmetically* reachable + locked
|
||
by tests, but the
|
||
±1.0 opponent factor is still dead because `opp_rank_stat` is underivable from
|
||
ESPN (see open item 6). Live boards remain **B/C only**.
|
||
**→ The "A-RATED" marketing hold STANDS** (`AccuracyBadge` correctly falls
|
||
through to "MODEL · X% HIT"; `TopSignals` self-hides). Do not promote A-rated.
|
||
- **EV IS OVERCONFIDENT AND UNVALIDATED.** First captured values include
|
||
**+62.1% / +61% / +56.9%** — real markets don't offer that. Cause: `p_win`
|
||
clamps at `PROB_CEIL` 0.95 off ~10 games. **Hero v2 ranks on `ev_pct`, so it
|
||
picks the MOST overconfident read.** Do not surface EV as truth until
|
||
calibration runs on the harness.
|
||
|
||
## 🟢 OFF-BOX BACKUP — CLOSED. FULL ROUND TRIP PROVEN (2026-07-20 05:25 UTC)
|
||
|
||
**Root cause (Kev):** the vyndr-backup key was only in Hetzner's PROJECT ssh-key
|
||
store, never in the Storage Box's own `.ssh/authorized_keys`. Installed there →
|
||
the box now offers `publickey,password` (it previously offered an EMPTY auth
|
||
list, which is what "SSH not enabled on this box" looks like).
|
||
|
||
**Proven end-to-end:**
|
||
- `POST /api/internal/backup/run` → `ok:true`, **`offbox_ok:true`**, exit 0, 31s.
|
||
- Output confirms both env questions: **`ssh key: base64-decoded`** (the key is
|
||
correct base64, not a mangled raw paste) and destination
|
||
**`u635423@...:vyndr-backups/`** — no leading-slash bug.
|
||
- **File VERIFIED on the box, not inferred from exit 0:**
|
||
`GET /api/internal/backup/offbox` (rsync `--list-only` through the pinned host
|
||
key) returned `vyndr-20260720-051158.dump`, **833,917 bytes, 05:12:28 UTC** —
|
||
byte-identical size to the local dump, timestamp from that run.
|
||
|
||
**Hardening shipped with it:**
|
||
- **Host key STATICALLY PINNED.** `ssh-keyscan` matched the out-of-band
|
||
fingerprint `SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM`;
|
||
`scripts/storagebox_known_hosts` ships to the container and the script uses
|
||
`StrictHostKeyChecking=yes` + `UserKnownHostsFile=<pin>`. `accept-new`
|
||
(trust-on-first-use) is GONE; a missing pin file REFUSES the push. Tests
|
||
assert the weakenings never return.
|
||
- **Remote dir guaranteed** — `--mkpath`, else remote `mkdir -p`. Without it a
|
||
push into a missing parent can write the dump AS the directory name: one file,
|
||
overwritten nightly, reading as "backups exist" while retaining exactly one.
|
||
- **A failed required push now PAGES** at urgent (was "low/deferred") and the run
|
||
reports `offbox_ok:false`. Exit code still tracks ON-BOX durability so a good
|
||
on-box dump never raises a false total-failure alarm.
|
||
|
||
### ✅ ROUND TRIP CLOSED — pulled back off-box and rebuilt into a live DB (05:25 UTC)
|
||
The full insurance chain is now PROVEN, not assumed: dump → validated → pushed
|
||
off-box → verified on the box → **pulled back down from the box → restored into a
|
||
live Postgres**.
|
||
- **Pulled FROM the Storage Box** (not the local copy) with the in-session key
|
||
through the pinned host key (`StrictHostKeyChecking=yes`, never bypassed):
|
||
`vyndr-20260720-051158.dump`, 833,917 bytes — byte-identical to what was pushed.
|
||
- **Restored into scratch Postgres 17** (docker): 715 archive objects,
|
||
**42 public tables**, `ledger_entries` with all **27 columns** and real rows
|
||
(spot-checked: Kelsey Mitchell points o22.5 B −105 hit).
|
||
- **ASSERTION PASSED — `ledger_entries` restored = 645, live = 645** (target was
|
||
≥ 645). `model_snapshots` also restored 100/100, so the new retention store is
|
||
covered by backups from day one.
|
||
- Scratch DB torn down; the pulled copy deleted; **both dumps still on the box**
|
||
(nothing rotated out, cron untouched).
|
||
|
||
**🔴 OPERATIONAL GOTCHA FOUND DURING THE RESTORE — record this in any DR runbook:**
|
||
the dump is written by **pg_dump 17** (Supabase is PG 17.6), so **pg_restore 16
|
||
CANNOT read it** — it fails with `unsupported version (1.16) in file header`.
|
||
The first restore attempt failed for exactly this reason. **Restore with PG 17+
|
||
tooling** (`postgres:17-alpine` works). Restoring into vanilla Postgres also logs
|
||
`errors ignored on restore: 12` — Supabase-specific roles/extensions that do not
|
||
exist locally; harmless, the data restores completely.
|
||
|
||
## Open items — with status
|
||
|
||
| Item | Status | Note |
|
||
|---|---|---|
|
||
| **Settlement: 0 pushes / 470 settled** | ✅ **CLOSED 2026-07-29 — premise re-verified, NOT a bug** | The ruling's basis STILL HOLDS: the feed is **100% half-numbers**. 0 whole-number lines in **117,970 captured market lines** (5 books, 16 markets, `book`+`sharp`, Jul 20–29), 6,050 `model_snapshots` (incl. refusals), 1,141 ledger rows, 173 `lock_lines` (today). All 992 settled actuals are INTEGERS → **smallest actual-vs-line gap across all 992 = 0.5**, the arithmetic minimum. Expected pushes = **exactly 0 (P=0)**, not chance. Push branch sound + tested; CHECK constraints accept `'push'`. See the verify block above. |
|
||
| **~28 props/day never settle** | 🔴 OPEN, unstarted | Jul 17 MLB 86 graded/57 settled; Jul 18 103/75. Cause undiagnosed. |
|
||
| **Model-version contamination** | 🟠 PERMANENT, mitigate | `ledger_entries` mixes pre/post-2026-07-19-fix grades with no marker; eras cannot be separated retroactively. **Any backtest/accuracy claim off existing ledger history MUST treat the fix boundary as a hard cutoff.** `model_snapshots` stamps `model_version`+`code_sha` so it can't recur. |
|
||
| **A-grade unreachable in prod** | 🔴 OPEN | `opp_rank_stat` null; ESPN team endpoint has no defensive metric at all. Marketing hold stands. |
|
||
| **EV overconfident** | 🟠 OPEN | Needs calibration before it drives any surface. Hero v2 already ranks on it. |
|
||
| **`edge_pct` broken scale (U-deg pt 2)** | 🔴 OPEN — **DIAGNOSED 2026-07-29, fork reported, not chosen** | Re-verified: **44.0% over cap-40 (502/1141)**, not 51.5%; **worst 620 exact**. MLB median 60 / max 620 (86% of MLB rows are 0.5 lines); **WNBA sane** (median 12, 0% >100). Cause = **small-denominator explosion**, NOT units. **Only 3 surfaces RENDER it** (not 13): GradeResultCard (uncapped, auth-gated), DeskShowcase (already honest "—"), SoccerGradeResult (uncapped, orphan). **It DRIVES the dashboard top-10 sort** (operative on 97.3% of rows) + the Desk alt-ladder. `ev_pct` **DOES** have a live render site (PriceTriplet) — starved, not unwired. **Recommendation: HIDE + re-point the sort at p_win.** See the diagnosis block above. |
|
||
| **CLV broken (C4)** | 🔴 OPEN | `closing_line == locked_line` on ~95% of rows. BEAT CLOSE suppressed. **CLV ledger stays PRIVATE until backtest-proven.** |
|
||
| **Consistency CV floor** | 🟠 STOPGAP | `CONSISTENCY_MIN_MEAN=4` leaves a ±1.0 dead for MLB low-count stats. Real fix = index-of-dispersion classifier; needs a backtest first. |
|
||
|
||
### 🔑 THREE KEYS TO ROTATE (all currently valid)
|
||
1. **Storage Box password** — Kev plans a fresh reset; do it before/with the SSH fix.
|
||
2. **`VYNDR_INTERNAL_KEY`** — was pasted in plaintext into a chat transcript
|
||
(2026-07-19) and is in the local `.env`. It gates snapshot triggers, backup
|
||
runs, settle passes and the newsletter send. **Rotate: Coolify env + local
|
||
`.env` together.**
|
||
3. **GitHub PAT in the `origin` remote** — plaintext in `.git/config`, leaked
|
||
since ~Session 31. Deploys run off `gitea` so it's inert, but it is a live
|
||
credential. Revoke on GitHub; the remote can simply be deleted.
|
||
|
||
## Next queued work (in order)
|
||
|
||
1. **Backtest harness** — reads `model_snapshots`. **Needs ~2 weeks of history**;
|
||
as of now it holds ONE night. Its honest first output is *"cannot validate
|
||
yet"* and that is the expected, correct result. Build the instrument now, let
|
||
data accumulate.
|
||
2. **Settlement-correctness audit** — the two bugs above. Highest-value
|
||
correctness work: everything downstream trusts hit/miss.
|
||
3. **A-grade / opponent-strength sourcing** — MLB is **solved and verified**:
|
||
`statsapi` `teams/stats?group=pitching` returns all 30 teams free with
|
||
`era/whip/avg/slg/ops/HR9`, and it's *stat-specific*, which beats what we were
|
||
fetching. NBA/WNBA need `stats.nba.com` or derived allowed-points. Implement
|
||
behind the **source-adapter pattern** (health check that PAGES on empty —
|
||
tonight's ESPN crash was invisible for exactly that reason).
|
||
4. **Metrics engine** — archetype-baselined proprietary metrics
|
||
(`specs/model-train.md` §2C). **Gated on the harness**: every metric family
|
||
ships only if it backtests better than without it.
|
||
|
||
**Governing docs:** `specs/model-train.md` (value engine + phases + north star) ·
|
||
`specs/snapshot-retention.md` (Phase 2 report + schema) ·
|
||
`specs/audit-data/grade-collapse.md` · `specs/audit-data/gate-simulation.md`.
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 80 — S7 PRICED-LINE NUDGE COMPLETED (freshness) (2026-07-22)
|
||
|
||
## 🔴 REPORT-FIRST corrected the order's premise
|
||
Order claimed "S7 is a shell that doesn't update per selection." **Not true** —
|
||
`pricedForSelection` is a `useMemo([pricedIndex, selectedPlayer, stat])` and
|
||
`setSelectedPlayer/setStat` fire on every pick → chips already update per
|
||
selection (S79 verification stood). **The real gap was FRESHNESS**: the fetch
|
||
depended on `[sport]` ONLY, so `pricedIndex` was fetched once per sport-change,
|
||
never refreshed → a page open across a pricing-cron boundary (5 UTC hours)
|
||
surfaced hour-stale lines. That was the ONLY defect, and the only thing fixed.
|
||
|
||
## Fix (display-layer only, one file)
|
||
- `refreshPriced()` re-fetches when the held snapshot is older than
|
||
**`PRICED_STALE_MS` (30s = the /api/snapshot cache)** at the moment of use —
|
||
**on selection change AND on window focus**. Long-open page never shows a
|
||
stale line. Sport change clears first (no old-sport flash).
|
||
- **Reversible gate:** whole nudge behind `PRICED_NUDGE_ENABLED = true`. False →
|
||
empties the surfaced set → S6 link-only fallback. The instant revert lever.
|
||
|
||
## Stale-tap (unchanged, already safe)
|
||
Scan submit re-fetches the live snapshot server-side, so a stale chip tapped →
|
||
real triplet (still priced) OR honest empty (rotated away). Never lands on
|
||
nothing.
|
||
|
||
## Phase 4 — every case induced live
|
||
| Case | Result |
|
||
|---|---|
|
||
| (a) priced selection | `Sam Huff hits → [u0.5 @+105]` (exact) |
|
||
| (b) change selection | `Tsung Che Cheng hits → [o0.5 @-140]` — updates |
|
||
| (c) none / off-slate | `[]` / `[]` |
|
||
| (d) tap → triplet | `book +105 · fair +120` — real |
|
||
| (e) stale-tap (8.5) | no market → **degrades to empty** |
|
||
| (f) freshness | refresh-on-stale + focus wired, fingerprint live in chunk |
|
||
|
||
## Verdict — S7 FULLY WIRED (not a shell)
|
||
Correct per-selection · updates on change · fresh on long-open · tap reaches a
|
||
real triplet · stale-tap degrades honestly. **Scan A byte-identical**
|
||
(GradeResultCard/PriceTriplet/gradeAdapter/valueState/both scan routes/
|
||
pricedLines unchanged). Independently revertible via the flag.
|
||
|
||
## Method note
|
||
Scanner page is auth-gated → data-correctness induced via the exact functions
|
||
the `useMemo` runs + the `/api/scan` endpoint the tap hits (the load-bearing
|
||
path), not the rendered gated page. Freshness proven by code fingerprint + unit
|
||
tests asserting the exact refresh logic.
|
||
|
||
## Next open (unchanged)
|
||
edge_pct/ev_pct mislabel = **already checked S(prev), no genuine mislabel** —
|
||
one metric-CHOICE flagged (GradeResultCard "EDGE" = stat-gap %, tied to the
|
||
deferred ranking decision). · No settled `p_win` yet · 3 key rotations
|
||
(PAT stripped locally; Kev must REVOKE on GitHub + rotate internal key + Storage
|
||
Box password).
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 79 — READ-CARD NO-MARKET EMPTY STATE + PRICED-LINE SURFACER (2026-07-22)
|
||
The S78 diagnosis (join works, marketless scans rightly self-hide) → made the
|
||
absence LEGIBLE without fabricating a market.
|
||
|
||
## Phase 0 gate (all clear)
|
||
- **Design-check:** built to the triplet's own REFUSAL honesty language + the
|
||
existing `EmptyState` path-forward pattern — NOT freelanced.
|
||
- **Reachability:** scanner already fetches games/odds/search; snapshot is ONE
|
||
more public 30s-cached fetch per sport.
|
||
- **Staleness = inherent guard:** every scan re-validates the market
|
||
server-side at submit, so a surfaced line that rotates → honest empty state on
|
||
tap, never a vanishing triplet.
|
||
|
||
## What shipped
|
||
- **`web/src/lib/pricedLines.js`** (pure, tested) — `indexPricedLines` keeps
|
||
ONLY snapshot rows with a real `book_odds`, keyed by exact `nameKey|stat`.
|
||
Wrong stat → []; off-slate → []; nothing suggested/interpolated/nearest.
|
||
- **`web/src/components/vyndr/NoMarketState.tsx`** — honest empty state: no
|
||
market numbers of its own; real priced lines as one-tap chips OR a link to the
|
||
board. Path forward, not a wall.
|
||
- **`scan/page.tsx`** — snapshot fetch + "PRICED TONIGHT" help chips under the
|
||
free-typed line input (help, not restriction — any prop still scannable) +
|
||
the empty state when `book_odds==null || fair_odds==null`. Tapping a chip
|
||
pre-fills + re-scans (market re-resolved server-side).
|
||
|
||
## 🔴 REVERSIBILITY (the design constraint) — VERIFIED
|
||
`GradeResultCard.tsx`, `PriceTriplet.tsx`, `gradeAdapter.js`, `valueState.js`,
|
||
`routes/scan.js` and the Next `api/scan/route.ts` are **BYTE-IDENTICAL** (a test
|
||
asserts none reference the new UI). All new behavior is in 2 added files + one
|
||
additive scan-page block. **Scan A untouched; free-tier model-lock intact.**
|
||
|
||
## Proven live
|
||
- Surfacer on live snapshot: `Sam Huff hits` → exact match returns
|
||
`[o0.5? no — u0.5 @105]`; wrong stat → `[]`; Aaron Judge (off-slate) → `[]`.
|
||
- Scan A (Sam Huff hits 0.5 under): `book 105 · fair 120` → triplet renders,
|
||
empty state hidden.
|
||
- Scan B (Sam Huff hits **3.5** over): grade D, no market → **empty state shows**
|
||
+ surfacer offers the real o0.5 line.
|
||
|
||
## Still open (unchanged; NOT this order)
|
||
- No settled `p_win` yet → adjuster measurement not begun.
|
||
- env_multiplier on a ledger row pending next MLB lock.
|
||
- mappingHarness nomination flow · A/D emission · EV overconfidence ·
|
||
edge_pct surfaces · 3 key rotations (PAT stripped locally S-runbook; Kev must
|
||
still REVOKE on GitHub + rotate internal key + Storage Box password).
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 77 — DORMANT ADJUSTERS WIRED LIVE (2026-07-21)
|
||
The four-order plumbing gap is CLOSED. park + weather + platoon now reach live
|
||
grades. Pure input-wiring — **the 5 adjuster files are byte-identical** (verified
|
||
`git diff --quiet`).
|
||
|
||
## Phase 0 gate — all three inputs AVAILABLE at build
|
||
| Input | Availability | Source |
|
||
|---|---|---|
|
||
| Venue | **ALWAYS** 15/15 | schedule game object |
|
||
| First-pitch | **ALWAYS** | schedule `gameTime` |
|
||
| Opposing-SP hand | **declared 30/30 this afternoon** (morning locks precede it → platoon absent, correct) | probable `pitcherId` → statsapi hand |
|
||
Batter hand (`statcast.bats`) + MLBAM id were already on each grade.
|
||
|
||
## What was built
|
||
**`src/services/environmentContext.js`** (the wiring; adjusters untouched).
|
||
Fetches ONCE per snapshot: schedule, probable pitchers, ONE batched handedness
|
||
call, ONE Open-Meteo forecast per home park, batter splits per hitter.
|
||
`PARK_COORDS` (30 parks) lives here as public geometry (like the dome list), NOT
|
||
in an adjuster. Best-effort throughout — any failure → that prop archetype-only,
|
||
never a pipeline break.
|
||
**`attachChallenger` is now async** + takes `contextFor(grade)` → `{environment,
|
||
matchup}` → `adjust()` (which already accepted both).
|
||
|
||
## Independent attribution (each of 4 measurable separately)
|
||
Ledger columns `env_multiplier / env_park_base / env_weather_mod /
|
||
env_weather_state` + `challenger_adjustments` jsonb keeping archetype /
|
||
environment / matchup as SEPARATE entries. Not one blended delta.
|
||
|
||
## Combined bound (tested)
|
||
Coors slugger + wind-out + favourable platoon, all at once → **δ < 0.12**. Every
|
||
layer capped, total nudge clamped. Stacks lean, never re-forecast.
|
||
|
||
## ✅ PASS CONDITION MET — env_multiplier non-null (was 0 for 4 orders)
|
||
Proven END-TO-END on the live slate (buildContext + adjust against real feeds):
|
||
`venues_with_weather 15 · opp_declared 30 · pitchers_with_hand 30`.
|
||
| Park | env | compose |
|
||
|---|---|---|
|
||
| PHI | **1.179** | park 1.09 × wx 1.082 → **δ +0.041** |
|
||
| COL/ATL | >1 | hitter parks |
|
||
| CLE 0.929 · KC 0.937 | <1 | pitcher parks |
|
||
| **TOR/MIL** | park only | `wx 1 [dome_na]` — weather stands down, **park still applies** |
|
||
|
||
## 🔴 THE ONE HONEST GAP — not yet on a LEDGER ROW
|
||
Immutability refused to rewrite today's already-locked props (`ignoreDuplicates`),
|
||
so re-inducing wrote 0 new rows and **`env_multiplier` is still 0 on ledger rows**.
|
||
The wiring is PROVEN to produce the value (above); it lands on a row at the **next
|
||
NEW MLB lock**. Verify then:
|
||
```sql
|
||
select count(*) filter (where env_multiplier is not null) with_env,
|
||
count(*) filter (where challenger_adjustments::text like '%matchup%') with_platoon
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
## Point-in-time / no lookahead
|
||
Weather = FORECAST for first pitch fetched at projection time. Platoon = split
|
||
entering the game. Neither reads settle-time.
|
||
|
||
## Still open (unchanged, NOT this order)
|
||
- **No settled row carries `p_win`** — all adjusters still unmeasurable until
|
||
tonight's locks settle.
|
||
- Read-card triplet gap · mappingHarness nomination flow · A/D emission ·
|
||
EV overconfidence · zero-push bug · edge_pct surfaces · 3 key rotations.
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 76 — PLATOON SPLITS (regressed) (2026-07-21)
|
||
**Built + tested + deployed. Not reaching live props** (same plumbing gap).
|
||
|
||
## Gates — both clear; one was ALREADY closed
|
||
1. **Splits** = statsapi pull, 1 call/hitter:
|
||
`people/{id}/stats?stats=statSplits&sitCodes=vl,vr&group=hitting`.
|
||
2. **🔴 CORRECTION — batter handedness is NOT pending.** S69's Tier-A join did
|
||
it: `statcast_aggregates.bats` = **604/604** (210 L · 327 R · 67 S). The S69
|
||
STATE note said pending; the data says done.
|
||
3. **Point-in-time holds** if fetched before first pitch (a split queried this
|
||
afternoon can't contain tonight). A HISTORICAL backtest would use
|
||
season-final and leak → **clean measurement is forward-accruing.**
|
||
|
||
## 🔑 THE SPINE — regression IS the feature
|
||
`regressed = (PA·observed + K·prior) / (PA + K)`, **K = 600 PA**, prior = the
|
||
hitter's **OWN** blended rate (not the league's). The question is *"is he
|
||
DIFFERENT vs this hand than he normally is"* → his own line is the null, so a
|
||
hitter with no evidence gets **nothing**.
|
||
Weights: `30 PA → 4.8% · 130 → 17% · 400 → 40% · 1000 → 63%`.
|
||
|
||
### The make-or-break test (both halves, tested)
|
||
| Same .310 vs LHP | weight | multiplier | δ |
|
||
|---|---|---|---|
|
||
| **30 PA** | 4.8% | 1.011 | **0.003** |
|
||
| **400 PA** | 40% | 1.096 | **0.023** |
|
||
**8× apart.** A test asserts the ratio stays **> 5** — if the regression ever
|
||
breaks, the suite fails instead of projections quietly drifting onto noise.
|
||
|
||
## Real data — most hitters get NOTHING, and that's correct
|
||
| Player | vs L | vs R | vs LHP mult |
|
||
|---|---|---|---|
|
||
| Josh Bell (S) | .259 / 126 PA | .248 / 265 PA | **1.005** (17% wt) |
|
||
| Aaron Judge (R) | .246 / 76 PA | .248 / 185 PA | **0.999** (11% wt) |
|
||
Bell *looks* like a platoon splitter until the sample speaks. Neither is
|
||
material. **Expect most hitters to get ~nothing — honest output, not failure.**
|
||
|
||
## Honest-absent — 5 distinct routes, all exactly 1.0
|
||
`batter_hand_absent` · `pitcher_hand_absent` · `splits_absent` ·
|
||
`not_applicable` (stat) · missing SIDE → falls back to prior (weight 0), never 0.
|
||
|
||
## INDEPENDENT of the environment — deliberate
|
||
Park+weather compose into ONE coefficient (both describe the stadium). Platoon
|
||
rides its **own** `matchup` slot with its own label — entangling them would make
|
||
both **harder to attribute** when the instrument scores them. Mirrored on under,
|
||
capped ±15%, **inverted for strikeouts** (higher K rate → higher K prop).
|
||
|
||
## 🔴 THE PLUMBING GAP — now FOUR orders old
|
||
`snapshotService` still doesn't attach **venue + first-pitch + opposing-pitcher
|
||
hand** per prop and pass `environment` / `matchup` into `attachChallenger`.
|
||
**Park, weather AND platoon all adjust NOTHING live.** Everything else is ready.
|
||
|
||
## 🔴 And still: no settled row carries `p_win`
|
||
All seven adjusters remain unmeasurable.
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
## Recommendation (repeated, now urgent)
|
||
**Stop adding adjusters.** The next order should be ONLY: (a) attach venue +
|
||
first-pitch + opposing SP hand to props, (b) pass environment/matchup through,
|
||
(c) confirm `p_win` lands on a settled row. The measurement stack is now far
|
||
deeper than the plumbing feeding it.
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 75 — WEATHER MODULATION (2026-07-21)
|
||
**Built + tested + deployed. NOT reaching live props** (same venue gap as park).
|
||
`effective_environment = park_base × weather_mod`.
|
||
|
||
## Phase 0 — both feeds free + keyless
|
||
statsapi `/venues?hydrate=location` → lat/lon for every park, ONE call.
|
||
Open-Meteo hourly (temp / wind speed / wind direction), **available hours before
|
||
first pitch** = when we project. Verified live.
|
||
|
||
## 🔑 THE SPINE — two weather values, NEVER crossed
|
||
| Value | Purpose | Lands in |
|
||
|---|---|---|
|
||
| **FORECAST** at projection time | **drives** the live adjustment **AND is what the instrument measures** | `ledger_entries.wx_forecast` + `env_*` |
|
||
| **ACTUAL** at game time | future SELF-DERIVED weather factors ONLY | `game_context.wx_*` |
|
||
Using the actual to measure tonight = **scoring ourselves on information we
|
||
didn't have**. The actual comes from Open-Meteo's **ARCHIVE** endpoint — asking
|
||
the *forecast* endpoint after the fact returns a re-forecast, not what happened.
|
||
|
||
## 🔴 WIND SIGN — the easiest error in the file
|
||
`wind_direction` is METEOROLOGICAL (where it comes FROM). Blowing OUT to centre
|
||
= arriving FROM the opposite bearing → **`(from + 180) % 360`** before comparing
|
||
to the CF bearing. **Getting this backwards inverts every wind adjustment.**
|
||
Pinned by a test on straight-out / straight-in / crosswind.
|
||
`CF_BEARING` (30 parks) is **public geometry**, same class as the dome list. A
|
||
park missing from it gets **no wind effect** (keeps temperature).
|
||
|
||
## Three honest do-nothing states — all 1.0, distinct
|
||
- **`dome_na`** — weather doesn't apply; **the PARK factor STILL does** (verified:
|
||
domed venue keeps its sub-1.0 base while weather stands down).
|
||
- **`forecast_absent`** — none available for this park/time.
|
||
- **`sub_threshold`** — real forecast below a meaningful bar (a 0.3% nudge on a
|
||
light breeze is false precision).
|
||
Plus `not_applicable` for stats weather says nothing about (strikeouts).
|
||
|
||
## Conservative + ledger-tunable
|
||
Every magnitude is an env var (`WX_WIND_PER_MPH` .006 · `WX_TEMP_PER_DEG` .003 ·
|
||
thresholds 5 mph / 8°F), total **capped ±12%**. Nothing asserted — a nominated
|
||
challenger that earns its place or is cut.
|
||
|
||
## Induced (Wrigley, CF bearing 32°)
|
||
`wind FROM 212° @15mph → 15mph straight OUT · wx 1.12 × park 1.06 = 1.187 ·
|
||
δ +0.043` · under mirrors · **pitcher HR-allowed moves WITH hitter HR** (both are
|
||
P(over) on a ball leaving the park) · wind IN → 0.955 · calm/dome/no-forecast all
|
||
→ 1.0 by three different routes, park base still applying.
|
||
|
||
## ⚠️ ORDER CORRECTION (recorded)
|
||
The order described wind-out as helping the hitter and **hurting** the pitcher's
|
||
HR-allowed as "opposite sides." In PROP terms both go the **same** way — the
|
||
HR-allowed OVER is also more likely. **The sign lives in the STAT**, as already
|
||
established for park. Implementation follows that, not the phrasing.
|
||
|
||
## 🔴 STILL OPEN (unchanged)
|
||
1. **Neither park nor weather reaches live props** — `snapshotService` still
|
||
doesn't attach venue+first-pitch per prop and pass `environment` into
|
||
`attachChallenger`. Everything else is ready; this one link is missing.
|
||
2. **No settled row carries `p_win`** — all six adjusters remain unmeasurable.
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 74 — PARK BASE (public, pluggable) + GAME-LEVEL CAPTURE (2026-07-21)
|
||
|
||
## Phase 0 — the grain is cheaper than expected
|
||
Settle sees a player's game-log line (date + teams), **never venue or totals**.
|
||
But **ONE statsapi schedule call per game DATE** returns every game that day with
|
||
venue + linescore + scoringPlays. **Not per game, not per prop.** Verified: 15
|
||
games, 1 call.
|
||
|
||
## Public base — the ingestion was ALREADY DONE
|
||
`src/data/parkFactors.js` (static FanGraphs, S15) **is** the public base.
|
||
`src/services/parkBase.js` converts its 100-index to the composable multiplier
|
||
(COL 128 → **1.28**) instead of ingesting a second copy.
|
||
|
||
**🔴 THE COMMODITY LINE IS IN CODE, not just a comment.** Every resolution
|
||
carries `provenance`:
|
||
- public → `proprietary: false`, *"Commodity: a public number. Not a VYNDR
|
||
derivation."*
|
||
- derived → `proprietary: true` — **reserved**, and only earned once it beats
|
||
the public base on the instrument.
|
||
**Never market "proprietary park factors" about the public base.**
|
||
|
||
**Honest-absent where even PUBLIC is thin** (`UNSTABLE_PUBLIC`: ATH/OAK/SAC/LV —
|
||
relocated club, temporary venue). Extend that set whenever a park opens or a
|
||
club moves; a missing entry means a confident factor for a stadium with no data.
|
||
|
||
## SOURCE-PLUGGABLE — the architectural point
|
||
`resolveParkBase()` is the ONLY accessor; public and derived return **identical
|
||
shapes**, so callers never branch on source. Self-derived swaps into the same
|
||
slot with nothing downstream rewritten. **A derived source with no factor
|
||
returns `absent` — it does NOT silently fall back to public** (that would turn a
|
||
commodity number into a proprietary claim).
|
||
|
||
## Game-level capture — LIVE
|
||
`game_context` (migration 035), **game grain, deduped on `game_id`, never copied
|
||
onto prop rows**. Named-purpose fields ONLY: venue → park factors · runs → run
|
||
environment · HR totals → HR factors. Nothing else stored.
|
||
- **Finals only** (an in-progress total is not a result).
|
||
- **No scoring plays → HR is `null`, not 0.**
|
||
- HR from scoringPlays is COMPLETE — every HR scores ≥ the batter.
|
||
- Runs nightly beside the settle pass (yesterday's date).
|
||
|
||
**Live fingerprint:** `POST /api/internal/gamectx/2026-07-20` → **15 games, 15
|
||
venues, 15 with HR totals, 0 duplicates.**
|
||
|
||
## 📅 ACCRUAL TARGET — "later" is a number
|
||
**150 home games/venue** (≈81/season) → **~1.9 seasons** before a self-derived
|
||
factor can be NOMINATED against the public base.
|
||
Live now: **15 games · 15 venues · 0 ready.** Track:
|
||
`GET /api/internal/gamectx/accrual`.
|
||
|
||
## Induced (live)
|
||
COL HR **+0.061** hitter · **+0.061** pitcher HR-allowed (same park, sign lives
|
||
in the STAT) · mirrored on under · SF negative · **TB weather_na=true with its
|
||
factor STILL applying** · ATH **absent** · strikeouts untouched.
|
||
|
||
## ⚠️ STILL OPEN (unchanged from S73)
|
||
**Park does not yet reach live props** — `snapshotService` still doesn't attach a
|
||
venue per prop and pass `environment` into `attachChallenger`. `parkBase` +
|
||
`parkMultiplier` + the challenger's `environment` slot are all ready; the
|
||
prop→venue lookup is the missing link.
|
||
**And still blocking ALL measurement: no settled row carries `p_win`.**
|
||
|
||
## Next
|
||
1. Attach venue → prop → `environment` (closes the park loop).
|
||
2. **Weather** — multiplies `park_base`; domes already `weather_na`; add
|
||
weather-actuals to `game_context` for later self-derived modulation.
|
||
3. Self-derived park factors at ~1.9 seasons of accrual.
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 73 — PARK FACTORS (derived, composable) (2026-07-21)
|
||
**Derivation + composition BUILT and proven. PIPELINE WIRING INCOMPLETE — see
|
||
the gap.** Layer 3 Step 4.
|
||
|
||
## 🔴 PHASE 0 — A STATIC PARK TABLE ALREADY EXISTED (nearly built a duplicate)
|
||
`src/data/parkFactors.js` — **static FanGraphs table, Session 15**, consumed by
|
||
`computeFeatures.js:374`. I almost created a second source of truth.
|
||
- **NEW derived service = `src/services/parkFactors.js`** (distinct path).
|
||
- **Checked:** the static table reaches `computeFeatures` but **NOT
|
||
`probabilityEstimator`** → it does **not** currently touch `p_win`.
|
||
- **This reframes the order:** if the champion ever does see park, adding it to
|
||
the challenger is DOUBLE-COUNTING — precisely what the S72 harness catches.
|
||
Park therefore ships as a **nominated challenger to be TESTED**, not an
|
||
assumed win.
|
||
|
||
## Derivation (ours, from real games)
|
||
statsapi `schedule?hydrate=linescore,venue,scoringplays` — **one call per date
|
||
range**. Every HR scores ≥ the batter, so **HR totals are fully recoverable from
|
||
scoringPlays**. Derived from **5,055 real games, 2022–2025**:
|
||
|
||
| Park | HR base | RUN base |
|
||
|---|---|---|
|
||
| Dodger Stadium | **1.106** | 1.023 |
|
||
| Coors Field | 1.029 | **1.099** |
|
||
| Oracle Park | **0.923** | 0.969 |
|
||
| PNC Park | **0.917** | 0.994 |
|
||
|
||
**18 parks present · 18 honestly absent** (below the 150-game regime floor).
|
||
*(Sample was a partial pull — month-chunked with `-31` end dates drops days — so
|
||
absences are inflated. Re-derive with proper date ranges before trusting counts.)*
|
||
|
||
## COMPOSABLE — the architectural point
|
||
Emits a **multiplier ~1.0**, never an additive nudge, because weather must
|
||
modulate it next: `effective = park_base × weather_mod × …`. Additive terms
|
||
don't compose (5% park + 8% wind = 1.05×1.08, **not** +13%). The challenger
|
||
converts via `Math.log(env)` into the same log-odds space the archetype nudges
|
||
use. A test multiplies a placeholder weather term on to prove the shape composes
|
||
**with no rearchitecting**.
|
||
|
||
## Honest states (three, distinct)
|
||
`absent` (thin → adjust nothing) · `present` (adjust) · **`weather_na`** (dome —
|
||
**the park factor STILL APPLIES**; the flag exists so next order's weather does
|
||
nothing there). **N/A ≠ absent** — conflating them either drops a valid park
|
||
factor or applies wind indoors.
|
||
|
||
## Directional by prop-owner
|
||
`home_runs` and `home_runs_allowed` both key off `hr_base` **in the same
|
||
direction** — the sign lives in the STAT, not the park.
|
||
|
||
## Structural breaks + shrinkage
|
||
A deviating season starts a new regime **only if the NEXT season confirms** (one
|
||
odd year = noise). Only post-break seasons used. Factors **regress toward
|
||
neutral by sample size**; fine conditioning (park×hand, park×batted-ball) stays
|
||
**unavailable** until its own larger floor (400 games).
|
||
|
||
## 🔴 THE GAP — park is NOT yet reaching live props
|
||
`adjust()` accepts `environment`, and `parkMultiplier` produces it — but
|
||
**nothing in `snapshotService` looks up a venue per prop and passes it**. Props
|
||
carry `game_time` (via gameBinder) but **no venue**. Remaining wire-up:
|
||
1. Attach venue to each prop (statsapi schedule already returns it).
|
||
2. Persist derived factors (a table or cached blob) + a refresh job.
|
||
3. Pass `environment` into `attachChallenger`.
|
||
**Until then park adjusts nothing live** — the derivation and composition are
|
||
proven, the plumbing is not.
|
||
|
||
## Next
|
||
1. Wire venue → prop → park lookup (the gap above).
|
||
2. **Weather modulation** — composes by multiplying `park_base`; domes already
|
||
flagged `weather_na`.
|
||
3. Still blocking all measurement: **no settled row carries `p_win` yet**.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 72 — TIER-1 MAPPINGS LIVE + TIER-2 HARNESS (2026-07-21)
|
||
|
||
## 🔴 PHASE 0 GATE — NO HISTORICAL OOS IS POSSIBLE (and why)
|
||
`statcast_aggregates` is **overwritten nightly by design** (Layer 1 = full
|
||
re-pull upsert) → season-TO-DATE only, **no point-in-time history**. Classifying
|
||
a player for a Jul-15 game with today's aggregate feeds in Jul 15–21 →
|
||
**look-ahead leakage**, making any "out-of-sample" verdict worthless.
|
||
**→ The harness reads the archetype vector RETAINED at grade time (S70
|
||
instrument) and runs FORWARD-ACCRUAL.** Reported, not worked around.
|
||
*(If point-in-time archetypes are ever wanted, Layer 1 must retain a dated
|
||
snapshot per refresh — a real schema change, not a query.)*
|
||
|
||
## Canonical names — asserted, not assumed
|
||
Every mapping references the axis keys the classifier **emits**; a test walks
|
||
both maps against `BATTER_AXES`/`PITCHER_AXES`. **A non-existent key would look
|
||
wired and never fire** — that's the failure this guards.
|
||
|
||
## TIER 1 — LIVE (tautological, directional, MLB-only)
|
||
| Axis (canonical) | Labels | Stat | Dir |
|
||
|---|---|---|---|
|
||
| `strikeout` | PUNCHOUT / WHIFF | strikeouts | ↑ |
|
||
| `ground_ball` | SINKER / SEAM | home_runs_allowed | ↓ |
|
||
| `fly_ball` | FLY BALL / ELEVATOR | home_runs_allowed | ↑ |
|
||
| `control` | SURGEON ARM / PINPOINT | walks_allowed | ↓ |
|
||
| `power` | SLUGGER / BOMBER | total_bases, home_runs | ↑ |
|
||
| `contact` | TECHNICIAN / SURGEON | hits ↑, strikeouts ↓ | |
|
||
| `patience` | GRINDER / SNIPER | walks | ↑ |
|
||
|
||
Average player untouched; UNDER mirrors OVER exactly.
|
||
**SPEED (BURNER→SB) is HONESTLY ABSENT** — SB never reached the aggregate store,
|
||
so Layer 2 shelved the axis. `BATTER_MAP.stolen_bases = {}`, not invented.
|
||
|
||
## TIER-2 HARNESS — `src/services/mappingHarness.js`
|
||
Three hurdles a plausible-but-worthless mapping fails:
|
||
1. **MARGINAL, not correlation** — scores the archetype against the
|
||
**projection's RESIDUAL** (`outcome − p_win`). If the projection already
|
||
knows it → `REJECT (redundant)`. **This is the double-counting catch.**
|
||
2. **OOS split by DATE, never random** — rows from one game share pitcher/park/
|
||
lineup; a random split leaks.
|
||
3. **Direction validated from held-out data; a contradicted sign is REJECTED,
|
||
never silently flipped** (flipping = fitting noise).
|
||
Floors: `MIN_ROWS 200 · MIN_HELD_OUT 60 · MIN_EFFECT 0.02` → else `INSUFFICIENT`.
|
||
|
||
**Lifecycle encoded: `nominated → live → claimed`.** A survivor RUNS LIVE and is
|
||
measured; **only the quantified public claim waits.** Nothing sits dark.
|
||
|
||
## Real candidate, end-to-end
|
||
`GROUND-BALL → hits_allowed` → **INSUFFICIENT — 0 / 200 settled rows**, because
|
||
no settled row carries `p_win` yet (S70 instrument starts at the next new lock).
|
||
Correct and expected verdict.
|
||
|
||
## Gotcha worth keeping
|
||
My first synthetic fixture **aliased the carrier selector against the outcome
|
||
draw** and manufactured a 0.038 effect where zero was injected. The harness
|
||
rejected it correctly — but via the SIGN reason, not redundancy, which is how I
|
||
found the fixture bug. Draw now uses a coprime modulus. *Lesson: when a harness
|
||
gives the right verdict for the wrong reason, check the fixture.*
|
||
|
||
## Still blocking everything downstream
|
||
**No settled row carries `p_win`/`p_win_challenger` yet.** Tier-1 adjusts real
|
||
projections now, but neither it nor any Tier-2 candidate can be evaluated until
|
||
volume accrues. Verify:
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger,
|
||
count(*) filter (where challenger_delta <> 0) moved
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 71 — LAYER 3 STEP 2: CHALLENGER PROJECTION (2026-07-21)
|
||
**Code-complete + deployed + unit-proven. NOT yet observed writing a live row —
|
||
see the verification gap below.** Champion serves users UNCHANGED.
|
||
|
||
## What it is
|
||
`src/services/challengerProjection.js` — the champion's `p_win` **adjusted** by
|
||
the Layer-2 axis vector in **log-odds space** (a nudge can't cross 0/1 and means
|
||
the same at p=.5 and p=.9). Total capped at **0.45 log-odds** — a lean, not a
|
||
re-forecast. Every deviation is attributable: `challenger_adjustments` stores the
|
||
named axis + signed nudge. Only mechanically obvious stat↔axis pairs are mapped.
|
||
|
||
**Identical to champion by construction** where there's no signal — unremarkable
|
||
player, thin sample, unmapped stat, or no classification all return the champion
|
||
value byte-for-byte with a stated `reason`. The A/B therefore differs ONLY where
|
||
archetype-awareness could help or hurt.
|
||
|
||
**Isolation:** `adjust()` is pure; the champion field is read, never written; the
|
||
served snapshot payload is still `enriched` (untouched); a challenger throw is
|
||
caught so it can never break the pipeline it's measured inside. Statcast loads
|
||
ONCE per snapshot run — grade-time I/O stays at zero.
|
||
|
||
## Induced on real players (local, real data)
|
||
| Prop | Champion → Challenger | Via |
|
||
|---|---|---|
|
||
| Judge HR **over** | 0.42 → **0.447** | BOMBER +.22, WHIFF RISK −.11 |
|
||
| Judge HR **under** | 0.58 → **0.553** | exact mirror |
|
||
| Judge strikeouts | 0.55 → **0.55 (δ=0)** | WHIFF RISK +.11 / GRINDER −.11 **cancel** |
|
||
| Skubal Ks over | 0.60 → **0.702** | WHIFF, TRAPDOOR, CANNON aligned |
|
||
| Skubal hits-allowed | 0.50 → **0.392** | a K arm allows fewer hits |
|
||
| Bell / 12-PA sample | **unchanged** | `no_distinctive_axis` / `thin` |
|
||
|
||
## 🔴 VERIFICATION GAP — read before assuming this is recording
|
||
**Zero live rows carry `p_win` or `p_win_challenger`** (`latest_lock` still
|
||
`2026-07-21 02:58:45`). Three compounding reasons:
|
||
1. **Immutability refuses to backfill.** Every prop on the current slate was
|
||
locked BEFORE these columns existed; `ignoreDuplicates` correctly skips them.
|
||
**First live rows arrive on the next NEW lock** (new game date / new line).
|
||
2. **`ledgerWritten` OVER-REPORTS.** It counts `chunk.length`, not actual
|
||
inserts — 14 induced snapshots all reported "ledger 25" while inserting
|
||
**zero**. This misled my own verification; worth a one-line fix.
|
||
3. **`statcast_aggregates` is MLB-ONLY (1,354 rows; 0 wnba/nba).** I induced on
|
||
WNBA, where the challenger correctly no-ops for lack of mechanism data.
|
||
**The challenger currently applies to MLB only.**
|
||
|
||
**Verify with:**
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger,
|
||
count(*) filter (where challenger_delta <> 0) moved
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
Expect first non-zero after the next MLB slate locks.
|
||
|
||
## HELD
|
||
**Promotion** (challenger → serving) and **any superiority claim** stay held
|
||
until the ledger proves it on segmented volume. Baseline volume: blended n≈500
|
||
≈ 5 days; per-segment n≈200 ≈ 2–4 weeks.
|
||
|
||
## Next
|
||
1. Confirm first live champion+challenger rows after the next MLB lock.
|
||
2. Fix `ledgerWritten` to report real inserts.
|
||
3. Extend mechanism data beyond MLB (or accept MLB-only challenger).
|
||
4. Let volume accrue; then compare Brier/calibration segmented.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 70 — LAYER 3 STEP 0+1: THE MEASUREMENT INSTRUMENT (2026-07-21)
|
||
|
||
## 🔴 STEP 0 FINDING — WE HAVE BEEN FLYING WITHOUT AN INSTRUMENT
|
||
**"Is the projection calibrated / does it beat the market" has NEVER been
|
||
answerable.** Three separate causes:
|
||
- `model_snapshots` holds `p_win` but has **0 settled rows** (settle writes to
|
||
`ledger_entries` only) and only spans 07-19→07-20.
|
||
- `ledger_entries` settles 642 rows but carries **no probability at all**.
|
||
- `closing_captures` has the close but **no link to a result**.
|
||
|
||
The whole measurable universe was **35 rows**, recovered by a lossy in-memory
|
||
join. Harness agrees: `INSUFFICIENT_HISTORY, scored 35, joined 404`.
|
||
|
||
### The 35-row reading (NOT conclusive — reported for the record)
|
||
| | Brier ↓ | Directional |
|
||
|---|---|---|
|
||
| Our projection | **0.2682** | 18/35 |
|
||
| Market (de-vigged) | **0.2650** | 16/35 |
|
||
| Coin flip | 0.2500 | — |
|
||
|
||
**Both score worse than a coin flip; ours is marginally worse than the market.**
|
||
Mean model p **0.566** vs actual hit rate **0.486** → **systematically
|
||
overconfident ~8 pts** (same `PROB_CEIL` overconfidence flagged in EV).
|
||
At n=35 the SE is ~8.5 pts — **noise. We have no evidence the projection beats
|
||
the market, and none that it doesn't.** Every segment (max n=28) is honest-absent.
|
||
**Every settled row has archetype `(none)`** — Layer 2 postdates all of them.
|
||
|
||
## 🔴 THE PROJECTION IS NAIVE — verified, not assumed
|
||
`probabilityEstimator` = **game-log frequency over the line, 0.6/0.4 recency
|
||
blend, opponent nudge**. Zero archetype, zero mechanism data. Grep confirms.
|
||
|
||
## Step 1 — instrument wired (migration 033)
|
||
**The four-tuple now lands on ONE row** (`ledger_entries` — the row that
|
||
actually settles): `p_win` + `fair_prob_lock` + `archetype_vector` +
|
||
`projection_locked_at` at LOCK; `closing_prob` + `closing_captured_at` from
|
||
`closing_captures`; `outcome` at settle. Calibration and market-comparison
|
||
become plain SQL instead of a join that drops 90% of rows.
|
||
- **IMMUTABLE** — written once via the existing `ignoreDuplicates` upsert, never
|
||
re-derived at settle (that would measure a projection we never made).
|
||
- **VECTOR not label** — a text column cannot attribute "did archetype help?".
|
||
- **Honest-absent BOTH ways**: past game + no capture → `market_unavailable_reason`,
|
||
never an imputed line (calibration still scores); **future game → NOT marked**,
|
||
because a close can still arrive and premature absence is as dishonest as
|
||
imputation.
|
||
- **Closing coverage verified BEFORE reuse:** 83/83 = **100%** on graded props
|
||
since capture began. The 56,890 refusals are ungraded candidates, not our props.
|
||
|
||
## 🔴 TWO BUGS CAUGHT BY INDUCING
|
||
1. **Scheduler hook iterated `SPORTS`, which does not exist in that scope.**
|
||
Inside its try/catch it would throw `ReferenceError` every tick and silently
|
||
never run — **the instrument would have looked wired and captured nothing.**
|
||
Now `cadence.ALL_SPORTS`.
|
||
2. **`attachClosingProb` selected a `fair_prob` column `closing_captures` does
|
||
not have.** It stores raw `over_odds`/`under_odds` ON PURPOSE (S64) so the
|
||
de-vig runs later. First live run: **642 rows marked market-unavailable, 0
|
||
closes attached.** De-vig now runs via `devigTwoWay`. **Repair checked, not
|
||
assumed:** all 642 are pre-capture-era games, so the absence is TRUE; **0
|
||
capture-era rows wrongly marked.** The bug would have mis-marked every future row.
|
||
|
||
## ⚠️ THE BASELINE ACCRUES FORWARD — and does NOT backfill
|
||
Historical `p_win`/closes are gone. **Immutability means today's already-locked
|
||
rows never gain `p_win`** — the instrument starts recording at the **next NEW
|
||
lock** (verified: 0 new rows during induction because every prop was already
|
||
locked). Check with:
|
||
`select count(p_win) from ledger_entries where user_id is null;`
|
||
|
||
**Volume to a real baseline** (~100 settled/day): blended n≈500 → **~5 days**;
|
||
per-segment n≈200 → **2–4 weeks**.
|
||
|
||
## Layer 3 sequencing — RECOMMENDATION
|
||
The order proposed building archetype-awareness then proving it. **We cannot
|
||
prove anything yet**, so an adjuster shipped today is unmeasurable for weeks and
|
||
we would be back to shipping on plausibility. **Fix retention first (done), let
|
||
volume accrue, then build adjusters** — each measurable from its first settled row.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 69 — LAYER 2: MULTI-AXIS ARCHETYPES (2026-07-21)
|
||
Spec: **`docs/ARCHETYPE-AXES.md`** (every name accounted for). Layer 3 NOT built.
|
||
**Archetype still does not feed grades** — display + snapshot + future Layer 3.
|
||
|
||
## Tier-A joins — all three landed, each ONE free call
|
||
| Join | Before | After | Source |
|
||
|---|---|---|---|
|
||
| Handedness | batters **0%** | **1,354/1,354** | statsapi `/sports/1/players` (bats+throws+position) |
|
||
| True role | IP proxy (drifts) | **750/750** | statsapi stats `playerPool=ALL` (751 rows; default returns only ~57) |
|
||
| **Velo** | **53%** | **99%** (721/729) | wide `pitch-arsenals` matched BY PITCH TYPE |
|
||
**Role split (real usage):** starter 219 · reliever 365 · setup 126 · closer 40.
|
||
The velo fix is what makes CANNON/HOWITZER buildable instead of shelved.
|
||
|
||
## 🔬 AXIS INDEPENDENCE — MEASURED (467 batters / 531 pitchers)
|
||
**|r| ≥ 0.70 ⇒ one trait ⇒ COLLAPSED** (never show one trait as two archetypes):
|
||
`batter k%~whiff% +0.89` · `hard-hit%~exit velo +0.88` · `chase%~swing% +0.87` ·
|
||
`chase%~bb% −0.72` · `pitcher k%~whiff% +0.76` · `gb%~fb% −0.73`.
|
||
|
||
**The finding that mattered:** `pitcher velo ~ k% = +0.14`, `~whiff% = +0.07`,
|
||
`~gb% = +0.07`. **Velocity is NOT a proxy for missing bats** — a hard thrower
|
||
who misses no bats is a real distinct type, so it earns its own axis.
|
||
`pitcher k% ~ gb% = −0.10` → PUNCHOUT ⊥ SINKER, which IS the multi-axis thesis.
|
||
|
||
## Cut-lines
|
||
Real **p75 (distinctive) / p90 (elite)**, **per role where tails differ** even
|
||
when medians agree — reliever GB% p90 **54.1** vs starter **48.9**, both median
|
||
42.5. Floors PA≥50 / IP≥10.
|
||
|
||
## 🔴 THE FLEX FALLBACK IS DELETED (`archetypeService.js`)
|
||
`classify()` returned FLEX (mlb) / SHIELD (wnba) / CONNECTOR (nba) **at weight
|
||
1.0** when nothing scored — "could not classify" rendered as a confident
|
||
classification **with descriptive education copy**. 8 of 18 MLB players carried
|
||
it, and FLEX could never be *earned* (its only input, `utility`, had zero
|
||
writers). Every sport now does what MMA already did: **unclassified = absent.**
|
||
**Live fingerprint:** Alan Roden → `primary: null, blend: []`.
|
||
|
||
## Induced on real players
|
||
| Player | Blend |
|
||
|---|---|
|
||
| Skubal | STARTER · L · **WHIFF + SEAM + PINPOINT** (all elite) |
|
||
| Judge | **BOMBER + GRINDER + WHIFF RISK** (elite power, patient, whiffs) |
|
||
| Kwan | **SURGEON + SNIPER + SLASH**, power **absent** (0.4 barrel ≠ "low power") |
|
||
| **Josh Bell** (was DRIVER) | **empty** — "No standout profile" |
|
||
| **Alan Roden** (was FLEX @1.0, 21 PA) | **all axes absent** — "Not enough PA yet" |
|
||
| velo-less SP | keeps 3 axes; velocity+slot absent |
|
||
|
||
**`absent` (no data) is distinct from a null vector entry (measured, ordinary).**
|
||
|
||
## Gotchas
|
||
- `classify(sport, null)` **threw** — an explicit null does NOT trigger a default
|
||
parameter and every scorer dereferences its arg. Guarded in `classify`.
|
||
- Three tests asserted the fallback; updated to assert absence. One over-broad
|
||
edit briefly broke a test where a player has REAL stats and legitimately
|
||
classifies (Riley → BOMBER) — removing the FALLBACK is not removing
|
||
classification.
|
||
- `TeamHub.tsx` `|| 'zzz'` is a SORT key, not a display default — sorts
|
||
unclassified last. Correct; leave it.
|
||
|
||
## Layer 3 (HELD) reads the FULL vector
|
||
`classifyPlayer()` stores every axis; only the top ≤3 surface. Projection +
|
||
adjusters (park, weather, umpire, platoon) are the next arc, not built here.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 68 — LAYER 1: STATCAST MECHANISM DATA (2026-07-21)
|
||
Ingest + stay-current. **Layers 2 (archetypes) and 3 (projection) NOT built.**
|
||
Blueprint: **`docs/MECHANISM-DATA.md`** — the pattern every sport inherits.
|
||
|
||
## Phase 0 gate — both match rates 100%
|
||
- **Batters 40/40 · PITCHERS 66/66** (real rosters: CLE, DET, MIN, NYY, LAD),
|
||
joined by **MLBAM id** against the 713-pitcher Savant feed. Zero honest-absent
|
||
on identity — the join is an integer both systems use natively, and the
|
||
snapshot pipeline **already stores it** (`playerId`) per graded row.
|
||
- **pybaseball deliberately NOT used** — it is an MIT wrapper over the same CSV
|
||
URLs; adding it reintroduces a Python runtime in a stack where the existing
|
||
Python service is already offline. Direct axios + `savantAdapter`'s parser.
|
||
|
||
## What landed
|
||
`statcastAdapter` (6 free Savant feeds) → `statcastAggregateService` →
|
||
`statcast_aggregates` (migrations 030 + 031). **Live in prod:**
|
||
|
||
| | |
|
||
|---|---|
|
||
| Rows | **1,354** (604 batters · 750 pitchers) |
|
||
| Distinct players | 1,312 → **42 two-way players carry BOTH profiles** |
|
||
| Join rate | **1,354 / 1,354 · 0 unjoined** |
|
||
| Handedness | 677 pitchers (movement feed) |
|
||
| Sufficient / thin | 998 / 356 (PA≥50, IP≥10) |
|
||
| Arsenals | avg **4.4 pitches**, max 8, 679 with ≥3 |
|
||
| Table size | **5.3 MB** (vs 500 MB plan ceiling) |
|
||
| Pull time | ~5 s, all six feeds |
|
||
|
||
**Backfill and refresh are the SAME call** — full re-pull upserted on
|
||
`(sport, season, source_id, role)`. Idempotent + self-healing: **5 induced runs
|
||
→ still exactly 1,354 rows.** No incremental who-played bookkeeping to drift.
|
||
|
||
## 🔴 TWO REAL BUGS CAUGHT BY INDUCING, NOT BY REVIEW
|
||
1. **Two-way players broke the write.** PK `(sport, season, source_id)`
|
||
collapsed a player's batter and pitcher profiles into one key → one batch hit
|
||
the same row twice → `ON CONFLICT DO UPDATE command cannot affect row a
|
||
second time`, first chunk written, rest refused. **ROLE is in the key now**
|
||
(migration 031). Ohtani stores 419 PA / 17.2% barrel as a batter AND 85.2 IP
|
||
/ 33.7° arm angle / 52.1% GB as a pitcher — two real profiles, not one
|
||
invented player.
|
||
2. **Pitch mix was primary-pitch-only.** `pitch-movement` with an empty
|
||
`pitch_type` returns **ONE row per pitcher** (677 rows / ~700 pitchers), so a
|
||
five-pitch arsenal stored as one pitch — which would have classified every
|
||
pitcher as a one-pitch arm. Mix now comes from **`pitch-arsenal-stats`
|
||
(3,205 rows = pitcher × pitch)**; movement still supplies velo/break/
|
||
handedness folded onto the primary pitch. Skubal: 5 pitches, L, FF 96.7.
|
||
|
||
## Honesty rules (each has a test)
|
||
Absent metric → `null`, never 0 · thin sample **stored and flagged** (thin ≠
|
||
missing) · unjoined player stored with null `player_key`, joins later · **all
|
||
feeds empty → REFUSES to write** so a bad night can never blank a good table ·
|
||
`updated_at` on every row.
|
||
|
||
## Staleness — verified across simulated conditions
|
||
`0.1h → false · 26h → false · 50h → TRUE · 168h → TRUE · NEVER-BUILT → false`
|
||
(threshold 48h). **Never-built is deliberately NOT stale** — different
|
||
condition, different fix; paging on a fresh install teaches the operator to
|
||
ignore the alarm. The scheduler pages on a **failed run** AND on **silent
|
||
staleness** separately: a job that stops being scheduled never produces a
|
||
failure.
|
||
|
||
## Ops
|
||
Nightly at `STATCAST_HOUR_UTC` (default **11 UTC**, after every game is final);
|
||
kill switch `STATCAST=0`; gates `STATCAST_MIN_PA` / `MIN_IP` / `MAX_AGE_HOURS`.
|
||
**Induce:** `POST /api/internal/statcast/refresh` · **probe:**
|
||
`GET /api/internal/statcast/status`.
|
||
**⚠️ `/api/internal/*` is NOT reachable on `vyndr.app`** (no Next proxy, by
|
||
design) — use **`api.vyndr.app`**. Also: local `@supabase/supabase-js` egress
|
||
from the WSL box FAILS (`fetch failed`); the MCP tool works. Induce DB-writing
|
||
jobs on the server, not locally.
|
||
|
||
## Raw per-pitch — measured, and deliberately NOT stored
|
||
2026-07-19: **4,791 pitches · 119 cols · 3.26 MB · 680 B/row · 1.2 s** →
|
||
**~776k pitches/season = 503 MB CSV, 171 MB gzipped, ~0.85 GB in Postgres**
|
||
against a **500 MB free-plan ceiling** (DB is 26 MB today). Aggregate grain is
|
||
5.3 MB. Raw stays re-pullable from the free source if Layer 3 needs it; the
|
||
cold tier (Storage Box) is the two-tier home when it does.
|
||
|
||
## Commodity, not moat
|
||
Raw Statcast is public — anyone can pull it in a second. The edge is Layer 2
|
||
(which signals define an archetype, and where the boundaries sit), Layer 3
|
||
(projections), and the settled ledger that proves them. **Having the data is not
|
||
having an edge.**
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 67 — PRICE LAYER GATED + WIRED (2026-07-20)
|
||
|
||
## ⚠️ THE PREMISE WAS WRONG TWICE — read before trusting a "wire the fields" order
|
||
1. **There is no `/api/slate`.** The board reads `GET /api/snapshot/:sport`.
|
||
2. **The snapshot already carried the price fields.** `book_odds`, `fair_odds`,
|
||
`model_odds`, `p_win`, `ev_pct`, `value`, `takeable`, `fair_prob`,
|
||
`devig_method` were all live. Nothing needed adding.
|
||
3. **`fairLine` does not exist** (zero hits repo-wide). The Option-A
|
||
"derive from fairLine" instruction had no referent — `fair_odds` is already
|
||
a real de-vigged American price.
|
||
|
||
## Phase 0.5 gates — all PASS, verified numerically
|
||
- **Is it a price?** Yes. `fair_odds` = American odds from
|
||
`impliedProbToAmerican` inside `devigTwoWay`; `fair_prob` is the probability.
|
||
Both distinct from `line` (the stat threshold).
|
||
- **Numeric match?** **8/8 EXACT.** Recomputed `fair_odds`+`fair_prob`
|
||
independently from stored raw over/under prices; every value matched to the
|
||
integer and to 3dp.
|
||
- **Poison-independent on quarantine rows?** **YES, proven on the quarantined
|
||
cohort itself.** `devigTwoWay(over, under)` takes market prices only — no
|
||
model term is reachable — and the 8 rows recomputed above are all
|
||
`wrong_opponent_grade`. The poison is in the GRADE, not the price. So
|
||
**quarantine suppresses the MODEL leg only; the fair leg stands.**
|
||
|
||
## 🔴 THE REAL FINDING — the model price was PUBLIC
|
||
`GET /api/snapshot/:sport` is unauthenticated and was serving `model_odds`,
|
||
`p_win`, `ev_pct`, `value`, `takeable` to anonymous callers on **25/25 live
|
||
rows**. The Session-66 gate on `/api/analyze` was bypassed entirely.
|
||
**`/api/hero-prop` was a SECOND door** — it reads Redis directly via
|
||
`heroPropService`, so the first fix didn't cover it.
|
||
|
||
**`src/utils/snapshotGating.js` strips FIVE fields, not one** —
|
||
`model_odds` is the price, **`p_win` IS the price in another base**, and
|
||
**`ev_pct` is INVERTIBLE** (ev = f(p_win, book_odds); book_odds is public, so
|
||
ev hands the price over). `value`/`takeable` bound it. Market legs
|
||
(`book_odds`, `fair_odds`, `fair_prob`, `overround`, `devig_method`) stay on
|
||
**every** tier — **the fair leg is never the paywall**.
|
||
`src/utils/requestTier.js` resolves tier from a bearer token when present and
|
||
**FAILS CLOSED to `free`** on every error path, so a resolution failure can
|
||
only withhold, never leak. Responses now vary by entitlement → `/:sport` and
|
||
`/hero-prop` emit `private` Cache-Control for authenticated callers and both
|
||
Next proxies forward the bearer token.
|
||
|
||
## Read card — why it was empty, and the fix
|
||
A manual scan carries **no market** (`{player, stat, line, direction}`), so the
|
||
engine has no over/under to de-vig — `book_odds`/`fair_odds` were legitimately
|
||
absent. `lookupSnapshotPrices` in the Next `/api/scan` route recovers them from
|
||
the pre-graded snapshot (the same cache-only read it already did for locked
|
||
odds + team). Join is EXACT on **player + stat + line + side** (`fair_odds` is
|
||
side-specific) and returns nothing unless book AND fair are both present — a
|
||
user-chosen line the board never graded borrows no other line's price.
|
||
|
||
## 🔴 LIVE INDUCTION CAUGHT A SECOND CLASS OF BUG
|
||
With the leg stripped, `LiveHeroProp` forwarded everything to `PriceTriplet`
|
||
**except `model_price_locked`** → `deriveValueState` fell to the
|
||
missing-model-price branch and the landing page told every anonymous visitor
|
||
**"MODEL READ WITHHELD"** — quarantine's copy, which says our own data is
|
||
untrustworthy. Nothing was poisoned; it was the paywall.
|
||
**A paywall must never wear poison's face.** Unit tests and markup review both
|
||
passed this; only rendering the deployed page caught it.
|
||
|
||
## Live fingerprint (anonymous, cache-busted)
|
||
- Snapshot: model fields `[]`, market fields intact, `model_price_locked: true`
|
||
- Read card `/api/scan`: `book -150 · fair -126 · model null · locked true`
|
||
- Landing hero renders: **BOOK −153 · ◆FAIR −129 (amber) · MODEL [LOCK BAR]**,
|
||
verdict **"MODEL PRICE ON ANALYST"** + UNLOCK CTA
|
||
- `/dashboard`, `/ledger`, `/parlay` — 200, no triplet, unchanged. No consumer
|
||
reads `ev_pct`/`model_odds`/`p_win`, so the strip broke nothing.
|
||
|
||
## 📊 FAIR-LEG ABSENCE: 0.0%
|
||
636 graded rows: 636 with book, 636 with fair, **0 one-sided**. The hero number
|
||
is not a sometimes-number on current data.
|
||
|
||
## ⚠️ PRODUCT CALL PARKED FOR KEV
|
||
The landing hero is anonymous, so it now shows the LOCK instead of the full
|
||
triplet it showed this morning (`VALUE +21.1% VS FAIR`). That follows the
|
||
stated free-tier rule exactly, but trades a strong shop-window moment for gate
|
||
consistency. **Reversal is one line** — add `model_price: true` to the free
|
||
tier in `src/config/tiers.js`, or special-case the hero route.
|
||
|
||
## Still open
|
||
- **Quarantine flag is NOT available at snapshot build.** No service writes
|
||
`quarantine_reason` (manual SQL only); it's a retrospective judgement, so a
|
||
fresh row can never carry it. The component handles the state if it ever
|
||
arrives; wiring an always-null field was refused.
|
||
- **Slate row** still held — `lib/marketBreadth.js` computes LINES only.
|
||
- **Ledger / public profile** still out — no fair-odds column.
|
||
- Screenshots unavailable again (browser capture tool errored all session);
|
||
substituted computed-style readback per state.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 66 — PRICE-LAYER TOKENS + THE PRICE TRIPLET (2026-07-20)
|
||
Path 3: establish the token foundation, build the hero feature natively on it,
|
||
migrate nothing. Design source = `specs/design-reference/Vyndr Price
|
||
Triplet.dc.html` (ACT 01) + `HANDOFF.md`, committed in `f549422`.
|
||
|
||
## ⚠️ THE PHASE-0 FINDING — the token layer already existed
|
||
The order asked me to establish a token layer. **It was already there and
|
||
already correct.** `web/src/app/globals.css` `:root` carries HANDOFF's
|
||
surfaces / borders / text ramp / fonts / grade colours byte-for-byte (aligned
|
||
2026-07-16), and `web/src/lib/colorContract.js` already encodes green-is-edge-
|
||
only + glow-is-A-tier-only with an enforcing test. Building a second parallel
|
||
layer would have created two competing sources of truth, so this session
|
||
**EXTENDED** the existing one. If a future order says "establish tokens", check
|
||
globals.css first.
|
||
|
||
**Vehicle (confirmed against the stack, not assumed):** Tailwind v4 CSS-first
|
||
(`@import "tailwindcss"`, NO tailwind.config file) + inline `style={{}}`
|
||
components reading `var(--x)` — 2,305 inline-style sites and 1,916 `var(--…)`
|
||
reads. **CSS custom properties are the only vehicle this stack consumes
|
||
natively.** A TS token module or a Tailwind theme config would have fought it.
|
||
|
||
## What was ADDED (additive proof: 149 insertions, 0 deletions in existing files)
|
||
- **One genuinely new colour** — the priced-out blue `--priced-out #8fb2de`
|
||
(+ dim/tints/borders). It is the carve-out that lets a row say "there IS
|
||
edge" without green saying "take it".
|
||
- `--glow-a` (the A-tier glow finally tokenized — the law lived in
|
||
`gradeGlows()` with no value to hand back) and the fair-leg tints.
|
||
- **The LAWS are written INTO the token block** and asserted by
|
||
`tests/unit/priceTriplet.test.js`: green = takeable edge only · glow =
|
||
A-tier only · amber = caution + the fair leg · red = miss/negative only ·
|
||
blue = edge priced out · JetBrains Mono = all data. A test also asserts every
|
||
newly-declared name is NEW (zero collisions, zero overrides).
|
||
|
||
## The law layer — `web/src/lib/valueState.js`
|
||
ONE function (`deriveValueState`) answers "is this value?"; the component
|
||
renders what it returns and never re-derives a verdict.
|
||
- **VALUE requires `ev >= 2` AND a takeable price (−160…+200). Both.** Raw
|
||
positive EV is NOT value — +11.7% EV at −210 is a juiced price we won't call
|
||
a play. The constants MIRROR `src/config/valueEngine.js` and a test reads
|
||
BOTH files and fails on drift (the `playerName.js` pattern).
|
||
- Five states + a sixth display state: VALUE · PRICED_OUT · NO_EDGE ·
|
||
QUARANTINE · REFUSAL · NO_VERDICT_LOCKED (free tier — entitled data withheld,
|
||
which is NOT quarantine and NOT refusal).
|
||
- **Free tier gates at the WIRE**: `tiers.js` gains `model_price`,
|
||
`utils/tierGating` deletes `model_odds` and sets `model_price_locked`. Book
|
||
and fair pass through on EVERY tier — **the fair leg is never the paywall**.
|
||
|
||
## 🔴 INDUCTION CAUGHT A REAL BUG THAT REVIEW WOULD NOT HAVE
|
||
The "VS FAIR" figure compared **BOOK to fair**, so state 1 rendered
|
||
**"VALUE · −6.5% VS FAIR"** — a contradiction on screen, from code that read
|
||
fine and passed its tests. The design's own two worked examples pin the
|
||
formula: **MODEL minus FAIR in implied-probability PERCENTAGE POINTS**
|
||
(book +125 · fair +110 · model +98 → +2.9; book +118 · fair +104 · model +112
|
||
→ −1.8). `modelVsFair` now reproduces both exactly and a test locks them.
|
||
The figure renders **only when its sign agrees with the verdict** — a row that
|
||
clears the EV bar on the book price while our price sits level with fair leads
|
||
with the EV instead of a number that reads as a contradiction.
|
||
|
||
## Proof (induced in a real browser; computed styles, not markup)
|
||
| State | Rendered |
|
||
|---|---|
|
||
| VALUE | model + verdict `rgb(0,212,160)` — the ONLY green in the set |
|
||
| EDGE · NOT TAKEABLE | verdict `rgb(143,178,222)`, model leg **white, not green** |
|
||
| NO EDGE | grey verdict, stated at full voice |
|
||
| QUARANTINE | `MODEL —` withheld, **book + fair still render**, amber |
|
||
| REFUSAL | no legs, no gauge — "CAN'T PRICE THIS ONE" |
|
||
| FREE TIER | lock bar; book −120 + fair −104 still honest |
|
||
|
||
**Live landing hero (real `/api/hero-prop` data, not samples):** book −153 ·
|
||
fair −129 · model −343 · `VALUE +21.1% VS FAIR` · +28.1% EV.
|
||
**Mobile 390:** three legs hold at 117px each, right edge 378 < 390, no
|
||
horizontal overflow, fair no smaller than its neighbours.
|
||
**Post-deploy fingerprint:** tokens resolve in prod (`--priced-out: #8fb2de`,
|
||
`--glow-a`), triplet live on the landing hero.
|
||
**Additive proof:** `/dashboard` and `/ledger` return 200 with NO triplet and
|
||
were not touched — 0 deletions across every existing file.
|
||
|
||
*(Screenshots were not captured — the browser screenshot tool errored on every
|
||
attempt this session. Substituted stronger evidence: computed-style readback
|
||
per state and measured layout at a real 390px column.)*
|
||
|
||
## Scope held / deferred
|
||
- **Slate row HELD** — `lib/marketBreadth.js` computes LINES only (median book
|
||
line vs model value); it has no prices at all, so the triplet needs new
|
||
upstream fields there. Not a same-pass fit.
|
||
- **Ledger + public profile OUT** — `ledger_entries` has no fair-odds column
|
||
(`dclv_fair_lock` is 0/707). Needs a schema decision.
|
||
- **The rest of the redesign is a ranked migration backlog** — see the
|
||
Session-65 inventory. Nothing was migrated this order.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 65 — TWO TRUTH-LAW FIXES (2026-07-20, from Chrome's logged-out audit)
|
||
Both found by auditing the product as a STRANGER sees it. Neither touches the
|
||
triplet build; push scoring untouched.
|
||
|
||
## Doc drift corrected in this session
|
||
The header of this file was stale at `7ea0af2` and this file carried **no record
|
||
of the directional-CLV session** (`dcdad60` → `ca9ca34`: `src/services/
|
||
directionalClv.js`, the Analyst+Desk server-gated `ClvBadge`, and removal of the
|
||
legacy C4 CLV chip from five surfaces). `BUILD-STATE.md` is stale to April 2026 —
|
||
**`specs/STATE.md` is the authoritative state file**; treat BUILD-STATE as archive.
|
||
|
||
## FIX 1 — "CLV-VERIFIED" UN-CLAIMED ON THE PUBLIC PROFILE (copy + 2 render blocks)
|
||
`/u/[handle]` was titled **"CLV-verified record"** with the subhead **"closing-line
|
||
value included"** while **zero** closing-line value renders there. Verified live,
|
||
not assumed: `GET /api/profiles/vyndr` → `beat_close_pct: null` (gated behind
|
||
`clvCaptureReliable()` / `CLV_CAPTURE_RELIABLE`, which is unset while C4 is open).
|
||
A verification claim with nothing verified shown is an unearned claim, and it
|
||
leaked the standing CLV hold onto a surface that had never been audited.
|
||
|
||
**EIGHT instances found — two of them ONLY by the post-removal residual sweep:**
|
||
|
||
| # | Where | Was |
|
||
|---|---|---|
|
||
| 1 | `u/[handle]/page.tsx` | `<title>` / `og:title` / `twitter:title` — "CLV-verified record" |
|
||
| 2 | `u/[handle]/page.tsx` | meta + og + twitter description — "closing-line value included" |
|
||
| 3 | `u/[handle]/opengraph-image.tsx` | no-record fallback tagline |
|
||
| 4 | `u/[handle]/portrait/route.tsx` | no-record fallback tagline |
|
||
| 5 | `PublicProfile.tsx` | house subhead — "closing-line value included" |
|
||
| 6 | `PublicProfile.tsx` | `✓ CLV-VERIFIED` badge (gated, dark, one env flag from re-claiming) |
|
||
| **7** | `opengraph-image.tsx` | **eyebrow `CLV-VERIFIED RECORD · 30D`** — missed in the first report |
|
||
| **8** | `portrait/route.tsx` | **same eyebrow** — missed in the first report |
|
||
|
||
Instances 3 + 4 printed the verification claim **exactly when there is no record
|
||
at all** (the `rec == null` branch), i.e. the claim was loudest where the data was
|
||
most absent.
|
||
|
||
**Applied:** honest copy on 1–5 + 7–8 (settled record · wins and misses · nothing
|
||
curated); the badge (6a) and the `BEAT CLOSE` figure (6b) REMOVED from the public
|
||
profile, OG card and portrait card. **DISPLAY ONLY — `beat_close_pct`,
|
||
`clvCaptureReliable()`, `closing_captures` and the whole CLV data path are
|
||
untouched**, and the earned directional badge stays Analyst+Desk (never public).
|
||
The claim returns here when CLV genuinely renders here.
|
||
|
||
**Title double-append fixed and VERIFIED ON COMPOSED OUTPUT, not source:** the page
|
||
title carried its own `· VYNDR` while `layout.tsx`'s metadata template already
|
||
appends `%s · VYNDR` → live was `…@vyndr · VYNDR · VYNDR`. Built the app, served
|
||
it, and read the real HTML: `<title>Settled record — @vyndr · VYNDR</title>` —
|
||
exactly one suffix. Note `openGraph`/`twitter` titles **bypass** the template, so
|
||
they carry the brand suffix explicitly.
|
||
|
||
Both share images re-rendered locally (200 / image/png) and eyeballed: no overflow,
|
||
no clip, no residual CLV scaffold.
|
||
|
||
**POST-DEPLOY FINGERPRINT (live, logged-out, `f5156dd`):** `<title>Settled record
|
||
— @vyndr · VYNDR</title>` (one suffix), og/twitter title + description un-claimed,
|
||
**0 occurrences** of CLV-verified / BEAT CLOSE / "closing-line value included" in
|
||
the served HTML, and the live OG PNG re-rendered clean. Player pages fingerprinted
|
||
in Chrome across **MLB (Josh Bell, 1 graded prop — the exact row that produced
|
||
"74"), WNBA (A'ja Wilson, 2 props), NBA (Jayson Tatum, 0 props)**: FORM and MATCHUP
|
||
render `—` with intact layout and **zero console errors**; USAGE/REST still show
|
||
real values where they exist.
|
||
|
||
## FIX 2 — FORM WAS FABRICATED ON THE PLAYER PAGE (un-fabricate; steps 2+3 HELD)
|
||
Chrome saw Josh Bell as **FORM 74** (player page) and **Form 63** (generated read)
|
||
in one session. Traced both. **Neither feeds the grade** — proven, not assumed:
|
||
`engine1.gradeProp` runs at `analyzeViaEngine1.js:446`, `buildIntelFields` (which
|
||
computes form) decorates the already-graded object at `:468`; `engine1.js` never
|
||
reads a `form` key (only raw `l5_avg`/`l20_avg` **against the line**), and nothing
|
||
in `src/` reads `result.form`. Grade inputs are byte-identical before/after.
|
||
|
||
- **63 is real** — `70 + (l5_avg / l20_avg − 1) × 60`, per-prop, per-stat.
|
||
- **74 was manufactured** — `70 + 4 × (count of tonight's graded props)`, capped
|
||
92. **Nothing on the HTTP path ever sets `stats.form`**, so that fallback WAS
|
||
the live value. Confirmed empirically: Bell's payload carries `activeProps`
|
||
length **1** → 70 + 4 = **74**, exact.
|
||
- It **poisoned a second tile**: `MATCHUP` was `gradeFromForm(form)` — Bell's
|
||
"B" was `gradeFromForm(74)`, from the prop count, with **no opponent input on
|
||
that path at all**; the no-archetype branch hardcoded `'B'`. Both fabricated.
|
||
- **SYSTEMIC, not Bell-specific** — `buildIntel` is the unconditional path for
|
||
every player and every sport.
|
||
|
||
**Applied (absent renders absent):** FORM and MATCHUP now emit `kind:'plain'`
|
||
with `'—'` when there is no real value — `'plain'` deliberately, so the page
|
||
draws no progress bar and computes no width/colour off a null. `gradeFromForm`
|
||
deleted; the prop count is no longer passed into `buildIntel` at all.
|
||
`computeFormScore`'s hardcoded **75** (returned whenever the baseline was missing
|
||
or zero) now returns `undefined` → the card's self-hiding intel section drops the
|
||
row. Induced live across **MLB / NBA / WNBA**: all render `—` cleanly; real values
|
||
(`USAGE 3.6 AB/G`, `REST B2B`) still render — the fix is surgical to fabrications.
|
||
|
||
**HELD (needs a per-sport headline-stat design call, deliberately NOT built):** a
|
||
real player-level form metric + label disambiguation (page-level vs per-stat). It
|
||
lands as `stats.form` / `stats.matchup` and lights these tiles straight back up.
|
||
|
||
### Audit of the same file (order item 10) — reported, NOT changed
|
||
`toNum()` (`playerIntelService.js:18`) returns **0** for absent, feeding
|
||
`classifierInput`: `k_rate` (`pa > 0 ? … : 0`), `ip_per_start`, `k9`, `saves`.
|
||
These are **classifier inputs, not display** — a 0 there can misclassify an
|
||
archetype (0% K-rate reads as elite contact). Same family as the `Number(null)
|
||
=== 0` rule. `mlbLast10Rows` renders `${st.hits ?? 0}-${st.atBats ?? 0}`, which
|
||
prints a real-looking `0-0` line if a game-log row lacks its stat object.
|
||
**Ambiguous — changing classifier inputs shifts archetypes, so it needs its own
|
||
order.** Display rows elsewhere (`mlbSeasonRows`, ESPN season rows) correctly use
|
||
`?? '—'`.
|
||
|
||
|
||
---
|
||
|
||
## 🟢 MLB opp_rank_stat LIVE — computed, CONSUMED, retained (2026-07-20 16:51 UTC)
|
||
- **Consumption path verified BEFORE wiring** (a value the grader doesn't read
|
||
is a dead end): `featureCache.teamFeatures` sets `out.opp_rank_stat` →
|
||
`engine1.computeFactors` READS `features.opp_rank_stat` → fires
|
||
`weak_opponent_defense` (≥0.70) / `top_opponent_defense` (≤0.30).
|
||
- **Source:** `mlbStatsAdapter.getTeamPitchingStats` — all 30 teams in ONE free
|
||
unauthenticated statsapi call. Only fills when the ESPN path yields nothing
|
||
(always, for MLB), so WNBA is untouched.
|
||
- **CONSUMPTION PROVEN on the exact retained vector** (Josh Bell hits o0.5),
|
||
varying only `opp_rank_stat`:
|
||
|
||
| opp_rank_stat | grade | factors |
|
||
|---|---|---|
|
||
| 0.367 (as retained, mid) | **B-** | l5_hot_vs_line, l20_over_line, back_to_back |
|
||
| 0.98 (weak) | **B** | + **weak_opponent_defense** |
|
||
| 0.02 (tough) | **C** | + **top_opponent_defense**, away_vs_top5_defense |
|
||
| absent | **B-** | identical to mid — neither fires ✅ |
|
||
|
||
- **RETAINED** grade-time in `model_snapshots.features` (immutable, joinable).
|
||
|
||
### 📊 COVERAGE FIRST — then the grade read
|
||
**8 of 14 graded MLB rows carried a non-null `opp_rank_stat` (57 %); 6 null.**
|
||
The nulls are almost certainly unresolved `opponentAbbr` (`teamFeatures` returns
|
||
early without one), NOT the sample floor — but that is **undetermined without
|
||
container logs**, so it is reported as such rather than asserted.
|
||
|
||
**🔴 THE FACTOR FIRED ON ZERO ROWS THIS CYCLE.** All 8 non-null values were the
|
||
SAME mid-scale `0.367` (one opponent, neither threshold crossed:
|
||
`would_fire_weak 0`, `would_fire_tough 0`). **So this slate cannot test
|
||
A-emission at all** — the factor was present but inert. Any "A didn't emit"
|
||
reading from tonight is meaningless.
|
||
|
||
**Live MLB 11-step distribution (14 graded):** `C+ 4 · F 4 · B- 2 · C 2 · D 2`.
|
||
**No A.** Notably **F now emits (4 rows)** — the range has widened DOWNWARD, as
|
||
the restored ±1.0 makes Σδ ≤ −2.51 reachable for the first time.
|
||
|
||
**A-rated marketing stays HELD.** No thresholds changed. A genuine test needs a
|
||
slate with a weak-BAA opponent (≥0.70) so the +1.0 can actually lift a read.
|
||
|
||
### Thin-sample floor in production
|
||
No thin-sample opponent appeared on this slate, so the production null-floor path
|
||
was **not exercised live** — unit-tested only. Stated rather than implied.
|
||
|
||
## 🟢 HARNESS ARMED + CLOSING CAPTURE STARTED (2026-07-20 ~16:10 UTC)
|
||
Both verified by INDUCING the real code path — never by waiting for a slot.
|
||
|
||
**A — Backtest harness runs on OUR scheduler.** `snapshotScheduler` fires it at
|
||
`HARNESS_HOUR_UTC` (default 14), appends to `harness_results`, and
|
||
`opsWatch.harnessStaleAlarm` pages if it stops (a validator that isn't running
|
||
looks exactly like one that keeps passing). No external dependency.
|
||
**INDUCED live:** `verdict INSUFFICIENT_HISTORY, scored 35, joined 283,
|
||
snapshots 478` — the correct, passing output. `POST /api/internal/harness/run`.
|
||
|
||
**B — Closing capture (CAPTURE ONLY; no CLV metric, as ordered).**
|
||
- **C4 diagnosed:** `closing_line` is ONE overwritable field with no timestamp
|
||
and no provenance. `captureClosing` writes the current line and, on a prop
|
||
match failure, silently leaves the earlier value (= the lock) in place — so
|
||
"captured a real close" is indistinguishable from "never updated". It is
|
||
**92 % equal, not 100 %** (56 rows DID move): the defect is **provenance**.
|
||
- **`closing_captures` (migration 029)** — append-only, never overwritten, BOTH
|
||
raw side prices retained so the existing de-vig engine computes the fair
|
||
closing probability later. Sharp (`pinnacle`) vs book close kept distinct:
|
||
"beat the market" and "beat our book" are different questions.
|
||
- **JOIN drops `line` on purpose** — a close that MOVED off the graded line is
|
||
the whole point of CLV. Verified safe: all 164 identity groups have exactly
|
||
ONE line per `(sport, player_key, stat, side, game_date)`; zero ambiguity.
|
||
- **Wired into the intraday refresh** (already polls ~20 min through slate
|
||
hours) with a capture-rate alarm — a missed close is unrecoverable.
|
||
- **🔴 ODAWA IS NOT CAPTURABLE.** `odawa` exists only as a UI preference option
|
||
in onboarding/settings — no adapter, no `ALLOWED_BOOKS` entry, no feed. CLV
|
||
vs our own book cannot be measured until a real Odawa feed exists.
|
||
- **Refusals proven first (test-first):** unbound game_time, doubleheader
|
||
ambiguity, missed window, or a one-sided price → `missed_reason` and **NO
|
||
price**. Induced live: 1,076 rows correctly refused `one_sided_price`.
|
||
|
||
### ⚠️ MY OWN ERROR, CAUGHT AND REVERTED
|
||
To force a capture I widened the window to 1,200 min and **persisted 7,114
|
||
rows** — every one had a lead time of **392–602 minutes**, i.e. 6.5–10 hours
|
||
before lock. **Those were not closes.** Leaving them in a table named
|
||
`closing_captures` would have poisoned the first CLV computation with pseudo-
|
||
closes. **All 7,114 deleted; table verified back to 0.** Verification is now
|
||
done with `?dry=1`, which proves the pricing path (wnba: 486 rows priced, 0
|
||
missed) while writing nothing. The real 45-min window currently yields 0 rows
|
||
because no game is near lock — which is the correct answer, not a failure.
|
||
|
||
## 🟢 HEAL EXECUTED (Order 2, 2026-07-20 ~10:00 UTC) — record is clean
|
||
Rollback point first: `vyndr-20260720-093821.dump` (856,890 B) **verified ON the
|
||
Storage Box**, not just exit 0.
|
||
|
||
**⚠️ THE DRY-RUN OVERTURNED MY OWN DIAGNOSIS — read this before trusting the
|
||
Order 1.5 notes above.** Across every bindable row the stored `game_date`
|
||
matched a real game (**MIS-DATED: 0**). The players I cited as proof of
|
||
mis-dating were **genuine DNPs** on their true dates (Freeman 07-18;
|
||
Kwan/Hedges/Davis 07-17 — their teams played, they did not). **The original 64
|
||
`player_dnp` voids were CORRECT and my Order 1.5 revert undid correct work.**
|
||
|
||
| Axis | Result |
|
||
|---|---|
|
||
| Settlement scorer | **578/578 re-computed identical, 0 differ** — confirmed, no heal needed |
|
||
| Dates | 527 correct · **0 mis-dated** · 170 unbindable (my dry-run's parser, NOT proven damage) |
|
||
| Grade damage | **51 rows, not 303** — the rest hit the same opponent anyway (baseball SERIES effect) |
|
||
|
||
**Writes executed (each verified against the dry-run prediction):**
|
||
- **64 re-voided** (60 mlb + 4 wnba) — by the settle pass's positive-evidence
|
||
guard, not hand-picked ids. DNP voiding re-enabled deliberately because its
|
||
disable rested on a false premise.
|
||
- **25 ledger rows quarantined** `wrong_opponent_grade` (wnba 2026-07-18,
|
||
10 hit / 15 miss). They REMAIN real public settled results; only train/validate
|
||
excludes them.
|
||
- **200 `model_snapshots` rows quarantined** — the 4 pre-fix MLB cycles for
|
||
game_date 07-20 (06:00/06:02/06:03/07:25 UTC). Superseded by the correct
|
||
07:50/07:51 captures of the same props, so nothing is lost.
|
||
- **70 rows flagged** `doubleheader_game_identified`.
|
||
|
||
**TWO EXCLUSION SCOPES (migration 027) — deliberately separate:**
|
||
- `quarantine_reason` → grade untrustworthy. Row stays a real public result but
|
||
is **excluded from the model-record denominator** (`getModelAggregate`), like
|
||
void/unrecoverable.
|
||
- `analysis_flags` → valid for settlement AND the record; excluded only from
|
||
per-game/opponent analysis. **Never filtered from aggregates** (a test asserts
|
||
this, so a future edit can't silently collapse the two).
|
||
|
||
**Corrections to my own Phase 0 numbers, found while executing:**
|
||
- The 50 WNBA "poisoned moat rows" were **NOT poisoned** — captured 03:02 UTC
|
||
(23:02 ET 07-19) for game_date 07-19, so ESPN's "today" was the right day.
|
||
- The 166 doubleheader rows came from **loose team matching** and could not be
|
||
reconciled (a strict predicate found 121). I flagged only the **70** rows the
|
||
FEED itself marks (`G2…`/`(Game1)`) — and those are attributable, not
|
||
unattributed. The wider unattributed set needs a stricter pass; it is advisory
|
||
metadata that affects no number today.
|
||
|
||
**Final state:** 360 hit · 218 miss · 64 void · 55 pending · 25 quarantined
|
||
(697 public rows). 26 pending mlb 07-20 rows left untouched to self-heal on the
|
||
next cycle, as ordered.
|
||
|
||
## 🟢 GRADING DATE BINDING FIXED — wrong-opponent bug closed (2026-07-20 07:50 UTC)
|
||
`computeFeatures.lookupTodayGame` → **`lookupGameOnDate`**: it now sends
|
||
`?dates=YYYYMMDD` from the prop's BOUND game instead of taking ESPN's dateless
|
||
"today". Grading, settlement, retention and the ledger finally reference the
|
||
SAME game. **This was a MODEL-OUTPUT bug**: `opponentAbbr` feeds `opp_rank_stat`
|
||
(a ±1.0 factor) and `isHome` feeds `home_away` (+0.5).
|
||
|
||
**PROVEN against live ESPN, same instant:**
|
||
```
|
||
dateless "today" CLE→PIT NYY→LAD LAD→NYY (Jul 19 card)
|
||
bound to 2026-07-20 CLE→MIN NYY→PIT LAD→PHI (the real games)
|
||
bound to 2026-07-19 CLE→PIT NYY→LAD LAD→NYY (reproduces OLD exactly)
|
||
```
|
||
**Every opponent was wrong.** Measured at 07:5x UTC = **03:5x ET** — ESPN's
|
||
dateless scoreboard was STILL serving the previous day's card, so **the affected
|
||
window is wider than the 01:00/03:00 slots**.
|
||
|
||
- **Honest degradation:** no bound date → NO fallback lookup. Records
|
||
`no_bound_game_date`, leaves `opponentAbbr`/`isHome`/`gameId` null, and engine1
|
||
omits those factors rather than scoring a wrong matchup. Tests lock both ways.
|
||
- **Same class, also fixed:** Tank01 augmentation keyed its cache on TODAY's UTC
|
||
date; now uses the bound game date.
|
||
- **Audited:** no other dateless/"today" lookups remain in the feature path
|
||
(weather is current-conditions by venue; park/pace are static).
|
||
- Live cycles at 07:50/07:51 UTC wrote retention rows with `game_date 2026-07-20`
|
||
and `home_away` on 38/50 — games binding. MLB `opp_rank_stat` stays 0 for the
|
||
separate, known reason (ESPN's MLB team endpoint carries no defensive metric).
|
||
|
||
### 🔴 BLAST RADIUS (report-only — gates the heal)
|
||
**`ledger_entries` — grades by UTC hour (at-risk = late slots):**
|
||
|
||
| UTC hour | grades | settled results | risk |
|
||
|---|---|---|---|
|
||
| 00 | 32 | 29 | 🔴 at risk |
|
||
| 01 | 68 | 67 | 🔴 at risk |
|
||
| 03 | 152 | 103 | 🔴 at risk |
|
||
| 06 | 51 | 0 | 🔴 at risk (ESPN still on prior card at 02 ET) |
|
||
| 13/14/19/22 | 390 | 379 | ✅ same-ET-day, correct |
|
||
|
||
**At-risk: 303 grades (199 already settled with real results).** ~44 % of the
|
||
public record was graded in the affected window.
|
||
|
||
**`model_snapshots`:** 350 rows total — **50 carry `opp_rank_stat`** (the 03:02
|
||
UTC WNBA cycle) and are the poisoned moat rows; 190 carry `home_away`. Rows
|
||
captured 07:50+ are clean (post-fix).
|
||
|
||
**Recoverable vs quarantine (PLAN ONLY, nothing written):**
|
||
- **RECOVERABLE:** the correct opponent for a past date is still fetchable from
|
||
the date-pinned ESPN scoreboard, so a grade can be RE-DERIVED — but only where
|
||
the feature vector is reproducible. `model_snapshots` rows store the full
|
||
feature vector, so those are re-computable; **ledger rows do not** and would
|
||
need a full re-grade against historical inputs.
|
||
- **QUARANTINE:** any at-risk ledger grade that cannot be re-derived should be
|
||
marked computed-against-wrong-opponent and EXCLUDED from harness/validation —
|
||
never silently trusted. It stays in the public record as a real settled bet
|
||
(the result happened) but must not train or validate anything.
|
||
- **Scope note for the heal:** at-risk is defined by grading HOUR, not by proven
|
||
wrongness. Proving per-row wrongness requires re-deriving each row's opponent —
|
||
that IS the heal work.
|
||
|
||
## 🟢 game_date ROOT FIXED + 64 WRONG VOIDS REVERTED (2026-07-20 07:30 UTC)
|
||
- **ROOT (not DNP, as first diagnosed):** PropLine emits **no `commence_time`**
|
||
(grep-verified), so `ledgerService`'s `dateET(prop.game_time) ||
|
||
dateET(gradedTs)` always fell through to the **GRADE clock**. A 01:00/03:00
|
||
UTC snapshot is 21:00/23:00 ET the PREVIOUS day → tonight's props filed under
|
||
yesterday → settlement correctly found no game → Order 1's void logic
|
||
destroyed 64 real results.
|
||
- **FIX — `src/services/gameBinder.js`:** binds each prop to a scheduled game by
|
||
TEAMS across the ET window (grade date, +1, −1) and attaches the GAME'S own
|
||
time/date/id, in `snapshotService` **before** grading and the ledger write, so
|
||
ledger + retention + settlement all inherit one correct date.
|
||
**PROVEN against the live schedule at the failing hour:** a `2026-07-20T03:00Z`
|
||
grade (ET date 07-19, what the old code used) bound Steven Kwan to
|
||
**game_date 2026-07-20**, `mlb:2026-07-20:401816188`.
|
||
- **HARD CONTRACT:** unbindable → **nothing**. The grade-clock fallback is gone;
|
||
ledgerService SKIPS and counts such rows. Verified live: a prop with an
|
||
invented matchup returned `unresolved`, never a guessed date.
|
||
- **Doubleheaders** are flagged `ambiguous`, never attributed. Real case already
|
||
in data: `mlb:2026-07-11:MilwaukeeBrewers@PittsburghPirates(Game1)`.
|
||
- **Retention had the same bug** (rows dated by the snapshot clock) — fixed to
|
||
the bound game's ET date.
|
||
- **64 wrong voids REVERTED** (60 mlb + 4 wnba), scoped exactly to
|
||
`settlement_version=2 AND outcome='void' AND settlement_source='player_dnp'`.
|
||
- **LIVE PROOF of the safety fix:** the next induced settle pass returned
|
||
**`voided: 0, pending: 64`**. The old code voided all 64; the deployed code
|
||
refuses to claim DNP on an unconfirmed absence.
|
||
|
||
### 🔴 STOP-AND-REPORT — the date bug also reaches GRADING (not just settlement)
|
||
`computeFeatures.lookupTodayGame` calls the ESPN scoreboard with **no date
|
||
param** — it takes whatever ESPN calls "today". A prop graded at 23:00 ET can
|
||
therefore bind its `home_away` / `opponentAbbr` features to the WRONG game
|
||
(yesterday's finished one), which feeds `opp_rank_stat` and the home/away factor.
|
||
**Not fixed — flagged per the order's stop rule.** Blast radius:
|
||
- **Grading: AFFECTED** (feature correctness at late-ET slots — 01:00/03:00 UTC).
|
||
- **Settlement: FIXED** (this order).
|
||
- **Retention: FIXED** (this order).
|
||
- **Display/slate: NOT affected** — the board renders from the date-pinned
|
||
schedule and joins grades by player+stat, never by ledger `game_date`.
|
||
|
||
### 🔶 STILL OPEN — the reverted 64 cannot settle yet
|
||
Their `game_date` is still the OLD wrong value (Jul 17/18); the binder only
|
||
corrects rows written from now on. They will retry and age to `unrecoverable`
|
||
until their dates are corrected. **Correcting them is the first act of the heal
|
||
order** — healing against wrong dates would simply re-fail.
|
||
|
||
## 🟢 RETENTION CLOCK TICKS ON ITS OWN (2026-07-20 06:05 UTC)
|
||
Induced the REAL cron entrypoint rather than waiting for 14:00 UTC.
|
||
- **Cron capture needed NO wiring — verified in code.** `snapshotScheduler`'s
|
||
tick calls `runAll` = `snapshotService.runAllSnapshots` → `runSnapshot` per
|
||
sport → `gradeAndCacheSlate` → `onGraded` → retention. **The scheduled and
|
||
manual paths are the same function.** Nothing had been captured only because
|
||
no slot had fired since retention deployed (~02:55; slots 14/19/22/1/3).
|
||
- **PROVEN by induction:** `POST /api/internal/snapshot/all` (= `runAllSnapshots`,
|
||
the exact function the tick calls) wrote cycles at **06:03:23 mlb** and
|
||
**06:05:09 wnba**, 50 rows each, `cycle_hour_utc=6`, features 50/50.
|
||
*(The HTTP request gateway-timed-out at ~2 min; the job completed server-side —
|
||
verified in the DB, not from the response.)*
|
||
- **`archetype` FIXED — 50/50 on MLB** (was 0/100). Retention still COLLECTS at
|
||
grade time (only moment features exist) but now PERSISTS after enrichment via
|
||
`retentionService.mergeEnrichment`, which fills ONLY archetype/team/opponent
|
||
and never rewrites grade-time features or model output (test-locked).
|
||
- **🔶 `team`/`opponent` STILL NULL — upstream, not the merge.** The ENRICHED
|
||
grades themselves carry `team:null` (verified live: 0/25 with team, 25/25 with
|
||
archetype). `resolvePlayerStats` doesn't return `stats.team` for MLB, and
|
||
enrichment never sets `opponent` at all. Fixing that is an upstream change,
|
||
deliberately not made under this order.
|
||
- **WNBA archetype is null** — honest absence: WNBA stats don't resolve (Python
|
||
service offline), so there is nothing to classify from.
|
||
- **ZERO-WRITE ALARM ARMED** — `opsWatch.retentionZeroWriteAlarm` pages at
|
||
missed-snapshot severity when a slot GRADED props but retention wrote fewer
|
||
rows than the slate. `runSnapshot` now returns `retentionRows` so the tick can
|
||
evaluate it. Retention is best-effort by design, so a broken write is silent by
|
||
construction; this is the counterweight. Verified quiet on a healthy run
|
||
(gradeCount 25 → retentionRows 50).
|
||
- **INTRADAY DEFERRED (Phase 4).** `intradayRefreshService` calls
|
||
`analyzeViaEngine1` DIRECTLY — it never goes through `gradeAndCacheSlate`, so
|
||
there is no `onGraded` hook to reuse. Capturing it needs its own ctx +
|
||
snapshot_id + persist path and would ~5× row volume. Not small; not done.
|
||
- Outcome stamping still NOT implemented (depends on the settlement fix).
|
||
|
||
## 🟢 RETENTION LIVE — model_snapshots capturing (2026-07-20, migration 025 applied)
|
||
Phase 2 priority zero. **History now compounds from tonight.** Spec:
|
||
`specs/snapshot-retention.md`.
|
||
- **Why:** the ONLY model history was `ledger_entries` (640 rows / 6 game days)
|
||
and it stores NO model inputs. All ~30 other warehouse tables are empty. We
|
||
could score our own grades but could not replay a different model — the only
|
||
question a backtest answers.
|
||
- **`model_snapshots`** (migration 025, APPLIED): append-only, ONE ROW PER GRADED
|
||
PROP PER SIDE PER CYCLE. Unique on `(snapshot_id, player_key, stat, line,
|
||
side)` so a retried cycle can't duplicate. RLS on, service-role writes only.
|
||
- **VERIFIED CAPTURING (2026-07-20 03:02 UTC, 2 triggered cycles): 100 rows.**
|
||
MLB 50 (14 graded / **36 refused**), WNBA 50 (50 graded). `features`,
|
||
`grade_11`, `p_win`, `ev_pct` populated on 100% of graded rows; 1 model_version.
|
||
- **Refusals are captured — the ledger drops them entirely.** First-ever refusal
|
||
breakdown: `juiced_no_edge` 18, `rare_event_over_below_line` 13,
|
||
`insufficient_data` 5. **The MLB gate refused 36 of 50 sides (72%)** — whether
|
||
that costs edge is now measurable for the first time.
|
||
- Contract: retention is best-effort + injectable; `persist()` returns errors and
|
||
never throws. **A retention failure can never break a snapshot.**
|
||
- 🔶 **EV LOOKS OVERCONFIDENT — do not surface it as truth yet.** First captured
|
||
EVs include +62.1%, +61%, +56.9% — implausible for real markets. Cause is the
|
||
quantile estimator clamping `p_win` at 0.95 (`PROB_CEIL`) off ~10 games.
|
||
**Hero v2 already ranks by highest `ev_pct`, so it will select the MOST
|
||
overconfident read.** Calibration (C-cal) must run on the harness before EV
|
||
drives anything user-facing.
|
||
|
||
## 🔶 SETTLEMENT-CORRECTNESS AUDIT — two concrete findings (open, scheduled)
|
||
Logged from the Phase 2 history report; do NOT fix inside a retention build.
|
||
1. **ZERO pushes across 470 settled rows.** Implausible — hits/TB land exactly on
|
||
the number regularly. The exact-number push case is almost certainly
|
||
mis-scored (silently counted as hit or miss).
|
||
2. **~28 props/day never settle** (Jul 17 MLB 86 graded/57 settled; Jul 18
|
||
103/75). Cause undiagnosed.
|
||
Both are correctness bugs in the scoring that accuracy, ROI, CLV and every public
|
||
record depend on.
|
||
|
||
## 🔴 MODEL-VERSION CONTAMINATION in ledger_entries (permanent, mitigate by cutoff)
|
||
`ledger_entries` mixes pre- and post-2026-07-19-fix grades with **no version
|
||
marker**, and the eras cannot be separated retroactively. **Any backtest or
|
||
accuracy claim computed off existing ledger history must treat the fix boundary
|
||
as a hard cutoff.** `model_snapshots` stamps `model_version` + `code_sha` on every
|
||
row so this can never recur.
|
||
|
||
## 🟢 BACKUP — DURABLE ON-BOX, VERIFIED BY READ-BACK (2026-07-20, deployed)
|
||
The highest-severity open item is closed **on-box**. Off-box is a tracked follow-up.
|
||
- **Runs as CODE, not a cron somebody must install** (`src/backupScheduler.js`,
|
||
armed in `server.js`). Arming is opt-OUT: armed whenever `SUPABASE_DB_URL`
|
||
exists, `BACKUP_CRON=0` kills it, nightly `BACKUP_HOUR_UTC:BACKUP_MINUTE_UTC`
|
||
(default 03:10 UTC). The S62 design was opt-in and nobody ever opted in — that
|
||
is why the DB went unbacked for weeks. Deploy == installed now.
|
||
- **DURABLE:** `BACKUP_DIR=/app/backups` is a Coolify persistent volume, so dumps
|
||
survive redeploys. The container-ephemeral risk is closed.
|
||
- **VERIFIED END-TO-END (2026-07-20 02:31 UTC):** `POST /api/internal/backup/run`
|
||
→ exit 0 in 31s → `vyndr-20260720-023136.dump`, **821,549 bytes**, 698 archive
|
||
objects, ledger_entries present. `GET /api/internal/backup/verify` read the
|
||
rows back OUT of the archive with `pg_restore`: **643 rows in the dump vs 643
|
||
live — exact match**, DB 18 MB.
|
||
- The earlier figure of 604 was a moving target, not a discrepancy: the ledger
|
||
is append-only and 39 rows were written between the census and the dump
|
||
(2 manual regrades + the 01:00 UTC cron; newest lock 01:01:45 UTC).
|
||
604 + 39 = 643 reconciles exactly.
|
||
- **Permission bug found and fixed by the real run:** the first attempt failed
|
||
with `pg_dump: Permission denied` — the container runs as non-root `vyndr`
|
||
(uid 100/gid 101) and the mounted volume was root-owned. The Dockerfile now
|
||
creates + chowns `/app/backups`; the verify endpoint reports `uid`/`gid`/
|
||
`backup_dir_writable` so a mounted-but-unwritable volume is loud, not silent.
|
||
- **🔶 OFF-BOX IS DEFERRED, NOT DONE.** Storage Box SSH auth is not working yet,
|
||
so the rsync is gated behind `BACKUP_OFFBOX=1` (+ `BACKUP_REMOTE` +
|
||
`BACKUP_SSH_KEY`) and currently logs "DEFERRED". A failing push is a LOW
|
||
priority notice, never a failure — the durable on-box dump succeeded, and
|
||
crying wolf trains us to ignore backup alerts. **Remaining risk: a single
|
||
Hetzner-side loss takes the dumps with it.** To finish: install the generated
|
||
public key on the Storage Box, set `BACKUP_SSH_KEY` in Coolify (port 23 —
|
||
22 is SFTP-only), flip `BACKUP_OFFBOX=1`.
|
||
- **Not yet done:** a full restore into a scratch Postgres SERVER. The read-back
|
||
proves the archive CONTAINS the data; a server restore would additionally
|
||
prove schema/constraint restorability. It needs the dump off the container,
|
||
which is blocked on the same Storage Box auth.
|
||
|
||
## MODEL TRAIN — ARC 2 PREP: DEAD PROBABILITY LAYER REVIVED (2026-07-19, on main, DEPLOYED + fingerprinted)
|
||
Re-sequenced ahead of the gate work after the audit found the ranking signal did
|
||
not exist. Reports: `specs/audit-data/gate-simulation.md` (G-b/C-cal) +
|
||
`specs/audit-data/grade-collapse.md` (the big one).
|
||
- **p_win/ev_pct/kelly/model_odds/value were absent on 100% of live grades** —
|
||
`gameLogService` returns null for MLB by construction + the Python service is
|
||
offline, so `meta.gameLogs` was `[]` everywhere. EV, the Model Train's whole
|
||
ranking signal, had never been computed on a live prop; hero v2 always fell
|
||
through to its fallback; quarter-Kelly (sold on /pricing) never ran.
|
||
FIX: `featureCache.getStatRows`. **Fingerprinted live: p_win 0.523, ev_pct
|
||
-10.4, model_odds -109 — the value triplet is whole.**
|
||
- **Grade range restored ON MERIT** (rescaling permanently rejected):
|
||
`refreshTeamStats` wired into `runSnapshot` (it had ZERO callers), L20 made
|
||
symmetric (no negative branch existed), `game_count_in_7d` derived. D now emits
|
||
(WNBA 1/25, earned at p_win 0.365). **A-emission awaits the 01 UTC snapshot** —
|
||
`opp_rank_stat` needs that run. MARKETING HOLD on "A-RATED" copy until then.
|
||
- **Calibration guard:** consistency CV is NBA-tuned; `cv ~ 1/sqrt(mean)` made
|
||
every MLB stat `boom_bust`. Floored at `CONSISTENCY_MIN_MEAN=4`.
|
||
- Ledger truth (576 rows): only B/C ever emitted, 9 confidence values, ceiling 63;
|
||
past -400 was -7.7% ROI (the -400 floor was the whole win); flex -161..-250 is
|
||
our BEST band (+2.2%) so its EV gate ships DISABLED. Suite 276/3286, build 0.
|
||
|
||
## MODEL TRAIN — ARC 1 (ENGINE) SHIPPED (2026-07-19, on main `7a925f4`, pushed, NOT yet deploy-fingerprinted)
|
||
Spec written retroactively: `specs/model-train.md` (arc list + what's open). Goal of the
|
||
train: make "real opportunities at takeable prices" the ENGINE, not a display filter.
|
||
Arc 1 = steps 1–6, backend only. **No frontend consumes the new fields yet** — the
|
||
triplet/VALUE marker/refusal copy are Design's Session-2 surfaces.
|
||
- **De-vig** (`src/utils/devig.js`, NEW): two-way multiplicative de-vig → fair prob +
|
||
fair American price per side + `overround`. One side unpriced → fair is NULL, never
|
||
faked. Also `americanToImpliedProb` / `americanToDecimal` / `impliedProbToAmerican`.
|
||
- **EV** (`devig.evPct`): `ev_pct = modelProb × decimal − 1`, at the graded side's
|
||
ACTUAL price. This is the ranking signal now, replacing raw |model − consensus|.
|
||
- **Takeable / value gates** (`src/config/valueEngine.js`, NEW): `TAKEABLE_ODDS_CEILING`
|
||
**−160**, `TAKEABLE_ODDS_MAX` **+200**, `VALUE_EV_THRESHOLD` **2** (%) — all env-tunable.
|
||
`takeable` = price in band; `value` = takeable AND ev ≥ threshold. Strict null guards
|
||
(`Number(null)===0` would have made a missing price "takeable").
|
||
- **Hero v2** (`heroPropService`): the hero is the highest `ev_pct` among **takeable**
|
||
A/B reads — a huge model-vs-line gap on a −900 line is trivia, not an opportunity.
|
||
Falls back to most-recent-real graded read on an empty slate (unchanged).
|
||
- **Value triplet** on every graded read: `book_odds` · `fair_odds` · `model_odds`
|
||
(+ `fair_prob`, `overround`, `devig_method`). Wired inside `analyzeViaEngine1`'s
|
||
existing p_win/Kelly block; rides the `...result` spread through snapshot / hero /
|
||
scan. Documented in `BACKEND_HANDOFF.md`.
|
||
- **Test baseline: 276 suites / 3306 tests green** (was 274/3289), `next build` exit 0.
|
||
New suites: `devig.test.js`, `valueEngine.test.js` (+ `heroPropService.test.js` rewritten).
|
||
- **Knobs NOT in this arc** (asked about, confirmed absent from the codebase):
|
||
`EDGE_FLEX_WALL`, `HARD_JUICE_WALL`, `LADDER_ODDS_MAX`, `MIN_RUNG_PROBABILITY`.
|
||
The only price wall on disk is the flat `JUICE_ODDS_FLOOR` (−400) from the
|
||
rare-event work below — it is **not** edge-aware, and there is no −250 wall.
|
||
|
||
## RARE-EVENT UNDER SUPPRESSION (2026-07-19, on main)
|
||
Betting-logic audit: the CONSENSUS-vs-MODEL board flooded with "doubles u0.5 · MODEL 0.2 · +edge" fake reads (juiced rare-event unders). Report: the doubles projection is REAL per-player (MLB_LOG_FIELD doubles→doubles; values varied 0.03/0.16/0.2/0.22) — NOT a flat fallback; the issue is purely structural.
|
||
- **Config-driven** (`src/config/rareEventMarkets.js`): RARE_EVENT_STATS = doubles/triples/home_runs/stolen_bases, RARE_EVENT_LINE_MAX = 0.5.
|
||
- **Grade layer** (`analyzeViaEngine1`): rare-event UNDER at ≤0.5 → always refused (grade null + suppressed); rare-event OVER at ≤0.5 → refused unless projection > line (a 0.2-over-0.5 carries the SAME |edge| as the under, so it'd just take its board rank — refusing it is what actually clears the market). OVER with a genuine projection > line still grades.
|
||
- **Board layer** (`marketBreadth.collectBreadth`): drops null-model rows — a suppressed/ungraded prop can't rank a "MODEL —" placeholder onto the board.
|
||
- Suite 274/3289 green. Fingerprint: next MLB snapshot should have no rare-event u0.5 grades.
|
||
|
||
## BACKUP + FOUNDER CHECKOUT (2026-07-18, on main)
|
||
- **Task A — DB backup (SHIPPED; Kev cron+fingerprint on box)** `c2c43cd`: Dockerfile now has pg_dump/pg_restore/rsync; `backup-db.sh` validates every dump via `pg_restore --list` (must contain ledger_entries). Runs IN the API container (SUPABASE_DB_URL is there; WSL2 can't reach Supabase). Runbook = host cron `docker exec <api> sh /app/scripts/backup-db.sh`. **Mechanism fingerprint PASSED locally** (137 rows → dump → validate → restore → 137 rows). Kev: install the host cron + run the prod restore fingerprint.
|
||
- **Task B — founder checkout SEAT-GATED (SHIPPED)** `ccb9668`: `resolveCheckoutPrice` attaches the founder price while seats remain (< FOUNDER_SEATS_TOTAL, same `countFounderSeats()` truth as the meter), flips to standard at 100; meter shows SOLD OUT. Fixes "Claim a Founder Desk" charging $44.99 vs advertised $34.99. **payment_failed grace 48h→14d** (spans Stripe retries; revoke only on real cancel). Tested seat 0/99/100/null. Needs `STRIPE_PRICE_*_FOUNDER` set in prod for founder pricing to activate.
|
||
|
||
## SECURITY + PLUMBING FOLLOW-UP (2026-07-18, on main) — code shipped, migrations authored
|
||
Chrome ops session findings (Supabase advisor + Stripe). Code items shipped + fingerprinted; DB migrations authored for Kev to apply (no DB access from the box).
|
||
- **Item 0 (SHIPPED, fingerprinted)** `3b12c6c`: founder count = real active Stripe subs (`stripeService.countFounderSeats`), not DB tier fields. Phantom 1 (a comped profile) gone → prod now `available:false` (founder price IDs unset in prod env — set `STRIPE_PRICE_*_FOUNDER` to show the real 0).
|
||
- **Items 1,3,4,5 (AUTHORED — Kev applies in Supabase)** `78c1929`: `migrations/023_security_hardening.sql` (founder_pricing_seats → security_invoker; waitlist anon-INSERT-only; function search_path) + `migrations/024_anon_revoke_discoverability.sql` (revoke anon SELECT). Fingerprint = re-run Security Advisor.
|
||
- **Item 2 (SHIPPED — Kev sets env + cron)** `5f5c004`: `scripts/backup-db.sh` + `docs/BACKUP-RUNBOOK.md`. Needs `SUPABASE_DB_URL` (direct :5432 URI) in Coolify + the cron on the Hetzner box. Fingerprint = restore a dump + count ledger_entries.
|
||
- **Item 6 (SHIPPED — Kev activates portal in Stripe)** `889e862`: portal proxy + "Manage billing" button. Dunning verified: cancel-on-exhaustion → subscription.deleted → 48h grace → `gracePeriod.checkGracePeriod` revokes to free. Nuance: 48h grace on first payment_failed < Stripe's 2-week retry window (self-correcting).
|
||
- **Item 7 (SHIPPED)** `ae3cff9`: ESPN BET → theScore Bet (PENN) in books.js + ALLOWED_BOOKS; design HANDOFF flagged for the BookChip swap.
|
||
|
||
## TRUTH-EVERYWHERE PART 2 — COMPLETE (2026-07-18, on main, all 8 items fingerprinted)
|
||
Standing law: NO FABRICATED DATA ANYWHERE (UI, marketing, demos, examples). If it renders a number, it's from the DB or it doesn't render. Live data IS the better demo.
|
||
- **Item 1 — Verb law** (`66d52a9`): one verb, READ (never SCAN), swept every user-visible surface (FAB/nav/tier/ticker/upgrade). `tests/unit/verbLaw.test.js` enforces. Ticker "slate read" text flips at the next snapshot cron.
|
||
- **Item 2 — Founder counter** (`41fc2b9`): real active-paying-founder count from Stripe-synced user_profiles, cached; HIDES when unavailable. Live: **1/100** (was fabricated 47/100).
|
||
- **Item 3/4 — Reports**: full fabricated-data inventory + pricing claim audit (LIVE/PARTIAL). Alt-ladder/Kelly/correlation are LIVE in the engine; scratch-cascade + real-time-feed are PARTIAL.
|
||
- **Item 5 — Daily hero prop** (`9b9aab4`): live RULE = largest |model−consensus| gap, A/B only. Shows book line vs VYNDR model + real timestamp. Empty→most-recent-real; nothing→hidden. NO static Jokic fallback. `heroPropService` + `/api/hero-prop`. Live: Erica Wheeler 9.5 vs 14.2.
|
||
- **Item 6 — Desk showcase** (`cb3237c`): real alt-line ladder from the snapshot (edge guarded on small-line artifact); φ = real same-team correlation for two real legs or hidden; Kelly tile REMOVED (no odds → can't size honestly). `deskShowcaseService` + `/api/desk-showcase`. Live: Vandersloot ladder.
|
||
- **Item 7 — Accuracy → TRUTH** (`89a2977`): public 30D accuracy (`/api/accuracy` + `/api/ledger/accuracy`) now reads the CLEAN Postgres ledger aggregate (getModelAggregate/getAccuracyView, model_value>0), not the Redis outcome log. Number moved 67%→61%. **BEAT CLOSE suppressed** at source (`clvCaptureReliable()`) — see C4.
|
||
- **Item 8 — Blog** (`3b7a1f5`,`a8e383e`): wired the 5 REAL articles (content lives in `web/content/articles` — the runtime content root, NOT repo-root content/) with honest dates + markdown render + OG. Backdated 2026-03-22 orphan deleted + 301. Live: 5 articles serve.
|
||
|
||
### 🔴 C4 — CLV CAPTURE BROKEN (open, backend)
|
||
`captureClosing` re-records the LOCKED line as "closing" (closing_line == locked_line on the whole sample) → CLV is flat → BEAT CLOSE = a measured-wrong 0%. Suppressed on all public surfaces until fixed. Full write-up + fix plan: `specs/audit-data/clv-capture-broken.md`. Flip `CLV_CAPTURE_RELIABLE=1` after C4 verifies real closes.
|
||
|
||
## PHONE AUDIT PUNCH LIST — COMPLETE (2026-07-17, on main, deployed, fingerprinted)
|
||
Kev's 35-screenshot real-390px audit (signed-out + Desk signed-in). Worked top-down, fingerprinted the batch (compare "on current form" literal live in served chunk off `71e35e9`). Test baseline: 268 suites / 3239.
|
||
- **P0-1** blue-B/amber-C banished → token-derived grade badges (earlier batch).
|
||
- **P0-2** image pipeline (logos/headshots) `loading=eager` (earlier batch).
|
||
- **P0-3** player grouping/ladders across leaders/explore/ledger/breadth (`lib/playerGrouping.js`).
|
||
- **P0-4** mobile app bar collapsed to one line (ticker+heartbeat hidden <768, `MobileSyncClock`).
|
||
- **P1-5** 390px overflow containment (breadth rows stack, pitcher line wraps).
|
||
- **P1-6** (`ff0d3b1`) ONE freshness source: Slate "UPDATED" now reads the pipeline `refreshed_at` (the app-bar clock's field), not the client poll time — no more "0s ago" vs "SYNC 46:03" contradiction.
|
||
- **P1-7** (`d755b43`) flat edge board: an impossible `|edge|>40` is treated ABSENT at both data (can't rank a fake +140% over a real +8.4%) and display layers. Chips render now (P0-2). `EDGE_BOARD_SANE_MAX`.
|
||
- **P1-8** (`00537eb`) Compare verdict truth: no phantom "edge tonight" for July NBA(0) → "on current form"; cites what actually wins (scoring/boards/playmaking, not usage).
|
||
- **P2-9** (`77e8937`) leaderboard stat labels `web/src/lib/statAbbrev.js` (SB/ER/TB) + **FLAGGED** the backend grade degradation (below).
|
||
- **P2-10** (`71e35e9`) ledger read-card density (padding/margins tightened; archetype showDesc kept — it's inline, zero height cost).
|
||
|
||
### 🔴🔴 BACKEND FINDING — MLB GRADE PIPELINE DEGRADED (needs backend work, NOT frontend)
|
||
Diagnosed against live `/api/snapshot/mlb` (2026-07-17). Full write-up + do-not list: `specs/audit-data/mlb-grade-degradation.md`. Frontend guards (sane-edge cap, statAbbrev) are damage control, NOT the fix.
|
||
- **projection == 0 for 9/25** grades — the MLB feature path feeds 0 instead of refusing (S58 `insufficient_data`). This also breaks `computeEdge` → the broken edge_pct.
|
||
- **edge_pct on a broken scale** {20,60,100,140} — not a market %. Real prop edge is single-digit.
|
||
- **grade↔confidence mismatch on 25/25** — B shown at conf 55 (= the B- band per `grade_thresholds.json`); systematic one-sub-tier gap. The surfaced `confidence` is not the probability that derived the letter (likely the data-sufficiency penalty applied to display-only). This is the audit's "B at 45%" — NOT a frontend labeling or threshold bug; DO NOT re-letter/re-threshold on the frontend.
|
||
|
||
## DESIGN COMPLETION TRAIN — SUBSTANTIALLY SHIPPED (2026-07-16, on main, deployed) · ONE MASTER CHROME AUDIT PENDING
|
||
Goal: the live product matches Design's mockup at EVERY width. **HONESTY LAW: I cannot see rendered output (WSL2↔Chrome unreachable) — every visual change is "built to Design's spec, VISUALLY UNVERIFIED." Kev runs ONE master Chrome audit at the END (390px + desktop) against the running unverified list below; fixes propagate through tokens (surfaces inherit, don't fork).** Design package (3 revs) authoritative at `specs/design-reference/` (HANDOFF.md = entry point; `vyndr-system.html` desktop, `vyndr-mobile.html` 19 screens, `vyndr-landing.html`, `assets/glyphs/` 83 SVGs+MANIFEST).
|
||
SHIPPED + deployed:
|
||
- **Palette alignment** (`e0a26a8`,`c3d469e`,`ebaab00`): tokens → Design's EXACT ramp. **Grade colors corrected: B blue→WHITE #F0F0F0, C amber→GREY #B8BCC8, D #FF4757** (the old blue/amber violated DESIGN-SPEC's own "B neutral-bright, C muted"). Surfaces #0E0E14/#14141E, text #F0F0F0/#B8BCC8/#707080, borders #1E1E2A/#2A2A38. vyndrTokens + tests aligned. ShareCard/OG billboards grade-B white.
|
||
- **Glyph system** (`77d8fd6`,`771d8b0`): 44 archetypes render Design's REAL marks (83-mark set: 74 display + 9 classifier-legacy), front `lib/archetypes.js` + back `archetypeService.js` color-synced. ZERO generic placeholders left. Combat keeps its char-glyph namespace (separate test).
|
||
- **Entity layer VERIFIED already-complete** (DS0/Wave 2A/2B — the audit was stale): TeamLogo (real ESPN CDN, monogram fallback) on game cards/team hub; PlayerAvatar headshots MLB/NBA/WNBA + NFL/NHL resolver ready + soccer-monogram (blocked) + combat-no-photos; books via bookInfo/BookChip everywhere (fixed BookComparison holdout `1558710`); line-shopping surfaces (detectBestBook/MarketBreadth) self-deepen. No `$1M` copy.
|
||
- **Rev 3 matchup chips** (`f24f941`,`baf977f`,`bf8bf16`): `TeamChip` primitive (real TeamLogo at 10-12px + abbr, inherits row opacity) on StatStrip board rows, grade reveal, pitcher-vs. Game-card headers already had TeamLogo.
|
||
- **Mobile: M1a header collapse** (`a10227e`) · **M1b app-bar clock** (`7c7ab24`, Hybrid: wall clock rests, STALE reacts off refreshed_at) · **grade-reveal mobile** (`baf977f`, 74px hero) · **flat EDGE BOARD** (`8a24ac9` — Design's core mobile screen: NEW flat edge-ranked IA, `slateAdapter.flattenToEdgeBoard`+`MobileEdgeBoard`, shown <768px, cards ≥768px; threaded edge_pct through the pipeline) · **M4 structural locks** (`6bc9093`).
|
||
- **Key finding:** the 19 mobile screens were LARGELY already built (pages built to the SAME DS0-DS5/§7 contracts Design's mobile follows) → they already match structurally + stack on mobile. The genuinely-new work was the flat board (done) + narrow chip/hero polish.
|
||
REMAINING (all VISUAL, for the master audit): per-screen mobile polish of the ~17 already-built screens (blind, mostly Design-aligned) · M2 billboard crop-ratio verification (OG routes exist, colors fixed). Test baseline: 266 suites / 3214.
|
||
|
||
### 🔴 MASTER-AUDIT UNVERIFIED LIST (every visual change this train — Kev's one end-of-session audit)
|
||
1. **Palette** (biggest): grade B-white/C-grey/red-#FF4757 + full surface/text/border ramp — every surface inherits.
|
||
2. **Archetype marks + colors** — 44 archetypes render Design's real 83-mark set.
|
||
3. **Flat EDGE BOARD** at 390px — the core mobile screen (rows/opacity-ramp/chips/hero-edge%).
|
||
4. **Grade-reveal mobile** (74px hero + team chip) · **app-bar clock** (Hybrid) · **M1a header collapse**.
|
||
5. **TeamChip** on board rows / grade reveal / pitcher-vs (real logo at 10-12px).
|
||
6. **OG billboards** (ShareCard/player/u) grade-B white + text #F0F0F0.
|
||
|
||
## SHIPPED — HEARTBEAT + CADENCE + QUOTA-GUARD (2026-07-15, on main, deployed)
|
||
Three ships, each green + deployed. Standing authority is now PERMANENT full autonomy (three-gate floor: tests green, web build exit 0, post-deploy fingerprint of NEW code).
|
||
- **`7712f0a` Heartbeat honesty (fingerprint-verified live):** the "frozen snapshot / STALE 8h" was a MEASUREMENT bug, not a dead pipeline. `updated_at` = grade-LOCK time (advances 5×/day by design; grades never change in-game); the SYNC badge measured the 20-min `expected_interval_s` against it → structurally STALE. `refreshed_at` (the real 20-min intraday heartbeat) was written to Redis but NEVER serialized. FIX: full snapshot seeds `refreshed_at`, `/api/snapshot/summary` + GET `/:sport` expose it, `LiveLayer` badge measures from it. Fingerprint PROVED the pipeline was alive all along (`refreshed_at` advancing 20:20→20:40→21:00→22:00; the 22:00 full slot fired → scheduler healthy, re-armed fine after the 19:37 restart).
|
||
- **`4cd933d` + `2d413cf` Per-sport cadence (Job 1):** `src/config/sportCadence.js` (config-over-constants). REALITY-CORRECTED: mlb/nba/**wnba** all keep the full grid `14/19/22/1/3` (PropLine 9k/day = cheap; WNBA games span afternoon→late-evening, so 14 UTC/10am ET catches afternoon-game props — the earlier "drop it as waste" was wrong, corrected by observing 2 finished afternoon WNBA games on Jul 15). **Soccer alone is trimmed** to `14/19` + NO intraday — the genuine win, protecting the scarce odds-api key. Scheduler fires at HOURS_UTC and grades only `sportsForHour(h)`; only `intradaySports()` refresh.
|
||
- **`2d413cf` Quota guard (the real finding):** odds-api quota was EXHAUSTED 0/500 mid-cycle, unpaged. Root cause: `futuresService` (the only regular burner) used raw axios, BYPASSING the gateway → never hit `recordCall` (the ONLY place the WARN/BLOCK pager fires) → the counter reached 100% via silent header-sync. FIX: futures now routes through `gateway.fetch('odds-api', …)` (counted + blocked); a RESERVE floor (`ODDS_API_RESERVE`=50) refuses DISCRETIONARY calls (futures/soccer) while `remaining ≤ reserve` so they can NEVER starve MLB's essential prop-backup; `quotaTracker.syncFromHeaders` (authoritative) now ALSO fires the once-per-period alert; `POST /api/internal/quota/test-alert` test-fires the pager end-to-end. **Fingerprint caveat:** backend-internal, no unauthenticated public surface (`/api/internal/*` 401s for any path — auth runs before route-match); verify via authenticated test-alert (200=new/404=old) with the internal key.
|
||
- **odds-api 0/500 → BOTH soccer props AND futures are blocked on QUOTA, not code.** Kev's ruling: HOLD THE LINE (don't pay; zero-out-of-pocket). Soccer stays honest-empty; cadence is ready for the reset. Env: keep `API_FOOTBALL_KEY` (code reads it), delete inert `API_FOOTBALL_KEY2` (Kev handling). See memory [[odds-api-quota-soccer-blocked]]. Test baseline: 264 suites / 3195.
|
||
|
||
## SHIPPED — TRUTH-EVERYWHERE + OFFSEASON HUB TRAIN (on main, deployed)
|
||
Governed by `specs/truth-everywhere-train.md`. **The reframe that drove it: only MLB GRADED in prod (NBA/WNBA refused their whole slate — offline Python feature source, no fallback; soccer errored pre-grade). So grading, not settlement, was the first domino.** Standing deploy authority granted (green + fingerprint → merge + deploy, no ask). Waves DEPLOYED:
|
||
- **Wave 0 — NBA/WNBA GRADE (free ESPN gamelogs):** `espnStatsAdapter.getPlayerGameLog` feeds `featureCache` when the Python source is null → l5/l20 → props grade. **Verified live** (A'ja Wilson→B, Ionescu→C off real ESPN form). Endpoint `site.web.api.espn.com/apis/common/v3/sports/basketball/{nba|wnba}/athletes/{id}/gamelog`.
|
||
- **Resolver-hardening:** roster-index primary resolver (206 WNBA / 544 NBA) + `?season=` gamelog retry — the real fix for dual-league (WNBA+NCAA) players who got an empty gamelog (Collier/Brionna Jones now resolve). Broad coverage.
|
||
- **Wave 1 — NBA/WNBA SETTLE:** `defaultGetPlayerStats` routes nba/wnba to the ESPN gamelog in BOTH `outcomeService` + `ledgerService`; separate `NBA_BOX_KEY`/`NBA_COMBO` settle maps; final-honesty guard (never settle in-progress); `opsWatch` finals-gated zero-settle alarm + per-sport boot announce. accuracy:{sport}+by_tier light up automatically. **PROOF is next-day** (grades lock tonight → settle tomorrow → `accuracy:wnba`).
|
||
- **Wave 2 — NEVER-DARK HUB:** ESPN `/news` wire (`newsService` + `/api/news/:sport`) — **verified live (real headlines)**; `NewsWire`+`FuturesBoard` on `/explore` (offseason-aware, self-hide on empty, futures "TRACKED · NOT GRADED"). **KNOWN GAP: `/api/futures/:sport` returns `markets:[]` for ALL sports in prod** — code path looks correct (odds-api `/sports/{FUTURES_KEYS}/odds?markets=outrights`); needs a live odds-api probe (ODDS_API_KEY) to finish. Self-hides honestly meanwhile. `FUTURES_ENABLED=0` kill-switch; 12h TTL; 1 credit/refresh.
|
||
- **Also rode along:** `espnAthleteIndex` cross-sport headshot capture (headshot-coverage) — NBA/WNBA/NFL/NHL real headshots when in-season.
|
||
DEFERRED (locked decisions): **soccer** — grade+settle path, needs `API_FOOTBALL_KEY` (VALIDATED, Free 100/day, active to 2027 — set it in Coolify env `API_FOOTBALL_KEY`; code dormant until wired) + ESPN-soccer-free-first attempt; **MMA moneyline settle** (combat not yet in the grade/lock loop); **NFL/NHL** (dormant, in-season later); **futures debug** (odds-api probe). Test baseline now **263 suites / 3183**.
|
||
|
||
## SHIPPED — WIRING & DATA TRAIN (on main `3b1aa9f`, deployed 2026-07-13)
|
||
Governed by `specs/wiring-data-train.md` + `specs/combat-intelligence.md` + the global visual reference `specs/design-reference/vyndr-system.html` (build toward it; live wordmark kept). **All 6 waves shipped + deployed** (fingerprint-verified: /parlay, /fight, /u/vyndr live; house record 88-38/70% real). Real assets verified live (MLB/ESPN headshot CDNs, ESPN-MMA feed, Baseball Savant CSV).
|
||
- **Wave 1 — trust bugs:** billing renewal honest render (`billingDisplay.classifyRenewal` — no far-future placeholder); James Wood nameKey-collision fixed (`mlbStatsAdapter` teamHint disambiguation + streaks join-invariant); DeskShowcase "$1M terminal" → deadpan copy.
|
||
- **Wave 2 — sport-agnostic entity layer:** real player headshots threaded from ingestion (MLB MLBAM + NBA/WNBA ESPN athlete ids that were fetched-and-discarded) across slate/scan/hotlist/search/grade card; soccer = honest monogram (no free id); 8 self-authored SVG book wordmarks (`web/public/books/*.svg`, swappable for official art) + all 10 book keys resolve; team-logo abbr aliases. Storage: id on the `enriched` grade at `snapshotService` (zero new I/O).
|
||
- **Wave 3 — record by grade tier (Addition 2):** ONE shared `TierRecord` (`lib/tierRecord.js` + component) on dashboard + /u + ledger; per-tier W-L always, hit-% only at n≥20 per tier (gate stays in `getModelAggregate`).
|
||
- **Wave 4 — missing surfaces:** Outlook mode (grid never blank → yesterday receipts / tomorrow schedule); Market-Breadth median-consensus-vs-model strip (self-hides <2 books); Parlay Lab `/parlay` (slate-independent leg source); live Grade-Shift timeline (`GradeShift`).
|
||
- **Wave 5 — /u house-mode + arsenal:** house handle `vyndr` surfaces the real `user_id=NULL` public model record + per-tier calibration + 1080×1350 portrait/OG (user-handle privacy 404s stay byte-identical); Baseball Savant pitcher-arsenal (`savantAdapter`, free CSV, verified) → `PitcherArsenal` card, self-hides on absent.
|
||
- **Wave 6 — combat v1 (MMA):** ESPN-MMA fight cards + tale-of-the-tape + style-blend archetypes (sport-scoped `COMBAT_ARCHETYPES` — FINISHER color collided w/ soccer + tripped the ΔE gate, so kept separate) + odds-api ML/round-totals + style-edge verdict. NOT in the snapshot/settle loop; method/round/props + fighter photos + matchup-GRADE engine + ufcstats scraping all DEFERRED, flagged data-limited in-UI. New routes: `/fight/[id]`, `/parlay`, `/u/[handle]/portrait`.
|
||
POST-MERGE TODO: NBA/WNBA headshot coverage + combat depth need prod runtime verification; soccer headshots blocked on `API_FOOTBALL_KEY`; combat settlement + matchup-grade engine are the next combat sub-wave.
|
||
|
||
## SHIPPED — DESIGN TRAIN v2 (merged to main `e9c0a59`, deployed & fingerprint-verified live)
|
||
Governed by `specs/DESIGN-SPEC.md` v2 (the raised standard: entities render as
|
||
themselves, color contract, sub-200ms, screenshot-first billboards). **All six
|
||
sessions DS0–DS5 shipped: merged to main, pushed to gitea, Coolify auto-deploy
|
||
LANDED — verified live via `/pricing` DeskShowcase ("A $1M terminal") + `/u`
|
||
route serving the unified EmptyState. 238 suites / 2885 tests green,
|
||
`next build` exit 0.** Post-ship TODO: re-run the Chrome design audit against
|
||
prod v2 to confirm all 22 audit findings closed (env here can't screenshot —
|
||
WSL2↔Windows-Chrome localhost unreachable).
|
||
- **DS0 Entity Layer** (`24af247`): `teamMeta.js` (real logos+colors for 30 MLB /
|
||
30 NBA / 13 WNBA / 48 WC nations via ESPN CDNs), `TeamLogo` / `PlayerAvatar`
|
||
(team-colored monogram fallback, no gray silhouette) / `BookWordmark`, swapped
|
||
into GameCard/StatStrip/StreaksPanel/TeamHub.
|
||
- **DS1 Speed + Trust Bugs** (`1c681df`): React #418 hydration fix (mounted-flag
|
||
guard), layout-matched `Skeleton`s replace text-wall loaders, scan→ledger
|
||
persistence via `session.access_token`. **DS1 follow-up** (`cf91c04`): closed
|
||
the sb-token trust-bug CLASS — `lib/authToken.js` `currentAccessToken()` reads
|
||
the REAL Supabase session (not the OAuth-only `sb-token` key) across
|
||
profile/slip/dashboard/settings/tracker.
|
||
- **DS3 Color Contract** (`49a3323`): `lib/colorContract.js` — green = one meaning
|
||
(edge/A-tier/CTA), edge/CLV by sign, glow A/A+ only, archetype greens deduped.
|
||
Enforced by `colorContract.test.js` (fails on violation).
|
||
- **DS4 Billboards** (`45bafbc`): STREAKS row (length as mono hero), grade reveal
|
||
(sign-colored edge), CLV reframe (flat says so), `/u` public profile + OG.
|
||
- **DS2 Dashboard Slate Rebuild** (`fe294a5`): never-empty hero (falls back to
|
||
yesterday's A-tier settled receipts), one bold hero per card, pending-filler
|
||
collapse, ranked tonight-grades. `slateAdapter.js` engine + 21 tests.
|
||
- **DS5 Pricing + Motion + States** (`a18a3f3`): Desk $44.99 as the hero tier
|
||
(single green CTA, `DeskShowcase`), ticker → punctuated stillness (≥4.2s hold,
|
||
one idle motion), unified `EmptyState` (404 grammar) at TeamHub/game/ledger,
|
||
archetype glyphs propagated to streaks/ledger.
|
||
NOTE: WSL2↔Windows-Chrome localhost is unreachable in this env, so visual proof
|
||
was SSR-render + source-assertion tests + live prod fingerprint curls, not
|
||
pixel screenshots. The `design/v2-train` + per-DS worktree branches can be
|
||
pruned. Stale `origin` GitHub remote still carries the long-dead leaked PAT —
|
||
deploy runs off `gitea`, so it's inert, but rotate it when convenient.
|
||
|
||
|
||
|
||
## WHERE THINGS STAND
|
||
|
||
**Tests: 2757 / 229 suites, all green. Web build exit 0.** Baseline at the A1
|
||
board's start was 2398. Governing docs: `specs/VYNDR-NORTH-STAR.md` (v1.1) +
|
||
`specs/VOICE.md` (v1.1) + `specs/ROW-GRAMMAR.md` — read them before building
|
||
anything.
|
||
|
||
**The record is live and writing itself.** Day one (Jul 11): 24→25 MLB rows,
|
||
real locks (line/odds/book), closing capture every snapshot + every 20-min
|
||
intraday refresh, settlement each morning slot with signed CLV. No percentage
|
||
renders anywhere under n≥20 — enforced in `ledgerService.getModelAggregate`,
|
||
the single gate. WNBA/soccer rows pend honestly until Phase 4.5.
|
||
|
||
## THE THREE A1 TRAINS (all merged to main, all pushed)
|
||
|
||
| Train | Range | Carried |
|
||
|---|---|---|
|
||
| #1 | `4d2b27d → 219167e` | S0 verify · S1 promise audit (`PROMISE-AUDIT.md`: alt ladder + quarter-Kelly BUILT, analyst unlimited, "40+" claims made honest) · S2 compliance (/responsible-gambling rebuilt, /terms /privacy drafts, /methodology, 5 seed articles in `content/articles/`) · S3 affiliate (BOOK IT deep links organic-until-config-flip, best-price dots, ?ref= partner attribution) · S4 media engine (`mediaEngine` + executable VOICE lint, `/desk` arsenal + DATA BRIEF, Ghost DRAFTS-only publisher) · S5 viability (statsapi lineups → CONFIRMED/NOT-IN kills grades visibly; ESPN injury wire; Yesterday/Today/Tomorrow nav + Settle panel) · S7 newsletter (Listmonk capture + THE VYNDR REPORT assembly, env-gated) · S8 ops (settle alarm, 3-slot failure pager, quota/disk/mem, 9AM pulse — test alert delivered to the real ntfy channel) · S9 slip reader (tesseract.js OCR + per-book parsers, /slip page; synthetic-image acceptance passed) · S10 public profiles (/u/{handle}, private-by-default, no-existence-leak) · migrations 021/022 |
|
||
| #2 | `219167e → 14dc9cf` | S6 display (ROW-GRAMMAR.md + slot-order fixes, line sparklines from intraday history, last-10 ●●○ dots, CLV distribution on MODEL tab, ⌘K SearchModal + mobile search, landing LCP fixes) |
|
||
| #3 | `14dc9cf → ec5ee3d` | S11 live tracking (statsapi/ESPN live box lines → ON PACE / NEEDS N / HIT ✓ in the outcome slot, live games float up, "TRACKING — READ LOCKED PRE-GAME"; real-live acceptance: Harper 3/1.5 TB ▼8th → HIT ✓). Grades never change in-game. |
|
||
|
||
**Deploy status:** ALL LIVE as of `17fb981` (Jul 12 17:43 UTC). Fingerprints
|
||
confirmed in prod: `/api/live/mlb` serves JSON with `hasLive:true` (S11 live
|
||
tracking working — a real game was tracked at confirm time), `/api/health`
|
||
healthy, `/api/snapshot/summary` → `expected_interval_s:1200` (env var now set).
|
||
|
||
**P0 postmortem (`17fb981`):** the API silently served a 14h-old image
|
||
(`4d2b27d`) from Jul 11 ~20:00 to Jul 12 17:43. Root cause: the Dockerfile
|
||
never copied `content/`, and `mediaEngine.js` read `content/stark-lines.json`
|
||
with an unguarded module-load `readFileSync` → ENOENT at require time crashed
|
||
`app.js`'s eager route chain (app → routes/desk → deskService → mediaEngine)
|
||
→ boot failed → Coolify healthcheck rolled back. Every merge from S6 onward
|
||
was merged-but-not-running for 14h. Fixed: Dockerfile COPYs content/;
|
||
mediaEngine treats the file as optional garnish (loadStark try/catch → {});
|
||
`src/preflight.js` now prints `[preflight] OK|DEGRADED` at boot so a missing
|
||
file/env is legible instead of masked by a rollback. LESSON: any module-load
|
||
file read is a boot-crash risk — keep them lazy + try/catch, and the
|
||
Dockerfile must copy every dir the runtime reads.
|
||
|
||
## LIVE INFRA MAP (verified in prod)
|
||
|
||
- **VYNDR Web** — Next.js :3000, vyndr.app. **VYNDR API** — Express :3001,
|
||
api.vyndr.app (`/api/internal/*` behind `VYNDR_INTERNAL_KEY`).
|
||
- **Scheduler (in-process, API app, `SNAPSHOT_CRON=1`)** — snapshot slots
|
||
14,19,22,1,3 UTC; settle pass (outcomes + ledger, idempotent) runs FIRST
|
||
each slot; 20-min intraday odds refresh during slate hours
|
||
(`INTRADAY_REFRESH=0` kills); desk-ready ping after the day's first slot;
|
||
daily pulse 13:00 UTC; missed-cron + 3-slot-failure + quota pagers → ntfy
|
||
`vyndr-pipeline-kev2026`. Boot logs: `[snapshotScheduler] armed` +
|
||
`[settlement] armed`.
|
||
- **Data**: Redis (mastermind-cache) for all pipeline caches (SNAP_TTL 24h —
|
||
do NOT shorten; the 11h overnight gap killed morning settles once).
|
||
**Supabase = the ledger** (`ledger_entries` + `public_profiles` +
|
||
`user_profiles.partner_ref`; RLS everywhere, service-role-only writes;
|
||
migrations 001–022 applied, files in `supabase/migrations/`).
|
||
- **Deploys**: gitea `git.builtbykev.com/builtbykev/vyndr` → Coolify webhook
|
||
(HMAC secret fixed Jul 11; auto-deploy proven). GitHub `origin` remote is
|
||
dead-PAT hygiene debt only.
|
||
- **Free feeds** (zero out-of-pocket law): PropLine ×3 keys (9k/day capacity;
|
||
intraday uses ≤144/day), odds-api backup (quota-gated), statsapi.mlb.com,
|
||
ESPN site API. tesseract.js OCR self-hosted (traineddata caches on first
|
||
call — needs outbound network once).
|
||
|
||
## ENV VARS STILL PENDING ON KEV'S SIDE (Coolify, API app)
|
||
|
||
1. `SNAPSHOT_EXPECTED_INTERVAL=1200` — **confirmed absent** (fresh restart
|
||
still served the 18000 default). SYNC badge tracks intraday cadence once set.
|
||
2. `DESK_OWNERS=kevdevelops@gmail.com` — /desk 403s everyone until set.
|
||
3. `LISTMONK_URL` / `LISTMONK_USER` / `LISTMONK_TOKEN` / `LISTMONK_LIST_ID`
|
||
(all four together) — newsletter no-ops politely without them; box-side
|
||
setup in `docs/NEWSLETTER.md`.
|
||
4. Optional: `GHOST_URL` + `GHOST_ADMIN_API_KEY` (drafts), `PULSE_HOUR_UTC`.
|
||
|
||
**Entity placeholders** (S2, blocked on Kev's values): `[ENTITY NAME]`,
|
||
`[STATE OF FORMATION]`, `[ARBITRATION VENUE]`, `[CONTACT EMAIL]` in
|
||
`web/src/app/terms/page.tsx` (6×) + `web/src/app/privacy/page.tsx` (5×).
|
||
A test keeps unverified entity names out until replaced.
|
||
|
||
## BOOK COMPARISON — data layer shipped (per the build order)
|
||
|
||
**What shipped (backend + tooling; UI routing is the gated next step):**
|
||
- **Phase 1 — fenced, snapshot-locked per-book store.** `src/services/
|
||
bookPriceStore.js` captures per-book prices from the multi-book `props` array
|
||
BEFORE `dedupeProps` runs, keyed by `nameKey|stat_type`, written to
|
||
`bookprices:{sport}` at SNAP_TTL (24h) inside `snapshotService.runSnapshot`.
|
||
STRUCTURAL FENCE: it only reads `props` and writes its own key; nothing on the
|
||
grade path reads it. Proven byte-identical — the graded slate is unchanged with
|
||
or without the capture (`tests/unit/bookPriceStore.test.js`, incl. a grep test
|
||
asserting no grade-path file references `bookprices`).
|
||
- **Phase 2 — the crown is threshold-gated.** `scripts/measure-book-spread.js`
|
||
reports best-vs-worst same-line price spread (American cents AND implied-prob
|
||
points), book-count histogram, %-zero, pinnacle presence — per sport, never
|
||
pooled. PRE-REGISTERED: crown ships only if median ≥8¢ OR ≥2 implied-prob pts.
|
||
Runs post-deploy against real captured data (Redis degraded locally → exits 0).
|
||
- **Phase 3 (backend) — honest-absent + gated crown.** `bookComparisonService.
|
||
compareProp` now renders single-book / flat-market props with NO crown; the
|
||
crown fires only among ≥2 books at the SAME line with DIFFERING prices, and only
|
||
when `BOOK_CROWN_ENABLED=1` (default OFF until Phase 2 clears). `/api/books/
|
||
:sport` (a crown claim) returns [] while gated off; `/api/books/:sport/:player/
|
||
:stat` returns the honest grid always. Route repointed to read `bookprices:*`
|
||
(fallback: odds cache), matched by `nameKey`. Response `source` field is the
|
||
deploy fingerprint (`bookprices` = new store serving).
|
||
|
||
**What it revealed:** the premise ("BookComparison.tsx is a dead component") was
|
||
incomplete — the ENTIRE feature (service + `/api/books` routes + Next proxy) was
|
||
built and mounted, but NON-FUNCTIONAL end-to-end: the route fed FLAT odds-cache
|
||
rows to a comparator that expects GROUPED `{lines|books}`, so it always returned
|
||
empty. Per-book prices were retained only transiently (odds cache, ~1h, raw
|
||
names, grade-path INPUT); every grade-path persistence point collapses to one
|
||
book. Fixed by the snapshot-locked store + route repoint.
|
||
|
||
**Gated next step (NOT this order-turn):** route `BookComparison.tsx` onto the
|
||
prop card fed by the store, and design the crown treatment — AFTER Phase 2's
|
||
measurement returns real spread numbers on prod (you can't honestly design a
|
||
crown before you know it ships). `BOOK_CROWN_ENABLED` stays 0 until then.
|
||
HELD unchanged: dedupeProps, snapshot dedup, the selector, the grade, champion,
|
||
challengers, ranking, edge_pct/ev_pct. No push-to-book, no movement strip, no
|
||
pinnacle edge test (data persisted only).
|
||
|
||
## HONEST OPEN ITEMS
|
||
|
||
- **Phase 4.5 — WNBA settlement via ESPN box scores. DUE ~Jul 24** (hard
|
||
date). Until then WNBA/soccer/NBA ledger rows pend forever; accuracy = MLB only.
|
||
- **The record needs time**: no percentage anywhere before n≥20 settles, by
|
||
design. First percentages expected ~day 3–4 of MLB volume.
|
||
- **Slip reader**: acceptance was a clean synthetic image; needs one real
|
||
phone screenshot (DK) to validate OCR robustness.
|
||
- **Landing LCP**: S6's fixes are static-analysis-provable but unmeasured —
|
||
run PageSpeed mobile vs the 3.8s baseline; next suspect is the
|
||
LiveHeroProp post-fetch swap.
|
||
- **Partner report endpoint** returns real numbers only for signups AFTER
|
||
migration 021 (metadata path live; OAuth signups carry no ref — known gap
|
||
in `docs/PARTNERS.md`).
|
||
- **NBA Oct / NFL Sep readiness** + soccer end-to-end: `specs/vyndr-roadmap.md`.
|
||
- **Killed permanently** (founder ruling): AI chat assistant, light mode.
|
||
|
||
## OPERATING LAWS (never relax)
|
||
|
||
Zero out-of-pocket · absent beats wrong, refusal beats hollow · n≥20 before
|
||
any percentage · nothing auto-posts anywhere, ever · grades never change
|
||
in-game (revisions are public via `revised_from_grade`) · one meaning per
|
||
color, data in mono · no exclamation points, anywhere · the pipeline is the
|
||
only source of numbers, including in marketing (`mediaEngine` lint enforces).
|
||
|
||
— STATE as of ec5ee3d · the record is writing —
|