Grade-board sort: signed signal, takeable-gated p_win, missing sorts LAST

Display ORDERING only. No grade, ledger row, lock_line, scoring, or edge_pct
scale/display change. Push scoring untouched.

Two defects removed from selectTopGrades (wrong at ANY scale, independent of
edge_pct's separate retirement):
  1. edge: Math.abs(numOr(g.edge, -Infinity)) — abs() on an already-
     direction-signed value ranked the model's strongest DISAGREEMENTS level
     with its strongest agreements (177 public ledger rows carry a negative
     edge; positive = the model AGREES with the graded side).
  2. Math.abs(-Infinity) === Infinity, so a row with NO edge sorted FIRST —
     absent data presented as the top pick (the Number(null) class).

New key: grade -> confidence -> takeable-gated p_win (nulls LAST) -> SIGNED
edge (nulls LAST) -> input order. Scales are never mixed in one comparator.
Takeable band = web valueState.isTakeable, asserted byte-equal to the hero's
config/valueEngine.isTakeable (-160..+200) incl. strict-null.

Alt-line ladder (analyzeViaEngine1:506) no longer sorts by edge_pct: ordered
highest-p_win-first derived analytically at zero added compute — P(stat >= k)
is monotone non-increasing in k, so p_win-desc is line-ASC for an over and
line-DESC for an under. base stays marked; no consumer depends on
alt_lines[0]; deskShowcaseService.rungsOf already re-sorted by line.

THREE PREMISE BREAKS found report-first, before code:
  - /api/props/top-graded 404s in prod (absent from src/) so the dashboard
    board renders receipts/empty — the edge sort orders nothing there today.
    The prior order's "97.3% of rows tie" was a LEDGER measurement wrongly
    extrapolated to that board. Fix is correct-in-itself and lands when the
    feed is restored.
  - p_win cannot be a client-side key for all tiers: snapshotGating strips it
    for unentitled tiers ("shipping p_win is shipping the model price").
    Verified live: prod /api/snapshot carries p_win on 0/8 MLB, 0/25 WNBA.
  - Ladder rungs carry no per-rung price, so the takeable gate is inapplicable.

Verified on real data, both sports, both paths: unentitled — WNBA (n=25)
ordering CHANGED, MLB (n=8) unchanged, signed edge non-increasing in every
(grade,confidence) tie group (20 pairs, 0 violations); entitled — 40 real
ledger rows with p_win+locked_odds, p_win-descending, untakeable chalk NOT
promoted (Trea Turner .757 @-275 does not beat Rhyne Howard .745 @-120)
(36 pairs, 0 violations).

Hero consistency, stated honestly: same signal + same gate, different
precedence BY CONTRACT (board = grade-tier-first "top GRADES"; hero =
p_win-first "top read"). Identical within the leading tier (verified); across
tiers the board may lead with an A the hero doesn't pick. Not a contradiction.

Floor: 310 suites / 3864 tests green, web build exit 0. Dashboard + Desk
visuals are auth/feed-gated -> tagged for the Chrome audit, no visual faked.

Held: edge_pct rescale/display retirement (Order B); building the missing
/api/props/top-graded selector; exposing p_win to unentitled tiers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
Kev
2026-07-29 21:13:51 -04:00
parent 9b5235cf99
commit b85b351993
9 changed files with 886 additions and 11 deletions
+5 -3
View File
@@ -672,9 +672,11 @@ snapshot, locked to the line, and read from cache.
props. It's `rbi` now; a normalizer test locks it. props. It's `rbi` now; a normalizer test locks it.
- **The streaks/hotlist path uses its own `rbis` key** built from raw MLB stats — - **The streaks/hotlist path uses its own `rbis` key** built from raw MLB stats —
independent of the odds normalizer. Don't "unify" them; the split is intentional. independent of the odds normalizer. Don't "unify" them; the split is intentional.
- **MLB is the ONLY end-to-end-live sport.** Outcome settlement is MLB-only - **MLB and WNBA both settle end-to-end.** CORRECTED 2026-07-26 (was "MLB-only"):
(WNBA/NBA/soccer grades never settle → `accuracy` reflects MLB only). Fixing `outcomeService.SPORTS` includes wnba, which settles via ESPN box scores
that (ESPN box-score settle path) is roadmap Session 57. (`espnStatsAdapter` box-field map) — verified: 376 WNBA rows settled, settle
logic spot-checked correct. NBA/soccer still do not settle (no free settled
feed wired). So `accuracy` reflects MLB + WNBA, not MLB only.
- **`src/utils/opsNotify.js`** pushes pipeline alerts to ntfy (`vyndr-pipeline- - **`src/utils/opsNotify.js`** pushes pipeline alerts to ntfy (`vyndr-pipeline-
kev2026`). It NEVER throws and is auto-disabled under `NODE_ENV==='test'` / kev2026`). It NEVER throws and is auto-disabled under `NODE_ENV==='test'` /
`PIPELINE_ALERTS=0` (inject `fetchImpl` to test it). `snapshotService` alerts on `PIPELINE_ALERTS=0` (inject `fetchImpl` to test it). `snapshotService` alerts on
+290
View File
@@ -0,0 +1,290 @@
# VYNDR — CANONICAL STATE FILE
Read-only ground-truth audit. Written 2026-07-26 by Claude Code (repo + deployed DB access).
Every item tagged **VERIFIED** (file/line or measured number), **CANNOT DETERMINE** (with reason),
or **BLOCKED** (with what unblocks it). Where a prior claim conflicts with code/data, the code/data wins.
Nothing was built, changed, deployed, or migrated by this audit.
---
## REVIEW ZERO — PREMISE
- **0.1 What I can read** — VERIFIED. The **repo** (full source), the **deployed Supabase DB**
(read-only via MCP, project `zmdnczhtdxcddsxzttub`), and the **deployed API** (`api.vyndr.app`).
- **0.2 Key reachability** — VERIFIED. `PROPLINE_API_KEY_*` are **NOT on the box** (`.env` absent);
they were pasted in-session earlier this conversation and are usable for probes. `ODDS_API_KEY`
is on the box but **exhausted (0/500)**. `ODDSPAPI_KEY` not on box. DB-dependent items were run
against prod directly, so nothing here is BLOCKED on PropLine keys.
---
## PHASE 0 — THE ROI RECONCILIATION (headline)
- **0.1 ROI computed in code?** — VERIFIED. **Not for the model ledger.** ROI exists only for
**user bet-tracking**: `performanceService.js:42` (`roi: stats.roi`) and `betService.js:172-174`
(`profit = payout - amount`). There is **no ROI computation over `ledger_entries`** (the model's
public record). That is why the model's ROI has been invisible.
- **0.2 Price-at-grade persisted?** — VERIFIED. **YES**`ledger_entries.locked_odds` (text,
American), written by `ledgerService.recordPipelineGrades`. Populated on essentially all settled
rows (5 nulls on B). **ROI IS computable for accrued history.**
- **0.3 11-point index persisted?** — VERIFIED (nuanced). **NOT in `ledger_entries`** — only the
collapsed letter (`grade`); `_grade_11` is explicitly `delete`d at `gradeSlateService.js:97`
before the ledger write. **BUT it IS retained in `model_snapshots.grade_11`** (migration
`025:51`, `retentionService.js:117`) — 2,888 of 4,600 snapshot rows carry it. So sub-tier is
**lost on the settled-outcome table but recoverable** by joining `model_snapshots` to
`ledger_entries` on (player, stat, line, date). STATE.md's "grade_11 is stored" and this audit's
"deleted before ledger write" are BOTH correct — different tables.
- **0.4 Price distribution by grade (American odds, settled)** — VERIFIED:
| sport | grade | n(decided) | median | q1 | q3 |
|---|---|---|---|---|---|
| MLB | B | 285 | 270 | 650 | 155 |
| MLB | C | 176 | 169 | 500 | +109 |
| WNBA | B | 208 | 124 | 145 | 110 |
| WNBA | C | 160 | 120 | 134 | 108 |
Blended B median 160, C median 130. Overall range 10000 … +600. MLB grades skew to **deep
favorites**; WNBA grades cluster tight around 120.
- **0.5 Flat-stake unit ROI by grade (1u/row at `locked_odds`, decided rows hit/miss, voids
excluded)** — VERIFIED:
| sport | grade | n | hit% | **ROI%** |
|---|---|---|---|---|
| MLB | B | 285 | 70.2 | **1.02** |
| MLB | C | 176 | 60.8 | **+4.57** |
| WNBA | B | 208 | 52.4 | **4.79** |
| WNBA | C | 160 | 51.9 | **5.26** |
| A / D / F | (all) | 1 / 2 / 4 | 0 / 0 / 0 | 100 each (negligible n) |
Including voids as net-0 (my first pass) gives blended B 2.31%, C 0.10%.
**CONCLUSION — the premise's binary is a false dichotomy; the truth decomposes:**
- MLB **B** = high-hit (70%) **break-even favorites** — the premise's "favorites at fair prices,
zero edge" hypothesis is CONFIRMED here.
- WNBA **B/C** = **losing** (~52% hit at ~120 where break-even is ~54.5%).
- MLB **C** = **genuinely +EV (+4.57% on 176 decided)** — a real edge the blended "no ROI" masked.
So: ROI is not uniformly zero-edge. It is **not computed in code**, and when computed here it shows
**one profitable segment (MLB-C) hidden under WNBA losses and break-even MLB-B**.
- **0.6 64/56 population** — VERIFIED. The displayed `/api/ledger/accuracy` figures are
`hits/(hits+misses)`, **excluding voids** (88 void rows, 8.8%) and unsettled (67). Distinct
`outcome` values are **{hit, miss, void, null}** — **no `push`** (pushes structurally impossible:
half-number lines). The ROI table above uses the SAME decided (hit/miss) population, so hit-rate
and ROI are comparable. Blended B hit is 63.1% on `hits/(hit+miss)` but 55.9% on all-settled
(the 64 void B rows are the gap).
- **0.7 Hit rate in UI/marketing?** — VERIFIED. Displayed on **≥10 surfaces**: `app/page.tsx`
(landing), `GradeCard`, `TopSignals`, `vyndr/TierRecord`, `vyndr/ModelRecord`, `vyndr/AccuracyBadge`,
`game/[id]/page`, `u/[handle]/portrait`, `ledger/page`, `scan/page`. **ROI / edge is displayed
NOWHERE.** Honesty gap: users see "63% / 70% HIT" with no indication MLB-B is 1% and WNBA is 5%.
- **0.8 CLV computed / closing persisted?** — VERIFIED (with a broken-ness caveat). `closing_odds`
on **962/996** rows, `clv` on **842/996** — so CLV IS computed and closing IS persisted. BUT
`closingCapture.js:7-8` documents it as **effectively broken**: `closing_line == locked_line on
92% of rows` (only ~56 rows show real movement), because `captureClosing` overwrites the field
with the current feed on every snapshot and most props leave the feed near their lock. CLV exists
but is **largely degenerate (≈0)**.
---
## PHASE 1 — THE GRADED LINE
- **1.9 Selector location** — VERIFIED. Two stages: (a) `gradeSlateService.dedupeProps`
(`gradeSlateService.js:34`, called `:144`); (b) the `snapshotService` dedup at
`snapshotService.js:425-434`.
- **1.10 Reads over which set?** — VERIFIED. **One provider at a time** — PropLine-normalized rows
(primary) via `oddsService``recordDownstream``gradeAndCacheSlate`. Falls back to
odds-api / oddspapi only if PropLine fails (`oddsService` fallback chain). Not all providers merged.
- **1.11 dedupeProps exists?** — VERIFIED. **YES, it EXISTS** (`gradeSlateService.js:34`). Logic:
Set on key `` `${player}::${stat_type}::${line}` `` — keeps the **FIRST** row per (player, stat,
line), discards later rows at the same (player, stat, line) (i.e. other books at the same line),
caps at `limit`. Written **Session 32** (`f0c8b4f`, "Grades pipeline + NFL/NHL wiring"). This
settles the asserted/un-asserted question: **it is real, not imagined.**
- **1.12 snapshotService 411-434** — VERIFIED. Confirmed: dedup by `` `${nameKey}|${stat_type}` ``,
keeps the row with the **highest `confidence`**. `confidence` derives from the **grade letter**
(band-midpoint from `grade_thresholds.json`, `confidence_basis:'grade_band'`) — so "highest
confidence" == "highest grade letter". It carries no information beyond the letter.
- **1.13 consensus vs first-book** — VERIFIED. **NO consensus rule exists in code.** The graded line
is chosen by **first-book-at-each-line (dedupeProps) then highest-grade-across-lines
(snapshotService)**. The "book-agnostic consensus rule" description is a prior order's *proposal*,
never built. The "first book" description is CLOSER to correct. **Code wins: no consensus.**
- **1.14 5 books at 3 lines → which line?** — VERIFIED (walked). `normalizeProps` emits one row per
book → `dedupeProps` keeps the first book at each of the 3 distinct lines (3 survivors) →
`gradeAndCacheSlate` grades both sides of each → `snapshotService` keeps the **highest-grade** of
the 3. **The graded line = whichever of the 3 lines grades highest** (a best-grade-for-us
selection, active now that the feed is 27% MLB / 82% WNBA multi-book).
- **1.15 Consumers assuming one book/prop** — VERIFIED (partial list): the grade path
(`dedupeProps` collapses book multiplicity), `analyzeViaEngine1` (`book_odds`/`fair_prob` use the
single surviving row's odds), the snapshot GameCard overlay, `detectBestBook` (no-ops <2 books).
- **1.16 Test pinning graded line vs book-set change?** — VERIFIED: **NONE.** `dedupeProps` and
`snapshotService` have unit tests for dedup mechanics, but **no test pins graded-line stability
against a change in the book set** (a multi-book different-line scenario).
- **1.17 Ledger flag for pre/post book-set change?** — VERIFIED. **No dedicated flag.** `model_version`
stamps the model era and challenger-version columns exist, but **nothing distinguishes grades by
book-set basis.** A silent book-set change would not be visible on the ledger.
---
## PHASE 2 — THE FEED
- **2.18 Books-per-prop TODAY (2026-07-26, live `/api/odds`)** — VERIFIED, and it **CORRECTS the
prior "137/140 single-book, DK129/FD14" measurement — the feed has broadened:**
- **MLB**: n=322 → `{1 book: 234 (73%), 2: 51, 3: 26, 4: 10, 5: 1}`. Books: **draftkings 273,
betmgm 105, betrivers 49, pinnacle 30, fanduel 2.** So 73% single-book (mostly DK), 27%
multi-book, **5 books now present** (not DK-only).
- **WNBA**: n=163 → `{1 book: 29 (18%), 2 books: 134 (82%)}`. Books: **fanduel 149, draftkings 148.**
**WNBA is graded off a genuine 2-book feed (DK + FD)** — better multi-book coverage than MLB.
- **2.19 ALLOWED_BOOKS** — VERIFIED. `oddsNormalizer.js:9` = 11 books (draftkings, fanduel, betmgm,
caesars, fanatics, bet365, hardrockbet, pointsbet, betrivers, pinnacle, thescore), applied at
`:110`/`:198`/`:235`. **Drops nothing today** (every book in the live feed is on the list).
- **2.20 PropLine request** — VERIFIED. `proplineAdapter.js:128` `buildUrl = ${BASE}/${sportKey}/odds`;
`:151-152` params = `{ apiKey, markets }`. **No `regions`/`bookmakers` param is sent.** The feed
broadening (2.18) happens on PropLine's DEFAULT response, not a param we added.
- **2.21 PropLine docs / multi-book param** — CANNOT DETERMINE (not re-tested this order). Prior
research (WebSearch): PropLine advertises "13 books + 5 exchanges; every payload includes a
bookmakers array" and is The-Odds-API-compatible (which uses `regions`/`bookmakers`). Whether a
param unlocks the full set on our tier is **unconfirmed** — would need a keyed test with the param.
---
## PHASE 3 — THE CHALLENGER LEDGER
- **3.22 Row counts** — VERIFIED:
| challenger | total | settled | first row | note |
|---|---|---|---|---|
| arch-v1 | 164 | 128 | 2026-07-21 | 94 with non-zero delta |
| contact-v1 | 122 | 86 | 2026-07-23 | 107 non-null `p_win_contact` (15 abstained) |
| proj-v1 + proj-v1.1 | 76 + 46 = 122 | 86 | 2026-07-23 | 119 with `proj_point` |
**All MLB** (challengers are statcast-gated → MLB only; the 407 WNBA rows carry none). Public
ledger total = **996 rows** (929 settled, 67 unsettled). **The earlier "ZERO settled p_win /
measurement not begun" is now STALE — 86-128 settled per challenger.**
- **3.23 Population rate** — VERIFIED. 872 rows in the last 14 days across **12 active days** (2 days
had no rows). Challengers populate a fraction of rows (arch 164/996) — the rest are pre-deploy or
WNBA.
- **3.24 Idempotent-lock lag** — VERIFIED, **still structural.** The ledger upsert is
`ignoreDuplicates:true`, so a challenger's fields land ONLY on rows first written AFTER that
challenger's code deployed (arch 07-21, contact/proj 07-23); re-running a snapshot never
backfills challenger fields onto an already-locked row. Any challenger added later inherits the
same gap. **Settlement itself is healthy** (`stale_unsettled = 0`).
- **3.25 Silent-failure / abstain path** — VERIFIED: **none producing fake accrual.** contact-v1's
nulls are honest abstentions (thin/absent Statcast); proj-v1 projects 119/122; arch-v1's
non-moved rows are byte-identical-to-champion (no distinctive axis). No caught-throw-writes-null
path masquerading as accrual.
- **3.26 Void/DNP rate** — VERIFIED. **88 void (8.8%)** + 67 unsettled of 996. Matches the ~9% premise.
- **3.27 model_snapshots / archetype / opp_rank** — VERIFIED (partial). `model_snapshots` is LIVE:
**4,600 rows, latest 2026-07-26 22:01**, `grade_11` on 2,888. Archetype + `opp_rank_stat` were
verified live per-sport (MLB + WNBA) in prior orders; **CANNOT re-confirm per-sport freshness here
without additional queries** (not run to keep this pass bounded).
---
## PHASE 4 — WHAT IS ACTUALLY WIRED
- **4.28 SportsGameOdds wired?** — VERIFIED. **NOTHING.** No SGO reference in `src/`, `web/src/`,
or `.env`. Not wired, configured, committed, or deployed. (The audit that qualified it as a source
was report-only.)
- **4.29 Design surfaces live vs designed** — VERIFIED:
- **S2 (book comparison / crown / disagreement):** `BookComparison.tsx` EXISTS in
`web/src/components/` but is **NOT imported/routed anywhere (dead component)**. `BookChip` +
`BookWordmark` exist and are used. `MovementStrip`, `CrownBadge` → **NOT FOUND** (never built).
(Corrects an earlier order that said "BookComparison doesn't exist" — it exists, just unrouted.)
- **THE WIRE:** = the **daily newsletter/content format** (see 4.30). Newsletter engine built
(`newsletterService`); send is unscheduled (internal endpoint only).
- **S3 article media:** **NOT FOUND** (no article-media generator; `mediaEngine` only has wire/
share text templates).
- **S-2 Offseason hub / SeasonBoard:** **NOT FOUND** (never built).
- **System / Intelligence:** design files present in `specs/design-reference/`; live components partial.
- **4.30 What is THE WIRE?** — VERIFIED. It is **VYNDR's daily newsletter / content voice**, not a UI
surface: `newsletterService.js:219-220` (`THE WIRE — {date}`) + `mediaEngine.js:6,121`
(`MORNING WIRE`, `SIGNAL` deterministic templates). It's the editorial format for the daily report.
- **4.31 detectBestBook + LineSparkline** — VERIFIED. `slateAdapter.detectBestBook` returns a book
only when ≥2 books post the SAME line at differing prices (no-ops on 1 book, never marks a lone
price "best"). `StatStrip.LineSparkline` (`StatStrip.tsx:130`) renders only at **≥3 history
points**, returns `null` below. Both degrade honest-absent on single-book/shallow data.
- **4.32 Line history / closing overwrite** — VERIFIED. `history` (`{t,line}`, line-deduped, cap 24)
is shallow — most props sit at 1-2 flat points (line moves are rare; books move odds, not the
half-point line). Two distinct closing mechanisms: the **`closing_captures` TABLE is INSERTED /
appended** (`closingCapture.js:234`), but the **`ledger_entries.closing_line/closing_odds` is
OVERWRITTEN every snapshot** (`ledgerService.js:18-21` doc; overwrite is why CLV is degenerate — 0.8).
- **4.33 Migration drift** — VERIFIED, **still present.** Repo `supabase/migrations/` has
**001-022, 025, 030, 031, 032**. **MISSING: 023, 024, 026, 027, 028, 029** — applied to prod but
untracked (they carry the challenger columns `p_win_challenger`/`challenger_*`, `model_version`,
`env_*`, `archetype_vector` — confirmed live in prod). The repo does NOT reflect prod schema; check
`information_schema.columns` via MCP, not the repo, before schema work.
- **4.34 WNBA grading live?** — VERIFIED. **YES.** 407 WNBA ledger rows, **376 settled**, graded off
a 2-book (DK+FD) feed. WNBA carries no challenger rows (statcast-gated MLB-only).
---
## PHASE 5 — DOCTRINE AND THE BOARD
- **5.35 CLAUDE.md (read in full)** — VERIFIED. Standing rules & findings:
- **NO CODE WITHOUT A SPEC**; 5 quality gates; WSL2 heredoc rule (python3 for >10-line files);
update BUILD-STATE.md / BLOCKERS.md.
- **Data Semantics Rule:** VYNDR never generates lines/odds — market values are REAL captured book
numbers; only model_value/grade/edge are model output. `Number(null)===0` is the recurring
fabrication bug; use strict null guards.
- **Grade internals:** grade is an additive integer index (`engine1`, NEUTRAL_INDEX 3), moved by
flat ±1.0/±0.5 deltas; A needs sum ≥+4.5, D ≤−1.51. `confidence` is NOT a probability (grade-band
midpoint). `p_win` is the real signal. **NEVER rescale thresholds to mint A's (permanent founder
ruling).** **A-RATED marketing on hold** until a prod fingerprint shows real A grades.
- **Three stat_type whitelists must stay in sync** (analyze.js, scan.js, validation.py).
- **Snapshot pipeline** is the product model (scheduled grade → lock to line → read from cache);
on-demand "Read" retired. SNAP_TTL 24h.
- **DO-NOT-WIRE / DO-NOT-TOUCH:** Tank01 player props = empty (do not wire); ParlayAPI host dead;
`gameLogService.getGameLogs` returns null for MLB (a trap — use `featureCache.getStatRows`);
`mlbGrader.js` is dead code; the legacy `--grade-a` token alias block kept until consumers migrate.
- **Three separate MLB stat maps on purpose** (featureCache / outcomeService / liveTrackingService) —
do not merge. Settlement is MLB-only (WNBA/NBA/soccer never settle via that path — but WNBA IS
grading + settling per 4.34, so confirm the settle path).
- **5.36 Running task list / open items** — VERIFIED. Primary: **`specs/STATE.md`** (1,886 lines,
"STATE OF THE WORLD," CURRENT STATUS + OPEN ITEMS block, last dated 2026-07-22). Also
`BUILD-STATE.md`, `BLOCKERS.md`, `DECISIONS.md`, `AUTONOMY.md`, `PROMISE-AUDIT.md`, `ROADMAP.md`.
**Open threads I have been tracking across recent orders that a strategist chat may not have:**
(a) challenger measurement now HAS settled rows (arch 128 / contact 86 / proj 86) — promotion is a
future per-prop-type ledger decision; (b) the design-migration arc (Landing hero migrated; scanner
S6/S7 blue-channel reskin shipped; S2/THE WIRE/Offseason/article-media are GAP/unbuilt);
(c) multi-book: SGO qualified (report-only), nothing wired; (d) **ROI is uncomputed in code and
decomposes to MLB-C +4.57% / MLB-B 1% / WNBA 5%** (this file).
- **5.37 Spec / roadmap files** — VERIFIED (in `specs/`): `STATE.md` (running state),
`VYNDR-NORTH-STAR.md`, `DESIGN-SPEC.md`, `ROW-GRAMMAR.md` (row grammar law), `LIVE-TRACKING.md`,
`VOICE.md`, `model-train.md`, `phase-0-kill-the-lies.md`, `phase-1-truth-infrastructure.md`,
`propline-audit.md`, `feature-1-1…4-1` + `a1-s3/s7/s9/s10` feature specs, `combat-intelligence.md`,
`design-reference/` (the design bundle). Root docs: `CLAUDE.md`, `DECISIONS.md`, `ARCHITECTURE.md`,
`AUTONOMY.md`, `BACKEND_HANDOFF.md`, `BUILD-STATE.md`, `BLOCKERS.md`, `PROMISE-AUDIT.md`.
- **5.38 Standing decisions a new session might contradict** — VERIFIED: `DECISIONS.md`
(DECISION-001+ architecture log); CLAUDE.md's permanent rulings (never mint A's; A-RATED hold;
data-semantics; do-not-wire list); STATE.md open items (A does not emit in prod; EV overconfident/
unvalidated — hero ranks on ev_pct and picks the most overconfident read; grade_11 stored in
model_snapshots). `AUTONOMY.md` = the zero-touch loop trace.
---
## CLAIMS I WAS ASKED ABOUT THAT TURNED OUT TO BE FALSE OR UNSUPPORTED
1. **"the ledger reports … 597 settled with 'no ROI'"** — the population is now **929 settled / 996
total** (597 was an earlier snapshot); and **ROI IS computable** (`locked_odds` persisted) — it is
simply **not computed in the ledger code**. "No ROI" = no computation, not incomputable.
2. **"Either ROI is not computed, or the model is selecting heavy favorites at fair prices"** — the
binary is false; **both are partially true and it decomposes by sport/grade**: MLB-B = high-hit
break-even favorites (the hypothesis), WNBA = losing, **MLB-C = genuinely +4.57% EV**. Not
uniformly zero-edge.
3. **"the graded line is chosen by a book-agnostic consensus rule"** — **no consensus rule exists in
code.** It is first-book-at-line (`dedupeProps`) + highest-grade-across-lines (`snapshotService`).
4. **Prior "137/140 single-book, draftkings 129 / fanduel 14" (~98% single-book)** — **corrected**:
today MLB is 73% single-book with **5 books present** (DK/BetMGM/BetRivers/Pinnacle/FanDuel), and
**WNBA is 82% two-book (DK+FD)**. The feed broadened.
5. **"the 11-point index is lost / unrecoverable"** — **it is retained in `model_snapshots.grade_11`**
(2,888 rows); lost only from `ledger_entries`. Recoverable via join.
6. **"ZERO settled p_win yet / challenger measurement not begun"** (from prior challenger orders) —
**stale**: arch-v1 128, contact-v1 86, proj-v1 86 rows are now settled.
7. **"BookComparison doesn't exist, only BookChip"** (an earlier order) — **BookComparison.tsx
EXISTS**, it is just unrouted/dead.
8. **CLV framed as "held / not computed"****CLV IS computed** (842 rows) and closing IS persisted
(962), but it is **largely degenerate** (closing==locked on 92%) due to the overwrite — a
different problem than "not computed."
9. **WNBA implicitly treated as not-really-grading** (settlement described as MLB-only in CLAUDE.md) —
**WNBA IS grading AND settling** (376 settled). The doctrine note about MLB-only settlement is
contradicted by the data; the WNBA settle path should be confirmed.
---
*End of canonical state. Regenerate the measured numbers before citing them in a later session —
they move as the ledger accrues. Structural facts (file/line, schema, wiring) are stable until code changes.*
+162
View File
@@ -139,6 +139,11 @@ scorer, pipeline, or real feature touched. Updated HONEST cells:
## KNOWN HONESTY GAPS (not fixed this order — logged, not fabrication) ## KNOWN HONESTY GAPS (not fixed this order — logged, not fabrication)
- **Hit rate 59% (n=763) shown without ROI/CLV** — thin, not false. ROI/CLV surfacing is a later build. - **Hit rate 59% (n=763) shown without ROI/CLV** — thin, not false. ROI/CLV surfacing is a later build.
- **"0 pushes = mis-scoring" — RETIRED 2026-07-29 as a false alarm** (premise re-verified, report-only).
The displayed hit/miss denominators are NOT corrupted by a hidden push bug: the feed is still 100%
half-numbers (0 whole lines in 117,970 captured market lines / 6,050 snapshots / 1,141 ledger rows /
173 lock_lines), all 992 settled actuals are integers, and the smallest actual-vs-line gap in the
whole ledger is 0.5. Expected pushes = exactly 0. See the verdict block below.
- **CLV instrument REPAIRED 2026-07-28** (commit 6552281). Was: 59 usable closing_prob. Now: **406** (MLB 248, WNBA 158) — the collapse was `attachClosingProb`'s `.limit(50000)`/no-ORDER-BY read + write-once `market_unavailable`, NOT capture (95% per-prop coverage) or the join (0 key mismatches). **CLV finding, straight: MLB unders lag the close (mean 9.1 prob-pts, 74% lose); MLB overs +2.0; WNBA flat.** → the +4.57% MLB-C and over/under asymmetry are substantially stale-line artifacts. This UNBLOCKS the proof order (proj-v1.1), which gates promoting p_win/ev to served grades. - **CLV instrument REPAIRED 2026-07-28** (commit 6552281). Was: 59 usable closing_prob. Now: **406** (MLB 248, WNBA 158) — the collapse was `attachClosingProb`'s `.limit(50000)`/no-ORDER-BY read + write-once `market_unavailable`, NOT capture (95% per-prop coverage) or the join (0 key mismatches). **CLV finding, straight: MLB unders lag the close (mean 9.1 prob-pts, 74% lose); MLB overs +2.0; WNBA flat.** → the +4.57% MLB-C and over/under asymmetry are substantially stale-line artifacts. This UNBLOCKS the proof order (proj-v1.1), which gates promoting p_win/ev to served grades.
**Honest state after this order: "no KNOWN live fabrications" — not "provably none."** The audit was thorough (repo + prod), but absence of a claim of falsehood is not a proof of universal truth. **Honest state after this order: "no KNOWN live fabrications" — not "provably none."** The audit was thorough (repo + prod), but absence of a claim of falsehood is not a proof of universal truth.
@@ -191,3 +196,160 @@ Gates the champion's over-CLV signal (partial r=0.375, p≈0.003, n=62 takeable
# HERO RANKING FIX — 2026-07-29 (commit 41b86e3, deployed) # HERO RANKING FIX — 2026-07-29 (commit 41b86e3, deployed)
The landing/hero (matrix row 1) selection was silently broken: it ranked on `ev_pct`, which is NULL on served grades, and **`Number(null) === 0`** made every prop tie at EV 0 → the "top read" was the FIRST takeable A/B prop in cache order — **arbitrary, dressed as ranked** (prod served Kelsey Mitchell, the #6 read by p_win). FIXED: rank by the **champion's p_win** (the only promising edge signal) among A/B **takeable-priced** reads (`isTakeable` 160..+200, same band as the proof/audit); strict-null guard; takeable filter excludes chalk; **no backfill** → honest empty state when nothing qualifies. p_win is ranking-only (never exposed; the route strips it). Display-only — reads caches, writes to nothing. No proven-edge/+EV/best-bet claim, no CLV/ROI/edge number. This makes the champion's p_win a real (display) consumer for the first time. Fingerprint VERIFIED: hero is the max-p_win read across sports (WNBA A), not old code's first-in-order MLB pick (Schanuel 135); untakeable chalk excluded. Visual auth-gated → data fingerprint. The landing/hero (matrix row 1) selection was silently broken: it ranked on `ev_pct`, which is NULL on served grades, and **`Number(null) === 0`** made every prop tie at EV 0 → the "top read" was the FIRST takeable A/B prop in cache order — **arbitrary, dressed as ranked** (prod served Kelsey Mitchell, the #6 read by p_win). FIXED: rank by the **champion's p_win** (the only promising edge signal) among A/B **takeable-priced** reads (`isTakeable` 160..+200, same band as the proof/audit); strict-null guard; takeable filter excludes chalk; **no backfill** → honest empty state when nothing qualifies. p_win is ranking-only (never exposed; the route strips it). Display-only — reads caches, writes to nothing. No proven-edge/+EV/best-bet claim, no CLV/ROI/edge number. This makes the champion's p_win a real (display) consumer for the first time. Fingerprint VERIFIED: hero is the max-p_win read across sports (WNBA A), not old code's first-in-order MLB pick (Schanuel 135); untakeable chalk excluded. Visual auth-gated → data fingerprint.
---
# PUSH-SCORING PREMISE VERIFY — verdict 2026-07-29 (report-only, read-only)
Tested the standing ruling "push scoring is correct — do not touch." That ruling rested on
"100% half-number lines → pushes structurally impossible," which was true for the data it was
made on. If whole-number lines had entered the feed since, 0 pushes across settled rows would be
a real mis-scoring bug the ruling was shielding. **The premise HOLDS — the ruling stands.**
**Phase 1 — feed distribution, 4 independent populations, per sport AND per market (never blended):**
| population | what it covers | rows with a line | whole-number lines |
|---|---|---|---|
| `closing_captures` | raw captured market lines, 5 books, `book`+`sharp`, Jul 20-29 continuous | **117,970** | **0** |
| `model_snapshots` | every graded prop **incl. grader refusals** (not survivorship-filtered) | 6,050 | **0** |
| `ledger_entries` (public) | the settled public record, 11 markets | 1,141 | **0** |
| `lock_lines` | TODAY's lock-time per-book lines (freshest feed, migration 033) | 173 | **0** |
Per-market: MLB hits / doubles / rbi / total_bases / stolen_bases / runs / strikeouts / home_runs /
walks / earned_runs / outs / hits_allowed and WNBA points / rebounds / assists / threes — **every
market's min AND max line ends in `.5`** (e.g. MLB strikeouts 2.5-8.5, WNBA points 5.5-26.5, MLB
outs 3.5-19.5). No whole-number market is hiding inside a blended fraction.
**Phase 2 — the push branch would fire.** `outcomeService.js:151` `if (a === l) return 'push'`,
reached **after** `Number()` + `Number.isFinite` guards on both operands — a sound numeric compare,
not the `Number(null) === 0` string-vs-number class that hit the hero. It is the **single scoring
chokepoint** (`ledgerService.js:31` imports `settleResult`; no parallel hit/miss derivation exists
in `src/`), it is **unit-tested live** (`outcomeService.test.js:38`, `nbaSettlement.test.js:104`),
and both `ledger_entries.outcome` and `outcomes.result` CHECK constraints **include `'push'`** — a
real push would score, write, and persist end-to-end.
**Phase 2.6 — the decisive number.** Across 992 settled rows carrying an actual: **0 exact ties, 0
fractional actuals, and the smallest actual-vs-line gap is 0.5** — the arithmetic minimum between an
integer result and a half-number line.
**VERDICT: RULING HOLDS.** Expected push rate is **exactly 0 (P = 0), not "low"** — 0/992 is
*forced*, not chance. The "implausible" flag mistook an arithmetic impossibility for a suspicious
absence; the row closes honestly. Stale n corrected: the flag said 470 settled, it is now **1,097**
(593 hit / 399 miss / 105 void / 44 unsettled-today). Nothing modified — no scoring, settlement,
re-settle, or backfill.
**No latent bug either.** Because the branch is correct and covered, a whole-number market entering
later (NFL/NHL are code-wired but out of season; whole-number strikeout props exist at some books)
would be scored as a push automatically. The residual is a **monitoring** gap, not a scoring gap:
nothing alerts on the first whole-number line to enter the feed. Logged, not built.
---
# EDGE_PCT SCALE DIAGNOSIS — 2026-07-29 (report-only, read-only). Fork REPORTED, not chosen.
**What it is (0.1).** `analyzeViaEngine1.js:265-270``edge_pct = ((projection line) / line) × 100`,
signed by direction, where `projection = l5_avg ?? l20_avg ?? {stat}_per_90 ?? xg_per_90`.
**Independent of `p_win`** (so NOT tainted by the overconfidence that damns `ev_pct`) but it takes
**no price input at all**, so it cannot express a betting edge. **Arithmetically correct, MISLABELLED:**
honest as "% the projection differs from the line," **a lie at any scale as "EDGE."** Two independent
implementations — backend `edgePctFor` and `web/src/lib/gradeAdapter.js:25-31 computeEdge`; the grade
card renders the WEB one, so a backend-only fix would miss it.
**The cap (0.2).** `SANE_EDGE_MAX = 40` (`deskShowcaseService.js:31`: *"beyond this the (model-line)/line
value isn't a market edge"*), mirrored in `slateAdapter.js:613` and `MobileEdgeBoard.tsx:45`. A
self-declared plausibility bound from an earlier order, not a derived statistical one.
**Mechanism (0.3) = SMALL-DENOMINATOR EXPLOSION** — not units, not inversion, not a missing ×100.
`line` is the denominator and **86% of MLB rows (562/655) sit at line 0.5**. Max 620 = a ~3.6 projection
on a 0.5 line.
| population | n | >cap 40 | >100 | median | p95 | max | min |
|---|---|---|---|---|---|---|---|
| **MLB** | 655 | 65.8% (>50) | 13.6% | **60** | 180 | **620** | 86.7 |
| **WNBA** | 486 | 4.9% (>50) | **0%** | 12 | 49 | 77.8 | 51.7 |
| blended | 1,141 | **44.0% (502)** | 7.8% | — | — | 620 | — |
Per line (the proof): MLB 0.5 → 73.0% over cap, max 620 · MLB 1.5 → 43.8%, max 153 · WNBA 12.5 → 6.7%
· **WNBA 26.5 → max 1.9.** Matrix figures re-verified: **51.5% is stale → 44.0%; worst 620 is exact.**
**Shape: structurally broken for MLB, sane for WNBA** — and the scale is a *function of line size*, so
the metric is incomparable across markets **by construction**. No rescaling fixes that.
**Surfaces (Phase 2) — the "~13" count is NOT confirmed. Three surfaces RENDER it:**
| surface | live | access | role | user sees at 620 |
|---|---|---|---|---|
| `GradeResultCard.tsx:182,216,325` | YES (3 importers) | **auth-gated `/scan`****TAGGED FOR CHROME AUDIT** | display | **"+620% edge", raw + GREEN** |
| `DeskShowcase.tsx:40` | YES | **PUBLIC** `/pricing` | display | **"—"** (already honest) |
| `SoccerGradeResult.tsx:229` | YES | orphan `/soccer` (0 nav links, public by URL) | display | raw uncapped `X.X% edge` |
| `MobileEdgeBoard.tsx:47` | **DEAD** (0 importers) | — | sort+display | "—" (pulled in honesty pass) |
| `PropRow:45`, `GradeCard:32`, `ledger/page:40` | live components | — | **type-only, never rendered** | nothing |
| `contentTemplateService.js:164` | public `/api/content` | API | string | uncapped — **no page fetches it** |
**🔴 IT DRIVES TWO LIVE SORTS (the fork's load-bearing answer).**
1. `slateAdapter.selectTopGrades:469-471``grade → confidence → |edge| desc` → **dashboard TOP GRADES
top-10** (`dashboard/page.tsx:419`). **97.3% of rows (1110/1141) sit in a (date,sport,grade,confidence)
tie group of ≥2** (biggest 56), so |edge| is operative for essentially the whole slate — the **de facto
ordering** of that leaderboard.
2. `analyzeViaEngine1.js:506` — the Desk **alt-line ladder** is sorted by `edge_pct` desc.
**Two scale-INDEPENDENT defects inside that sort** (`Math.abs(numOr(g.edge, -Infinity))`): **(i) abs()**
on an already-direction-signed value ranks the model's strongest *disagreements* equal to its strongest
agreements (**177 negative-edge rows**: 58 B / 118 C / 1 F, worst 86.7); **(ii)** `Math.abs(-Infinity)
= Infinity` → a **missing edge sorts FIRST**. The `Number(null)` fabrication class again, new costume.
**Phase 4 correction — "nothing renders `ev_pct`" is WRONG.** `PriceTriplet.tsx:60,67,76` renders
`${pct(ev)} EV`, live and wired (scan → `gradeAdapter:143`). It shows nothing only because `ev_pct` is
NULL on served grades → `valueState.js:121` falls to **NO_MODEL honest-absent**. The right metric already
has a live honest render site, **starved of data, not unwired** — and the card's "EDGE" row sits exactly
where a price-aware number belongs.
**THE FORK (reported, not chosen).**
- **FIX** — dishonest: no rescaling turns a price-free projection gap into an edge (renaming, not fixing);
it silently re-ranks the dashboard top-10 (the hero-class bug just fixed); needs BOTH implementations.
- **HIDE** — cheap: 3 render sites, each already has a null branch (no layout breaks), and DeskShowcase
already proves the honest "—" pattern in-product. Not load-bearing for layout anywhere.
- **RECOMMENDED: HIDE the number and re-point the sort at `p_win`** — the hero order established p_win is
on 100% of recent ledger rows and is the only signal that survived an adversarial audit. Repairing a key
that is a 0.5-line artifact is not worth it. End state: **p_win ranks · ev_pct displays · edge_pct retires.**
The abs()/null-first sort defects deserve their own small order either way.
*Nothing changed: no edge_pct, scale, surface, sort, grade, ledger, or accruing edge touched.*
---
# GRADE-BOARD SORT FIX — 2026-07-29 (spec `specs/grade-board-sort.md`, shipped)
Display ORDERING only. Fixes two defects that were wrong at ANY scale, independent of edge_pct's
separate retirement (Order B, still held).
**Defects removed.** `selectTopGrades` ranked on `Math.abs(numOr(g.edge, -Infinity))`:
`abs()` on an already-direction-signed value ranked the model's strongest **disagreements** level with
its agreements (177 public ledger rows carry a negative edge); and `Math.abs(-Infinity) === Infinity`
made a **missing** signal sort **FIRST** — absent data as the top pick. Now: `grade → confidence →
takeable-gated p_win (nulls LAST) → SIGNED edge (nulls LAST) → input order`, scales never mixed.
The alt-line ladder (`analyzeViaEngine1:506`) no longer sorts by `edge_pct`; it is ordered
highest-p_win-first via the monotonic line rule (line-ASC for an over, line-DESC for an under) at zero
added compute.
**Three premise breaks found report-first.** (1) **`/api/props/top-graded` 404s in prod** — the
dashboard board's feed does not exist, so that board renders receipts/empty and the sort orders nothing
there today; the prior order's "97.3% of rows tie → the edge key decides the board" was a ledger
measurement wrongly extrapolated to it. (2) **p_win is stripped for unentitled tiers by design**
(`snapshotGating`, Session 67 — "shipping p_win is shipping the model price"); verified live, prod
`/api/snapshot` carries p_win on **0/8 MLB and 0/25 WNBA** grades, so the browser path uses the signed
edge and only entitled callers rank on p_win. (3) **Ladder rungs carry no per-rung price**, so the
hero's takeable gate is inapplicable there.
**Verified on real data, both sports, both paths.** Unentitled: WNBA (n=25) ordering CHANGED, MLB (n=8)
unchanged; signed edge non-increasing within every (grade,confidence) tie group — 20 pairs, 0
violations. Entitled: 40 real ledger rows with p_win+locked_odds — p_win-descending, untakeable chalk
not promoted, 36 pairs, 0 violations.
**Hero consistency, honestly:** same signal + same gate, different precedence by contract (board =
grade-tier-first "top GRADES"; hero = p_win-first "top read"). They agree exactly **within** the
leading tier (verified); across tiers the board may lead with an A the hero doesn't pick. Not a
contradiction — do not "fix" it by making the board ignore grade.
**Floor:** 310 suites / 3864 tests green, web build exit 0. Dashboard + Desk visuals are auth/feed-gated
→ tagged for the Chrome audit, no visual faked. **Held:** edge_pct rescale/display retirement, building
the missing `/api/props/top-graded` selector, exposing p_win to unentitled tiers.
+142 -2
View File
@@ -110,6 +110,146 @@
> hero is Kelsey Mitchell (WNBA A), the max-p_win read ACROSS sports, and untakeable chalk > hero is Kelsey Mitchell (WNBA A), the max-p_win read ACROSS sports, and untakeable chalk
> (Yainer Diaz 200, Altuve C) is excluded → new code confirmed serving. Visual is auth-gated > (Yainer Diaz 200, Altuve C) is excluded → new code confirmed serving. Visual is auth-gated
> (landing hero public, dashboard not) — data fingerprint used, `cf-cache-status: DYNAMIC`. > (landing hero public, dashboard not) — data fingerprint used, `cf-cache-status: DYNAMIC`.
> ## ✅ PUSH-SCORING PREMISE VERIFY 2026-07-29 (report-only): **RULING HOLDS — row closed**
> Tested whether the standing "push scoring is correct — do not touch" ruling still rests on a
> true premise, or whether whole-number lines had entered the feed since it was made (which
> would make 0 pushes a real mis-scoring bug the ruling was shielding). **The premise HOLDS.**
> 0.1 basis confirmed verbatim (`VYNDR-CANONICAL-STATE.md:69`): "*no `push`* (pushes structurally
> impossible: half-number lines)". **Phase 1 — the feed is STILL 100% half-numbers, on 4
> independent populations, per sport AND per market (no blend): `closing_captures` **117,970
> priced market lines → 0 whole** (5 books incl. `sharp`/pinnacle, 16 sport/stat groups, Jul 2029
> continuous); `model_snapshots` 6,050 → 0 (incl. grader refusals, so not survivorship);
> `ledger_entries` 1,141 public → 0 (11 markets: MLB hits/doubles/TB/SB/ER/HR/outs, WNBA pts/reb/
> ast/3s — every min AND max ends in `.5`); `lock_lines` 173 → 0 (TODAY's lock, freshest feed).
> **Phase 2** — push branch is `outcomeService.js:151` `if (a === l) return 'push'` **after**
> `Number()` + `Number.isFinite` guards on both operands → sound numeric compare, NOT the
> `Number(null)===0` string-vs-number class. It is the **single scoring chokepoint** (`ledgerService.js:31`
> imports `settleResult`; no parallel scorer derives hit/miss anywhere in `src/`), it is **unit-tested
> live** (`outcomeService.test.js:38` `settleResult('over',2,2)==='push'`, `nbaSettlement.test.js:104`),
> and both `ledger_entries.outcome` + `outcomes.result` CHECK constraints **include `'push'`** → a real
> push would score, write, and persist. **Phase 2.6 / the decisive number: 0 exact ties in 992 settled
> rows, 0 fractional actuals, and the SMALLEST actual-vs-line gap across all 992 rows is 0.5** — the
> arithmetic minimum between an integer result and a half-number line. **VERDICT: RULING HOLDS.**
> Expected push rate is **exactly 0 (P=0), not "low" — 0/992 is FORCED, not chance**, so the
> "implausible" flag is retired: it mistook an arithmetic impossibility for a suspicious absence.
> The open-items row is CLOSED honestly. Nothing was modified (no scoring, settlement, re-settle,
> or backfill). NOTE the stale n: the flag said 470 settled; it is now **1,097 settled** (593 hit /
> 399 miss / 105 void / 44 unsettled-today). **No latent bug either** — the branch is correct and
> covered, so if a whole-number market ever DOES enter (NFL/NHL are code-wired but out of season;
> whole-number K props exist at some books), it scores as a push automatically. Residual risk is
> a monitoring gap, not a scoring gap: nothing ALERTS on the first whole-number line.
> ## 🔬 EDGE_PCT SCALE DIAGNOSIS 2026-07-29 (report-only): **not a units bug — a mislabelled metric that DRIVES A LIVE SORT**
> **0.1 What it is.** `analyzeViaEngine1.js:265-270 edgePctFor()`:
> `signed = over ? (projection line) : (line projection)`; `Math.round((signed/line)*1000)/10`
> → **edge_pct = ((projection line) / line) × 100**, signed by direction, 1dp. `projection` =
> `projectionFor` = `l5_avg ?? l20_avg ?? {stat}_per_90 ?? xg_per_90` (`:252-257`). **It is
> INDEPENDENT of `p_win`** — so it is NOT tainted by the overconfidence that damns `ev_pct`. But it
> takes **NO price input at all**, so it cannot express edge in the betting sense. **Verdict:
> arithmetically correct, MISLABELLED. Honest as "% the model's projection differs from the line";
> a lie at ANY scale as "EDGE"** — which is exactly how every surface labels it. **TWO independent
> implementations exist** — backend `edgePctFor` AND `web/src/lib/gradeAdapter.js:25-31 computeEdge`
> (same formula); the grade card renders the WEB one, so a backend-only fix would not reach it.
> **0.2 The cap.** `SANE_EDGE_MAX = 40`, `deskShowcaseService.js:31`, comment: *"beyond this the
> (model-line)/line value isn't a market edge"*. Mirrored `slateAdapter.js:613 EDGE_BOARD_SANE_MAX=40`
> + `MobileEdgeBoard.tsx:45`. It is a self-declared plausibility bound from an earlier order, **not a
> derived statistical bound**.
> **0.3 Mechanism = SMALL-DENOMINATOR EXPLOSION** (not units, not inverted, not missing ×100 — the
> ×100 is present and correct). `line` is the denominator and **562 of 655 MLB rows (86%) sit at line
> 0.5**, where every 0.1 of projection is ±20 points. Max 620 = projection ≈3.6 on a 0.5 line. PROVEN
> per-line: MLB 0.5 → 73.0% over cap, max 620 · MLB 1.5 → 43.8%, max 153 · WNBA 12.5 → 6.7% · **WNBA
> 26.5 → max 1.9**. Monotone decay with line size.
> **PHASE 1 — matrix figures RE-VERIFIED, partly stale.** Blended over cap-40 = **502/1141 = 44.0%**
> (matrix said 51.5% — direction right, number stale). **Worst value 620 = EXACT match.** Per sport:
> **MLB n=655** — 13.6% >100, 65.8% >50, median **60** (already 1.5× the cap), p95 180, p99 238,
> min 86.7, max 620. **WNBA n=486 — 0% >100**, median 12, p95 49, max 77.8. **Shape: structurally
> broken for MLB, essentially SANE for WNBA.** Not a mild calibration — and note what that means:
> the scale is a FUNCTION OF LINE SIZE, so the metric is incomparable across markets **by
> construction**. No rescaling fixes that; only changing what the metric IS would.
> **PHASE 2 — the ~13-surface count is NOT confirmed. Only 3 surfaces RENDER it:**
> (a) **`GradeResultCard.tsx:182,216,325`** — LIVE (3 importers), **auth-gated `/scan`** →
> **TAGGED FOR THE CHROME AUDIT, no visual faked** — renders **"+620% edge" RAW and GREEN**, no cap;
> (b) **`DeskShowcase.tsx:40`** — LIVE, **PUBLIC** `/pricing` — **already honest**, shows "—"
> (service nulls >40); (c) **`SoccerGradeResult.tsx:229`** — uncapped, on the ORPHAN `/soccer`
> (0 nav links, public by URL). **Dead:** `MobileEdgeBoard` (0 importers, pulled in the honesty pass),
> `DemoScan` (0 importers). **TYPE-ONLY, never rendered:** `PropRow.tsx:45`, `GradeCard.tsx:32`,
> `ledger/page.tsx:40`. **API-only:** `contentTemplateService.js:164` emits an uncapped "+620% edge"
> string on public `/api/content` — **no page fetches it** (verified).
> **🔴 THE SORT ANSWER — YES, IT DRIVES TWO LIVE SORTS.** (1) `slateAdapter.selectTopGrades:469-471`
> sorts `grade → confidence → |edge| desc`, consumed by `dashboard/page.tsx:419` for the dashboard
> **TOP GRADES top-10**. **MEASURED: 97.3% of rows (1110/1141) sit in a (date,sport,grade,confidence)
> tie group of ≥2 (biggest 56)** → the |edge| key is operative for essentially the whole slate, so it
> is the **de facto ordering** of that leaderboard. (2) `analyzeViaEngine1.js:506` sorts the Desk
> alt-line ladder by `edge_pct` desc. **TWO SCALE-INDEPENDENT DEFECTS FOUND INSIDE THAT SORT:**
> `edge: Math.abs(numOr(g.edge, -Infinity))` — (i) **abs()** on an already-direction-signed value ranks
> the model's strongest DISAGREEMENTS equal to its strongest agreements (**177 ledger rows carry a
> negative edge**: 58 B / 118 C / 1 F, most negative 86.7); (ii) `Math.abs(-Infinity) = Infinity`, so
> a **MISSING edge sorts FIRST** — the `Number(null)` fabrication class again, in a new costume.
> **PHASE 4 — the matrix's "nothing renders `ev_pct`" is WRONG.** `PriceTriplet.tsx:60,67,76` renders
> `${pct(ev)} EV` and is LIVE + wired (scan → `gradeAdapter:143` → PriceTriplet). It shows nothing only
> because `ev_pct` is NULL on served grades, so `valueState.js:121` correctly falls to **NO_MODEL
> honest-absent**. So: **the right metric already has a live, honest render site starved of data** —
> and the card's "EDGE" row sits exactly where a price-aware number belongs.
> **PHASE 3 — THE FORK (reported, NOT chosen).** **FIX is not honest**: no rescaling turns a
> price-free projection gap into an edge — you would be renaming, not fixing; it silently re-ranks the
> dashboard top-10 (the hero-class bug just fixed); and it must land in TWO implementations.
> **HIDE is cheap**: only 3 render sites, every one already has a null branch (so no layout breaks),
> and **DeskShowcase already proves the honest "—" pattern in-product**. **RECOMMENDED: HIDE the
> number, and re-point the sort at `p_win`** (the hero order established p_win is on 100% of recent
> ledger rows and is the one signal that survived an adversarial audit) rather than repair a key that
> is a 0.5-line artifact. Coherent end state: **p_win ranks · ev_pct displays · edge_pct retires.**
> The abs()/null-first sort defects are worth their own small order regardless of the fork.
> ## 🔧 GRADE-BOARD SORT FIXED 2026-07-29 (spec `specs/grade-board-sort.md`): signed signal, missing sorts LAST
> Display ORDERING only — no grade, ledger, lock_line, scoring, or edge_pct scale/display change.
> **THREE PREMISE BREAKS found REPORT-FIRST, before code:** (1) **`/api/props/top-graded` returns 404
> in prod** — it does not exist in `src/` (only 3 axios *callers*), so the Next proxy catches →
> `{props:[]}` → the dashboard board renders `proofMode`/empty and **the edge sort orders nothing on
> that surface today**. CORRECTION to my prior order: its "97.3% of rows tie → the edge key decides the
> board" was a LEDGER measurement I extrapolated to this board — wrong; the board has no rows. The fix
> is still correct-in-itself and lands the moment the feed is restored. (2) **p_win CANNOT be a
> client-side sort key for all tiers** — `utils/snapshotGating.stripModelPrice` (Session 67) strips
> `p_win`/`ev_pct`/`model_odds`/`value`/`takeable` for unentitled tiers because *"shipping p_win is
> shipping the model price in a different base"*, and `selectTopGrades` runs in the BROWSER. Ranking
> there by p_win for everyone would REVERSE that gate. **VERIFIED LIVE: prod `/api/snapshot` returns
> p_win on 0/8 MLB and 0/25 WNBA grades** (stripped, as designed). (3) **The ladder cannot take the
> takeable gate** — rungs carry NO per-rung price (books price each line differently; we don't fetch
> them) and `isTakeable` is a property of price alone.
> **WHAT SHIPPED.** `selectTopGrades`: `grade → confidence → takeable-gated p_win (nulls LAST) →
> SIGNED edge (nulls LAST) → input order`. `Math.abs()` GONE — `edge` is signed by direction upstream
> so positive = the model AGREES; `|edge|` had been ranking the model's strongest DISAGREEMENTS level
> with its agreements (177 public ledger rows carry a negative edge). `Math.abs(-Infinity)=Infinity`
> GONE — a missing signal sorted FIRST (absent data as the top pick, the `Number(null)` class); it now
> sorts LAST and rows are never dropped. **Scales are never mixed** (p_win 0..1 vs edge % — 0.62 vs 62
> is not a comparison). Takeable band = `web/src/lib/valueState.isTakeable`, asserted by test to be
> byte-equal to the hero's `config/valueEngine.isTakeable` (160..+200) incl. strict-null.
> **Alt-line ladder** (`analyzeViaEngine1:506`): was `edge_pct desc` (and `Number(x)||0` collapsed
> absent edges to mid-pack); now ordered **highest-p_win-first derived ANALYTICALLY at zero added
> compute** — `P(stat ≥ k)` is monotone non-increasing in k, so p_win-desc is exactly line-ASC for an
> over and line-DESC for an under. `base` stays marked, so order never implies a recommendation; no
> consumer depends on `alt_lines[0]` (grepped), and `deskShowcaseService.rungsOf:40` already re-sorted
> by line anyway.
> **FINGERPRINTED ON REAL DATA (both sports, both paths).** Unentitled path, live prod snapshots:
> **WNBA (n=25) ordering CHANGED**, MLB (n=8) unchanged (its top edges were already positive);
> **within every (grade,confidence) tie group the signed edge is non-increasing — 20 adjacent pairs,
> 0 violations**. Entitled path, 40 real ledger rows carrying p_win+locked_odds: p_win-descending
> within tie groups, **untakeable chalk correctly NOT promoted** (Trea Turner p_win .757 @275 does
> not beat Rhyne Howard .745 @120) — **36 pairs, 0 violations**.
> **HERO CONSISTENCY — stated honestly, they are NOT identical and should not be.** Same signal, same
> gate, DIFFERENT precedence by contract: the board is `top GRADES` (grade-tier first), the hero is
> `top read` (p_win first). On the real rows the board leads with Angel Reese (A, p_win .555) while
> the hero picks Brionna Jones (B, p_win .90). **Within the leading tier they agree exactly (verified
> true).** That is a difference of question, not a contradiction — do not "fix" it by making the board
> ignore grade.
> **FLOOR: 310 suites / 3864 tests green, web build exit 0.** New `tests/unit/gradeBoardSort.test.js`
> locks: disagreement never outranks agreement, missing sorts LAST and stays present, null/'' edge is
> absent-not-zero, takeable p_win beats untakeable higher-p_win chalk, scales never mixed, the two
> takeable bands match, grade tier still dominates, and the ladder order for BOTH directions.
> **HELD:** edge_pct rescale/display retirement (Order B) · building the missing
> `/api/props/top-graded` server selector (that is what would make the board render at all) ·
> exposing p_win to unentitled tiers. **Dashboard + Desk visuals are auth/feed-gated → TAGGED FOR
> THE CHROME AUDIT, no visual faked.**
- **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows``closing_captures` (append-only, - **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows``closing_captures` (append-only,
provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221` provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221`
+ internal endpoint; `ledgerService.attachClosingProb``closing_prob` (de-vigs both raw sides, + internal endpoint; `ledgerService.attachClosingProb``closing_prob` (de-vigs both raw sides,
@@ -243,12 +383,12 @@ exist locally; harmless, the data restores completely.
| Item | Status | Note | | Item | Status | Note |
|---|---|---| |---|---|---|
| **Settlement: 0 pushes / 470 settled** | 🔴 OPEN, unstarted | Implausible — hits/TB land on the number regularly. Exact-number push almost certainly mis-scored as hit or miss. Corrupts every accuracy/ROI number. | | **Settlement: 0 pushes / 470 settled** | **CLOSED 2026-07-29 — premise re-verified, NOT a bug** | The ruling's basis STILL HOLDS: the feed is **100% half-numbers**. 0 whole-number lines in **117,970 captured market lines** (5 books, 16 markets, `book`+`sharp`, Jul 2029), 6,050 `model_snapshots` (incl. refusals), 1,141 ledger rows, 173 `lock_lines` (today). All 992 settled actuals are INTEGERS → **smallest actual-vs-line gap across all 992 = 0.5**, the arithmetic minimum. Expected pushes = **exactly 0 (P=0)**, not chance. Push branch sound + tested; CHECK constraints accept `'push'`. See the verify block above. |
| **~28 props/day never settle** | 🔴 OPEN, unstarted | Jul 17 MLB 86 graded/57 settled; Jul 18 103/75. Cause undiagnosed. | | **~28 props/day never settle** | 🔴 OPEN, unstarted | Jul 17 MLB 86 graded/57 settled; Jul 18 103/75. Cause undiagnosed. |
| **Model-version contamination** | 🟠 PERMANENT, mitigate | `ledger_entries` mixes pre/post-2026-07-19-fix grades with no marker; eras cannot be separated retroactively. **Any backtest/accuracy claim off existing ledger history MUST treat the fix boundary as a hard cutoff.** `model_snapshots` stamps `model_version`+`code_sha` so it can't recur. | | **Model-version contamination** | 🟠 PERMANENT, mitigate | `ledger_entries` mixes pre/post-2026-07-19-fix grades with no marker; eras cannot be separated retroactively. **Any backtest/accuracy claim off existing ledger history MUST treat the fix boundary as a hard cutoff.** `model_snapshots` stamps `model_version`+`code_sha` so it can't recur. |
| **A-grade unreachable in prod** | 🔴 OPEN | `opp_rank_stat` null; ESPN team endpoint has no defensive metric at all. Marketing hold stands. | | **A-grade unreachable in prod** | 🔴 OPEN | `opp_rank_stat` null; ESPN team endpoint has no defensive metric at all. Marketing hold stands. |
| **EV overconfident** | 🟠 OPEN | Needs calibration before it drives any surface. Hero v2 already ranks on it. | | **EV overconfident** | 🟠 OPEN | Needs calibration before it drives any surface. Hero v2 already ranks on it. |
| **`edge_pct` broken scale (U-deg pt 2)** | 🔴 OPEN | 51.5% of ledger rows exceed the sane cap; worst 620. 13 frontend surfaces render it; **nothing renders `ev_pct`**; it's the free-tier hook; it's written to the append-only `edge` column every cron. | | **`edge_pct` broken scale (U-deg pt 2)** | 🔴 OPEN **DIAGNOSED 2026-07-29, fork reported, not chosen** | Re-verified: **44.0% over cap-40 (502/1141)**, not 51.5%; **worst 620 exact**. MLB median 60 / max 620 (86% of MLB rows are 0.5 lines); **WNBA sane** (median 12, 0% >100). Cause = **small-denominator explosion**, NOT units. **Only 3 surfaces RENDER it** (not 13): GradeResultCard (uncapped, auth-gated), DeskShowcase (already honest "—"), SoccerGradeResult (uncapped, orphan). **It DRIVES the dashboard top-10 sort** (operative on 97.3% of rows) + the Desk alt-ladder. `ev_pct` **DOES** have a live render site (PriceTriplet) — starved, not unwired. **Recommendation: HIDE + re-point the sort at p_win.** See the diagnosis block above. |
| **CLV broken (C4)** | 🔴 OPEN | `closing_line == locked_line` on ~95% of rows. BEAT CLOSE suppressed. **CLV ledger stays PRIVATE until backtest-proven.** | | **CLV broken (C4)** | 🔴 OPEN | `closing_line == locked_line` on ~95% of rows. BEAT CLOSE suppressed. **CLV ledger stays PRIVATE until backtest-proven.** |
| **Consistency CV floor** | 🟠 STOPGAP | `CONSISTENCY_MIN_MEAN=4` leaves a ±1.0 dead for MLB low-count stats. Real fix = index-of-dispersion classifier; needs a backtest first. | | **Consistency CV floor** | 🟠 STOPGAP | `CONSISTENCY_MIN_MEAN=4` leaves a ±1.0 dead for MLB low-count stats. Real fix = index-of-dispersion classifier; needs a backtest first. |
+73
View File
@@ -0,0 +1,73 @@
# SPEC — Fix the grade-board sort (display ordering only)
**Status:** built 2026-07-29. Report-first found three premise breaks — see §2.
**Scope:** display ORDERING only. No grade, ledger row, lock_line, scoring, model,
or edge_pct scale/display change. Push scoring untouched.
## 1. The two defects (wrong at ANY scale, independent of edge_pct's retirement)
`web/src/lib/slateAdapter.js selectTopGrades` ranked on
`edge: Math.abs(numOr(g.edge, -Infinity))`:
1. **abs() on an already-signed value.** `edge` is signed BY DIRECTION upstream
(`edgePctFor`: `over ? projline : lineproj`), so **positive = the model AGREES
with the graded side**. Taking `|edge|` ranked the model's strongest
DISAGREEMENTS equal to its strongest agreements. 177 public ledger rows carry a
negative edge (58 B / 118 C / 1 F, worst 86.7).
2. **`Math.abs(-Infinity) === Infinity`** → a row with NO edge sorted **FIRST**.
Absent data presented as the top pick — the `Number(null)` fabrication class.
`src/services/intelligence/analyzeViaEngine1.js:506` sorted the Desk alt-line
ladder by `edge_pct` desc — the same price-free 0.5-line artifact.
## 2. REPORT-FIRST — three premise breaks found before building
- **B1. Site 1's board is fed by a 404.** `/api/props/top-graded` **does not exist**
in `src/` (only three axios *callers* reference it) and returns **404 in prod**.
The Next proxy catches → `{props: []}``topGrades = []``selectTopGrades([])`
→ the dashboard renders `proofMode` (yesterday's receipts) or honest empty copy.
**The edge sort orders nothing on that surface today.** CORRECTION to the prior
order: its "97.3% of rows tie → the edge key decides the board" was a LEDGER
population measurement extrapolated to this board; the board has no rows to order.
The fix is still correct-in-itself and lands the moment the feed is restored.
- **B2. p_win CANNOT be the client-side sort key.** `src/utils/snapshotGating.js`
(Session 67) strips `p_win`/`ev_pct`/`model_odds`/`value`/`takeable` for
unentitled tiers, with the explicit rationale *"shipping p_win is shipping the
price in a different base."* `selectTopGrades` runs in the BROWSER. Ranking there
by p_win for all tiers would REVERSE that gate. Implemented instead:
**p_win is used WHEN THE ROW CARRIES IT** (entitled tiers / server callers),
takeable-gated identically to the hero; unentitled tiers fall through to the
signed edge uniformly. Scales are NEVER mixed in one comparator.
- **B3. The ladder cannot take the takeable gate, and p_win order there is
ANALYTICALLY the line order.** Ladder rungs carry **no per-rung price** (books
price each line differently; we do not fetch them), and `isTakeable` is a
property of price alone → inapplicable per rung. And `P(stat ≥ k)` is monotone
non-increasing in `k`, so **p_win-desc ≡ line-asc for an over, line-desc for an
under**. Implemented as the direction-aware line order = exactly p_win-desc, at
zero added compute. (`deskShowcaseService.rungsOf:40` already re-sorted rungs by
line, discarding the edge order — so only `GradeResultCard` ever showed it.)
## 3. What ships
- `selectTopGrades`: `grade → confidence → takeable-gated p_win (nulls last) →
SIGNED edge (nulls last) → input order`. No `abs()`. Missing signal ranks LAST,
rows are never dropped.
- Alt-line ladder: ordered by highest p_win first, computed via the monotonic
line equivalence (direction-aware). No `edge_pct` in the sort.
- The takeable band is `web/src/lib/valueState.isTakeable` (160..+200), which
mirrors `src/config/valueEngine.isTakeable` — the hero's exact definition.
## 4. Acceptance criteria
1. A disagreement (negative-edge) prop does NOT outrank an agreement at equal
grade+confidence.
2. A missing-signal row sorts LAST and is still PRESENT.
3. A takeable p_win row outranks a higher-p_win UNTAKEABLE (chalk) row.
4. p_win and edge are never compared against each other.
5. Ladder rungs are ordered highest-p_win-first for both over and under.
6. No grade/ledger/lock_line/scoring write. Full suite green, web build exit 0.
## 5. Held (NOT this order)
edge_pct rescale or display retirement (Order B) · building the missing
`/api/props/top-graded` server selector · exposing p_win to unentitled tiers.
+17 -1
View File
@@ -503,7 +503,23 @@ async function analyzeViaEngine1(rawProp = {}) {
}, { edgePct: edgePctFor(features, { ...shiftedProp, stat_type: rawProp.stat_type, direction: prop.direction }) }); }, { edgePct: edgePctFor(features, { ...shiftedProp, stat_type: rawProp.stat_type, direction: prop.direction }) });
return { line: ln, grade: adapted.grade, edge_pct: adapted.edge_pct, base: ln === baseLine }; return { line: ln, grade: adapted.grade, edge_pct: adapted.edge_pct, base: ln === baseLine };
}) })
.sort((a, b) => (Number(b.edge_pct) || 0) - (Number(a.edge_pct) || 0)); // SORT FIX (2026-07-29, specs/grade-board-sort.md). Was
// `edge_pct desc` — a price-free (projline)/line artifact whose scale is
// a function of line size, so on a 0.5 line it explodes and ordered the
// ladder by nothing meaningful. `Number(x) || 0` also collapsed absent
// edges to 0 (mid-pack).
//
// Now ordered HIGHEST-p_win-FIRST, derived analytically at zero added
// compute: P(stat ≥ k) is monotone NON-INCREASING in k, so for an OVER
// p_win-desc is exactly line-ASC, and for an UNDER (p_win = 1 p_over)
// it is exactly line-DESC. Rungs carry NO per-rung price — books price
// each line differently and we do not fetch them — so the hero's
// takeable gate (a property of price alone) is inapplicable here; that
// is why this ranks on probability order only. The `base` rung stays
// marked, so ordering never implies a recommendation.
.sort((a, b) => (String(prop.direction || 'over').toLowerCase() === 'under'
? Number(b.line) - Number(a.line)
: Number(a.line) - Number(b.line)));
if (ladder.length > 1) legacy.alt_lines = ladder; if (ladder.length > 1) legacy.alt_lines = ladder;
} }
} catch { /* the ladder is additive — never breaks the read */ } } catch { /* the ladder is additive — never breaks the read */ }
+134
View File
@@ -0,0 +1,134 @@
/**
* Grade-board sort (specs/grade-board-sort.md) — display ORDERING only.
*
* Locks the two defects that were wrong at ANY scale:
* 1. abs() on an already-direction-signed edge ranked DISAGREEMENTS level with
* agreements.
* 2. Math.abs(-Infinity) === Infinity made a MISSING signal sort FIRST.
* Plus: the takeable gate matches the hero, and p_win/edge scales never mix.
*/
const adapter = require('../../web/src/lib/slateAdapter');
const { isTakeable } = require('../../web/src/lib/valueState');
const { isTakeable: backendIsTakeable } = require('../../src/config/valueEngine');
const names = (rows) => rows.map((r) => r.player);
describe('selectTopGrades — signed signal, missing sorts LAST', () => {
test('a DISAGREEMENT does not outrank an AGREEMENT at equal grade+confidence', () => {
// edge is signed by direction: positive = model agrees with the graded side.
// Old |edge| ranked -80 (strong disagreement) above +10 (mild agreement).
const grades = [
{ player: 'disagrees', grade: 'B', confidence: 50, edge: -80 },
{ player: 'agrees', grade: 'B', confidence: 50, edge: 10 },
];
expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['agrees', 'disagrees']);
});
test('a MISSING signal sorts LAST and is still PRESENT (never dropped)', () => {
const grades = [
{ player: 'no-signal', grade: 'B', confidence: 50 },
{ player: 'weak-but-real', grade: 'B', confidence: 50, edge: 0.1 },
{ player: 'negative-but-real', grade: 'B', confidence: 50, edge: -5 },
];
const out = names(adapter.selectTopGrades(grades, 10));
expect(out).toEqual(['weak-but-real', 'negative-but-real', 'no-signal']);
expect(out).toHaveLength(3); // present, not dropped
expect(out[out.length - 1]).toBe('no-signal');
});
test('null/empty-string edge is absent, NOT zero (Number(null) === 0 guard)', () => {
const grades = [
{ player: 'nullish', grade: 'B', confidence: 50, edge: null },
{ player: 'empty', grade: 'B', confidence: 50, edge: '' },
{ player: 'real-negative', grade: 'B', confidence: 50, edge: -1 },
];
// A real negative beats two absents; absents keep input order at the bottom.
expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['real-negative', 'nullish', 'empty']);
});
});
describe('selectTopGrades — takeable-gated p_win outranks edge, scales never mix', () => {
test('takeable p_win row outranks an UNTAKEABLE higher-p_win chalk row', () => {
const grades = [
{ player: 'chalk', grade: 'B', confidence: 50, p_win: 0.92, book_odds: -300 }, // untakeable
{ player: 'takeable', grade: 'B', confidence: 50, p_win: 0.61, book_odds: -120 },
];
expect(names(adapter.selectTopGrades(grades, 10))[0]).toBe('takeable');
});
test('p_win is preferred over edge, and a p_win row outranks an edge-only row', () => {
const grades = [
{ player: 'edge-only', grade: 'B', confidence: 50, edge: 300 },
{ player: 'has-pwin', grade: 'B', confidence: 50, p_win: 0.55, book_odds: 100 },
];
// 0.55 must NOT be compared against 300 — the p_win-bearing row wins on the
// earlier key instead (scales never mixed in one comparator).
expect(names(adapter.selectTopGrades(grades, 10))[0]).toBe('has-pwin');
});
test('among takeable p_win rows the HIGHER p_win wins', () => {
const grades = [
{ player: 'lower', grade: 'B', confidence: 50, p_win: 0.55, book_odds: -110 },
{ player: 'higher', grade: 'B', confidence: 50, p_win: 0.71, book_odds: -110 },
];
expect(names(adapter.selectTopGrades(grades, 10))[0]).toBe('higher');
});
test('gradedAt.odds is accepted as the price when book_odds is absent', () => {
const grades = [
{ player: 'via-gradedAt', grade: 'B', confidence: 50, p_win: 0.66, gradedAt: { odds: -115 } },
{ player: 'edge-only', grade: 'B', confidence: 50, edge: 5 },
];
expect(names(adapter.selectTopGrades(grades, 10))[0]).toBe('via-gradedAt');
});
test('the frontend takeable band MATCHES the backend hero gate exactly', () => {
for (const price of [-400, -160, -159, -110, 0, 100, 200, 201, 500]) {
expect(isTakeable(price)).toBe(backendIsTakeable(price));
}
// and the strict-null contract both sides
expect(isTakeable(null)).toBe(false);
expect(backendIsTakeable(null)).toBe(false);
});
test('grade tier still dominates every signal', () => {
const grades = [
{ player: 'B-strong', grade: 'B', confidence: 99, p_win: 0.99, book_odds: -110 },
{ player: 'A-weak', grade: 'A', confidence: 1, edge: -50 },
];
expect(names(adapter.selectTopGrades(grades, 10))[0]).toBe('A-weak');
});
});
describe('alt-line ladder — ordered highest-p_win-first via the monotonic line rule', () => {
// P(stat >= k) is monotone non-increasing in k, so p_win-desc is line-ASC for an
// over and line-DESC for an under. Assert the ordering the engine emits.
const ladderOrder = (direction, lines) => {
const rungs = lines.map((line) => ({ line }));
return rungs
.slice()
.sort((a, b) => (String(direction).toLowerCase() === 'under'
? Number(b.line) - Number(a.line)
: Number(a.line) - Number(b.line)))
.map((r) => r.line);
};
test('OVER: lowest line (highest p_win) first', () => {
expect(ladderOrder('over', [1.5, 0.5, 2.5, 1, 2])).toEqual([0.5, 1, 1.5, 2, 2.5]);
});
test('UNDER: highest line (highest p_win) first', () => {
expect(ladderOrder('under', [1.5, 0.5, 2.5, 1, 2])).toEqual([2.5, 2, 1.5, 1, 0.5]);
});
test('the engine sorts its ladder by that rule, NOT by edge_pct', () => {
const src = require('fs').readFileSync(
require('path').join(__dirname, '../../src/services/intelligence/analyzeViaEngine1.js'),
'utf8',
);
// the old key must be gone from the ladder sort
expect(src).not.toMatch(/sort\(\(a, b\) => \(Number\(b\.edge_pct\)/);
expect(src).toMatch(/Number\(b\.line\) - Number\(a\.line\)/);
});
});
+1 -1
View File
File diff suppressed because one or more lines are too long
+62 -4
View File
@@ -454,10 +454,62 @@ const numOr = (v, fallback) => {
return Number.isFinite(n) ? n : fallback; return Number.isFinite(n) ? n : fallback;
}; };
/** Strict numeric read — `Number(null) === 0` is the recurring fabrication bug. */
const strictNum = (v) => (v == null || v === '' ? null : (Number.isFinite(Number(v)) ? Number(v) : null));
/**
* The takeable-gated champion probability for one grade row, or null.
*
* Gate = `valueState.isTakeable` (160..+200), which mirrors
* `src/config/valueEngine.isTakeable` — the IDENTICAL band the hero ranks on
* (`heroPropService` HERO RULE v3), so the two ranking surfaces agree. The gate
* is mandatory: raw p_win crowns 300 chalk, which is not the product.
*
* p_win is ABSENT BY DESIGN for unentitled tiers — `utils/snapshotGating`
* strips it on the way out because "shipping p_win is shipping the model price
* in a different base" (Session 67). Those callers legitimately get null here
* and fall through to the signed edge.
*/
function takeablePWin(g) {
const { isTakeable } = require('./valueState');
const p = strictNum(g && g.p_win);
if (p == null) return null;
const price = strictNum(g && g.book_odds) ?? strictNum(g && g.gradedAt && g.gradedAt.odds);
if (price == null || !isTakeable(price)) return null;
return p;
}
/** Descending comparator that always sorts a null signal LAST (never first). */
function descNullsLast(a, b) {
if (a == null && b == null) return 0;
if (a == null) return 1;
if (b == null) return -1;
return b - a;
}
/** /**
* #13 — rank tonight's grades by TIER, then the VARYING signal so a leaderboard * #13 — rank tonight's grades by TIER, then the VARYING signal so a leaderboard
* of near-identical rows stops being noise: confidence desc, then |edge| desc, * of near-identical rows stops being noise. Drops gradeless rows. Returns at
* stable by input order. Drops gradeless rows. Returns at most `limit`. * most `limit`.
*
* SORT FIX (2026-07-29, `specs/grade-board-sort.md`). The old key was
* `edge: Math.abs(numOr(g.edge, -Infinity))`, carrying two defects that were
* wrong at ANY scale — independent of edge_pct's separate retirement:
*
* 1. abs() ON AN ALREADY-SIGNED VALUE. `edge` is signed BY DIRECTION upstream
* (`edgePctFor`: over ? projline : lineproj), so POSITIVE = the model
* AGREES with the graded side. Ranking |edge| put the model's strongest
* DISAGREEMENTS level with its strongest agreements — the board promoted
* reads the model contradicts (177 public ledger rows carry a negative edge).
* 2. `Math.abs(-Infinity) === Infinity`, so a row with NO edge sorted FIRST.
* Absent data presented as the top pick — the `Number(null)` class again.
*
* Now: takeable-gated champion p_win first (so this board and the hero agree),
* then the SIGNED edge, each descending with MISSING SORTING LAST. Rows are
* never dropped for a missing signal — they rank at the bottom.
*
* SCALES ARE NEVER MIXED: p_win (0..1) is compared only against p_win, edge (%)
* only against edge. Comparing 0.62 against 62 is not a comparison.
*/ */
function selectTopGrades(grades, limit = 10) { function selectTopGrades(grades, limit = 10) {
const arr = (Array.isArray(grades) ? grades : []).filter((g) => g && g.grade); const arr = (Array.isArray(grades) ? grades : []).filter((g) => g && g.grade);
@@ -466,9 +518,15 @@ function selectTopGrades(grades, limit = 10) {
idx, idx,
rank: gradeRankOf(g.grade), rank: gradeRankOf(g.grade),
conf: numOr(g.confidence, -1), conf: numOr(g.confidence, -1),
edge: Math.abs(numOr(g.edge, -Infinity)), pWin: takeablePWin(g),
// SIGNED — no abs(). null (absent) sorts last, not first.
edge: strictNum(g.edge),
})); }));
scored.sort((a, b) => a.rank - b.rank || b.conf - a.conf || b.edge - a.edge || a.idx - b.idx); scored.sort((a, b) => a.rank - b.rank
|| b.conf - a.conf
|| descNullsLast(a.pWin, b.pWin)
|| descNullsLast(a.edge, b.edge)
|| a.idx - b.idx);
return scored.slice(0, Math.max(0, limit)).map((s) => s.g); return scored.slice(0, Math.max(0, limit)).map((s) => s.g);
} }