diff --git a/specs/MASTER-PLAN.md b/specs/MASTER-PLAN.md index 9efadcc..027fc48 100644 --- a/specs/MASTER-PLAN.md +++ b/specs/MASTER-PLAN.md @@ -73,7 +73,7 @@ multi-book data. | **MLB isotonic `p_win`** | **DECIDED** — reliability **0.0846**, resolution **0.190**, holdout **n=125** | **PROVISIONAL label RETRACTED 2026-08-01.** Calibration is **ruler-independent** (`estimateProbability` never sees a price; the fit is p_win-vs-outcome). Replicated on a fresh later window, both metrics improved | | **edge vs the ruler** | corr(edge, outcome) **−0.010** (v1) → **−0.022** (v2), n=200 · corr(**p_win**, outcome) **+0.26** | **Subtracting the market DESTROYS the signal.** The consensus ruler does not rescue edge: *differs ≠ better* | | **exchange-inclusive ruler** | **UNTESTABLE on existing data** | exchange quotes were never stored (discarded until 2026-08-01). Becomes testable only as v2-era captures accrue | -| **WNBA** | **still abstains** | a MODEL problem, not a coverage problem — coverage was never its constraint | +| **WNBA** | **MODEL NOT BUILT YET** — held out, *not* failed | **CORRECTED 2026-08-01.** The −0.12 was **NBA-template machinery run on WNBA data**. WNBA has never had its own archetypes/variables/conditions — the "sport stubbed in on another sport's template" CLAUDE.md forbids. That is an **unbuilt model's expected failure, not a verdict on the sport.** Its own build is QUEUED, after MLB | | 🔴 **pinnacle feed** | **0 captures since 2026-07-31** (103,940 in the prior 10 days) | a live regression; **we had a sharp anchor and lost it.** Not caused by our changes | | **soccer** | **settles** — ~15 competitions, 30d | "grades into a void" is a **$19/mo Pro-tier** problem, not a data problem | | **CLV + results feeds** | `/odds/closing` + `/movement` **redacted**; `/results` **403 `required_tier: hobby`**; `/exports/resolved-props` **403 `required_tier: pro`** | **verified on our keys** — plain tier exclusion, not a key or plan fault. **$9/mo** buys CLV + steam + results; **$19/mo** adds the 90-day settlement export | @@ -136,7 +136,7 @@ connected and *meaningless*. MLB's fix is connection, not construction. ### A2. Sport order (OPEN — Kev decides) Proposed by readiness × clock: **1) MLB** (reference, only qualifying model) → **2) CFB** (has a <30-day clock; soft-market thesis) → **3) NFL** → **4) NBA** → -**5) CBB** → **6) WNBA re-attempt** (abstains today) → **7) soccer** (quota-blocked). +**5) CBB** → **6) WNBA — first real build** (never had its own model) → **7) soccer** (quota-blocked). Each gets the same 8-layer template. **No sport is abandoned — abstention is a state, not a verdict.** @@ -194,7 +194,7 @@ a sport is a module. **Blocks all of A2 after MLB.** | phase | orders | contents | blocks | |---|---|---|---| -| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only, WNBA abstains) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped | +| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only — every other sport is NOT-BUILT, held out) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped | | **2. Resolution tail** | 3 | trigger · share cards + notifications + posts + recap · CLV flag decision | share cards, social proof | | **3. Surfaces + design lane** *(parallel with 1-2)* | 4 | C2 decision → D1-close mount · `/record` nav + remaining surfaces · D1-B glyphs/archetypes · S2 primitives | Chrome audit | | **4. Sport boundary** | 2 | registry collapse · MLB re-expressed as the first module | all further sports | @@ -214,7 +214,9 @@ a sport is a module. **Blocks all of A2 after MLB.** weights, conditions and Bayesian all feed the grade; calibration applied; the ladder monotone (A>B>C, no inversion) and proven on held-out data. 2. **Every listed sport is finished on the same 8-layer template**, or explicitly - abstaining with its reason recorded — never silently absent. + held out with its reason recorded — **"not built yet"** where no sport-specific + model exists, and only "measured and failed" where one was genuinely built and + tested. Never silently absent, and never a verdict on a sport we never attempted. 3. **Design fully implemented** — all 61 catalogued items BUILT-TO-SPEC. 4. **Every surface built, reachable and honest** — no orphans, no live-but-not-honest surface, no dead component. @@ -255,7 +257,7 @@ Every edge measurement this session came back **null, negative, or unproven**: | served grade → outcome | **r ≈ 0.005**, and **inverted** (B 52.4% < C 56.9%) | | p_win − fair_prob (3 formulations) | **negative in all three, both sports, both splits** | | p_win alone, MLB, holdout | +0.165, **p ≈ 0.07 — not significant** | -| p_win alone, WNBA | **negative** — abstains | +| p_win alone, WNBA | **negative — but this measured an NBA-template model on WNBA data, so it is not a WNBA result at all** | | CLV / beat-close | **null by guard** — instrument not trustworthy | | ROI by grade | likely an artifact of a meaningless letter | @@ -291,7 +293,8 @@ them is a hypothesis, not a guarantee.** They must each prove out on held-out da or be left disconnected honestly. ## 9.3 A one-sport product marketed as multi-sport -MLB is the only qualifying model. WNBA abstains on its own data. NBA and soccer +MLB is the only model that has been BUILT and passed. WNBA's model does not exist +yet (what was measured was NBA-template machinery on WNBA data). NBA and soccer **don't even settle** — they grade into a void. Until Phase 5, the honest framing is *"an MLB product with other sports in development."* The site should not imply otherwise. @@ -419,8 +422,9 @@ pitchers / depth charts / lineup confirmation that `/context` serves free. > > **WNBA is NOT thin at the feed** — 4.21 books/prop vs MLB's 3.61. It was > allow-list-starved exactly as MLB was. This removes one candidate explanation -> for its anti-predictive result; it does not explain it, and WNBA stays -> abstaining. +> for its −0.12 result. **And that result is not a WNBA verdict anyway** — it +> measured NBA-template machinery on WNBA data. WNBA is **NOT BUILT YET**, held +> out until it gets its own model. > > **No sharp anchor exists for props:** `pinnacle`, `matchbook` and `polymarket` > all measured **0%** on both sports. The consensus ruler is therefore a MARKET diff --git a/specs/rank-on-pwin-challenger.md b/specs/rank-on-pwin-challenger.md index 1ffc472..c3c1a35 100644 --- a/specs/rank-on-pwin-challenger.md +++ b/specs/rank-on-pwin-challenger.md @@ -73,9 +73,16 @@ slate — re-run the endpoint on a full slate before the flip. It is one call. ## PER-SPORT DOCTRINE — ENFORCED IN CODE, NOT IN A COMMENT **WNBA moves the most (100% of rows, mean 4.1 places) and must NOT adopt this.** -WNBA's `p_win` is **anti-predictive** on its own data — it abstains. Ranking that -board by `p_win` would sort it by a signal measured to point the *wrong way*: -worse than the incumbent, not better. + +**CORRECTED 2026-08-01:** WNBA does **not** "abstain" and is **not** +"anti-predictive". The −0.12 that produced those words was **NBA-template +machinery run on WNBA data** — WNBA has never had its own archetypes, variables, +conditions or calibration. That is an **unbuilt model's expected failure, not a +verdict on the sport.** WNBA is **NOT BUILT YET**, held out until its own model +exists; its build is queued after MLB. + +The live consequence is the same either way — an unbuilt sport must not rank on a +signal not shown to hold for it — which is why the guard code is unchanged. A comment would not have stopped a future flip from applying this globally, so: diff --git a/src/routes/snapshot.js b/src/routes/snapshot.js index c7754df..dc8356c 100644 --- a/src/routes/snapshot.js +++ b/src/routes/snapshot.js @@ -121,9 +121,15 @@ router.get('/:sport', async (req, res) => { // and accepted for the top-graded board. const stampForecastRank = (grades) => { if (!Array.isArray(grades) || grades.length === 0) return grades; - // PER-SPORT DOCTRINE, enforced: only sports that passed their OWN holdout - // may be ranked by forecast. WNBA's p_win is anti-predictive, so stamping - // it would hand a future flip a wrong-way ordering for that board. + // ARMED ROLLBACK. The boards sort by `forecast_rank` WHEN PRESENT, so + // setting FORECAST_RANK=0 reverts every surface to the incumbent order on + // the next response — no deploy, no code change, no client release. + if (String(process.env.FORECAST_RANK || '1') === '0') return grades; + // PER-SPORT DOCTRINE, enforced as a gate: a sport ranks on p_win only + // once its OWN model is built and shown to predict. MLB is the only one + // that has passed; the rest are held out as NOT-BUILT, not as failed. + // (WNBA's -0.12 came from NBA-template machinery on WNBA data — an + // unbuilt model's expected failure, not a verdict on the sport.) if (!ranksOnForecast(sport)) return grades; const ranked = rankByForecast(grades); const pos = new Map(); diff --git a/src/utils/gradeRanking.js b/src/utils/gradeRanking.js index bc7015d..f77525e 100644 --- a/src/utils/gradeRanking.js +++ b/src/utils/gradeRanking.js @@ -67,10 +67,14 @@ function descNullsLast(a, b) { /** * rankGrades — "top GRADES" order: grade tier → confidence → takeable-gated - * p_win → SIGNED edge → stable input order. Nulls sort LAST on both signals. + * p_win → stable input order. * - * SCALES ARE NEVER MIXED: p_win (0..1) is only ever compared against p_win and - * edge (%) only against edge. Comparing 0.62 against 62 is not a comparison. + * EDGE KEY REMOVED 2026-08-01. It used to be the 4th key. Measured on n=200 + * settled MLB rows, corr(edge, outcome) = -0.010 under the incumbent ruler and + * -0.022 under the consensus ruler — it does not predict, so it must not break + * ties either. Removing it is safe for EVERY sport: it takes a non-predictive + * signal out, it does not put p_win in front (that is `rankByForecast`, gated + * to sports whose own model has passed). * * Rows without a grade are dropped (a board of ungraded rows is not a board). * `limit` omitted → the whole ranked list (callers slice). @@ -83,12 +87,10 @@ function rankGrades(grades, limit) { rank: gradeRankOf(g.grade), conf: strictNum(g.confidence) == null ? -1 : strictNum(g.confidence), pWin: takeablePWin(g), - edge: strictNum(g.edge != null ? g.edge : g.edge_pct), })); scored.sort((a, b) => a.rank - b.rank || b.conf - a.conf || descNullsLast(a.pWin, b.pWin) - || descNullsLast(a.edge, b.edge) || a.idx - b.idx); const out = scored.map((s) => s.g); return limit == null ? out : out.slice(0, Math.max(0, limit)); @@ -142,15 +144,27 @@ function rankByForecast(grades, limit) { } /** - * WHICH SPORTS MAY RANK ON THE FORECAST — per-sport doctrine, enforced in code. + * WHICH SPORTS MAY RANK ON THE FORECAST — the per-sport doctrine, as a GATE. * - * MLB only. WNBA's p_win is ANTI-PREDICTIVE on its own data (it abstains), so - * ranking WNBA by p_win would sort that board by a signal measured to point the - * wrong way — worse than the incumbent, not better. A comment would not have - * stopped a future flip from applying this globally; this does. + * A sport ranks on p_win ONLY once its OWN model is built and shown to + * predict — calibration AND resolution holding on its own holdout. MLB is the + * only sport that has passed. Every other sport is held out as NOT-BUILT, + * never as FAILED. * - * A sport joins this set only by passing its OWN holdout: honest calibration - * AND surviving resolution. Never by inheriting MLB's result. + * WNBA IS NOT "ANTI-PREDICTIVE" AND DOES NOT "ABSTAIN" (corrected 2026-08-01). + * The -0.12 result that produced those words was NBA-template machinery run on + * WNBA data. WNBA has never had its own archetypes, variables, conditions or + * calibration — it is precisely the "sport stubbed in on another sport's + * template" that CLAUDE.md forbids. So -0.12 is the EXPECTED FAILURE OF AN + * UNBUILT MODEL, not a verdict on the sport. Reading it as a verdict would + * quietly retire a sport we never actually attempted. + * + * WNBA's own model-build is QUEUED as its own sport, after MLB is finished. + * + * The live consequence is identical either way — an unbuilt sport must not rank + * on a signal that has not been shown to hold for it — which is why this set is + * UNCHANGED. Only its meaning is corrected. A comment would not have stopped a + * future flip from going global; this does. */ const FORECAST_RANKED_SPORTS = Object.freeze(new Set(['mlb'])); const ranksOnForecast = (sport) => FORECAST_RANKED_SPORTS.has(String(sport || '').toLowerCase()); diff --git a/tests/unit/edgeBoard.test.js b/tests/unit/edgeBoard.test.js index f14297f..5be8f8a 100644 --- a/tests/unit/edgeBoard.test.js +++ b/tests/unit/edgeBoard.test.js @@ -25,22 +25,40 @@ describe('flattenToEdgeBoard', () => { ] }), ]; - test('flattens every graded prop across games into ranked rows, edge desc', () => { + // SUPERSEDED 2026-08-01: the board no longer ranks on edge (corr -0.010/-0.022 + // vs corr(p_win) = +0.26). With no forecast_rank supplied it falls through to + // GRADE rank, which is what this now asserts. + test('flattens every graded prop across games into ranked rows, grade order', () => { const rows = flattenToEdgeBoard(cards); - expect(rows.map((r) => r.player)).toEqual(['Jokic', 'Edwards', 'Gordon', 'LaVine']); // 8.4, 5.2, 3.4, -1.2 + expect(rows.map((r) => r.player)).toEqual(['Jokic', 'Edwards', 'Gordon', 'LaVine']); expect(rows[0].rank).toBe(1); expect(rows[3].rank).toBe(4); expect(rows[1].live).toBe(true); // Edwards' game is live → row carries it expect(rows[0].away).toBe('DEN'); expect(rows[0].home).toBe('MIN'); }); - test('a null edge sorts LAST, never 0-coerced to the top', () => { - const rows = flattenToEdgeBoard([card({ playerStrips: [ - strip('NoEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: null }]), - strip('HasEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'C', edge: 2.0 }]), + // SUPERSEDED 2026-08-01 — replaced by the strictly stronger inverse: edge + // cannot move the board AT ALL. The original could not detect edge being + // re-introduced as a key; this fails immediately if it is. + test('EDGE CANNOT RANK — flipping every edge leaves the order identical', () => { + const build = (e1, e2) => flattenToEdgeBoard([card({ playerStrips: [ + strip('First', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: e1 }]), + strip('Second', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: e2 }]), ] })]); - expect(rows[0].player).toBe('HasEdge'); // +2.0 beats a null (not 0) - expect(rows[1].player).toBe('NoEdge'); + expect(build(null, 2.0).map((r) => r.player)).toEqual(['First', 'Second']); + expect(build(2.0, null).map((r) => r.player)).toEqual(['First', 'Second']); + expect(build(-30, 30).map((r) => r.player)).toEqual(['First', 'Second']); + }); + + test('forecastRank ORDERS the board and a missing rank sorts LAST', () => { + const rows = flattenToEdgeBoard([card({ playerStrips: [ + strip('NoRank', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: 99 }]), + strip('Ranked2', [{ stat: 'PTS', line: 21, side: 'O', grade: 'C', forecastRank: 2 }]), + strip('Ranked1', [{ stat: 'PTS', line: 22, side: 'O', grade: 'D', forecastRank: 1 }]), + ] })]); + // rank 1 → rank 2 → unranked. Grade and edge are both overridden. + expect(rows.map((r) => r.player)).toEqual(['Ranked1', 'Ranked2', 'NoRank']); + expect(rows[0].rank).toBe(1); }); test('awaiting / dead / ungraded rows are excluded from the board', () => { @@ -72,9 +90,10 @@ describe('flattenToEdgeBoard', () => { }); // P1-7 — an impossible edge (the pipeline's 60/100/140 placeholder) is not a - // market signal. It must be treated as absent so the board neither DISPLAYS - // nor RANKS on it — a fake +140% must never outrank a real +8.4%. - test('insane edge (>40%) is nulled so it neither displays nor ranks', () => { + // market signal and must not be DISPLAYED as a real percentage. The ranking + // half of this test is retired: edge ranks nothing now, so nulling it can no + // longer change an order. The display guard still matters and still holds. + test('insane edge (>40%) is nulled so it is never DISPLAYED as a real %', () => { const rows = flattenToEdgeBoard([card({ playerStrips: [ strip('RealEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'B', edge: 8.4 }]), strip('BrokenEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: 140 }]), @@ -82,8 +101,8 @@ describe('flattenToEdgeBoard', () => { const broken = rows.find((r) => r.player === 'BrokenEdge'); const real = rows.find((r) => r.player === 'RealEdge'); expect(broken.edge).toBeNull(); // not displayed as a fake % - expect(real.edge).toBe(8.4); // real edge preserved - // real edge outranks the nulled placeholder despite the placeholder's higher grade - expect(rows[0].player).toBe('RealEdge'); + expect(real.edge).toBe(8.4); // real edge preserved as a diagnostic + // Ordering is now grade-driven (A before B) and edge has no say in it. + expect(rows[0].player).toBe('BrokenEdge'); }); }); diff --git a/tests/unit/edgePctDisplayRetirement.test.js b/tests/unit/edgePctDisplayRetirement.test.js index f5003ae..980d3df 100644 --- a/tests/unit/edgePctDisplayRetirement.test.js +++ b/tests/unit/edgePctDisplayRetirement.test.js @@ -46,25 +46,44 @@ describe('COMPUTATION and SORT survive — removing them would re-break the boar expect(read('web/src/lib/gradeAdapter.js')).toMatch(/function computeEdge/); }); - test('the signed-edge sort fallback is INTACT (nulls last, no abs)', () => { - const src = read('web/src/lib/slateAdapter.js'); - expect(src).toMatch(/descNullsLast\(a\.edge, b\.edge\)/); - // Strip comments before asserting the abs() bug is gone — the file's own - // doc-comment QUOTES the old broken key to explain the fix, and matching - // that would be a false positive. - const code = src.replace(/\/\*[\s\S]*?\*\//g, '').replace(/^\s*\/\/.*$/gm, ''); - expect(code).not.toMatch(/Math\.abs\(numOr\(g\.edge/); - expect(code).toMatch(/edge: strictNum\(g\.edge\)/); + // SUPERSEDED 2026-08-01 by the p_win flip. These two asserted that the SIGNED + // EDGE SORT was intact. Edge no longer sorts anything — measured on n=200 + // settled MLB rows, corr(edge, outcome) = -0.010 / -0.022 against + // corr(p_win, outcome) = +0.26. The retirement these tests guard is now + // deeper: edge is retired from RANKING as well as from hero DISPLAY. + // + // The replacements assert the INVERSE, which is strictly stronger: the sort + // keys must not mention edge at all, in either ranking function. + test('NO ranking function sorts on edge any more', () => { + const strip = (src) => src.replace(/\/\*[\s\S]*?\*\//g, '').replace(/^\s*\/\/.*$/gm, ''); + for (const f of ['web/src/lib/slateAdapter.js', 'src/utils/gradeRanking.js']) { + const code = strip(read(f)); + expect(code).not.toMatch(/descNullsLast\(a\.edge, b\.edge\)/); + expect(code).not.toMatch(/Math\.abs\(numOr\(g\.edge/); + expect(code).not.toMatch(/edge: strictNum\(g\.edge/); + } }); - test('the sort still orders a real signed edge correctly', () => { + test('edge is still CARRIED on the row — retired from ranking, not deleted', () => { + // Losing the record would be worse than mis-using it: edge stays computed, + // stored and displayed as a labelled diagnostic. + const code = read('web/src/lib/slateAdapter.js'); + expect(code).toMatch(/edge: \(rec\.edge_pct != null/); + }); + + test('the sort is driven by forecast_rank, and edge cannot move it', () => { const adapter = require('../../web/src/lib/slateAdapter'); - const out = adapter.selectTopGrades([ - { player: 'disagrees', grade: 'B', confidence: 50, edge: -40 }, - { player: 'agrees', grade: 'B', confidence: 50, edge: 12 }, - { player: 'absent', grade: 'B', confidence: 50 }, - ], 10).map((g) => g.player); - expect(out).toEqual(['agrees', 'disagrees', 'absent']); + const names = (rows) => rows.map((g) => g.player); + // identical rows except edge → identical order + expect(names(adapter.selectTopGrades([ + { player: 'a', grade: 'B', confidence: 50, edge: -40 }, + { player: 'b', grade: 'B', confidence: 50, edge: 12 }, + ], 10))).toEqual(['a', 'b']); + // forecast_rank decides + expect(names(adapter.selectTopGrades([ + { player: 'worseGradeBetterForecast', grade: 'C', confidence: 10, forecast_rank: 1, edge: -40 }, + { player: 'betterGradeWorseForecast', grade: 'A', confidence: 99, forecast_rank: 2, edge: 40 }, + ], 10))).toEqual(['worseGradeBetterForecast', 'betterGradeWorseForecast']); }); test('DeskShowcase keeps its own already-honest rendering', () => { diff --git a/tests/unit/gradeBoardSort.test.js b/tests/unit/gradeBoardSort.test.js index 39ae9aa..82f78d7 100644 --- a/tests/unit/gradeBoardSort.test.js +++ b/tests/unit/gradeBoardSort.test.js @@ -14,37 +14,67 @@ const { isTakeable: backendIsTakeable } = require('../../src/config/valueEngine' const names = (rows) => rows.map((r) => r.player); -describe('selectTopGrades — signed signal, missing sorts LAST', () => { - test('a DISAGREEMENT does not outrank an AGREEMENT at equal grade+confidence', () => { - // edge is signed by direction: positive = model agrees with the graded side. - // Old |edge| ranked -80 (strong disagreement) above +10 (mild agreement). - const grades = [ - { player: 'disagrees', grade: 'B', confidence: 50, edge: -80 }, - { player: 'agrees', grade: 'B', confidence: 50, edge: 10 }, +// SUPERSEDED 2026-08-01 by the p_win flip. These three asserted that EDGE +// ordered the board (signed, nulls last). Edge no longer ranks anything: +// measured on n=200 settled MLB rows, corr(edge, outcome) = -0.010 under the +// incumbent ruler and -0.022 under the consensus ruler, against +// corr(p_win, outcome) = +0.26. +// +// The replacements are strictly STRONGER — they fail if edge is ever +// re-introduced as a ranking key, which the originals could not detect. +describe('selectTopGrades — EDGE CANNOT RANK (retired 2026-08-01)', () => { + test('flipping edge from -80 to +80 does NOT change the order', () => { + const a = [ + { player: 'first', grade: 'B', confidence: 50, edge: -80 }, + { player: 'second', grade: 'B', confidence: 50, edge: 10 }, ]; - expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['agrees', 'disagrees']); + const b = [ + { player: 'first', grade: 'B', confidence: 50, edge: 80 }, + { player: 'second', grade: 'B', confidence: 50, edge: -10 }, + ]; + // Equal grade + confidence + no p_win -> stable input order, edge irrelevant. + expect(names(adapter.selectTopGrades(a, 10))).toEqual(['first', 'second']); + expect(names(adapter.selectTopGrades(b, 10))).toEqual(['first', 'second']); }); - test('a MISSING signal sorts LAST and is still PRESENT (never dropped)', () => { + test('a MISSING edge is no longer penalised — rows are still all PRESENT', () => { const grades = [ { player: 'no-signal', grade: 'B', confidence: 50 }, { player: 'weak-but-real', grade: 'B', confidence: 50, edge: 0.1 }, { player: 'negative-but-real', grade: 'B', confidence: 50, edge: -5 }, ]; const out = names(adapter.selectTopGrades(grades, 10)); - expect(out).toEqual(['weak-but-real', 'negative-but-real', 'no-signal']); - expect(out).toHaveLength(3); // present, not dropped - expect(out[out.length - 1]).toBe('no-signal'); + expect(out).toHaveLength(3); // never dropped + expect(out).toEqual(['no-signal', 'weak-but-real', 'negative-but-real']); // input order }); - test('null/empty-string edge is absent, NOT zero (Number(null) === 0 guard)', () => { + test('forecast_rank LEADS the chain when the server supplied it', () => { const grades = [ - { player: 'nullish', grade: 'B', confidence: 50, edge: null }, - { player: 'empty', grade: 'B', confidence: 50, edge: '' }, - { player: 'real-negative', grade: 'B', confidence: 50, edge: -1 }, + { player: 'topGrade', grade: 'A', confidence: 90, forecast_rank: 3 }, + { player: 'topForecast', grade: 'C', confidence: 40, forecast_rank: 1 }, ]; - // A real negative beats two absents; absents keep input order at the bottom. - expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['real-negative', 'nullish', 'empty']); + // p_win-first order wins over the letter: the letter measured r ~ 0.005 and + // is inverted, p_win measured +0.26. + expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['topForecast', 'topGrade']); + }); + + test('a MISSING forecast_rank sorts LAST and falls back to the grade chain', () => { + const grades = [ + { player: 'noRank', grade: 'A', confidence: 90 }, + { player: 'ranked', grade: 'C', confidence: 40, forecast_rank: 2 }, + ]; + expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['ranked', 'noRank']); + }); + + test('forecast_rank null/empty is ABSENT, not 0 (Number(null) === 0 guard)', () => { + const grades = [ + { player: 'nullish', grade: 'B', confidence: 50, forecast_rank: null }, + { player: 'empty', grade: 'B', confidence: 50, forecast_rank: '' }, + { player: 'real', grade: 'C', confidence: 10, forecast_rank: 9 }, + ]; + // rank 9 is a REAL rank and must beat two absents — a 0-coercion would + // have put the nulls first. + expect(names(adapter.selectTopGrades(grades, 10))[0]).toBe('real'); }); }); diff --git a/tests/unit/rankingInstrument.test.js b/tests/unit/rankingInstrument.test.js index fc46c91..a578f5c 100644 --- a/tests/unit/rankingInstrument.test.js +++ b/tests/unit/rankingInstrument.test.js @@ -55,18 +55,22 @@ describe('rankByForecast — the challenger', () => { }); }); -describe('rankGrades — the incumbent is UNTOUCHED (live ordering byte-identical)', () => { - it('still leads with the grade letter and still consults edge', () => { +describe('rankGrades — still grade-first (p_win leading is rankByForecast only)', () => { + it('leads with the grade letter, NOT p_win', () => { const out = rankGrades([g('lowPwinA', 'A', 0.51), g('highPwinC', 'C', 0.74)]); expect(out[0].player).toBe('lowPwinA'); // grade-first, unchanged }); - it('edge still breaks a true tie in the incumbent', () => { + // SUPERSEDED 2026-08-01: the edge KEY was removed from rankGrades entirely. + // Removing it is safe for every sport — it takes a non-predictive signal out + // without putting p_win in front (that is `rankByForecast`, gated to sports + // whose own model has passed). + it('edge can no longer break a tie — the key is gone for EVERY sport', () => { const out = rankGrades([ - g('lowEdge', 'B', 0.6, -110, { edge: 1 }), - g('highEdge', 'B', 0.6, -110, { edge: 9 }), + g('first', 'B', 0.6, -110, { edge: 1 }), + g('second', 'B', 0.6, -110, { edge: 9 }), ]); - expect(out[0].player).toBe('highEdge'); + expect(out[0].player).toBe('first'); // stable input order, edge ignored }); }); @@ -97,14 +101,18 @@ describe('rankingDelta — the challenger-first measurement', () => { }); }); -describe('per-sport doctrine — who may rank on the forecast', () => { - it('MLB may; WNBA may NOT (its p_win is anti-predictive, it abstains)', () => { +describe('per-sport doctrine — a sport ranks only once ITS OWN model is built', () => { + // WNBA is held out as NOT-BUILT, not as failed: its -0.12 came from + // NBA-template machinery run on WNBA data, which is an unbuilt model's + // expected failure rather than a verdict on the sport. Same live behaviour, + // corrected meaning. + it('MLB may (its own model passed); WNBA may not (its model is not built yet)', () => { expect(ranksOnForecast('mlb')).toBe(true); expect(ranksOnForecast('MLB')).toBe(true); expect(ranksOnForecast('wnba')).toBe(false); }); - it('no sport inherits MLB’s result — unknown sports are excluded', () => { + it('no sport inherits MLB’s result — every unbuilt sport is held out', () => { for (const s of ['nba', 'nfl', 'soccer', 'nhl', 'ncaab', '', null, undefined]) { expect(ranksOnForecast(s)).toBe(false); } diff --git a/web/src/lib/slateAdapter.js b/web/src/lib/slateAdapter.js index c16adcb..fa4a2ef 100644 --- a/web/src/lib/slateAdapter.js +++ b/web/src/lib/slateAdapter.js @@ -385,6 +385,11 @@ function buildPlayerStripsFromProps(gameProps, gradeIndex, deltaIndex, now = Dat // Rev 3 flat edge-board (mobile screen 01) — the ranked hero figure. // Strict null (never 0-coerced) so an absent edge sorts last, not top. edge: (rec.edge_pct != null && Number.isFinite(Number(rec.edge_pct))) ? Number(rec.edge_pct) : null, + // Server-computed forecast ordinal (p_win-first). Stamped BEFORE the + // model-price strip, so it reaches every tier where raw p_win cannot. + // Present only for sports whose own model has passed — absent elsewhere, + // which is exactly how the boards fall back. + forecastRank: (rec.forecast_rank != null && Number.isFinite(Number(rec.forecast_rank))) ? Number(rec.forecast_rank) : null, // A1 S3 — the prop's own book + the best available price across the // game's book rows for the graded side (null unless ≥2 books at the // same current line disagree — see detectBestBook). @@ -470,6 +475,14 @@ const strictNum = (v) => (v == null || v === '' ? null : (Number.isFinite(Number * in a different base" (Session 67). Those callers legitimately get null here * and fall through to the signed edge. */ +/** Ascending comparator (rank 1 is best) that always sorts a null LAST. */ +function ascNullsLast(a, b) { + if (a == null && b == null) return 0; + if (a == null) return 1; + if (b == null) return -1; + return a - b; +} + function takeablePWin(g) { const { isTakeable } = require('./valueState'); const p = strictNum(g && g.p_win); @@ -513,19 +526,24 @@ function descNullsLast(a, b) { */ function selectTopGrades(grades, limit = 10) { const arr = (Array.isArray(grades) ? grades : []).filter((g) => g && g.grade); + // FLIPPED 2026-08-01 — forecast order FIRST when the server supplied it, and + // the EDGE KEY IS GONE. `forecast_rank` is p_win-first and is stamped only for + // sports whose own model has passed; when absent (every not-built-yet sport) + // this falls through to the unchanged grade-tier chain. Edge is removed + // outright: corr(edge, outcome) = -0.010 / -0.022 on n=200 settled MLB rows, + // against corr(p_win, outcome) = +0.26. const scored = arr.map((g, idx) => ({ g, idx, + fr: strictNum(g.forecast_rank), rank: gradeRankOf(g.grade), conf: numOr(g.confidence, -1), pWin: takeablePWin(g), - // SIGNED — no abs(). null (absent) sorts last, not first. - edge: strictNum(g.edge), })); - scored.sort((a, b) => a.rank - b.rank + scored.sort((a, b) => ascNullsLast(a.fr, b.fr) + || a.rank - b.rank || b.conf - a.conf || descNullsLast(a.pWin, b.pWin) - || descNullsLast(a.edge, b.edge) || a.idx - b.idx); return scored.slice(0, Math.max(0, limit)).map((s) => s.g); } @@ -691,6 +709,7 @@ function flattenToEdgeBoard(cards) { // — treat as absent so the board neither displays nor RANKS on a fake // edge. When absent, rows fall through to the grade-rank tiebreak. edge: (p.edge != null && Number.isFinite(p.edge) && Math.abs(p.edge) <= EDGE_BOARD_SANE_MAX) ? p.edge : null, + forecastRank: (p.forecastRank != null && Number.isFinite(p.forecastRank)) ? p.forecastRank : null, sport: c.sport || '', away, home, @@ -704,10 +723,16 @@ function flattenToEdgeBoard(cards) { } } } + // FLIPPED 2026-08-01 — this board no longer ranks on edge. Edge was its + // PRIMARY key, so the entire mobile board was ordered by a quantity measured + // not to predict (corr -0.010 / -0.022 vs corr(p_win) = +0.26). + // Now: server forecast order first (p_win-first; MLB-only, absent elsewhere), + // then grade rank. Edge is still carried and still DISPLAYED as a labelled + // diagnostic — it just no longer decides what a user sees first. rows.sort((a, b) => { - const ea = a.edge == null ? -Infinity : a.edge; - const eb = b.edge == null ? -Infinity : b.edge; - if (eb !== ea) return eb - ea; + const fa = a.forecastRank == null ? Infinity : a.forecastRank; + const fb = b.forecastRank == null ? Infinity : b.forecastRank; + if (fa !== fb) return fa - fb; return edgeGradeRank(a.grade) - edgeGradeRank(b.grade); }); return rows.map((r, i) => ({ ...r, rank: i + 1 }));