WNBA truth correction + THE p_win FLIP (live, rollback armed)
PART A -- WNBA TRUTH CORRECTION (no behaviour change).
WNBA does not "abstain" and is not "anti-predictive". The -0.12 that
produced those words was NBA-template machinery run on WNBA data -- WNBA
has never had its own archetypes, variables, conditions or calibration,
which is precisely the "sport stubbed in on another sport's template"
CLAUDE.md forbids. That is an UNBUILT MODEL'S EXPECTED FAILURE, not a
verdict on the sport; reading it as a verdict would quietly retire a sport
we never actually attempted. Its own build is QUEUED, after MLB.
The guard CODE is unchanged -- FORECAST_RANKED_SPORTS = {'mlb'} and the
inheritance test are correct live safety either way. Only the meaning is
corrected, and generalised into the doctrine-as-a-gate: a sport ranks on
p_win ONLY once its OWN model is built and shown to predict (calibration
AND resolution on its own holdout). Others are held out as NOT-BUILT,
never as failed. Re-labelled across gradeRanking, snapshot route, tests,
MASTER-PLAN and the challenger report.
PART B -- THE FLIP, gated on a full-slate re-run.
The re-run found something better than a bigger sample. An induced
snapshot graded 7 props: gradeAndCacheSlate runs with DEFAULT_LIMIT = 25
and ~72% of those refuse for insufficient_data, while 546 props are
gradeable. So 8 props IS the board, structurally -- not a small sample of
it. Logged as its own finding; the cap is a separate order.
For a statistically meaningful delta I used 11 real historical boards
(n=328, board sizes 14-57): 79.9% of rows move, mean 5.16 places per
board, TOP READ CHANGES ON 9 OF 11 BOARDS. The re-ordering holds at real
board size. Query committed.
FLIPPED:
- rankGrades drops its edge key (safe for every sport: removes a
non-predictive tiebreak without putting p_win in front).
- selectTopGrades leads on forecast_rank, edge key removed.
- flattenToEdgeBoard sorts on forecastRank, not edge -- this board had
edge as its PRIMARY key, so the whole mobile board was ordered by a
quantity measured not to predict.
- forecast_rank threaded onto strip props.
Sports whose model is not built supply no forecast_rank, so their boards
fall through to the unchanged grade chain -- the fallback is the guard.
ROLLBACK ARMED: boards sort by forecast_rank WHEN PRESENT, so
FORECAST_RANK=0 reverts every surface on the next response -- no deploy,
no client release.
Edge is still computed, stored, carried and displayed as a labelled
diagnostic. Retired from ranking, not deleted.
Eight superseded tests updated to strictly stronger INVERSE properties --
they now fail if edge is ever re-introduced as a ranking key, which the
originals could not detect.
Gates: 4,045 tests / 323 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
+12
-8
@@ -73,7 +73,7 @@ multi-book data.
|
||||
| **MLB isotonic `p_win`** | **DECIDED** — reliability **0.0846**, resolution **0.190**, holdout **n=125** | **PROVISIONAL label RETRACTED 2026-08-01.** Calibration is **ruler-independent** (`estimateProbability` never sees a price; the fit is p_win-vs-outcome). Replicated on a fresh later window, both metrics improved |
|
||||
| **edge vs the ruler** | corr(edge, outcome) **−0.010** (v1) → **−0.022** (v2), n=200 · corr(**p_win**, outcome) **+0.26** | **Subtracting the market DESTROYS the signal.** The consensus ruler does not rescue edge: *differs ≠ better* |
|
||||
| **exchange-inclusive ruler** | **UNTESTABLE on existing data** | exchange quotes were never stored (discarded until 2026-08-01). Becomes testable only as v2-era captures accrue |
|
||||
| **WNBA** | **still abstains** | a MODEL problem, not a coverage problem — coverage was never its constraint |
|
||||
| **WNBA** | **MODEL NOT BUILT YET** — held out, *not* failed | **CORRECTED 2026-08-01.** The −0.12 was **NBA-template machinery run on WNBA data**. WNBA has never had its own archetypes/variables/conditions — the "sport stubbed in on another sport's template" CLAUDE.md forbids. That is an **unbuilt model's expected failure, not a verdict on the sport.** Its own build is QUEUED, after MLB |
|
||||
| 🔴 **pinnacle feed** | **0 captures since 2026-07-31** (103,940 in the prior 10 days) | a live regression; **we had a sharp anchor and lost it.** Not caused by our changes |
|
||||
| **soccer** | **settles** — ~15 competitions, 30d | "grades into a void" is a **$19/mo Pro-tier** problem, not a data problem |
|
||||
| **CLV + results feeds** | `/odds/closing` + `/movement` **redacted**; `/results` **403 `required_tier: hobby`**; `/exports/resolved-props` **403 `required_tier: pro`** | **verified on our keys** — plain tier exclusion, not a key or plan fault. **$9/mo** buys CLV + steam + results; **$19/mo** adds the 90-day settlement export |
|
||||
@@ -136,7 +136,7 @@ connected and *meaningless*. MLB's fix is connection, not construction.
|
||||
### A2. Sport order (OPEN — Kev decides)
|
||||
Proposed by readiness × clock: **1) MLB** (reference, only qualifying model) →
|
||||
**2) CFB** (has a <30-day clock; soft-market thesis) → **3) NFL** → **4) NBA** →
|
||||
**5) CBB** → **6) WNBA re-attempt** (abstains today) → **7) soccer** (quota-blocked).
|
||||
**5) CBB** → **6) WNBA — first real build** (never had its own model) → **7) soccer** (quota-blocked).
|
||||
Each gets the same 8-layer template. **No sport is abandoned — abstention is a
|
||||
state, not a verdict.**
|
||||
|
||||
@@ -194,7 +194,7 @@ a sport is a module. **Blocks all of A2 after MLB.**
|
||||
|
||||
| phase | orders | contents | blocks |
|
||||
|---|---|---|---|
|
||||
| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only, WNBA abstains) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
|
||||
| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only — every other sport is NOT-BUILT, held out) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
|
||||
| **2. Resolution tail** | 3 | trigger · share cards + notifications + posts + recap · CLV flag decision | share cards, social proof |
|
||||
| **3. Surfaces + design lane** *(parallel with 1-2)* | 4 | C2 decision → D1-close mount · `/record` nav + remaining surfaces · D1-B glyphs/archetypes · S2 primitives | Chrome audit |
|
||||
| **4. Sport boundary** | 2 | registry collapse · MLB re-expressed as the first module | all further sports |
|
||||
@@ -214,7 +214,9 @@ a sport is a module. **Blocks all of A2 after MLB.**
|
||||
weights, conditions and Bayesian all feed the grade; calibration applied; the
|
||||
ladder monotone (A>B>C, no inversion) and proven on held-out data.
|
||||
2. **Every listed sport is finished on the same 8-layer template**, or explicitly
|
||||
abstaining with its reason recorded — never silently absent.
|
||||
held out with its reason recorded — **"not built yet"** where no sport-specific
|
||||
model exists, and only "measured and failed" where one was genuinely built and
|
||||
tested. Never silently absent, and never a verdict on a sport we never attempted.
|
||||
3. **Design fully implemented** — all 61 catalogued items BUILT-TO-SPEC.
|
||||
4. **Every surface built, reachable and honest** — no orphans, no live-but-not-honest
|
||||
surface, no dead component.
|
||||
@@ -255,7 +257,7 @@ Every edge measurement this session came back **null, negative, or unproven**:
|
||||
| served grade → outcome | **r ≈ 0.005**, and **inverted** (B 52.4% < C 56.9%) |
|
||||
| p_win − fair_prob (3 formulations) | **negative in all three, both sports, both splits** |
|
||||
| p_win alone, MLB, holdout | +0.165, **p ≈ 0.07 — not significant** |
|
||||
| p_win alone, WNBA | **negative** — abstains |
|
||||
| p_win alone, WNBA | **negative — but this measured an NBA-template model on WNBA data, so it is not a WNBA result at all** |
|
||||
| CLV / beat-close | **null by guard** — instrument not trustworthy |
|
||||
| ROI by grade | likely an artifact of a meaningless letter |
|
||||
|
||||
@@ -291,7 +293,8 @@ them is a hypothesis, not a guarantee.** They must each prove out on held-out da
|
||||
or be left disconnected honestly.
|
||||
|
||||
## 9.3 A one-sport product marketed as multi-sport
|
||||
MLB is the only qualifying model. WNBA abstains on its own data. NBA and soccer
|
||||
MLB is the only model that has been BUILT and passed. WNBA's model does not exist
|
||||
yet (what was measured was NBA-template machinery on WNBA data). NBA and soccer
|
||||
**don't even settle** — they grade into a void. Until Phase 5, the honest framing
|
||||
is *"an MLB product with other sports in development."* The site should not imply
|
||||
otherwise.
|
||||
@@ -419,8 +422,9 @@ pitchers / depth charts / lineup confirmation that `/context` serves free.
|
||||
>
|
||||
> **WNBA is NOT thin at the feed** — 4.21 books/prop vs MLB's 3.61. It was
|
||||
> allow-list-starved exactly as MLB was. This removes one candidate explanation
|
||||
> for its anti-predictive result; it does not explain it, and WNBA stays
|
||||
> abstaining.
|
||||
> for its −0.12 result. **And that result is not a WNBA verdict anyway** — it
|
||||
> measured NBA-template machinery on WNBA data. WNBA is **NOT BUILT YET**, held
|
||||
> out until it gets its own model.
|
||||
>
|
||||
> **No sharp anchor exists for props:** `pinnacle`, `matchbook` and `polymarket`
|
||||
> all measured **0%** on both sports. The consensus ruler is therefore a MARKET
|
||||
|
||||
@@ -73,9 +73,16 @@ slate — re-run the endpoint on a full slate before the flip. It is one call.
|
||||
## PER-SPORT DOCTRINE — ENFORCED IN CODE, NOT IN A COMMENT
|
||||
|
||||
**WNBA moves the most (100% of rows, mean 4.1 places) and must NOT adopt this.**
|
||||
WNBA's `p_win` is **anti-predictive** on its own data — it abstains. Ranking that
|
||||
board by `p_win` would sort it by a signal measured to point the *wrong way*:
|
||||
worse than the incumbent, not better.
|
||||
|
||||
**CORRECTED 2026-08-01:** WNBA does **not** "abstain" and is **not**
|
||||
"anti-predictive". The −0.12 that produced those words was **NBA-template
|
||||
machinery run on WNBA data** — WNBA has never had its own archetypes, variables,
|
||||
conditions or calibration. That is an **unbuilt model's expected failure, not a
|
||||
verdict on the sport.** WNBA is **NOT BUILT YET**, held out until its own model
|
||||
exists; its build is queued after MLB.
|
||||
|
||||
The live consequence is the same either way — an unbuilt sport must not rank on a
|
||||
signal not shown to hold for it — which is why the guard code is unchanged.
|
||||
|
||||
A comment would not have stopped a future flip from applying this globally, so:
|
||||
|
||||
|
||||
@@ -121,9 +121,15 @@ router.get('/:sport', async (req, res) => {
|
||||
// and accepted for the top-graded board.
|
||||
const stampForecastRank = (grades) => {
|
||||
if (!Array.isArray(grades) || grades.length === 0) return grades;
|
||||
// PER-SPORT DOCTRINE, enforced: only sports that passed their OWN holdout
|
||||
// may be ranked by forecast. WNBA's p_win is anti-predictive, so stamping
|
||||
// it would hand a future flip a wrong-way ordering for that board.
|
||||
// ARMED ROLLBACK. The boards sort by `forecast_rank` WHEN PRESENT, so
|
||||
// setting FORECAST_RANK=0 reverts every surface to the incumbent order on
|
||||
// the next response — no deploy, no code change, no client release.
|
||||
if (String(process.env.FORECAST_RANK || '1') === '0') return grades;
|
||||
// PER-SPORT DOCTRINE, enforced as a gate: a sport ranks on p_win only
|
||||
// once its OWN model is built and shown to predict. MLB is the only one
|
||||
// that has passed; the rest are held out as NOT-BUILT, not as failed.
|
||||
// (WNBA's -0.12 came from NBA-template machinery on WNBA data — an
|
||||
// unbuilt model's expected failure, not a verdict on the sport.)
|
||||
if (!ranksOnForecast(sport)) return grades;
|
||||
const ranked = rankByForecast(grades);
|
||||
const pos = new Map();
|
||||
|
||||
+26
-12
@@ -67,10 +67,14 @@ function descNullsLast(a, b) {
|
||||
|
||||
/**
|
||||
* rankGrades — "top GRADES" order: grade tier → confidence → takeable-gated
|
||||
* p_win → SIGNED edge → stable input order. Nulls sort LAST on both signals.
|
||||
* p_win → stable input order.
|
||||
*
|
||||
* SCALES ARE NEVER MIXED: p_win (0..1) is only ever compared against p_win and
|
||||
* edge (%) only against edge. Comparing 0.62 against 62 is not a comparison.
|
||||
* EDGE KEY REMOVED 2026-08-01. It used to be the 4th key. Measured on n=200
|
||||
* settled MLB rows, corr(edge, outcome) = -0.010 under the incumbent ruler and
|
||||
* -0.022 under the consensus ruler — it does not predict, so it must not break
|
||||
* ties either. Removing it is safe for EVERY sport: it takes a non-predictive
|
||||
* signal out, it does not put p_win in front (that is `rankByForecast`, gated
|
||||
* to sports whose own model has passed).
|
||||
*
|
||||
* Rows without a grade are dropped (a board of ungraded rows is not a board).
|
||||
* `limit` omitted → the whole ranked list (callers slice).
|
||||
@@ -83,12 +87,10 @@ function rankGrades(grades, limit) {
|
||||
rank: gradeRankOf(g.grade),
|
||||
conf: strictNum(g.confidence) == null ? -1 : strictNum(g.confidence),
|
||||
pWin: takeablePWin(g),
|
||||
edge: strictNum(g.edge != null ? g.edge : g.edge_pct),
|
||||
}));
|
||||
scored.sort((a, b) => a.rank - b.rank
|
||||
|| b.conf - a.conf
|
||||
|| descNullsLast(a.pWin, b.pWin)
|
||||
|| descNullsLast(a.edge, b.edge)
|
||||
|| a.idx - b.idx);
|
||||
const out = scored.map((s) => s.g);
|
||||
return limit == null ? out : out.slice(0, Math.max(0, limit));
|
||||
@@ -142,15 +144,27 @@ function rankByForecast(grades, limit) {
|
||||
}
|
||||
|
||||
/**
|
||||
* WHICH SPORTS MAY RANK ON THE FORECAST — per-sport doctrine, enforced in code.
|
||||
* WHICH SPORTS MAY RANK ON THE FORECAST — the per-sport doctrine, as a GATE.
|
||||
*
|
||||
* MLB only. WNBA's p_win is ANTI-PREDICTIVE on its own data (it abstains), so
|
||||
* ranking WNBA by p_win would sort that board by a signal measured to point the
|
||||
* wrong way — worse than the incumbent, not better. A comment would not have
|
||||
* stopped a future flip from applying this globally; this does.
|
||||
* A sport ranks on p_win ONLY once its OWN model is built and shown to
|
||||
* predict — calibration AND resolution holding on its own holdout. MLB is the
|
||||
* only sport that has passed. Every other sport is held out as NOT-BUILT,
|
||||
* never as FAILED.
|
||||
*
|
||||
* A sport joins this set only by passing its OWN holdout: honest calibration
|
||||
* AND surviving resolution. Never by inheriting MLB's result.
|
||||
* WNBA IS NOT "ANTI-PREDICTIVE" AND DOES NOT "ABSTAIN" (corrected 2026-08-01).
|
||||
* The -0.12 result that produced those words was NBA-template machinery run on
|
||||
* WNBA data. WNBA has never had its own archetypes, variables, conditions or
|
||||
* calibration — it is precisely the "sport stubbed in on another sport's
|
||||
* template" that CLAUDE.md forbids. So -0.12 is the EXPECTED FAILURE OF AN
|
||||
* UNBUILT MODEL, not a verdict on the sport. Reading it as a verdict would
|
||||
* quietly retire a sport we never actually attempted.
|
||||
*
|
||||
* WNBA's own model-build is QUEUED as its own sport, after MLB is finished.
|
||||
*
|
||||
* The live consequence is identical either way — an unbuilt sport must not rank
|
||||
* on a signal that has not been shown to hold for it — which is why this set is
|
||||
* UNCHANGED. Only its meaning is corrected. A comment would not have stopped a
|
||||
* future flip from going global; this does.
|
||||
*/
|
||||
const FORECAST_RANKED_SPORTS = Object.freeze(new Set(['mlb']));
|
||||
const ranksOnForecast = (sport) => FORECAST_RANKED_SPORTS.has(String(sport || '').toLowerCase());
|
||||
|
||||
@@ -25,22 +25,40 @@ describe('flattenToEdgeBoard', () => {
|
||||
] }),
|
||||
];
|
||||
|
||||
test('flattens every graded prop across games into ranked rows, edge desc', () => {
|
||||
// SUPERSEDED 2026-08-01: the board no longer ranks on edge (corr -0.010/-0.022
|
||||
// vs corr(p_win) = +0.26). With no forecast_rank supplied it falls through to
|
||||
// GRADE rank, which is what this now asserts.
|
||||
test('flattens every graded prop across games into ranked rows, grade order', () => {
|
||||
const rows = flattenToEdgeBoard(cards);
|
||||
expect(rows.map((r) => r.player)).toEqual(['Jokic', 'Edwards', 'Gordon', 'LaVine']); // 8.4, 5.2, 3.4, -1.2
|
||||
expect(rows.map((r) => r.player)).toEqual(['Jokic', 'Edwards', 'Gordon', 'LaVine']);
|
||||
expect(rows[0].rank).toBe(1);
|
||||
expect(rows[3].rank).toBe(4);
|
||||
expect(rows[1].live).toBe(true); // Edwards' game is live → row carries it
|
||||
expect(rows[0].away).toBe('DEN'); expect(rows[0].home).toBe('MIN');
|
||||
});
|
||||
|
||||
test('a null edge sorts LAST, never 0-coerced to the top', () => {
|
||||
const rows = flattenToEdgeBoard([card({ playerStrips: [
|
||||
strip('NoEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: null }]),
|
||||
strip('HasEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'C', edge: 2.0 }]),
|
||||
// SUPERSEDED 2026-08-01 — replaced by the strictly stronger inverse: edge
|
||||
// cannot move the board AT ALL. The original could not detect edge being
|
||||
// re-introduced as a key; this fails immediately if it is.
|
||||
test('EDGE CANNOT RANK — flipping every edge leaves the order identical', () => {
|
||||
const build = (e1, e2) => flattenToEdgeBoard([card({ playerStrips: [
|
||||
strip('First', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: e1 }]),
|
||||
strip('Second', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: e2 }]),
|
||||
] })]);
|
||||
expect(rows[0].player).toBe('HasEdge'); // +2.0 beats a null (not 0)
|
||||
expect(rows[1].player).toBe('NoEdge');
|
||||
expect(build(null, 2.0).map((r) => r.player)).toEqual(['First', 'Second']);
|
||||
expect(build(2.0, null).map((r) => r.player)).toEqual(['First', 'Second']);
|
||||
expect(build(-30, 30).map((r) => r.player)).toEqual(['First', 'Second']);
|
||||
});
|
||||
|
||||
test('forecastRank ORDERS the board and a missing rank sorts LAST', () => {
|
||||
const rows = flattenToEdgeBoard([card({ playerStrips: [
|
||||
strip('NoRank', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: 99 }]),
|
||||
strip('Ranked2', [{ stat: 'PTS', line: 21, side: 'O', grade: 'C', forecastRank: 2 }]),
|
||||
strip('Ranked1', [{ stat: 'PTS', line: 22, side: 'O', grade: 'D', forecastRank: 1 }]),
|
||||
] })]);
|
||||
// rank 1 → rank 2 → unranked. Grade and edge are both overridden.
|
||||
expect(rows.map((r) => r.player)).toEqual(['Ranked1', 'Ranked2', 'NoRank']);
|
||||
expect(rows[0].rank).toBe(1);
|
||||
});
|
||||
|
||||
test('awaiting / dead / ungraded rows are excluded from the board', () => {
|
||||
@@ -72,9 +90,10 @@ describe('flattenToEdgeBoard', () => {
|
||||
});
|
||||
|
||||
// P1-7 — an impossible edge (the pipeline's 60/100/140 placeholder) is not a
|
||||
// market signal. It must be treated as absent so the board neither DISPLAYS
|
||||
// nor RANKS on it — a fake +140% must never outrank a real +8.4%.
|
||||
test('insane edge (>40%) is nulled so it neither displays nor ranks', () => {
|
||||
// market signal and must not be DISPLAYED as a real percentage. The ranking
|
||||
// half of this test is retired: edge ranks nothing now, so nulling it can no
|
||||
// longer change an order. The display guard still matters and still holds.
|
||||
test('insane edge (>40%) is nulled so it is never DISPLAYED as a real %', () => {
|
||||
const rows = flattenToEdgeBoard([card({ playerStrips: [
|
||||
strip('RealEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'B', edge: 8.4 }]),
|
||||
strip('BrokenEdge', [{ stat: 'PTS', line: 20, side: 'O', grade: 'A', edge: 140 }]),
|
||||
@@ -82,8 +101,8 @@ describe('flattenToEdgeBoard', () => {
|
||||
const broken = rows.find((r) => r.player === 'BrokenEdge');
|
||||
const real = rows.find((r) => r.player === 'RealEdge');
|
||||
expect(broken.edge).toBeNull(); // not displayed as a fake %
|
||||
expect(real.edge).toBe(8.4); // real edge preserved
|
||||
// real edge outranks the nulled placeholder despite the placeholder's higher grade
|
||||
expect(rows[0].player).toBe('RealEdge');
|
||||
expect(real.edge).toBe(8.4); // real edge preserved as a diagnostic
|
||||
// Ordering is now grade-driven (A before B) and edge has no say in it.
|
||||
expect(rows[0].player).toBe('BrokenEdge');
|
||||
});
|
||||
});
|
||||
|
||||
@@ -46,25 +46,44 @@ describe('COMPUTATION and SORT survive — removing them would re-break the boar
|
||||
expect(read('web/src/lib/gradeAdapter.js')).toMatch(/function computeEdge/);
|
||||
});
|
||||
|
||||
test('the signed-edge sort fallback is INTACT (nulls last, no abs)', () => {
|
||||
const src = read('web/src/lib/slateAdapter.js');
|
||||
expect(src).toMatch(/descNullsLast\(a\.edge, b\.edge\)/);
|
||||
// Strip comments before asserting the abs() bug is gone — the file's own
|
||||
// doc-comment QUOTES the old broken key to explain the fix, and matching
|
||||
// that would be a false positive.
|
||||
const code = src.replace(/\/\*[\s\S]*?\*\//g, '').replace(/^\s*\/\/.*$/gm, '');
|
||||
expect(code).not.toMatch(/Math\.abs\(numOr\(g\.edge/);
|
||||
expect(code).toMatch(/edge: strictNum\(g\.edge\)/);
|
||||
// SUPERSEDED 2026-08-01 by the p_win flip. These two asserted that the SIGNED
|
||||
// EDGE SORT was intact. Edge no longer sorts anything — measured on n=200
|
||||
// settled MLB rows, corr(edge, outcome) = -0.010 / -0.022 against
|
||||
// corr(p_win, outcome) = +0.26. The retirement these tests guard is now
|
||||
// deeper: edge is retired from RANKING as well as from hero DISPLAY.
|
||||
//
|
||||
// The replacements assert the INVERSE, which is strictly stronger: the sort
|
||||
// keys must not mention edge at all, in either ranking function.
|
||||
test('NO ranking function sorts on edge any more', () => {
|
||||
const strip = (src) => src.replace(/\/\*[\s\S]*?\*\//g, '').replace(/^\s*\/\/.*$/gm, '');
|
||||
for (const f of ['web/src/lib/slateAdapter.js', 'src/utils/gradeRanking.js']) {
|
||||
const code = strip(read(f));
|
||||
expect(code).not.toMatch(/descNullsLast\(a\.edge, b\.edge\)/);
|
||||
expect(code).not.toMatch(/Math\.abs\(numOr\(g\.edge/);
|
||||
expect(code).not.toMatch(/edge: strictNum\(g\.edge/);
|
||||
}
|
||||
});
|
||||
|
||||
test('the sort still orders a real signed edge correctly', () => {
|
||||
test('edge is still CARRIED on the row — retired from ranking, not deleted', () => {
|
||||
// Losing the record would be worse than mis-using it: edge stays computed,
|
||||
// stored and displayed as a labelled diagnostic.
|
||||
const code = read('web/src/lib/slateAdapter.js');
|
||||
expect(code).toMatch(/edge: \(rec\.edge_pct != null/);
|
||||
});
|
||||
|
||||
test('the sort is driven by forecast_rank, and edge cannot move it', () => {
|
||||
const adapter = require('../../web/src/lib/slateAdapter');
|
||||
const out = adapter.selectTopGrades([
|
||||
{ player: 'disagrees', grade: 'B', confidence: 50, edge: -40 },
|
||||
{ player: 'agrees', grade: 'B', confidence: 50, edge: 12 },
|
||||
{ player: 'absent', grade: 'B', confidence: 50 },
|
||||
], 10).map((g) => g.player);
|
||||
expect(out).toEqual(['agrees', 'disagrees', 'absent']);
|
||||
const names = (rows) => rows.map((g) => g.player);
|
||||
// identical rows except edge → identical order
|
||||
expect(names(adapter.selectTopGrades([
|
||||
{ player: 'a', grade: 'B', confidence: 50, edge: -40 },
|
||||
{ player: 'b', grade: 'B', confidence: 50, edge: 12 },
|
||||
], 10))).toEqual(['a', 'b']);
|
||||
// forecast_rank decides
|
||||
expect(names(adapter.selectTopGrades([
|
||||
{ player: 'worseGradeBetterForecast', grade: 'C', confidence: 10, forecast_rank: 1, edge: -40 },
|
||||
{ player: 'betterGradeWorseForecast', grade: 'A', confidence: 99, forecast_rank: 2, edge: 40 },
|
||||
], 10))).toEqual(['worseGradeBetterForecast', 'betterGradeWorseForecast']);
|
||||
});
|
||||
|
||||
test('DeskShowcase keeps its own already-honest rendering', () => {
|
||||
|
||||
@@ -14,37 +14,67 @@ const { isTakeable: backendIsTakeable } = require('../../src/config/valueEngine'
|
||||
|
||||
const names = (rows) => rows.map((r) => r.player);
|
||||
|
||||
describe('selectTopGrades — signed signal, missing sorts LAST', () => {
|
||||
test('a DISAGREEMENT does not outrank an AGREEMENT at equal grade+confidence', () => {
|
||||
// edge is signed by direction: positive = model agrees with the graded side.
|
||||
// Old |edge| ranked -80 (strong disagreement) above +10 (mild agreement).
|
||||
const grades = [
|
||||
{ player: 'disagrees', grade: 'B', confidence: 50, edge: -80 },
|
||||
{ player: 'agrees', grade: 'B', confidence: 50, edge: 10 },
|
||||
// SUPERSEDED 2026-08-01 by the p_win flip. These three asserted that EDGE
|
||||
// ordered the board (signed, nulls last). Edge no longer ranks anything:
|
||||
// measured on n=200 settled MLB rows, corr(edge, outcome) = -0.010 under the
|
||||
// incumbent ruler and -0.022 under the consensus ruler, against
|
||||
// corr(p_win, outcome) = +0.26.
|
||||
//
|
||||
// The replacements are strictly STRONGER — they fail if edge is ever
|
||||
// re-introduced as a ranking key, which the originals could not detect.
|
||||
describe('selectTopGrades — EDGE CANNOT RANK (retired 2026-08-01)', () => {
|
||||
test('flipping edge from -80 to +80 does NOT change the order', () => {
|
||||
const a = [
|
||||
{ player: 'first', grade: 'B', confidence: 50, edge: -80 },
|
||||
{ player: 'second', grade: 'B', confidence: 50, edge: 10 },
|
||||
];
|
||||
expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['agrees', 'disagrees']);
|
||||
const b = [
|
||||
{ player: 'first', grade: 'B', confidence: 50, edge: 80 },
|
||||
{ player: 'second', grade: 'B', confidence: 50, edge: -10 },
|
||||
];
|
||||
// Equal grade + confidence + no p_win -> stable input order, edge irrelevant.
|
||||
expect(names(adapter.selectTopGrades(a, 10))).toEqual(['first', 'second']);
|
||||
expect(names(adapter.selectTopGrades(b, 10))).toEqual(['first', 'second']);
|
||||
});
|
||||
|
||||
test('a MISSING signal sorts LAST and is still PRESENT (never dropped)', () => {
|
||||
test('a MISSING edge is no longer penalised — rows are still all PRESENT', () => {
|
||||
const grades = [
|
||||
{ player: 'no-signal', grade: 'B', confidence: 50 },
|
||||
{ player: 'weak-but-real', grade: 'B', confidence: 50, edge: 0.1 },
|
||||
{ player: 'negative-but-real', grade: 'B', confidence: 50, edge: -5 },
|
||||
];
|
||||
const out = names(adapter.selectTopGrades(grades, 10));
|
||||
expect(out).toEqual(['weak-but-real', 'negative-but-real', 'no-signal']);
|
||||
expect(out).toHaveLength(3); // present, not dropped
|
||||
expect(out[out.length - 1]).toBe('no-signal');
|
||||
expect(out).toHaveLength(3); // never dropped
|
||||
expect(out).toEqual(['no-signal', 'weak-but-real', 'negative-but-real']); // input order
|
||||
});
|
||||
|
||||
test('null/empty-string edge is absent, NOT zero (Number(null) === 0 guard)', () => {
|
||||
test('forecast_rank LEADS the chain when the server supplied it', () => {
|
||||
const grades = [
|
||||
{ player: 'nullish', grade: 'B', confidence: 50, edge: null },
|
||||
{ player: 'empty', grade: 'B', confidence: 50, edge: '' },
|
||||
{ player: 'real-negative', grade: 'B', confidence: 50, edge: -1 },
|
||||
{ player: 'topGrade', grade: 'A', confidence: 90, forecast_rank: 3 },
|
||||
{ player: 'topForecast', grade: 'C', confidence: 40, forecast_rank: 1 },
|
||||
];
|
||||
// A real negative beats two absents; absents keep input order at the bottom.
|
||||
expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['real-negative', 'nullish', 'empty']);
|
||||
// p_win-first order wins over the letter: the letter measured r ~ 0.005 and
|
||||
// is inverted, p_win measured +0.26.
|
||||
expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['topForecast', 'topGrade']);
|
||||
});
|
||||
|
||||
test('a MISSING forecast_rank sorts LAST and falls back to the grade chain', () => {
|
||||
const grades = [
|
||||
{ player: 'noRank', grade: 'A', confidence: 90 },
|
||||
{ player: 'ranked', grade: 'C', confidence: 40, forecast_rank: 2 },
|
||||
];
|
||||
expect(names(adapter.selectTopGrades(grades, 10))).toEqual(['ranked', 'noRank']);
|
||||
});
|
||||
|
||||
test('forecast_rank null/empty is ABSENT, not 0 (Number(null) === 0 guard)', () => {
|
||||
const grades = [
|
||||
{ player: 'nullish', grade: 'B', confidence: 50, forecast_rank: null },
|
||||
{ player: 'empty', grade: 'B', confidence: 50, forecast_rank: '' },
|
||||
{ player: 'real', grade: 'C', confidence: 10, forecast_rank: 9 },
|
||||
];
|
||||
// rank 9 is a REAL rank and must beat two absents — a 0-coercion would
|
||||
// have put the nulls first.
|
||||
expect(names(adapter.selectTopGrades(grades, 10))[0]).toBe('real');
|
||||
});
|
||||
});
|
||||
|
||||
|
||||
@@ -55,18 +55,22 @@ describe('rankByForecast — the challenger', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('rankGrades — the incumbent is UNTOUCHED (live ordering byte-identical)', () => {
|
||||
it('still leads with the grade letter and still consults edge', () => {
|
||||
describe('rankGrades — still grade-first (p_win leading is rankByForecast only)', () => {
|
||||
it('leads with the grade letter, NOT p_win', () => {
|
||||
const out = rankGrades([g('lowPwinA', 'A', 0.51), g('highPwinC', 'C', 0.74)]);
|
||||
expect(out[0].player).toBe('lowPwinA'); // grade-first, unchanged
|
||||
});
|
||||
|
||||
it('edge still breaks a true tie in the incumbent', () => {
|
||||
// SUPERSEDED 2026-08-01: the edge KEY was removed from rankGrades entirely.
|
||||
// Removing it is safe for every sport — it takes a non-predictive signal out
|
||||
// without putting p_win in front (that is `rankByForecast`, gated to sports
|
||||
// whose own model has passed).
|
||||
it('edge can no longer break a tie — the key is gone for EVERY sport', () => {
|
||||
const out = rankGrades([
|
||||
g('lowEdge', 'B', 0.6, -110, { edge: 1 }),
|
||||
g('highEdge', 'B', 0.6, -110, { edge: 9 }),
|
||||
g('first', 'B', 0.6, -110, { edge: 1 }),
|
||||
g('second', 'B', 0.6, -110, { edge: 9 }),
|
||||
]);
|
||||
expect(out[0].player).toBe('highEdge');
|
||||
expect(out[0].player).toBe('first'); // stable input order, edge ignored
|
||||
});
|
||||
});
|
||||
|
||||
@@ -97,14 +101,18 @@ describe('rankingDelta — the challenger-first measurement', () => {
|
||||
});
|
||||
});
|
||||
|
||||
describe('per-sport doctrine — who may rank on the forecast', () => {
|
||||
it('MLB may; WNBA may NOT (its p_win is anti-predictive, it abstains)', () => {
|
||||
describe('per-sport doctrine — a sport ranks only once ITS OWN model is built', () => {
|
||||
// WNBA is held out as NOT-BUILT, not as failed: its -0.12 came from
|
||||
// NBA-template machinery run on WNBA data, which is an unbuilt model's
|
||||
// expected failure rather than a verdict on the sport. Same live behaviour,
|
||||
// corrected meaning.
|
||||
it('MLB may (its own model passed); WNBA may not (its model is not built yet)', () => {
|
||||
expect(ranksOnForecast('mlb')).toBe(true);
|
||||
expect(ranksOnForecast('MLB')).toBe(true);
|
||||
expect(ranksOnForecast('wnba')).toBe(false);
|
||||
});
|
||||
|
||||
it('no sport inherits MLB’s result — unknown sports are excluded', () => {
|
||||
it('no sport inherits MLB’s result — every unbuilt sport is held out', () => {
|
||||
for (const s of ['nba', 'nfl', 'soccer', 'nhl', 'ncaab', '', null, undefined]) {
|
||||
expect(ranksOnForecast(s)).toBe(false);
|
||||
}
|
||||
|
||||
@@ -385,6 +385,11 @@ function buildPlayerStripsFromProps(gameProps, gradeIndex, deltaIndex, now = Dat
|
||||
// Rev 3 flat edge-board (mobile screen 01) — the ranked hero figure.
|
||||
// Strict null (never 0-coerced) so an absent edge sorts last, not top.
|
||||
edge: (rec.edge_pct != null && Number.isFinite(Number(rec.edge_pct))) ? Number(rec.edge_pct) : null,
|
||||
// Server-computed forecast ordinal (p_win-first). Stamped BEFORE the
|
||||
// model-price strip, so it reaches every tier where raw p_win cannot.
|
||||
// Present only for sports whose own model has passed — absent elsewhere,
|
||||
// which is exactly how the boards fall back.
|
||||
forecastRank: (rec.forecast_rank != null && Number.isFinite(Number(rec.forecast_rank))) ? Number(rec.forecast_rank) : null,
|
||||
// A1 S3 — the prop's own book + the best available price across the
|
||||
// game's book rows for the graded side (null unless ≥2 books at the
|
||||
// same current line disagree — see detectBestBook).
|
||||
@@ -470,6 +475,14 @@ const strictNum = (v) => (v == null || v === '' ? null : (Number.isFinite(Number
|
||||
* in a different base" (Session 67). Those callers legitimately get null here
|
||||
* and fall through to the signed edge.
|
||||
*/
|
||||
/** Ascending comparator (rank 1 is best) that always sorts a null LAST. */
|
||||
function ascNullsLast(a, b) {
|
||||
if (a == null && b == null) return 0;
|
||||
if (a == null) return 1;
|
||||
if (b == null) return -1;
|
||||
return a - b;
|
||||
}
|
||||
|
||||
function takeablePWin(g) {
|
||||
const { isTakeable } = require('./valueState');
|
||||
const p = strictNum(g && g.p_win);
|
||||
@@ -513,19 +526,24 @@ function descNullsLast(a, b) {
|
||||
*/
|
||||
function selectTopGrades(grades, limit = 10) {
|
||||
const arr = (Array.isArray(grades) ? grades : []).filter((g) => g && g.grade);
|
||||
// FLIPPED 2026-08-01 — forecast order FIRST when the server supplied it, and
|
||||
// the EDGE KEY IS GONE. `forecast_rank` is p_win-first and is stamped only for
|
||||
// sports whose own model has passed; when absent (every not-built-yet sport)
|
||||
// this falls through to the unchanged grade-tier chain. Edge is removed
|
||||
// outright: corr(edge, outcome) = -0.010 / -0.022 on n=200 settled MLB rows,
|
||||
// against corr(p_win, outcome) = +0.26.
|
||||
const scored = arr.map((g, idx) => ({
|
||||
g,
|
||||
idx,
|
||||
fr: strictNum(g.forecast_rank),
|
||||
rank: gradeRankOf(g.grade),
|
||||
conf: numOr(g.confidence, -1),
|
||||
pWin: takeablePWin(g),
|
||||
// SIGNED — no abs(). null (absent) sorts last, not first.
|
||||
edge: strictNum(g.edge),
|
||||
}));
|
||||
scored.sort((a, b) => a.rank - b.rank
|
||||
scored.sort((a, b) => ascNullsLast(a.fr, b.fr)
|
||||
|| a.rank - b.rank
|
||||
|| b.conf - a.conf
|
||||
|| descNullsLast(a.pWin, b.pWin)
|
||||
|| descNullsLast(a.edge, b.edge)
|
||||
|| a.idx - b.idx);
|
||||
return scored.slice(0, Math.max(0, limit)).map((s) => s.g);
|
||||
}
|
||||
@@ -691,6 +709,7 @@ function flattenToEdgeBoard(cards) {
|
||||
// — treat as absent so the board neither displays nor RANKS on a fake
|
||||
// edge. When absent, rows fall through to the grade-rank tiebreak.
|
||||
edge: (p.edge != null && Number.isFinite(p.edge) && Math.abs(p.edge) <= EDGE_BOARD_SANE_MAX) ? p.edge : null,
|
||||
forecastRank: (p.forecastRank != null && Number.isFinite(p.forecastRank)) ? p.forecastRank : null,
|
||||
sport: c.sport || '',
|
||||
away,
|
||||
home,
|
||||
@@ -704,10 +723,16 @@ function flattenToEdgeBoard(cards) {
|
||||
}
|
||||
}
|
||||
}
|
||||
// FLIPPED 2026-08-01 — this board no longer ranks on edge. Edge was its
|
||||
// PRIMARY key, so the entire mobile board was ordered by a quantity measured
|
||||
// not to predict (corr -0.010 / -0.022 vs corr(p_win) = +0.26).
|
||||
// Now: server forecast order first (p_win-first; MLB-only, absent elsewhere),
|
||||
// then grade rank. Edge is still carried and still DISPLAYED as a labelled
|
||||
// diagnostic — it just no longer decides what a user sees first.
|
||||
rows.sort((a, b) => {
|
||||
const ea = a.edge == null ? -Infinity : a.edge;
|
||||
const eb = b.edge == null ? -Infinity : b.edge;
|
||||
if (eb !== ea) return eb - ea;
|
||||
const fa = a.forecastRank == null ? Infinity : a.forecastRank;
|
||||
const fb = b.forecastRank == null ? Infinity : b.forecastRank;
|
||||
if (fa !== fb) return fa - fb;
|
||||
return edgeGradeRank(a.grade) - edgeGradeRank(b.grade);
|
||||
});
|
||||
return rows.map((r, i) => ({ ...r, rank: i + 1 }));
|
||||
|
||||
Reference in New Issue
Block a user