WNBA truth correction + THE p_win FLIP (live, rollback armed)
PART A -- WNBA TRUTH CORRECTION (no behaviour change).
WNBA does not "abstain" and is not "anti-predictive". The -0.12 that
produced those words was NBA-template machinery run on WNBA data -- WNBA
has never had its own archetypes, variables, conditions or calibration,
which is precisely the "sport stubbed in on another sport's template"
CLAUDE.md forbids. That is an UNBUILT MODEL'S EXPECTED FAILURE, not a
verdict on the sport; reading it as a verdict would quietly retire a sport
we never actually attempted. Its own build is QUEUED, after MLB.
The guard CODE is unchanged -- FORECAST_RANKED_SPORTS = {'mlb'} and the
inheritance test are correct live safety either way. Only the meaning is
corrected, and generalised into the doctrine-as-a-gate: a sport ranks on
p_win ONLY once its OWN model is built and shown to predict (calibration
AND resolution on its own holdout). Others are held out as NOT-BUILT,
never as failed. Re-labelled across gradeRanking, snapshot route, tests,
MASTER-PLAN and the challenger report.
PART B -- THE FLIP, gated on a full-slate re-run.
The re-run found something better than a bigger sample. An induced
snapshot graded 7 props: gradeAndCacheSlate runs with DEFAULT_LIMIT = 25
and ~72% of those refuse for insufficient_data, while 546 props are
gradeable. So 8 props IS the board, structurally -- not a small sample of
it. Logged as its own finding; the cap is a separate order.
For a statistically meaningful delta I used 11 real historical boards
(n=328, board sizes 14-57): 79.9% of rows move, mean 5.16 places per
board, TOP READ CHANGES ON 9 OF 11 BOARDS. The re-ordering holds at real
board size. Query committed.
FLIPPED:
- rankGrades drops its edge key (safe for every sport: removes a
non-predictive tiebreak without putting p_win in front).
- selectTopGrades leads on forecast_rank, edge key removed.
- flattenToEdgeBoard sorts on forecastRank, not edge -- this board had
edge as its PRIMARY key, so the whole mobile board was ordered by a
quantity measured not to predict.
- forecast_rank threaded onto strip props.
Sports whose model is not built supply no forecast_rank, so their boards
fall through to the unchanged grade chain -- the fallback is the guard.
ROLLBACK ARMED: boards sort by forecast_rank WHEN PRESENT, so
FORECAST_RANK=0 reverts every surface on the next response -- no deploy,
no client release.
Edge is still computed, stored, carried and displayed as a labelled
diagnostic. Retired from ranking, not deleted.
Eight superseded tests updated to strictly stronger INVERSE properties --
they now fail if edge is ever re-introduced as a ranking key, which the
originals could not detect.
Gates: 4,045 tests / 323 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
+12
-8
@@ -73,7 +73,7 @@ multi-book data.
|
||||
| **MLB isotonic `p_win`** | **DECIDED** — reliability **0.0846**, resolution **0.190**, holdout **n=125** | **PROVISIONAL label RETRACTED 2026-08-01.** Calibration is **ruler-independent** (`estimateProbability` never sees a price; the fit is p_win-vs-outcome). Replicated on a fresh later window, both metrics improved |
|
||||
| **edge vs the ruler** | corr(edge, outcome) **−0.010** (v1) → **−0.022** (v2), n=200 · corr(**p_win**, outcome) **+0.26** | **Subtracting the market DESTROYS the signal.** The consensus ruler does not rescue edge: *differs ≠ better* |
|
||||
| **exchange-inclusive ruler** | **UNTESTABLE on existing data** | exchange quotes were never stored (discarded until 2026-08-01). Becomes testable only as v2-era captures accrue |
|
||||
| **WNBA** | **still abstains** | a MODEL problem, not a coverage problem — coverage was never its constraint |
|
||||
| **WNBA** | **MODEL NOT BUILT YET** — held out, *not* failed | **CORRECTED 2026-08-01.** The −0.12 was **NBA-template machinery run on WNBA data**. WNBA has never had its own archetypes/variables/conditions — the "sport stubbed in on another sport's template" CLAUDE.md forbids. That is an **unbuilt model's expected failure, not a verdict on the sport.** Its own build is QUEUED, after MLB |
|
||||
| 🔴 **pinnacle feed** | **0 captures since 2026-07-31** (103,940 in the prior 10 days) | a live regression; **we had a sharp anchor and lost it.** Not caused by our changes |
|
||||
| **soccer** | **settles** — ~15 competitions, 30d | "grades into a void" is a **$19/mo Pro-tier** problem, not a data problem |
|
||||
| **CLV + results feeds** | `/odds/closing` + `/movement` **redacted**; `/results` **403 `required_tier: hobby`**; `/exports/resolved-props` **403 `required_tier: pro`** | **verified on our keys** — plain tier exclusion, not a key or plan fault. **$9/mo** buys CLV + steam + results; **$19/mo** adds the 90-day settlement export |
|
||||
@@ -136,7 +136,7 @@ connected and *meaningless*. MLB's fix is connection, not construction.
|
||||
### A2. Sport order (OPEN — Kev decides)
|
||||
Proposed by readiness × clock: **1) MLB** (reference, only qualifying model) →
|
||||
**2) CFB** (has a <30-day clock; soft-market thesis) → **3) NFL** → **4) NBA** →
|
||||
**5) CBB** → **6) WNBA re-attempt** (abstains today) → **7) soccer** (quota-blocked).
|
||||
**5) CBB** → **6) WNBA — first real build** (never had its own model) → **7) soccer** (quota-blocked).
|
||||
Each gets the same 8-layer template. **No sport is abandoned — abstention is a
|
||||
state, not a verdict.**
|
||||
|
||||
@@ -194,7 +194,7 @@ a sport is a module. **Blocks all of A2 after MLB.**
|
||||
|
||||
| phase | orders | contents | blocks |
|
||||
|---|---|---|---|
|
||||
| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only, WNBA abstains) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
|
||||
| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only — every other sport is NOT-BUILT, held out) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
|
||||
| **2. Resolution tail** | 3 | trigger · share cards + notifications + posts + recap · CLV flag decision | share cards, social proof |
|
||||
| **3. Surfaces + design lane** *(parallel with 1-2)* | 4 | C2 decision → D1-close mount · `/record` nav + remaining surfaces · D1-B glyphs/archetypes · S2 primitives | Chrome audit |
|
||||
| **4. Sport boundary** | 2 | registry collapse · MLB re-expressed as the first module | all further sports |
|
||||
@@ -214,7 +214,9 @@ a sport is a module. **Blocks all of A2 after MLB.**
|
||||
weights, conditions and Bayesian all feed the grade; calibration applied; the
|
||||
ladder monotone (A>B>C, no inversion) and proven on held-out data.
|
||||
2. **Every listed sport is finished on the same 8-layer template**, or explicitly
|
||||
abstaining with its reason recorded — never silently absent.
|
||||
held out with its reason recorded — **"not built yet"** where no sport-specific
|
||||
model exists, and only "measured and failed" where one was genuinely built and
|
||||
tested. Never silently absent, and never a verdict on a sport we never attempted.
|
||||
3. **Design fully implemented** — all 61 catalogued items BUILT-TO-SPEC.
|
||||
4. **Every surface built, reachable and honest** — no orphans, no live-but-not-honest
|
||||
surface, no dead component.
|
||||
@@ -255,7 +257,7 @@ Every edge measurement this session came back **null, negative, or unproven**:
|
||||
| served grade → outcome | **r ≈ 0.005**, and **inverted** (B 52.4% < C 56.9%) |
|
||||
| p_win − fair_prob (3 formulations) | **negative in all three, both sports, both splits** |
|
||||
| p_win alone, MLB, holdout | +0.165, **p ≈ 0.07 — not significant** |
|
||||
| p_win alone, WNBA | **negative** — abstains |
|
||||
| p_win alone, WNBA | **negative — but this measured an NBA-template model on WNBA data, so it is not a WNBA result at all** |
|
||||
| CLV / beat-close | **null by guard** — instrument not trustworthy |
|
||||
| ROI by grade | likely an artifact of a meaningless letter |
|
||||
|
||||
@@ -291,7 +293,8 @@ them is a hypothesis, not a guarantee.** They must each prove out on held-out da
|
||||
or be left disconnected honestly.
|
||||
|
||||
## 9.3 A one-sport product marketed as multi-sport
|
||||
MLB is the only qualifying model. WNBA abstains on its own data. NBA and soccer
|
||||
MLB is the only model that has been BUILT and passed. WNBA's model does not exist
|
||||
yet (what was measured was NBA-template machinery on WNBA data). NBA and soccer
|
||||
**don't even settle** — they grade into a void. Until Phase 5, the honest framing
|
||||
is *"an MLB product with other sports in development."* The site should not imply
|
||||
otherwise.
|
||||
@@ -419,8 +422,9 @@ pitchers / depth charts / lineup confirmation that `/context` serves free.
|
||||
>
|
||||
> **WNBA is NOT thin at the feed** — 4.21 books/prop vs MLB's 3.61. It was
|
||||
> allow-list-starved exactly as MLB was. This removes one candidate explanation
|
||||
> for its anti-predictive result; it does not explain it, and WNBA stays
|
||||
> abstaining.
|
||||
> for its −0.12 result. **And that result is not a WNBA verdict anyway** — it
|
||||
> measured NBA-template machinery on WNBA data. WNBA is **NOT BUILT YET**, held
|
||||
> out until it gets its own model.
|
||||
>
|
||||
> **No sharp anchor exists for props:** `pinnacle`, `matchbook` and `polymarket`
|
||||
> all measured **0%** on both sports. The consensus ruler is therefore a MARKET
|
||||
|
||||
@@ -73,9 +73,16 @@ slate — re-run the endpoint on a full slate before the flip. It is one call.
|
||||
## PER-SPORT DOCTRINE — ENFORCED IN CODE, NOT IN A COMMENT
|
||||
|
||||
**WNBA moves the most (100% of rows, mean 4.1 places) and must NOT adopt this.**
|
||||
WNBA's `p_win` is **anti-predictive** on its own data — it abstains. Ranking that
|
||||
board by `p_win` would sort it by a signal measured to point the *wrong way*:
|
||||
worse than the incumbent, not better.
|
||||
|
||||
**CORRECTED 2026-08-01:** WNBA does **not** "abstain" and is **not**
|
||||
"anti-predictive". The −0.12 that produced those words was **NBA-template
|
||||
machinery run on WNBA data** — WNBA has never had its own archetypes, variables,
|
||||
conditions or calibration. That is an **unbuilt model's expected failure, not a
|
||||
verdict on the sport.** WNBA is **NOT BUILT YET**, held out until its own model
|
||||
exists; its build is queued after MLB.
|
||||
|
||||
The live consequence is the same either way — an unbuilt sport must not rank on a
|
||||
signal not shown to hold for it — which is why the guard code is unchanged.
|
||||
|
||||
A comment would not have stopped a future flip from applying this globally, so:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user