# MLB CALIBRATION RE-RUN vs THE CONSENSUS RULER **Date:** 2026-08-01 · **MEASURE-ONLY** · no tier spend · live path byte-identical · queries committed (`scripts/ruler-comparison.sql`, `scripts/pwin-timeforward.sql`). --- ## VERDICT UP FRONT **Mandate 1's premise does not hold, and I have to say so before reporting numbers against it.** > **The p_win calibration is RULER-INDEPENDENT. Re-deriving `fair_prob` against the > consensus ruler cannot change a single number in the isotonic result.** Two independent confirmations: 1. **`estimateProbability({ gameLogs, line, statType, features })`** (`src/services/intelligence/probabilityEstimator.js:54`) — **no market price, no book odds, no `fair_prob` anywhere in its inputs.** `p_win` is computed from game logs against the line. 2. **The calibration query fits `p_win` against OUTCOMES.** Reliability (predicted vs actual hit rate) and resolution (does higher `p_win` hit more often) are both `p_win`-vs-outcome measures. `fair_prob` cannot enter either. **So there is nothing to re-fit.** The ruler changes **edge** (`p_win − fair_prob`), **CLV**, and **takeable** — not calibration. ### I have to retract my own label **I declared the MLB isotonic result PROVISIONAL "because it was measured against the bent ruler." That was wrong** — I over-applied the ruler caveat to a measurement the ruler never touched. **The isotonic result was never contaminated.** It moves from **PROVISIONAL → DECIDED**, not by re-running, but because the gate I attached to it does not apply. It stands on its own terms. --- ## WHAT I RAN INSTEAD (the genuinely ruler-dependent question) **Does swapping the incumbent single-book ruler for a median consensus rescue EDGE?** That question *is* ruler-dependent, and it is the one that came back negative in three formulations. **Timing held constant** — both rulers read at CLOSE. A lock-time reconstruction is impossible at usable n (**only 43** settled rows join `lock_lines` with ≥2 two-sided books), and pairing a lock-time incumbent with a close-time consensus would confound *when* with *what*. ### Result — MLB, n=200 settled rows, 2026-07-21 → 07-30 | measure | value | |---|---:| | base rate | 0.525 | | **mean \|ruler gap\|** | **0.0085** (0.85 prob points) | | mean signed ruler gap | −0.0004 | | **corr(edge **v1 single-book**, outcome)** | **−0.0101** | | **corr(edge **v2 consensus**, outcome)** | **−0.0220** | | **corr(`p_win` alone, outcome)** | **+0.2598** | | edge v1: winners / losers | 0.0281 / 0.0315 | | edge v2: winners / losers | 0.0265 / 0.0339 | **The consensus ruler does not rescue edge.** Both rulers give a correlation indistinguishable from zero, and both are very slightly *negative* — higher edge associates marginally with *losing*, under either ruler. ### The number that actually matters **`p_win` alone correlates +0.26 with outcome. Edge — `p_win` minus the market — correlates −0.01 to −0.02.** **Subtracting the market price destroys the signal.** The model has genuine forecasting signal; the market-relative transformation removes it. That is a direct empirical vindication of the identity now at the top of CLAUDE.md: *market edge is a byproduct, never the success criterion.* Here it is not merely a poor criterion — **it is a strictly worse instrument than the raw forecast.** --- ## CALIBRATION REFRESH (ruler-independent, but n has grown) Worth running for a different reason: the prior report used **n=119** holdout. MLB now has **250** settled rows with `p_win`. Time-forward, earlier fits, later proves. | split | n | buckets | reliability (mean \|dev\|) | resolution (corr) | base rate | window | |---|---:|---:|---:|---:|---:|---| | train | 125 | 9 | 0.0951 | 0.2557 | 0.528 | 07-21 → 07-26 | | **holdout** | **125** | **9** | **0.0846** | **0.1902** | 0.560 | 07-26 → 07-30 | **Both hold on a fresh, later window — and both improved** versus the prior report (reliability 0.0939 → **0.0846**; resolution 0.123 → **0.190**). Resolution degrades train → holdout (0.256 → 0.190), which is expected and honest; it stays clearly positive rather than crushing to base rate. **This is independent replication on data the earlier fit never saw.** It is the strongest evidence to date that MLB `p_win` is honest *and* still ranks. --- ## THE LIMITATION THAT PREVENTS A FULL VERDICT **The ruler I could test is not the ruler that was built.** `closing_captures` contains **only MODEL books** — draftkings, betmgm, betrivers, fanduel, pinnacle. **Exchange quotes were never stored**, because `normalizeProps` discarded them until 2026-08-01. Mean books per prop in the historical join: **1.97**. So this measured a **US-books-median vs one-US-book** ruler change — which is small by construction (0.85 prob points), and nothing like the exchange-inclusive consensus, where the live delta showed **p90 +10 points and 17% of props moving ≥5 points**. > **The exchange-inclusive consensus ruler is UNTESTED and UNTESTABLE on existing > data at any n.** It becomes testable only as v2-era captures accumulate — which > starts now, because the display widening went live yesterday. **I am not forcing a verdict on it.** Per Mandate 4's third outcome: **inconclusive on holdout n.** --- ## VERDICT, PER MANDATE 4 | question | verdict | |---|---| | Does the consensus ruler improve the **calibration**? | **Question does not apply** — calibration is ruler-independent (proven in code and in the query). | | Is the MLB isotonic result still provisional? | **No — DECIDED.** My PROVISIONAL label was over-applied and is retracted. | | Does the consensus ruler rescue **EDGE** (US-books version, n=200)? | **DIFFERS BUT NOT BETTER.** −0.010 → −0.022. Rulers differ ≠ new ruler better. | | Does the **exchange-inclusive** ruler rescue edge? | **INCONCLUSIVE — n-blocked.** Untestable on existing data; the quotes were never stored. | | MLB `p_win` honest **and** ranking, held out? | **YES, replicated on fresh data.** reliability 0.0846, resolution 0.190, n=125. | | WNBA | **Abstains, unchanged.** Not re-litigated. | **Nothing is promoted. Nothing is flipped.** `CURRENT_RULER_VERSION` remains `v1_first_book`; the model still consumes `MODEL_BOOKS` only. **No promotion is scoped**, because the trigger condition — "new ruler is better" — was not met on the testable portion and could not be evaluated on the rest. --- ## 🔴 SEPARATE FINDING — A LIVE FEED REGRESSION While establishing feasibility I found this, and it is unrelated to the order: | date | pinnacle MLB captures | other books | |---|---:|---:| | 07-26 | 15,008 | 115,638 | | 07-29 | 10,406 | 82,080 | | 07-30 | 4,022 | 63,602 | | **07-31** | **0** | 47,606 | | **08-01** | **0** | 45,092 | **Pinnacle prop coverage went to zero on 2026-07-31 and has not returned**, while every other book continued normally. 103,940 captures over the preceding 10 days, then a clean cliff. **This also corrects an Order Zero claim of mine.** I reported "pinnacle 0% — no sharp anchor exists in our feed." That was accurate *for the day I measured* but wrong as a general statement: **Pinnacle was in the MLB props feed until 07-30**, with 17,090 two-sided captures. (`line_type='sharp'` is a label applied in `closingCapture.js` via `SHARP_BOOKS`, not a separate provider — I checked.) **We had a sharp anchor and lost it two days ago.** Whether that is a PropLine change, a regional restriction, or a bug is unknown and worth one question to the provider. It is not caused by anything in this session's work — the widening only adds books. ## TAGS **VERIFIED:** calibration is ruler-independent (code + query) · consensus ruler does not rescue edge at n=200 · MLB p_win replicates honest+ranking on a fresh holdout (n=125) · pinnacle coverage stopped 07-31. **CANNOT DETERMINE:** whether the exchange-inclusive consensus ruler improves edge — the quotes were never stored, so it is untestable at any n until v2-era captures accumulate. **RETRACTED:** my own "MLB isotonic is PROVISIONAL pending the ruler fix" — the ruler never touched it.