PHASE 1 (measured on the live prod feed with the real key): - WNBA is NOT thin at the feed -- 4.21 books/prop vs MLB's 3.61. It was allow-list-starved exactly as MLB was. This removes one candidate explanation for its anti-predictive result; it does not explain it, and WNBA stays abstaining. - We cannot see 64.8% of the MLB slate at all (zero admitted books). - Exchanges are real (smarkets 27%, novig 22%, kalshi 15% on MLB) but pinnacle, matchbook and polymarket measured 0% on BOTH sports. There is no sharp anchor for player props. The consensus is a MARKET consensus, not a SHARP one -- recorded as a permanent limitation, not a milestone. - DFS is the trap, quantified: prizepicks covers 82% of MLB props, the highest in the feed. Admitting it "for breadth" would have looked like the biggest available win. Permanently excluded. - Endpoints: /context WORKS and is FREE (umpire, roof, pitcher handedness, lineup confirmation -- richer than what we hand-built). /odds/closing and /movement are REDACTED (full structure, zero prices). /results and /exports/resolved-props are 403. - The $19/mo question is answered: soccer IS graded, ~15 competitions in 30 days (MLS 41k, Liga MX 15k, Brasileirao 12k, UCL/Europa/Conference). Our "soccer grades into a void" is a Pro-tier problem, not a data problem. NBA is absent because it is July -- seasonal, not inferable either way. PHASE 2 delta, corrected: MLB mean +1.50 pts, median 0, p90 +10.0, 17.0% of comparable props move >=5 pts, one-directional (the incumbent prices the over below the exchange-inclusive consensus). WNBA symmetric and tight. The median prop does not move -- the change is a right-skewed minority. That the rulers DIFFER is established; that the new one is BETTER is not, and that is the re-run. PHASE 2 item 6: ledger_entries.ruler_version applied to prod, 1,384 existing rows backfilled to v1_first_book (a statement of fact -- every row to date was produced by the first-book rule). ledgerService stamps CURRENT_RULER_VERSION on new rows. Never pool edge or CLV across it. Repo migration numbering lags prod; 025_ledger_ruler_version.sql records the DDL for review. PHASE 3: MLB isotonic p_win remains PROVISIONAL -- calibrated against v1_first_book, does not promote until re-run on the consensus ruler. NOT LIVE, deliberately: ALLOWED_BOOKS unchanged, served slate byte-identical, CURRENT_RULER_VERSION still v1_first_book, no live path calls consensusRuler. Gates: 4,022 tests passed / 322 suites; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
12 KiB
Order Zero — CONSENSUS RULER + ENDPOINT VERIFICATION
Date: 2026-08-01 · Phase 1 measured against the live prod feed with the real key; Phase 2 built challenger-first — the live ruler is byte-identical.
Gates: tests 4,022 passed / 322 suites · next build exit 0 · keyed
verification run on prod · challenger delta recorded.
PHASE 1 — KEYED VERIFICATION
Run via GET /api/internal/propline-verify (internal-key gated, read-only, no
cache/ledger/grade writes). Breadth reuses proplineAdapter.fetchRaw — the
exact live request — so this measures what the pipeline actually receives.
1.1 WNBA is NOT thin at the feed. It was starved by our own allow-list.
MLB 3,138 props / 20 events · WNBA 160 props / 4 events. Never pooled.
| MLB | WNBA | |
|---|---|---|
| books/prop in the feed | 3.61 | 4.21 |
books/prop after ALLOWED_BOOKS |
0.57 | 1.20 |
| props INVISIBLE (zero admitted books) | 2,033 (64.8%) | 46 (28.8%) |
WNBA has BETTER feed breadth than MLB (4.21 vs 3.61). The answer to the order's question is unambiguous: WNBA was allow-list-starved, exactly as MLB was — not thin at source.
That does not by itself explain WNBA's anti-predictive result, and I am not going to claim it does. It removes one candidate explanation ("no market to measure against") and leaves the ruler bend as a live hypothesis to re-test after the fix. Nothing here promotes WNBA out of abstention.
The headline number for MLB is not the mean — it is that we cannot see 64.8% of the MLB slate at all.
1.2 Exchange reality — PARTIALLY confirmed, and two candidates are dead
Reference-book presence, as a share of props:
| book | MLB | WNBA | |
|---|---|---|---|
smarkets |
27.1% | 0% | exchange |
novig |
22.2% | 46.3% | exchange |
kalshi |
14.9% | 8.8% | prediction market |
bovada |
9.2% | 16.3% | offshore, PropLine's own no-vig fallback |
pinnacle |
0% | 0% | the classic sharp anchor |
matchbook |
0% | 0% | exchange |
polymarket |
0% | 0% | prediction market |
Confirmed: exchanges do carry real props — smarkets 873 quotes, novig
730, kalshi 1,369 on MLB. The hypothesis was not empty.
Three candidates are dead for player props: pinnacle, matchbook,
polymarket returned zero on both sports. Pinnacle carries 144k markets
globally — just not these. We have no sharp anchor for props. Twice
measured, both sports.
1.3 The coverage curve — the number that decides the ruler
Strict eligibility: ≥2 reference books posting BOTH sides at the SAME line. One-sided quotes cannot be de-vigged; two books at different lines are not the same market.
| reference policy | MLB n≥2 | WNBA n≥2 |
|---|---|---|
exchange_only |
14.3% | 5.6% |
exchange_plus_sharp (+pinnacle, bovada) |
14.3% | 5.6% |
exchange_plus_us (+DK/FD/MGM/BR) |
28.1% | 46.3% |
takeable_only |
12.6% | 30.6% |
Adding pinnacle and matchbook/polymarket changes nothing — they are 0%.
Of props we can actually GRADE (those with an admitted book), consensus is available on 80.5% of MLB (923/1,146) and 54.9% of WNBA (79/144). That is the operative number: the 15% figure is diluted by props we cannot see at all.
1.4 DFS is the trap, quantified
prizepicks covers 82.3% of MLB props and sleeper 68.1% of WNBA — the
highest coverage in the entire feed. Admitting DFS "for breadth" would have
looked like the single biggest improvement available and would have bent the
ruler in a new direction. Fixed-payout, deliberately shaded, flat juice.
Permanently excluded from all pricing.
1.5 Endpoint verification — works / partial / no, FOR OUR TIER
| endpoint | verdict | evidence |
|---|---|---|
/v1/sports |
✅ WORKS | 54 sports, all active |
/v1/markets/resolution-summary |
✅ WORKS | 2,085,840 props graded in 30d |
/events/{id}/context (MLB) |
✅ WORKS, FREE | venue, roof_type, is_indoor, both probable pitchers + handedness, lineup_confirmed, home-plate umpire, weather |
/events/{id}/context (WNBA) |
⚠️ PARTIAL | 404 — MLB-only today, as documented |
/events/{id}/odds/closing |
⚠️ PARTIAL — REDACTED | 200 with 2,102 outcomes and 0 prices + upgrade_url |
/events/{id}/movement |
⚠️ PARTIAL — REDACTED | 2,102 outcomes, 0 prices, steam: [] |
/events/{id}/results |
❌ NO | 403 |
/exports/resolved-props |
❌ NO | 403 on all four sport keys |
A correction to my own first pass. I initially classified /odds/closing
and /movement as WORKS because the bodies were non-empty. They are
structure-only redactions — full shape, zero prices, plus an upgrade_url.
detectRedaction now counts actual prices and downgrades works→partial, and a
test locks it. A non-empty body is not proof of access.
1.6 The $19/mo question — answered, with one honest gap
/exports/resolved-props is 403 for us, so we cannot pull settlements. But
/markets/resolution-summary proves they exist to be bought (30 days):
| sport | props graded | events |
|---|---|---|
baseball_mlb |
1,798,850 | 363 |
basketball_wnba |
152,960 | 72 |
| soccer_mls | 41,444 | 35 |
| soccer_liga_mx | 15,037 | 21 |
| soccer_brasileirao / world_cup / UCL / Europa / Conference / … | ~60,000 combined | ~340 |
mma_ufc |
2,738 | 39 |
| tennis, golf, rugby | ~19,000 | ~1,150 |
Soccer settles. ~15 competitions graded in the last 30 days. Our standing "soccer grades into a void" is a Pro-tier ($19/mo) problem, not a data availability problem.
NBA is absent — and that is seasonal, not a coverage gap. It is July;
/sports lists basketball_nba as active. I cannot confirm NBA settlement
until the season starts, and I am not going to infer it.
PHASE 2 — THE CONSENSUS RULER (challenger-first, NOT live)
2.1 The three-way split — src/config/bookRoles.js
One allow-list was answering two different questions: "can we show this?" and "can we price against this?" Conflating them is what bent the ruler.
- TAKEABLE — a user can actually bet here. Best price, line shopping.
- REFERENCE — may price the ruler; never surfaced as a place to bet.
Selected policy:
exchange_plus_us. - EXCLUDED — DFS pick'em + offshore/intl. Permanently barred from pricing.
Deliberate deviation, stated plainly. The order said remove the 6 phantom
entries. I kept them. They returned zero PropLine quotes, but PropLine is
not our only provider — the odds-api backup path may carry Caesars/bet365/etc.
A book that never appears is never matched, so keeping it costs nothing;
deleting it risks silently dropping real books on the backup with no upside.
Recorded in PHANTOM_ON_PROPLINE rather than enacted as a deletion. Reverse
this if you'd rather.
The honest limitation, recorded in the config itself: including soft US books makes this a MARKET consensus, not a SHARP one. It is a large improvement on one arbitrary book. It is not the sharp closing line the category's leaders are measured against — and §1.2 says we cannot build that one, because no sharp book prices these props.
2.2 The rule — src/services/consensusRuler.js
- Group by the SAME LINE. Different lines are never averaged.
- TWO-SIDED REFERENCE quotes only. A one-sided price cannot be de-vigged.
- MEDIAN of per-book de-vigged
fair_prob— median, so one stale exchange cannot drag the ruler. - n≥2, or fall back to single-book and label it (
source: 'single_book', stamped v1). Never silently mixed — a column holding both is two rulers wearing one name. - No reference quote at all →
null,source: 'none'. Never fabricated.
2.3 The challenger delta — and a correction I caught mid-run
My first delta run was wrong and I stopped it. I modelled the incumbent as first-row-wins over the raw feed and it reported that an EXCLUDED book was "the market" on 69% of MLB prop-lines, prizepicks alone at 47%.
That is false. normalizeProps applies ALLOWED_BOOKS before
dedupeProps runs, so DFS books never reach the incumbent. The allow-list, for
all the coverage it costs, does keep DFS out of the ruler. Overstating the
incumbent's badness would have been as dishonest as understating it — and more
persuasive. Corrected, tested, and re-run.
Corrected results (allowedBooks respected — incumbent is 100% admitted
books, no DFS):
| MLB | WNBA | |
|---|---|---|
| prop-line groups | 6,422 | 853 |
| gradeable (has an admitted book) | 1,146 (17.8%) | 144 (16.9%) |
| consensus available | 995 | 87 |
| comparable (both rulers) | 923 | 79 |
| incumbent books | betmgm 586 · dk 538 · fd 16 · br 6 | dk 90 · fd 54 |
Delta (consensus − incumbent, probability points, over side):
| MLB | WNBA | |
|---|---|---|
| mean | +1.50 | −0.29 |
| median | 0 | 0 |
| p10 / p90 | −1.6 / +10.0 | −2.0 / +1.6 |
| median abs / p90 abs | 0.5 / 11.95 | 0.55 / 2.5 |
| disagree ≥2 pts | 24.8% | 20.3% |
| disagree ≥5 pts | 17.0% | 2.5% |
Read it honestly: the median prop does not move at all. The change is concentrated in a right-skewed minority — on MLB, 17% of comparable props move by 5+ probability points, and the skew is one-directional (p90 +10.0 vs p10 −1.6): the incumbent systematically prices the OVER below the exchange-inclusive consensus. WNBA is nearly symmetric and much tighter.
What this does and does not establish. It establishes the two rulers genuinely differ, on a minority of props, in a consistent direction. It does not establish that the new ruler is better — that requires re-running the edge and CLV measurements against it, which is the next order.
2.4 ruler_version stamped at the boundary — APPLIED
ledger_entries.ruler_version (migration recorded as
migrations/025_ledger_ruler_version.sql; applied to prod). 1,384 existing
rows backfilled to v1_first_book — a statement of fact, not an assumption:
every row to date was produced by the first-book rule. 629 of them carry a
fair_prob_lock. New rows stamp CURRENT_RULER_VERSION via ledgerService.
NEVER pool edge or CLV across differing ruler_version values. The
denominator changed; the numbers are not the same measurement. This is the same
no-pooling boundary as model_version, and it is arithmetically forced.
PHASE 3 — CONSEQUENCE
MLB isotonic p_win remains PROVISIONAL. It was calibrated against the
v1_first_book ruler. It does not promote until re-run against the consensus
ruler, after this delta is reviewed. Recorded in MASTER-PLAN.
WHAT IS AND IS NOT LIVE
LIVE: the internal verification endpoint; bookRoles.js and
consensusRuler.js as libraries; the ruler_version column and stamp.
NOT LIVE — deliberately: ALLOWED_BOOKS is unchanged, so the served
slate is byte-identical. CURRENT_RULER_VERSION is still v1_first_book. No
live path calls consensusRuler. No grade, cache or slate changed.
THE DECISION THIS TEES UP
Widening ALLOWED_BOOKS to the reference set would make 64.8% of the MLB
slate visible for the first time — that is the single largest product change
available, and it is a config edit.
But it is not free: those props would grade against a market we have never
measured our model on, and every edge number reset to v2. That is the next
order's decision, not this one's — and it should be taken after the ruler
re-run, not before it.
OPEN
- Ruler policy sign-off —
exchange_plus_usselected off the curve. Confirm. - Pro tier ($19/mo) — buys settlement for soccer (proven), CLV via
/odds/closing, steam via/movement, and/results. Genuinely cheap relative to what it unblocks. - The 6 phantom books — kept, against the order's wording, with reasoning above. Reverse if you disagree.
/contextis free and richer than what we hand-built — umpire, roof, pitcher handedness, lineup confirmation. Wiring it is a later order.