STOPPED at the T0 gate as instructed. Nothing fixed, no recalibration applied, no grade touched. T1-T4 deliberately not run. T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS. Condition 1 (mean |predicted-actual| > 0.05): MLB ~0.094, WNBA ~0.139. Condition 2 (monotonic slope): over-confidence GROWS with the prediction — MLB +0.034 -> +0.043 -> +0.084 -> +0.190 -> +0.189; WNBA +0.044 -> +0.109 -> +0.349. Worst cases: MLB predicted 0.842 actual 0.652 (n=23), predicted 0.917 actual 0.727 (n=11); WNBA predicted 0.730 actual 0.381 (n=21). WHY THIS EXPLAINS THE INVERSION, mechanically: p_win is over-stated and the overstatement SCALES with p_win, so p_win - fair_prob_lock is largest exactly where p_win is most inflated. Those props hit less than claimed, so the edge measure correlates negatively. The market was never the problem — fair_prob_lock is not a bent ruler, the thing subtracted from it is. It also explains why p_win ALONE still carries signal (+0.23 MLB): rank survives miscalibration, differences do not. This independently reconfirms the 2026-07-26 calibration finding (+0.02 at p<.5 -> +0.19 at p>=.8) on a newer, larger population, so it is structural rather than sampling noise. PART 0: P0a — only the GRADED side's fair prob is stored (fair_prob_lock; no opposite-side field), so T1's two-side-sum check cannot run and must use the stated no-vig recompute fallback. P0b — projection_locked_at exists as a timestamptz so T2 is potentially runnable, but distinctness from lock time was NOT verified because T0 gated it. Two cautions recorded before Part 2 runs: the top MLB buckets where the error is worst hold n=23 and n=11, so a flexible per-bucket correction would fit noise — isotonic with pooling or single-parameter Platt is safer; and calibration fixes magnitudes, so if the market is genuinely better the repaired edge may still land at ~0, which would be the honest ceiling and gets reported rather than graded around. Query committed at scripts/grade-calibration-t0.sql. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
4.4 KiB
GRADE DIAGNOSTIC v2 — PART 0 + T0 (report; STOPPED at the T0 gate)
2026-07-31. Nothing fixed, no recalibration applied, no grade touched.
🔴 T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS — p_win IS MISCALIBRATED
Per the order: "if T0 fires, THAT is the root — the market comparison is premature." It fired, so T1–T4 were not run and no fix was attempted.
The calibration curve (n=442 decided rows with p_win, buckets with n>=10)
| sport | predicted | actual | over-confidence | n |
|---|---|---|---|---|
| MLB | 0.352 | 0.286 | +0.066 | 28 |
| MLB | 0.462 | 0.513 | −0.051 | 39 |
| MLB | 0.545 | 0.511 | +0.034 | 47 |
| MLB | 0.639 | 0.596 | +0.043 | 52 |
| MLB | 0.751 | 0.667 | +0.084 | 39 |
| MLB | 0.842 | 0.652 | +0.190 | 23 |
| MLB | 0.917 | 0.727 | +0.189 | 11 |
| WNBA | 0.460 | 0.512 | −0.052 | 41 |
| WNBA | 0.544 | 0.500 | +0.044 | 74 |
| WNBA | 0.646 | 0.537 | +0.109 | 41 |
| WNBA | 0.730 | 0.381 | +0.349 | 21 |
Condition 1 — mean |predicted − actual| > 0.05: FIRES. MLB mean |dev| ≈ 0.094; WNBA ≈ 0.139. Both well past the 0.05 threshold.
Condition 2 — monotonic over/under-confidence slope: FIRES. Over-confidence grows with the prediction: MLB +0.034 → +0.043 → +0.084 → +0.190 → +0.189; WNBA +0.044 → +0.109 → +0.349. The model is roughly honest near a coin flip and badly over-stated at the top.
WHY THIS MECHANICALLY EXPLAINS THE INVERSION
This is the causal story the previous order was missing, and the data supports it end to end:
p_winis over-stated, and the overstatement SCALES with p_win.p_win − fair_prob_lockis therefore largest exactly where p_win is most inflated — the "biggest edges" are the most over-confident predictions.- Those props hit less than claimed (0.842 → 0.652; 0.917 → 0.727; WNBA 0.730 → 0.381), so the edge measure correlates negatively.
The market was never the problem. fair_prob_lock is not a bent ruler — the
thing being subtracted from it is. That also explains why p_win alone still
carries signal (+0.23 MLB): the ranking is partly right even though the
magnitudes are wrong. Rank survives miscalibration; differences do not.
This independently reconfirms the 2026-07-26 diagnosis (+0.02 at p<.5 → +0.19 at p>=.8) on a newer, larger population — so it is a persistent structural
property, not sampling noise.
PART 0 — SCHEMA PRE-CHECK (both premises resolved)
- P0a — only the GRADED side's fair prob is stored.
ledger_entrieshasfair_prob_lock(graded side),proj_book_implied,dclv_fair_lock/close— but no opposite-side fair prob. So T1's two-side-sum de-vig check cannot run; it must use the stated fallback (independent no-vig recompute fromlocked_odds) if T1 is ever reached. - P0b —
projection_locked_atEXISTS as atimestamptz, so T2 is potentially runnable, but whether it is genuinely distinct from lock time was not verified (T0 gated it). CANNOT DETERMINE until T2 is authorised.
WHAT THIS MEANS FOR PART 2 (not done — awaiting your go)
The named cause is T0: calibrate p_win. The order's prescribed fix applies —
isotonic or Platt on a training split, proven on a held-out/forward split, then
re-measure the market-relative edge on the repaired instrument.
Two honest cautions I want on the record before that runs:
- Sample is thin for a calibration map. The top MLB buckets — where the error is worst and the fix matters most — hold n=23 and n=11. Fitting a correction there risks fitting noise. A monotone (isotonic) fit with pooling, or a single-parameter Platt scaling, is far safer than a flexible per-bucket map.
- Recalibration may not resurrect the edge. Calibrating fixes the magnitudes;
if the market is genuinely better than the model,
p_win_calibrated − fair_prob_lockcan still land at ~0. That would be the honest ceiling, and per the order it gets reported plainly rather than graded around.
TAGS
VERIFIED: T0 fires on both pre-registered conditions; the monotone over-confidence slope; P0a (single-side fair prob only). CANNOT DETERMINE: P0b timestamp distinctness (T0 gated T2). BLOCKED by the gate, deliberately: T1-T4 and all of Part 2/3 — the root is named, and the order forbids fixing before reporting.
Query committed: scripts/grade-calibration-t0.sql.