report: grade diagnostic T0 — p_win is MISCALIBRATED, and it explains the inversion
STOPPED at the T0 gate as instructed. Nothing fixed, no recalibration applied, no grade touched. T1-T4 deliberately not run. T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS. Condition 1 (mean |predicted-actual| > 0.05): MLB ~0.094, WNBA ~0.139. Condition 2 (monotonic slope): over-confidence GROWS with the prediction — MLB +0.034 -> +0.043 -> +0.084 -> +0.190 -> +0.189; WNBA +0.044 -> +0.109 -> +0.349. Worst cases: MLB predicted 0.842 actual 0.652 (n=23), predicted 0.917 actual 0.727 (n=11); WNBA predicted 0.730 actual 0.381 (n=21). WHY THIS EXPLAINS THE INVERSION, mechanically: p_win is over-stated and the overstatement SCALES with p_win, so p_win - fair_prob_lock is largest exactly where p_win is most inflated. Those props hit less than claimed, so the edge measure correlates negatively. The market was never the problem — fair_prob_lock is not a bent ruler, the thing subtracted from it is. It also explains why p_win ALONE still carries signal (+0.23 MLB): rank survives miscalibration, differences do not. This independently reconfirms the 2026-07-26 calibration finding (+0.02 at p<.5 -> +0.19 at p>=.8) on a newer, larger population, so it is structural rather than sampling noise. PART 0: P0a — only the GRADED side's fair prob is stored (fair_prob_lock; no opposite-side field), so T1's two-side-sum check cannot run and must use the stated no-vig recompute fallback. P0b — projection_locked_at exists as a timestamptz so T2 is potentially runnable, but distinctness from lock time was NOT verified because T0 gated it. Two cautions recorded before Part 2 runs: the top MLB buckets where the error is worst hold n=23 and n=11, so a flexible per-bucket correction would fit noise — isotonic with pooling or single-parameter Platt is safer; and calibration fixes magnitudes, so if the market is genuinely better the repaired edge may still land at ~0, which would be the honest ceiling and gets reported rather than graded around. Query committed at scripts/grade-calibration-t0.sql. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
@@ -0,0 +1,13 @@
|
||||
-- T0 — CALIBRATION OF p_win (gates T1-T4). specs/grade-diagnostic-t0.md
|
||||
-- Pre-registered firing thresholds: mean |predicted-actual| > 0.05 across
|
||||
-- buckets, OR a monotonic over/under-confidence slope. BOTH fired.
|
||||
with r as (
|
||||
select sport, p_win::numeric as p, (outcome='hit')::int as won
|
||||
from ledger_entries
|
||||
where user_id is null and outcome in ('hit','miss') and p_win is not null
|
||||
), b as (
|
||||
select sport, width_bucket(p, 0.1, 1.0, 9) as bkt, p, won from r
|
||||
)
|
||||
select sport, bkt, count(*) n,
|
||||
avg(p) predicted, avg(won) actual, avg(p)-avg(won) over_confidence
|
||||
from b group by sport, bkt having count(*) >= 10 order by sport, bkt;
|
||||
Reference in New Issue
Block a user