report: grade diagnostic T0 — p_win is MISCALIBRATED, and it explains the inversion
STOPPED at the T0 gate as instructed. Nothing fixed, no recalibration applied, no grade touched. T1-T4 deliberately not run. T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS. Condition 1 (mean |predicted-actual| > 0.05): MLB ~0.094, WNBA ~0.139. Condition 2 (monotonic slope): over-confidence GROWS with the prediction — MLB +0.034 -> +0.043 -> +0.084 -> +0.190 -> +0.189; WNBA +0.044 -> +0.109 -> +0.349. Worst cases: MLB predicted 0.842 actual 0.652 (n=23), predicted 0.917 actual 0.727 (n=11); WNBA predicted 0.730 actual 0.381 (n=21). WHY THIS EXPLAINS THE INVERSION, mechanically: p_win is over-stated and the overstatement SCALES with p_win, so p_win - fair_prob_lock is largest exactly where p_win is most inflated. Those props hit less than claimed, so the edge measure correlates negatively. The market was never the problem — fair_prob_lock is not a bent ruler, the thing subtracted from it is. It also explains why p_win ALONE still carries signal (+0.23 MLB): rank survives miscalibration, differences do not. This independently reconfirms the 2026-07-26 calibration finding (+0.02 at p<.5 -> +0.19 at p>=.8) on a newer, larger population, so it is structural rather than sampling noise. PART 0: P0a — only the GRADED side's fair prob is stored (fair_prob_lock; no opposite-side field), so T1's two-side-sum check cannot run and must use the stated no-vig recompute fallback. P0b — projection_locked_at exists as a timestamptz so T2 is potentially runnable, but distinctness from lock time was NOT verified because T0 gated it. Two cautions recorded before Part 2 runs: the top MLB buckets where the error is worst hold n=23 and n=11, so a flexible per-bucket correction would fit noise — isotonic with pooling or single-parameter Platt is safer; and calibration fixes magnitudes, so if the market is genuinely better the repaired edge may still land at ~0, which would be the honest ceiling and gets reported rather than graded around. Query committed at scripts/grade-calibration-t0.sql. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
@@ -0,0 +1,13 @@
|
||||
-- T0 — CALIBRATION OF p_win (gates T1-T4). specs/grade-diagnostic-t0.md
|
||||
-- Pre-registered firing thresholds: mean |predicted-actual| > 0.05 across
|
||||
-- buckets, OR a monotonic over/under-confidence slope. BOTH fired.
|
||||
with r as (
|
||||
select sport, p_win::numeric as p, (outcome='hit')::int as won
|
||||
from ledger_entries
|
||||
where user_id is null and outcome in ('hit','miss') and p_win is not null
|
||||
), b as (
|
||||
select sport, width_bucket(p, 0.1, 1.0, 9) as bkt, p, won from r
|
||||
)
|
||||
select sport, bkt, count(*) n,
|
||||
avg(p) predicted, avg(won) actual, avg(p)-avg(won) over_confidence
|
||||
from b group by sport, bkt having count(*) >= 10 order by sport, bkt;
|
||||
@@ -0,0 +1,87 @@
|
||||
# GRADE DIAGNOSTIC v2 — PART 0 + T0 (report; STOPPED at the T0 gate)
|
||||
2026-07-31. Nothing fixed, no recalibration applied, no grade touched.
|
||||
|
||||
# 🔴 T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS — p_win IS MISCALIBRATED
|
||||
|
||||
Per the order: *"if T0 fires, THAT is the root — the market comparison is
|
||||
premature."* It fired, so T1–T4 were **not** run and no fix was attempted.
|
||||
|
||||
## The calibration curve (n=442 decided rows with p_win, buckets with n>=10)
|
||||
|
||||
| sport | predicted | actual | **over-confidence** | n |
|
||||
|---|---|---|---|---|
|
||||
| MLB | 0.352 | 0.286 | +0.066 | 28 |
|
||||
| MLB | 0.462 | 0.513 | −0.051 | 39 |
|
||||
| MLB | 0.545 | 0.511 | +0.034 | 47 |
|
||||
| MLB | 0.639 | 0.596 | +0.043 | 52 |
|
||||
| MLB | 0.751 | 0.667 | +0.084 | 39 |
|
||||
| MLB | **0.842** | **0.652** | **+0.190** | 23 |
|
||||
| MLB | **0.917** | **0.727** | **+0.189** | 11 |
|
||||
| WNBA | 0.460 | 0.512 | −0.052 | 41 |
|
||||
| WNBA | 0.544 | 0.500 | +0.044 | 74 |
|
||||
| WNBA | 0.646 | 0.537 | +0.109 | 41 |
|
||||
| WNBA | **0.730** | **0.381** | **+0.349** | 21 |
|
||||
|
||||
**Condition 1 — mean |predicted − actual| > 0.05: FIRES.**
|
||||
MLB mean |dev| ≈ **0.094**; WNBA ≈ **0.139**. Both well past the 0.05 threshold.
|
||||
|
||||
**Condition 2 — monotonic over/under-confidence slope: FIRES.**
|
||||
Over-confidence **grows with the prediction**: MLB +0.034 → +0.043 → +0.084 →
|
||||
**+0.190 → +0.189**; WNBA +0.044 → +0.109 → **+0.349**. The model is roughly
|
||||
honest near a coin flip and badly over-stated at the top.
|
||||
|
||||
# WHY THIS MECHANICALLY EXPLAINS THE INVERSION
|
||||
|
||||
This is the causal story the previous order was missing, and the data supports it
|
||||
end to end:
|
||||
|
||||
1. `p_win` is over-stated, and **the overstatement SCALES with p_win**.
|
||||
2. `p_win − fair_prob_lock` is therefore **largest exactly where p_win is most
|
||||
inflated** — the "biggest edges" are the most over-confident predictions.
|
||||
3. Those props hit **less** than claimed (0.842 → 0.652; 0.917 → 0.727; WNBA
|
||||
0.730 → 0.381), so the edge measure correlates **negatively**.
|
||||
|
||||
**The market was never the problem.** `fair_prob_lock` is not a bent ruler — the
|
||||
thing being subtracted from it is. That also explains why **p_win alone still
|
||||
carries signal** (+0.23 MLB): the *ranking* is partly right even though the
|
||||
*magnitudes* are wrong. Rank survives miscalibration; differences do not.
|
||||
|
||||
**This independently reconfirms the 2026-07-26 diagnosis** (`+0.02 at p<.5 →
|
||||
+0.19 at p>=.8`) on a newer, larger population — so it is a persistent structural
|
||||
property, not sampling noise.
|
||||
|
||||
# PART 0 — SCHEMA PRE-CHECK (both premises resolved)
|
||||
|
||||
- **P0a — only the GRADED side's fair prob is stored.** `ledger_entries` has
|
||||
`fair_prob_lock` (graded side), `proj_book_implied`, `dclv_fair_lock/close` — but
|
||||
**no opposite-side fair prob**. So T1's two-side-sum de-vig check **cannot run**;
|
||||
it must use the stated fallback (independent no-vig recompute from `locked_odds`)
|
||||
if T1 is ever reached.
|
||||
- **P0b — `projection_locked_at` EXISTS** as a `timestamptz`, so T2 is *potentially*
|
||||
runnable, but whether it is genuinely distinct from lock time was **not verified**
|
||||
(T0 gated it). **CANNOT DETERMINE** until T2 is authorised.
|
||||
|
||||
# WHAT THIS MEANS FOR PART 2 (not done — awaiting your go)
|
||||
|
||||
The named cause is **T0: calibrate `p_win`**. The order's prescribed fix applies —
|
||||
isotonic or Platt on a training split, **proven on a held-out/forward split**, then
|
||||
re-measure the market-relative edge on the *repaired* instrument.
|
||||
|
||||
Two honest cautions I want on the record before that runs:
|
||||
|
||||
1. **Sample is thin for a calibration map.** The top MLB buckets — where the error
|
||||
is worst and the fix matters most — hold **n=23 and n=11**. Fitting a correction
|
||||
there risks fitting noise. A monotone (isotonic) fit with pooling, or a
|
||||
single-parameter Platt scaling, is far safer than a flexible per-bucket map.
|
||||
2. **Recalibration may not resurrect the edge.** Calibrating fixes the *magnitudes*;
|
||||
if the market is genuinely better than the model, `p_win_calibrated −
|
||||
fair_prob_lock` can still land at ~0. **That would be the honest ceiling**, and
|
||||
per the order it gets reported plainly rather than graded around.
|
||||
|
||||
# TAGS
|
||||
VERIFIED: T0 fires on both pre-registered conditions; the monotone over-confidence
|
||||
slope; P0a (single-side fair prob only). CANNOT DETERMINE: P0b timestamp
|
||||
distinctness (T0 gated T2). **BLOCKED by the gate, deliberately: T1-T4 and all of
|
||||
Part 2/3 — the root is named, and the order forbids fixing before reporting.**
|
||||
|
||||
Query committed: `scripts/grade-calibration-t0.sql`.
|
||||
Reference in New Issue
Block a user