Files
vyndr/specs/grade-diagnostic-t0.md
builtbykev 249b3e8235 report: grade diagnostic T0 — p_win is MISCALIBRATED, and it explains the inversion
STOPPED at the T0 gate as instructed. Nothing fixed, no recalibration applied,
no grade touched. T1-T4 deliberately not run.

T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS.

Condition 1 (mean |predicted-actual| > 0.05): MLB ~0.094, WNBA ~0.139.
Condition 2 (monotonic slope): over-confidence GROWS with the prediction —
MLB +0.034 -> +0.043 -> +0.084 -> +0.190 -> +0.189; WNBA +0.044 -> +0.109 ->
+0.349. Worst cases: MLB predicted 0.842 actual 0.652 (n=23), predicted 0.917
actual 0.727 (n=11); WNBA predicted 0.730 actual 0.381 (n=21).

WHY THIS EXPLAINS THE INVERSION, mechanically: p_win is over-stated and the
overstatement SCALES with p_win, so p_win - fair_prob_lock is largest exactly
where p_win is most inflated. Those props hit less than claimed, so the edge
measure correlates negatively. The market was never the problem —
fair_prob_lock is not a bent ruler, the thing subtracted from it is. It also
explains why p_win ALONE still carries signal (+0.23 MLB): rank survives
miscalibration, differences do not.

This independently reconfirms the 2026-07-26 calibration finding (+0.02 at p<.5
-> +0.19 at p>=.8) on a newer, larger population, so it is structural rather
than sampling noise.

PART 0: P0a — only the GRADED side's fair prob is stored (fair_prob_lock;
no opposite-side field), so T1's two-side-sum check cannot run and must use the
stated no-vig recompute fallback. P0b — projection_locked_at exists as a
timestamptz so T2 is potentially runnable, but distinctness from lock time was
NOT verified because T0 gated it.

Two cautions recorded before Part 2 runs: the top MLB buckets where the error is
worst hold n=23 and n=11, so a flexible per-bucket correction would fit noise —
isotonic with pooling or single-parameter Platt is safer; and calibration fixes
magnitudes, so if the market is genuinely better the repaired edge may still land
at ~0, which would be the honest ceiling and gets reported rather than graded
around.

Query committed at scripts/grade-calibration-t0.sql.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 20:12:20 -04:00

88 lines
4.4 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GRADE DIAGNOSTIC v2 — PART 0 + T0 (report; STOPPED at the T0 gate)
2026-07-31. Nothing fixed, no recalibration applied, no grade touched.
# 🔴 T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS — p_win IS MISCALIBRATED
Per the order: *"if T0 fires, THAT is the root — the market comparison is
premature."* It fired, so T1T4 were **not** run and no fix was attempted.
## The calibration curve (n=442 decided rows with p_win, buckets with n>=10)
| sport | predicted | actual | **over-confidence** | n |
|---|---|---|---|---|
| MLB | 0.352 | 0.286 | +0.066 | 28 |
| MLB | 0.462 | 0.513 | 0.051 | 39 |
| MLB | 0.545 | 0.511 | +0.034 | 47 |
| MLB | 0.639 | 0.596 | +0.043 | 52 |
| MLB | 0.751 | 0.667 | +0.084 | 39 |
| MLB | **0.842** | **0.652** | **+0.190** | 23 |
| MLB | **0.917** | **0.727** | **+0.189** | 11 |
| WNBA | 0.460 | 0.512 | 0.052 | 41 |
| WNBA | 0.544 | 0.500 | +0.044 | 74 |
| WNBA | 0.646 | 0.537 | +0.109 | 41 |
| WNBA | **0.730** | **0.381** | **+0.349** | 21 |
**Condition 1 — mean |predicted actual| > 0.05: FIRES.**
MLB mean |dev| ≈ **0.094**; WNBA ≈ **0.139**. Both well past the 0.05 threshold.
**Condition 2 — monotonic over/under-confidence slope: FIRES.**
Over-confidence **grows with the prediction**: MLB +0.034 → +0.043 → +0.084 →
**+0.190 → +0.189**; WNBA +0.044 → +0.109 → **+0.349**. The model is roughly
honest near a coin flip and badly over-stated at the top.
# WHY THIS MECHANICALLY EXPLAINS THE INVERSION
This is the causal story the previous order was missing, and the data supports it
end to end:
1. `p_win` is over-stated, and **the overstatement SCALES with p_win**.
2. `p_win fair_prob_lock` is therefore **largest exactly where p_win is most
inflated** — the "biggest edges" are the most over-confident predictions.
3. Those props hit **less** than claimed (0.842 → 0.652; 0.917 → 0.727; WNBA
0.730 → 0.381), so the edge measure correlates **negatively**.
**The market was never the problem.** `fair_prob_lock` is not a bent ruler — the
thing being subtracted from it is. That also explains why **p_win alone still
carries signal** (+0.23 MLB): the *ranking* is partly right even though the
*magnitudes* are wrong. Rank survives miscalibration; differences do not.
**This independently reconfirms the 2026-07-26 diagnosis** (`+0.02 at p<.5 →
+0.19 at p>=.8`) on a newer, larger population — so it is a persistent structural
property, not sampling noise.
# PART 0 — SCHEMA PRE-CHECK (both premises resolved)
- **P0a — only the GRADED side's fair prob is stored.** `ledger_entries` has
`fair_prob_lock` (graded side), `proj_book_implied`, `dclv_fair_lock/close` — but
**no opposite-side fair prob**. So T1's two-side-sum de-vig check **cannot run**;
it must use the stated fallback (independent no-vig recompute from `locked_odds`)
if T1 is ever reached.
- **P0b — `projection_locked_at` EXISTS** as a `timestamptz`, so T2 is *potentially*
runnable, but whether it is genuinely distinct from lock time was **not verified**
(T0 gated it). **CANNOT DETERMINE** until T2 is authorised.
# WHAT THIS MEANS FOR PART 2 (not done — awaiting your go)
The named cause is **T0: calibrate `p_win`**. The order's prescribed fix applies —
isotonic or Platt on a training split, **proven on a held-out/forward split**, then
re-measure the market-relative edge on the *repaired* instrument.
Two honest cautions I want on the record before that runs:
1. **Sample is thin for a calibration map.** The top MLB buckets — where the error
is worst and the fix matters most — hold **n=23 and n=11**. Fitting a correction
there risks fitting noise. A monotone (isotonic) fit with pooling, or a
single-parameter Platt scaling, is far safer than a flexible per-bucket map.
2. **Recalibration may not resurrect the edge.** Calibrating fixes the *magnitudes*;
if the market is genuinely better than the model, `p_win_calibrated
fair_prob_lock` can still land at ~0. **That would be the honest ceiling**, and
per the order it gets reported plainly rather than graded around.
# TAGS
VERIFIED: T0 fires on both pre-registered conditions; the monotone over-confidence
slope; P0a (single-side fair prob only). CANNOT DETERMINE: P0b timestamp
distinctness (T0 gated T2). **BLOCKED by the gate, deliberately: T1-T4 and all of
Part 2/3 — the root is named, and the order forbids fixing before reporting.**
Query committed: `scripts/grade-calibration-t0.sql`.