Files
vyndr/specs/grade-diagnostic-t0.md
T
builtbykev 249b3e8235 report: grade diagnostic T0 — p_win is MISCALIBRATED, and it explains the inversion
STOPPED at the T0 gate as instructed. Nothing fixed, no recalibration applied,
no grade touched. T1-T4 deliberately not run.

T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS.

Condition 1 (mean |predicted-actual| > 0.05): MLB ~0.094, WNBA ~0.139.
Condition 2 (monotonic slope): over-confidence GROWS with the prediction —
MLB +0.034 -> +0.043 -> +0.084 -> +0.190 -> +0.189; WNBA +0.044 -> +0.109 ->
+0.349. Worst cases: MLB predicted 0.842 actual 0.652 (n=23), predicted 0.917
actual 0.727 (n=11); WNBA predicted 0.730 actual 0.381 (n=21).

WHY THIS EXPLAINS THE INVERSION, mechanically: p_win is over-stated and the
overstatement SCALES with p_win, so p_win - fair_prob_lock is largest exactly
where p_win is most inflated. Those props hit less than claimed, so the edge
measure correlates negatively. The market was never the problem —
fair_prob_lock is not a bent ruler, the thing subtracted from it is. It also
explains why p_win ALONE still carries signal (+0.23 MLB): rank survives
miscalibration, differences do not.

This independently reconfirms the 2026-07-26 calibration finding (+0.02 at p<.5
-> +0.19 at p>=.8) on a newer, larger population, so it is structural rather
than sampling noise.

PART 0: P0a — only the GRADED side's fair prob is stored (fair_prob_lock;
no opposite-side field), so T1's two-side-sum check cannot run and must use the
stated no-vig recompute fallback. P0b — projection_locked_at exists as a
timestamptz so T2 is potentially runnable, but distinctness from lock time was
NOT verified because T0 gated it.

Two cautions recorded before Part 2 runs: the top MLB buckets where the error is
worst hold n=23 and n=11, so a flexible per-bucket correction would fit noise —
isotonic with pooling or single-parameter Platt is safer; and calibration fixes
magnitudes, so if the market is genuinely better the repaired edge may still land
at ~0, which would be the honest ceiling and gets reported rather than graded
around.

Query committed at scripts/grade-calibration-t0.sql.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 20:12:20 -04:00

4.4 KiB
Raw Blame History

GRADE DIAGNOSTIC v2 — PART 0 + T0 (report; STOPPED at the T0 gate)

2026-07-31. Nothing fixed, no recalibration applied, no grade touched.

🔴 T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS — p_win IS MISCALIBRATED

Per the order: "if T0 fires, THAT is the root — the market comparison is premature." It fired, so T1T4 were not run and no fix was attempted.

The calibration curve (n=442 decided rows with p_win, buckets with n>=10)

sport predicted actual over-confidence n
MLB 0.352 0.286 +0.066 28
MLB 0.462 0.513 0.051 39
MLB 0.545 0.511 +0.034 47
MLB 0.639 0.596 +0.043 52
MLB 0.751 0.667 +0.084 39
MLB 0.842 0.652 +0.190 23
MLB 0.917 0.727 +0.189 11
WNBA 0.460 0.512 0.052 41
WNBA 0.544 0.500 +0.044 74
WNBA 0.646 0.537 +0.109 41
WNBA 0.730 0.381 +0.349 21

Condition 1 — mean |predicted actual| > 0.05: FIRES. MLB mean |dev| ≈ 0.094; WNBA ≈ 0.139. Both well past the 0.05 threshold.

Condition 2 — monotonic over/under-confidence slope: FIRES. Over-confidence grows with the prediction: MLB +0.034 → +0.043 → +0.084 → +0.190 → +0.189; WNBA +0.044 → +0.109 → +0.349. The model is roughly honest near a coin flip and badly over-stated at the top.

WHY THIS MECHANICALLY EXPLAINS THE INVERSION

This is the causal story the previous order was missing, and the data supports it end to end:

  1. p_win is over-stated, and the overstatement SCALES with p_win.
  2. p_win fair_prob_lock is therefore largest exactly where p_win is most inflated — the "biggest edges" are the most over-confident predictions.
  3. Those props hit less than claimed (0.842 → 0.652; 0.917 → 0.727; WNBA 0.730 → 0.381), so the edge measure correlates negatively.

The market was never the problem. fair_prob_lock is not a bent ruler — the thing being subtracted from it is. That also explains why p_win alone still carries signal (+0.23 MLB): the ranking is partly right even though the magnitudes are wrong. Rank survives miscalibration; differences do not.

This independently reconfirms the 2026-07-26 diagnosis (+0.02 at p<.5 → +0.19 at p>=.8) on a newer, larger population — so it is a persistent structural property, not sampling noise.

PART 0 — SCHEMA PRE-CHECK (both premises resolved)

  • P0a — only the GRADED side's fair prob is stored. ledger_entries has fair_prob_lock (graded side), proj_book_implied, dclv_fair_lock/close — but no opposite-side fair prob. So T1's two-side-sum de-vig check cannot run; it must use the stated fallback (independent no-vig recompute from locked_odds) if T1 is ever reached.
  • P0b — projection_locked_at EXISTS as a timestamptz, so T2 is potentially runnable, but whether it is genuinely distinct from lock time was not verified (T0 gated it). CANNOT DETERMINE until T2 is authorised.

WHAT THIS MEANS FOR PART 2 (not done — awaiting your go)

The named cause is T0: calibrate p_win. The order's prescribed fix applies — isotonic or Platt on a training split, proven on a held-out/forward split, then re-measure the market-relative edge on the repaired instrument.

Two honest cautions I want on the record before that runs:

  1. Sample is thin for a calibration map. The top MLB buckets — where the error is worst and the fix matters most — hold n=23 and n=11. Fitting a correction there risks fitting noise. A monotone (isotonic) fit with pooling, or a single-parameter Platt scaling, is far safer than a flexible per-bucket map.
  2. Recalibration may not resurrect the edge. Calibrating fixes the magnitudes; if the market is genuinely better than the model, p_win_calibrated fair_prob_lock can still land at ~0. That would be the honest ceiling, and per the order it gets reported plainly rather than graded around.

TAGS

VERIFIED: T0 fires on both pre-registered conditions; the monotone over-confidence slope; P0a (single-side fair prob only). CANNOT DETERMINE: P0b timestamp distinctness (T0 gated T2). BLOCKED by the gate, deliberately: T1-T4 and all of Part 2/3 — the root is named, and the order forbids fixing before reporting.

Query committed: scripts/grade-calibration-t0.sql.