Files
vyndr/specs/pwin-recalibration-holdout.md
T
builtbykev 6d87d7a33c report: p_win recalibration holdout — MLB qualifies on isotonic, WNBA abstains
Measure-only. p_win not flipped live, no grade rebuilt, no calibrator deployed.
Per the doctrine, MLB and WNBA were fitted, selected and judged as SEPARATE
models — and they reach opposite verdicts. No global instrument was fitted.

METHOD: time-forward split per sport (earlier fits, later proves). Both
instruments fitted on TRAIN only — single-parameter Platt and
isotonic-with-pooling. Inputs p_win + outcome only; no market field, no closing
value, no lookahead. Nothing about edge/CLV/beat-the-close enters any pass/fail
line.

MEASUREMENT CORRECTION made mid-run: the first pass reported mean|p - outcome|
(~0.46-0.51), which is NOT calibration — it is noise-dominated individual error
on 0/1 rows and would have made every instrument look identical. Reliability is
only meaningful on BUCKETS (bucket mean predicted vs bucket actual rate,
n-weighted), the metric T0 used. All reported numbers use the corrected metric.

HOLDOUT RELIABILITY (lower better): MLB n=119/4 buckets — raw 0.1038, Platt
0.1120, ISOTONIC 0.0939. WNBA n=93/3 buckets — raw 0.1322, Platt 0.0491,
isotonic 0.0667.
HOLDOUT RESOLUTION: MLB raw 0.1388 -> Platt 0.1284 -> isotonic 0.1225.
WNBA raw -0.1201 -> Platt +0.1269 -> isotonic +0.0322.
Fitted Platt: MLB a=-0.381 b=+0.705; WNBA a=+0.040 b=-0.081.

MLB QUALIFIES, MODESTLY — instrument selected BY HOLDOUT, not assumed: isotonic
beats both raw and Platt, and Platt actually made MLB worse. Reliability improves
0.1038 -> 0.0939 (~10% relative, real but modest) and resolution SURVIVES
(0.1388 -> 0.1225, not crushed). Both Mandate-3 conditions hold.

WNBA ABSTAINS — its Platt result is the best number in the report and is REJECTED
as a fake win. The fitted slope is b = -0.081, negative and near zero, so
sigmoid(0.040 - 0.081*logit p) is nearly constant at ~0.51 for every input: it
"calibrates" by discarding the prediction and emitting the base rate, which is
exactly the failure Mandate 3 pre-registered. Its apparent resolution gain
(-0.120 -> +0.127) is the sign flip, not skill — it would serve the opposite of
its own forecast, fitted on n~96 of anti-signal. Isotonic says the same quietly
(resolution collapses to +0.032).

HONEST CEILING: MLB is a usable-but-unimpressive forecaster (holdout resolution
~0.12, reliability ~0.094, n=119); WNBA has no honest forecast today. Holdout n
and bucket counts (4 and 3) suffice to reject WNBA and prefer isotonic for MLB,
NOT to certify a letter ladder, and the T0 pathology is reduced rather than cured.

CANNOT DETERMINE: per-archetype calibration (Mandate 3d) — bucket n falls below
the reporting floor once split by sport AND archetype on 442 rows.

Queries committed at scripts/pwin-calibration-holdout.sql.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 21:26:21 -04:00

4.9 KiB
Raw Blame History

RECALIBRATE p_win — HOLDOUT REPORT (measure-only; NOTHING DEPLOYED)

2026-07-31. p_win not flipped live, no grade rebuilt, no calibrator shipped. Per PART A doctrine: MLB and WNBA are treated as SEPARATE MODELS throughout — fit separately, selected separately, judged separately, and they reach opposite verdicts. No global instrument was fitted.

METHOD

Time-forward split PER SPORT (earlier half fits, later half proves — never the same rows). Both instruments fitted on TRAIN only: single-parameter Platt (p_cal = sigmoid(a + b·logit p), a/b from a weighted fit of TRAIN bucket empirical-logit on bucket mean-logit) and isotonic-with-pooling (TRAIN bucket rate applied to HOLDOUT by bucket, min 8 per bucket). Inputs: p_win + outcome only — no market field, no closing value, no lookahead.

⚠️ A MEASUREMENT CORRECTION I MADE MID-RUN

My first pass reported mean|p outcome| (~0.460.51). That is not calibration — it is individual-level error on 0/1 outcomes, dominated by irreducible noise, and it would have made every instrument look identical and useless. Reliability is only meaningful on buckets: bucket mean predicted vs bucket actual rate, n-weighted — the same metric T0 used. All numbers below are the corrected, bucketed metric.

RESULTS ON HOLDOUT

Reliability (n-weighted |predicted actual| across holdout buckets; lower better)

sport holdout n buckets raw Platt isotonic
MLB 119 4 0.1038 0.1120 0.0939
WNBA 93 3 0.1322 0.0491 0.0667

Resolution (does the ordering survive? Pearson r with outcome, holdout)

sport raw Platt isotonic
MLB 0.1388 0.1284 0.1225
WNBA 0.1201 +0.1269 +0.0322

Fitted Platt parameters: MLB a=0.381, b=+0.705 · WNBA a=+0.040, b=0.081.


VERDICT — MLB: QUALIFIES, MODESTLY. INSTRUMENT = ISOTONIC.

Selected by holdout, not assumed (the order was right to forbid assuming Platt): isotonic 0.0939 beats raw 0.1038 and beats Platt 0.1120 — Platt made MLB worse.

  • (a) Reliability: improves. 0.1038 → 0.0939, ~10% relative. Real but modest — this does not make an 84%-that-hits-65% into a truthful forecast; it shaves the exaggeration.
  • (b) Resolution: SURVIVES. 0.1388 → 0.1225. Slightly lower (expected — a monotone map preserves rank; the Pearson dip is the non-linearity), and critically it is not crushed to zero. High p_win still hits more than low.
  • Both (a) and (b) hold → MLB qualifies to be graded in a later, separately gated order.

🔴 VERDICT — WNBA: ABSTAINS. ITS PLATT "WIN" IS A FAKE WIN.

WNBA's Platt reliability (0.0491) is the best number in this report and must be rejected, exactly per Mandate 3.

Why it is fake: the fitted slope is b = 0.081 — NEGATIVE and near zero. So sigmoid(0.040 0.081·logit p) is nearly constant at ≈0.51 for every input. It "calibrates" by discarding the prediction and emitting the base rate. That is precisely the failure Mandate 3 pre-registered: "a model calibrated by crushing everything to the base rate has NO resolution → FAKE win → reject."

The apparent resolution improvement (0.120 → +0.127) is the sign flip, not skill: a negative slope inverts the ordering, so the model would be serving the opposite of its own forecast — fitted on n≈96 of anti-signal. Isotonic tells the same story more quietly (resolution collapses to +0.032, i.e. no ordering).

WNBA cannot be made honest AND predictive on this data. It abstains. Per Mandate 4 that is two models — one ready, one not — not one model half-working.


THE HONEST CEILING, STATED PLAINLY

MLB's forecast can be made less dishonest and keeps its (modest) ordering: holdout resolution ≈ 0.12, reliability ≈ 0.094 on n=119. That is a usable-but-unimpressive forecaster, not a sharp one. WNBA has no honest forecast at all today.

Cautions carried forward: holdout n is 119/93 with only 4 and 3 usable buckets — enough to reject WNBA and to prefer isotonic for MLB, not enough to certify a letter ladder. And the T0 pathology is only reduced, not cured: the top MLB buckets remain the thinnest and the most exaggerated.

Nothing about market edge, CLV, or beating the close appears in any pass/fail line above — as the order required. The instruments were judged only on whether the forecast tells the truth and still ranks.

TAGS

VERIFIED: per-sport time-forward holdout results; isotonic beats Platt and raw on MLB; WNBA's Platt win is a base-rate collapse with a negative slope. CANNOT DETERMINE: per-archetype calibration (Mandate 3d) — bucket n falls below the reporting floor once split by sport AND archetype on 442 rows. DEPLOYED: NOTHING. Queries committed at scripts/pwin-calibration-holdout.sql.