Files
vyndr/scripts/derive-lodo-test.js
T
builtbykev ced40421ed Audit the LODO instrument: it cannot evaluate any stat, and both prior
FAILs were false

PHASE 0 — the gate at 1f40014 was mine and was an incoherent pair. A 1-SE
informativeness bar with a ZERO-reversal rule: at exactly 1 SE a stable
stat's drop reverses with prob Phi(-1)=0.1587, so on four informative
drops P(>=1 reversal | perfectly stable) = 1 - 0.8413^4 = 0.50. It failed
stable stats half the time by construction. And the pooled n*=70
mis-credited EVERY stat -- too low for hits (own 77) and runs (81), too
high for total_bases (60) and rbi (54).

PHASE 1, blind. Per-stat (g, sigma_row): hits -0.01288/0.11251, TB
-0.01380/0.10680, rbi -0.00884/0.06459, runs -0.00902/0.08080. All four
clear z=1.96 at full n, so none is NO-EFFECT. Committed k=1 with per-stat
n* and a binomial cutoff holding FP at 0.004-0.031.

THE FINDING THAT DOMINATES: the test has no power. Against a strong
instability (date-to-date SD equal to the effect) it detects a failure
1.4%-9.3% of the time, and across every k from 1.0 to 2.0 the best any
stat reaches is 0.337. A gate that cannot fail cannot pass, so
LODO_POWER_FLOOR=0.50 makes UNTESTABLE structural -- "could not test" can
never read as "passed".

PHASE 2/3 cold, at each stat's OWN n*:

  hits  5 informative, 0 reversals, cutoff 2, power 0.093  UNTESTABLE
  TB    5 informative, 0 reversals, cutoff 2, power 0.093  UNTESTABLE
  rbi   4 informative, 1 reversal,  cutoff 2, power 0.045  UNTESTABLE
  runs  3 informative, 2 reversals, cutoff 2, power 0.014  UNTESTABLE

Setting the power floor aside entirely, NOT ONE STAT EXCEEDS ITS CUTOFF.

PHASE 4 — rbi's FAIL was false, as the order suspected. So was RUNS' --
which the order did not anticipate, having classified it DATE-DRIVEN on a
244-row reversal; two reversals in three drops does not clear a cutoff of
2. TB's PASS was vacuous: the test could not have failed it. hits' own n*
is LARGER than the pooled one (77 vs 70), and it remains untestable.

PHASE 5 — deploy basis is now the date-clustered CI alone:

  hits  CI [-0.0139,-0.0097], 4 date clusters   relabelled ci_only
  TB    CI [-0.0061,-0.0045], 2 date clusters   RELABELLED, kept
  rbi   CI [-0.0092,-0.0010], 2 date clusters   NEWLY DEPLOYED
  runs  no fittable map at its split            REFUSE, no CI either

Every deployed stat carries calibration_basis ci_only_lodo_untestable and
auto-demotion is the SOLE stability guard, not a backstop to a passed
test. Stated plainly: those intervals rest on 2-4 date clusters, which is
thin, and it is now the only support. rbi gains chainAcross stackability;
its bands rebuilt on p_win_calibrated (425 rows) are every-archetype
base_rate. runs is queued for the low-param calibrator for the ordinary
reason -- no fittable map -- not on the date-driven finding, which was an
artefact.

PHASE 6 — the deploy set was set by a coin-flip-power ruler; it is now set
by a per-stat power-coherent pre-committed test whose first act was to
report that it cannot evaluate anything. The audit was permitted to wound
the live deploy and did: total_bases lost its LODO claim. Standing
question unchanged -- 18 archetype slots across three deployed stats, every
one a single band indistinguishable from base rate.

Blind ordering held. p_win never mutated. No Bonferroni slot. Counter and
frozen clusters verified file-by-file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 23:01:57 -04:00

199 lines
8.4 KiB
JavaScript
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
#!/usr/bin/env node
'use strict';
/**
* PHASE 1 — derive a COHERENT LODO test, blind to reversals.
*
* ── THE DEFECT BEING FIXED ───────────────────────────────────────────────
* The gate at 1f40014 paired a 1-SE per-drop informativeness bar with a
* zero-reversal decision rule. Those two are incoherent. At exactly 1 SE, a
* genuinely STABLE stat's drop reverses with probability Phi(-1) = 0.159, so on
* four informative drops the chance of at least one reversal is
* 1 - 0.841^4 = 0.50. The rule failed stable stats half the time by construction.
*
* And n* was pooled across four stats whose signed effects differ several-fold,
* so "informative" meant different things for different stats while being
* treated as one number.
*
* ── THE FIX ──────────────────────────────────────────────────────────────
* The two halves have to be chosen together:
*
* informative bar n*_k = k^2 * (sigma_row / |g|)^2 PER STAT
* decision rule FAIL iff reversals > c, where under stability
* R ~ Binomial(D, Phi(-k)) and c is the smallest cutoff
* with P(R > c) <= 0.05
*
* `g` is the mean SIGNED per-row improvement — the quantity whose sign a
* reversal flips. Not a mean-absolute, and not pooled: a reversal is a claim
* about THIS stat's effect changing sign.
*
* THIS SCRIPT PRINTS NO REVERSAL AND NO VERDICT. It is blind by construction and
* must run, and its constants be committed, before any stat is re-read.
*
* SUPABASE_URL=... node scripts/derive-lodo-test.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** |g| must clear this many SE at the stat's full n or there is no effect to test. */
const EFFECT_Z = 1.96;
/** Target false-positive rate for the whole per-stat test. */
const TARGET_FP = 0.05;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Standard normal CDF (AbramowitzStegun 7.1.26 via erf). */
function normCdf(z) {
const t = 1 / (1 + 0.2316419 * Math.abs(z));
const d = 0.3989422804014327 * Math.exp(-z * z / 2);
const p = d * t * (0.319381530 + t * (-0.356563782 + t * (1.781477937 + t * (-1.821255978 + t * 1.330274429))));
return z >= 0 ? 1 - p : p;
}
const binomPmf = (n, k, p) => {
let logC = 0;
for (let i = 0; i < k; i += 1) logC += Math.log(n - i) - Math.log(i + 1);
return Math.exp(logC + k * Math.log(p) + (n - k) * Math.log(1 - p));
};
/** P(R > c) for R ~ Binomial(n, p). */
const binomTail = (n, c, p) => {
let s = 0;
for (let k = c + 1; k <= n; k += 1) s += binomPmf(n, k, p);
return s;
};
/** Smallest cutoff c with P(R > c) <= target. */
function cutoffFor(D, p, target) {
for (let c = 0; c <= D; c += 1) if (binomTail(D, c, p) <= target) return { cutoff: c, fp: binomTail(D, c, p) };
return { cutoff: D, fp: 0 };
}
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
const perStat = {};
const dateSizes = {};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b));
if (v === null) continue;
const over = v > L;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
const map = cal.fitIsotonic(rows.map((r) => ({ p: r.p, won: r.won })));
if (!map) { perStat[stat] = { n: rows.length, fittable: false }; continue; }
const d = [];
for (const r of rows) {
const pc = cal.applyIsotonic(map, r.p);
if (knownNumber(pc) === null) continue;
d.push((pc - r.won) ** 2 - (r.p - r.won) ** 2);
}
const g = mean(d);
const sigma = Math.sqrt(d.reduce((s, x) => s + (x - g) ** 2, 0) / (d.length - 1));
const seFull = sigma / Math.sqrt(d.length);
// Date sizes are sample STRUCTURE, not outcomes — safe to read here.
const sizes = new Map();
for (const r of rows) sizes.set(r.date, (sizes.get(r.date) || 0) + 1);
dateSizes[stat] = [...sizes.values()].sort((a, b) => b - a);
perStat[stat] = {
n: d.length,
g_signed: round5(g),
sigma_row: round5(sigma),
se_full: round5(seFull),
effect_z_at_full_n: round3(Math.abs(g) / seFull),
improves: g < 0,
no_effect: Math.abs(g) / seFull < EFFECT_Z,
fittable: true,
};
}
// ── Choose k jointly. Blind: uses only (g, sigma) and date SIZES. ──
const kTable = [];
for (const k of [1.0, 1.25, 1.5, 1.75, 2.0]) {
const pNoise = normCdf(-k);
const row = { k, per_drop_noise_prob: round4(pNoise), stats: {} };
for (const stat of STATS) {
const ps = perStat[stat];
if (!ps || !ps.fittable) continue;
const nStar = Math.ceil(k * k * (ps.sigma_row / Math.abs(ps.g_signed)) ** 2);
const D = (dateSizes[stat] || []).filter((n) => n >= nStar).length;
const { cutoff, fp } = D > 0 ? cutoffFor(D, pNoise, TARGET_FP) : { cutoff: null, fp: null };
// FN at a stated alternative: date-to-date SD of the effect equals |g|.
const pAlt = D > 0 ? normCdf(-Math.abs(ps.g_signed) / Math.sqrt(ps.g_signed ** 2 + (ps.sigma_row ** 2) / nStar)) : null;
const fn = D > 0 && cutoff !== null ? 1 - binomTail(D, cutoff, pAlt) : null;
row.stats[stat] = {
n_star: nStar, informative_drops: D, cutoff, fp: fp === null ? null : round4(fp),
fn_at_tau_equals_g: fn === null ? null : round4(fn),
};
}
kTable.push(row);
}
console.log(JSON.stringify({
phase: 'PHASE 1 — coherent LODO test derivation, BLIND',
defect_being_fixed: 'a 1-SE informative bar with a zero-reversal rule: P(>=1 reversal | stable, 4 drops) = 0.50',
per_stat_effect: perStat,
date_sizes: dateSizes,
k_selection_table: kTable,
target_fp: TARGET_FP,
blind: 'no reversal, no verdict, no reversing date referenced anywhere in this output',
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round3 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 1000) / 1000);