Power-derive the LODO threshold: hits restored through the gate, rbi/runs
routed as date-driven PHASE 0 — threshold derived BLIND, before any stat was re-read. A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count. Per-row Brier difference d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single stat's verdict could shape the threshold deciding it: pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175 n* = (0.09816/0.01175)^2 = 69.8 -> 70 The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect this removes, and why the previous verdict moved with the number. Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it equals the constant, so it cannot drift from its own justification. The derivation script prints no stat verdict, no date and no reversal. PHASE 1 — LODO at n*, applied cold: hits 5 informative drops, 0 reversals PASS total_bases 4 informative drops, 0 reversals PASS rbi reverses 2026-08-01 (n=99) FAIL runs reverses 08-01 (n=86), 08-05 (244) FAIL hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across 123-272 row dates, favourite sign holding on every testable drop. THIS IS THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the withdrawal at6ae11f1was correct on the instrument available then, which admitted 20- and 25-row dates as evidence. Nothing about hits changed; the threshold stopped being chosen. PHASE 2 — both failures are DATE-DRIVEN, not underpowered. Every reversal sits above n*=70 (99, 86, 244), so no threshold and no further accrual rescues either: isotonic is fitting day-structure. Routed to the low-parameter calibrator queue (Platt/beta), not built here. PHASE 3 — CALIBRATION_DEPLOYED is now ['hits','total_bases'], frozen and tested, both PROVISIONAL with auto-demotion armed and the >=40 date-cluster promotion bar unchanged. hits stackability for chain.chainAcross is RESTORED, and the record shows it returned through the powered gate rather than by fiat. hits bands rebuilt on p_win_calibrated (765 eval rows): every archetype still one band, still base_rate -- calibrated YES, proven-per-archetype NO. PHASE 4 logged: the deploy set is now set by a power-derived, pre-committed, tested constant rather than an operator-chosen number. At6ae11f1that rule moved the live path AGAINST the operator; it has now moved it back on the same evidence because the instrument changed. Both directions are the rule working. And calibrated p_win separates within archetype no better than raw across 13 archetype slots on two deployed stats -- per-archetype separation will come from proven factors or not at all. p_win never mutated; no Bonferroni slot consumed; counter and frozen clusters verified byte-identical file by file. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
@@ -11,18 +11,31 @@
|
||||
const snapshotService = require('../../src/services/snapshotService');
|
||||
|
||||
describe('the deployed set is LODO-gated', () => {
|
||||
it('serves total_bases — it passed at every held-size threshold', () => {
|
||||
it('serves the stats that passed LODO at the powered threshold', () => {
|
||||
expect(snapshotService.CALIBRATION_DEPLOYED).toContain('total_bases');
|
||||
// hits was withdrawn at 6ae11f1 under a hand-chosen threshold that admitted
|
||||
// 20-row dates as evidence, and is restored here because the powered
|
||||
// instrument finds no reversal. Through the gate, not around it.
|
||||
expect(snapshotService.CALIBRATION_DEPLOYED).toContain('hits');
|
||||
});
|
||||
|
||||
it('does NOT serve hits, rbi or runs — each fails LODO', () => {
|
||||
// hits reverses when 2026-07-22 or 2026-07-26 is dropped; rbi on 2026-08-01;
|
||||
// runs on 2026-08-01 and 2026-08-05.
|
||||
for (const stat of ['hits', 'rbi', 'runs']) {
|
||||
it('does NOT serve rbi or runs — both fail on dates ABOVE the threshold', () => {
|
||||
// These are DATE-DRIVEN failures, not underpowered ones: rbi reverses on a
|
||||
// 99-row date and runs on 86- and 244-row dates. No threshold rescues them.
|
||||
for (const stat of ['rbi', 'runs']) {
|
||||
expect(snapshotService.CALIBRATION_DEPLOYED).not.toContain(stat);
|
||||
}
|
||||
});
|
||||
|
||||
it('every deployed stat cleared the power-derived threshold, not a chosen one', () => {
|
||||
const { LODO_MIN_HELD_ROWS, LODO_THRESHOLD_BASIS } = require('../../src/services/model/calibrationRegistry');
|
||||
expect(LODO_MIN_HELD_ROWS).toBe(70);
|
||||
expect(LODO_THRESHOLD_BASIS.derived_blind).toBe(true);
|
||||
// The reversing dates that keep rbi/runs out are all at or above it, so
|
||||
// their exclusion cannot be an artefact of the threshold.
|
||||
for (const n of [99, 86, 244]) expect(n).toBeGreaterThanOrEqual(LODO_MIN_HELD_ROWS);
|
||||
});
|
||||
|
||||
it('is frozen, so a stat cannot be added at runtime without a code change', () => {
|
||||
expect(Object.isFrozen(snapshotService.CALIBRATION_DEPLOYED)).toBe(true);
|
||||
expect(() => { snapshotService.CALIBRATION_DEPLOYED.push('hits'); }).toThrow();
|
||||
|
||||
Reference in New Issue
Block a user