Read integrity, as-of context, and the shadow matchup resolve (A1-A7)

Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Kev
2026-08-11 22:49:56 -04:00
parent 387ae4d54e
commit f61ec6b391
49 changed files with 4874 additions and 308 deletions
+54
View File
@@ -0,0 +1,54 @@
'use strict';
/**
* MODEL VERSION — ONE source, because two were not the same.
*
* ── WHAT WENT WRONG ──────────────────────────────────────────────────────
* `ledgerService` and `retentionService` both read `process.env.MODEL_VERSION`,
* and both carried a HARDCODED DEFAULT — but different ones:
*
* ledgerService.js:57 … || 'engine1@2026-07-20'
* retentionService.js:59 … || 'engine1@2026-08-07-fullwindow'
*
* Production does not set the env var, so the two tables took different
* defaults. The 2026-08-07 champion repair bumped one and left the other behind.
* Result: `model_snapshots` carried the repaired marker on 34,128 rows while
* `ledger_entries` carried the OLD marker on all 15,739 — including rows graded
* by the repaired champion.
*
* That is not a cosmetic mismatch. `reAuditEligibility.isEligible` requires
* `model_version === REPAIRED_CHAMPION_VERSION`, so applied to the ledger it
* returned ZERO eligible rows and ZERO eligible dates, permanently, for all four
* pending measurements. The accrual clock read "blocked" when the real state was
* "mis-stamped".
*
* ── THE RULE ─────────────────────────────────────────────────────────────
* There is exactly one place a model version may be declared. Anything that
* stamps a row imports it from here. A default that lives next to its consumer
* will drift from the other consumer, and the drift is invisible because both
* sides look locally correct.
*
* ── WHAT THIS DOES NOT DO ────────────────────────────────────────────────
* It does not re-stamp history. Rows already written keep the marker they were
* written with — rewriting them would destroy the one record of which forecast
* actually produced them, which is the thing the marker exists to preserve. Old
* rows stay honestly old; a backfill is a separate, explicit decision.
*/
/** The version every NEW row is stamped with. Override with MODEL_VERSION. */
const MODEL_VERSION = process.env.MODEL_VERSION || 'engine1@2026-08-07-fullwindow';
/**
* Rows at or after this marker were produced by the repaired champion and are
* eligible for forward re-audit.
*
* Deliberately NOT `MODEL_VERSION`: the eligibility bar is a fixed historical
* fact about a specific repair, and tying it to "whatever we stamp today" would
* make every future version silently re-qualify itself.
*/
const REPAIRED_CHAMPION_VERSION = 'engine1@2026-08-07-fullwindow';
/** The marker that preceded the repair — kept so old rows are nameable. */
const PRE_REPAIR_VERSION = 'engine1@2026-07-20';
module.exports = { MODEL_VERSION, REPAIRED_CHAMPION_VERSION, PRE_REPAIR_VERSION };
+7
View File
@@ -101,8 +101,15 @@ async function gradeBestSide(grade, prop, sport, opts = {}) {
// previously computed downstream of the grade it should inform.
const factorContext = typeof opts.factorContext === 'function'
? opts.factorContext(prop, sport) : null;
// FIX A6 — the SHADOW pair. `factor_context` (above) is the LIVE context and
// is unchanged; these two let the engine compute what the factors WOULD do
// with the join keys, and freeze it. Neither reaches the served forecast.
const matchupKeys = typeof opts.matchupKeys === 'function'
? opts.matchupKeys(prop, sport) : null;
const base = {
factor_context: factorContext,
factor_context_resolver: typeof opts.factorContext === 'function' ? opts.factorContext : null,
matchup_keys: matchupKeys,
player: prop.player,
stat_type: prop.stat_type,
line: prop.line,
@@ -567,6 +567,27 @@ async function analyzeViaEngine1(rawProp = {}) {
pOver = adj.p_adjusted;
factorTrace = { multiplier: adj.multiplier, applied: adj.applied, skipped: adj.skipped, p_before: adj.p_base };
}
// ── FIX A5 — FREEZE THE INPUTS, NOT THE OUTPUT ──────────────────
// Captured from the SAME context object the factor was just handed, in
// the same pass, so the recorded value cannot drift from the used one.
// Purely additive: it reads `rawProp.factor_context` and writes a field.
// `pOver` above is untouched by this block, so the served grade cannot
// move — a freeze records, it does not compute.
try {
const ffz = require('../model/factorFreeze');
// ── FIX A6 — SHADOW RESOLVE ──────────────────────────────────
// With the join keys the factors CAN fire. We compute what they
// WOULD apply and FREEZE it. `pOver` is not touched by this block:
// turning them on for real is a separate, gated decision.
let shadow = null;
const keys = rawProp.matchup_keys || null;
if (keys && (keys.opponent || keys.opposing_pitcher)
&& typeof rawProp.factor_context_resolver === 'function') {
const shadowCtx = rawProp.factor_context_resolver(rawProp, rawProp.sport, keys);
if (shadowCtx) shadow = ffz.shadowBlock(shadowCtx, hf.hitsFactorMultiplier(shadowCtx), keys);
}
legacy.factor_inputs = ffz.freeze(rawProp.factor_context, rawProp, shadow);
} catch { /* recording must never break a grade either */ }
} catch { /* a factor must never break the grade */ }
}
+5 -1
View File
@@ -54,7 +54,11 @@ const SETTLE_ATTEMPT_CAP = Number(process.env.SETTLE_ATTEMPT_CAP || 4);
// from originals and the harness can filter by how a row was scored.
const SETTLEMENT_VERSION = Number(process.env.SETTLEMENT_VERSION || 2);
const settleSource = require('./settleSource');
const MODEL_ERA_VERSION = process.env.MODEL_VERSION || 'engine1@2026-07-20';
// FIX A3 — ONE SOURCE. This used to default to 'engine1@2026-07-20' while
// retentionService defaulted to the repaired version, so the same env var
// produced two different stamps and no ledger row ever carried the marker
// `reAuditEligibility` requires. Never re-declare a version default here.
const { MODEL_VERSION: MODEL_ERA_VERSION } = require('../config/modelVersion');
// Order Zero — which fair-probability RULER produced fair_prob_lock. It is the
// denominator of every edge and CLV number, so a value computed under one ruler
// is not the same measurement as one computed under another. Never pool across
+45 -14
View File
@@ -29,6 +29,7 @@
*/
const cal = require('./calibration');
const { paginate } = require('../../utils/safePaginate');
const { knownNumber } = require('../../utils/known');
/** Default: hold out the most recent quarter of history to certify on. */
@@ -100,24 +101,54 @@ function build(settled, opts = {}) {
* Injectable for tests; returns null rather than a permissive fallback, because
* "no calibrator" must mean "nothing is stackable", not "pass the raw numbers
* through".
*
* ── FIX A1 (2026-08-09) — THIS READ WAS 24.8% CORRUPT ────────────────────
* It walked pages with `.range()` and NO ORDER BY, and swallowed a query error
* as end-of-data. Measured on production: 617 of 2,490 rows returned twice and
* an equal number never returned, while `rows.length` matched the server count
* exactly. A map fitted on that is fitted on a sample where a fifth of the
* history is double-weighted and another fifth is absent.
*
* Both defects are now `safePaginate`: a stable ORDER BY on the unique `id`
* (an order on a non-unique column is not a fix — ties still scramble), plus a
* thrown error instead of a silent stop.
*
* A THROW IS NOT A NULL HERE, AND THE DIFFERENCE IS LOAD-BEARING. `null` means
* "not enough settled history to fit" — a real, expected state. A throw means
* "the read failed", which used to be indistinguishable from the first. It
* propagates so a caller can never mistake a broken read for thin history and
* quietly serve a map built on a fragment.
*/
async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) {
if (!sb) return null;
const cutoff = before || new Intl.DateTimeFormat('en-CA', {
function todayEt() {
return new Intl.DateTimeFormat('en-CA', {
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
}).format(new Date());
const rows = [];
for (let from = 0; ; from += 1000) {
const { data, error } = await sb.from('ledger_entries')
.select('p_win, outcome, game_date, quarantine_reason')
}
/**
* The ROW LOAD, exported so the read-integrity harness can measure THE REAL
* FUNCTION rather than a re-declaration of its query.
*
* That distinction is the whole point of the harness: a spec that restates the
* query would verify the restatement, and a flag saying "this one is fixed now"
* would verify a comment. The harness calls this.
*/
async function loadSettledRows(sb, { sport = 'mlb', stat = 'hits', before = null } = {}) {
const cutoff = before || todayEt();
return paginate(
() => sb.from('ledger_entries')
.select('id, p_win, outcome, game_date, quarantine_reason')
.eq('sport', sport).is('user_id', null).eq('stat', stat)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null)
.lt('game_date', cutoff) // STRICTLY before — the whole point
.range(from, from + 999);
if (error || !data || data.length === 0) break;
rows.push(...data);
if (data.length < 1000) break;
}
.lt('game_date', cutoff), // STRICTLY before — the whole point
{ key: 'id', pageSize: 1000, label: `calibrationService.fromLedger(${sport}/${stat})` },
);
}
async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) {
if (!sb) return null;
const cutoff = before || todayEt();
const rows = await loadSettledRows(sb, { sport, stat, before: cutoff });
const clean = rows
.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'))
.map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, date: String(r.game_date) }));
@@ -125,4 +156,4 @@ async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...
return built ? { ...built, cutoff } : null;
}
module.exports = { build, fromLedger, HOLDOUT_FRACTION, MIN_FIT };
module.exports = { build, fromLedger, loadSettledRows, HOLDOUT_FRACTION, MIN_FIT };
+164
View File
@@ -0,0 +1,164 @@
'use strict';
/**
* factorFreeze — record WHAT THE FACTOR READ, at the moment it read it.
*
* Spec: specs/read-integrity-harness.md §10
*
* ── WHY ──────────────────────────────────────────────────────────────────
* A4 established that a re-audit joining today's context to a past grade is
* contamination, and fixed it with an as-of cutoff. But an as-of join still
* RECONSTRUCTS: it asks "what would the context have been" rather than "what did
* the grader actually see". Those differ whenever a context table was refreshed
* late, backfilled, or corrected. The only way a row becomes independently
* checkable is if it carries its own inputs.
*
* ── INPUTS, NEVER OUTPUTS ────────────────────────────────────────────────
* This freezes the RAW values the factor consumed — spray shares, the opposing
* team's positional OAA, the pitcher's hard-hit rate, the platoon split counts,
* handedness — and deliberately NOT the resulting multiplier. A stored
* multiplier is unfalsifiable: it can only be compared to itself. Stored inputs
* can be re-run through `hitsFactors` and the answer checked, which is what
* makes an audit an audit. `recompute()` below exists precisely so that check is
* one call.
*
* ── IT RECORDS ABSENCE AS ABSENCE ────────────────────────────────────────
* Measured on 596 real graded hits props (2026-08-10): `positionOaa`, `throws`
* and `pitcherHardHit` resolved on ZERO of them, because nothing in the pipeline
* ever sets `prop.opponent` or `prop.opposing_pitcher` — the two join keys the
* resolver needs. All three factors were skipped on 596/596 and the multiplier
* was exactly 1 every time.
*
* So today this vector mostly freezes nulls, and that is the correct behaviour:
* it is the evidence that the factors are wired but inert. `available` records
* what COULD have been joined from the row, so a later order can measure what
* plumbing those keys would change before changing it. `available` is never fed
* to a factor — reading it here would silently alter the served grade.
*/
const { knownNumber } = require('../../utils/known');
/** Version the shape, so a later change is distinguishable from a null. */
const FREEZE_VERSION = 'hits-factors@1';
const num = (v) => knownNumber(v);
/** The spray shares a defence read consumes — raw, not the multiplier. */
function sprayInputs(spray) {
if (!spray) return null;
return {
as_of: spray.as_of_date || null,
pull_gb: num(spray.pull_gb), straight_gb: num(spray.straight_gb), oppo_gb: num(spray.oppo_gb),
pull_air: num(spray.pull_air), straight_air: num(spray.straight_air), oppo_air: num(spray.oppo_air),
};
}
/** Platoon split counts — the raw PAs and hits, so severity is re-derivable. */
function platoonInputs(splits) {
if (!splits) return null;
const side = (s) => (s ? { pa: num(s.pa), ab: num(s.atBats), hits: num(s.hits) } : null);
return { vl: side(splits.vl), vr: side(splits.vr) };
}
/**
* Build the frozen vector from the SAME context object the factor was handed.
*
* @param {object} ctx the resolved factor context (exactly what hitsFactors got)
* @param {object} prop the graded prop, for the join keys as they stood
*/
/**
* FIX A6 — the SHADOW would-fire block.
*
* With the join keys resolved, the factors CAN compute. This records what they
* WOULD have applied, namespaced under `would_fire`, and it is never added to
* the served forecast. It is evidence for a later, gated decision — not a grade.
*
* `multiplier` is stored here (unlike the inputs rule) because the shadow inputs
* are stored ALONGSIDE it under `shadow_inputs`, so it stays recomputable and
* checkable. A number you cannot re-derive is the thing that rule forbids.
*/
function shadowBlock(shadowCtx, applied, keys) {
if (!shadowCtx || !applied) return null;
return {
keys: {
opponent: (keys && keys.opponent) || null,
opposing_pitcher: (keys && keys.opposing_pitcher) || null,
source: (keys && keys.source) || null,
refused: (keys && keys.refused) || null,
},
multiplier: applied.multiplier,
factors_fired: applied.factors_fired,
per_factor: applied.applied.map((a) => ({ factor: a.factor, multiplier: a.multiplier })),
skipped: applied.skipped.map((sk) => sk.factor),
shadow_inputs: {
position_oaa: shadowCtx.positionOaa || null,
throws: shadowCtx.throws || null,
pitcher_hard_hit: knownNumber(shadowCtx.pitcherHardHit),
},
};
}
function freeze(ctx, prop = {}, shadow = null) {
if (!ctx) return null;
return {
v: FREEZE_VERSION,
as_of: ctx.as_of || null,
// ── the join keys, as the RESOLVER saw them ──
// Null here is the finding, not an omission: nothing sets these, so every
// opponent-side and pitcher-side input below is null as a consequence.
join: {
opponent: prop.opponent || prop.opp_team || null,
opposing_pitcher: prop.opposing_pitcher || null,
},
// ── the raw inputs each factor consumed ──
spray: sprayInputs(ctx.spray),
position_oaa: ctx.positionOaa || null,
bats: ctx.bats || null,
throws: ctx.throws || null,
pitcher_hard_hit: num(ctx.pitcherHardHit),
platoon: platoonInputs(ctx.platoonSplits),
// Carried from the graded row, never invented (A4).
archetype: ctx.archetype || null,
// ── WHAT WAS AVAILABLE BUT NOT USED ──
// Recorded so a later order can measure the effect of plumbing the join keys
// BEFORE it changes a served number. Never read by a factor.
available: {
team: prop.team || null,
game_id: prop.game_id || null,
game_date: prop.game_date || null,
},
// SHADOW ONLY (A6). Never read by the served grade.
would_fire: shadow || null,
};
}
/**
* Re-run the factors from a FROZEN vector.
*
* This is the whole point of freezing inputs rather than a multiplier: an audit
* can recompute and compare. Returns the same shape `hitsFactorMultiplier` does.
*/
function recompute(frozen, deps = {}) {
if (!frozen) return null;
const hf = deps.hitsFactors || require('./hitsFactors');
return hf.hitsFactorMultiplier({
spray: frozen.spray,
positionOaa: frozen.position_oaa,
bats: frozen.bats,
throws: frozen.throws,
pitcherHardHit: frozen.pitcher_hard_hit,
platoonSplits: frozen.platoon
? {
vl: frozen.platoon.vl ? { pa: frozen.platoon.vl.pa, atBats: frozen.platoon.vl.ab, hits: frozen.platoon.vl.hits } : null,
vr: frozen.platoon.vr ? { pa: frozen.platoon.vr.pa, atBats: frozen.platoon.vr.ab, hits: frozen.platoon.vr.hits } : null,
}
: null,
});
}
module.exports = { freeze, recompute, shadowBlock, FREEZE_VERSION, sprayInputs, platoonInputs };
+87 -21
View File
@@ -14,10 +14,38 @@
* Every load is best-effort: a missing table yields an empty index, the factor
* finds nothing readable, and the forecast is served unadjusted. A factor layer
* must never be able to break the pipeline it rides in.
*
* ── AS-OF (Fix A4) ───────────────────────────────────────────────────────
* These tables are dated snapshots — one row per entity per day. Without a
* cutoff, `latestBy` returns TODAY'S row, so re-auditing a 2026-08-07 grade
* would join a profile built from games played AFTER it. That is contamination
* through a perfectly clean read, and it is the reason a re-audit could not be
* trusted even once A1/A2b made every reader return the rows it thinks.
*
* `opts.asOf` bounds every dated read to `as_of_date <= asOf` and then takes the
* latest WITHIN that bound.
*
* REFUSAL OVER RECONSTRUCTION. If an entity has no row at or before `asOf`, its
* factor input is null and the factor does not apply — the base rate is left
* untouched. It NEVER falls back to the nearest or the latest row: a substituted
* row is a plausible wrong value wearing a date, which is worse than an absence
* because nothing downstream can see it.
*
* THE LIVE PATH IS UNTOUCHED. With no `asOf` the code below runs exactly as it
* did — same tables, same filters, same indexes — so the served grade cannot
* move. Only an explicitly dated call takes the audit path.
*
* STATCAST IS THE EXCEPTION THAT PROVES THE RULE. `statcast_aggregates` is
* upserted in place and keeps ONE as-of date, so it cannot answer an as-of
* question at all; the dated path reads `statcast_history` instead. Verified
* equivalent at the head: on 2026-08-09 the two agree on all 1,414 rows with 0
* differences, so switching sources does not itself move a number.
*/
const { knownNumber } = require('../../utils/known');
const { nameKey } = require('../../utils/playerName');
const { paginate } = require('../../utils/safePaginate');
const { uniqueKeyFor } = require('../../utils/tableKeys');
/** Rows a factor table must have before we trust it at all. */
const MIN_ROWS = 1;
@@ -32,20 +60,18 @@ const MIN_ROWS = 1;
* wiring fault has worn the costume of an honest absence, so the error is now
* surfaced rather than swallowed.
*/
async function page(sb, table, select, orderBy, apply) {
const out = [];
for (let from = 0; ; from += 1000) {
const q = apply ? apply(sb.from(table).select(select)) : sb.from(table).select(select);
const { data, error } = await q.order(orderBy, { ascending: true }).range(from, from + 999);
if (error) throw new Error(`${table}: ${error.message}`);
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < 1000) break;
}
return out;
async function page(sb, table, select, apply) {
return paginate(() => (apply ? apply(sb.from(table).select(select)) : sb.from(table).select(select)),
{ key: uniqueKeyFor(table), pageSize: 1000, label: `hitsFactorContext:${table}` });
}
/** Keep the most recent dated row per key. */
/**
* Keep the most recent dated row per key.
*
* With an as-of bound applied at the query, every row here is already at or
* before the cutoff, so "most recent" IS "the as-of row". An entity with no row
* in the bounded set simply never enters the map — which is the refusal.
*/
function latestBy(rows, keyFn, dateFn) {
const m = new Map();
for (const r of rows) {
@@ -63,13 +89,25 @@ function latestBy(rows, keyFn, dateFn) {
*/
async function build(sb, opts = {}) {
if (!sb) return null;
// null === LIVE. Only an explicitly dated call takes the audit path, so the
// served grade runs the same code it ran before this parameter existed.
const asOf = opts.asOf ? String(opts.asOf) : null;
const dated = (q) => (asOf ? q.lte('as_of_date', asOf) : q);
let spray; let defense; let platoon; let statcast;
try {
[spray, defense, platoon, statcast] = await Promise.all([
page(sb, 'batter_spray', '*', 'player_key', (q) => q.eq('sport', 'mlb')),
page(sb, 'team_defense', '*', 'team', (q) => q.eq('sport', 'mlb')),
page(sb, 'platoon_splits', '*', 'player_key', (q) => q.eq('sport', 'mlb')),
page(sb, 'statcast_aggregates', 'player_key, source_id, role, bats, throws, hard_hit_pct', 'player_key', (q) => q.eq('sport', 'mlb')),
page(sb, 'batter_spray', '*', (q) => dated(q.eq('sport', 'mlb'))),
page(sb, 'team_defense', '*', (q) => dated(q.eq('sport', 'mlb'))),
page(sb, 'platoon_splits', '*', (q) => dated(q.eq('sport', 'mlb'))),
// `statcast_aggregates` holds ONE as-of date (upserted in place), so it
// cannot answer an as-of question; the dated path reads the retained
// history instead.
asOf
? page(sb, 'statcast_history', 'player_key, source_id, role, bats, throws, hard_hit_pct, as_of_date',
(q) => q.eq('sport', 'mlb').lte('as_of_date', asOf))
: page(sb, 'statcast_aggregates', 'player_key, source_id, role, bats, throws, hard_hit_pct',
(q) => q.eq('sport', 'mlb')),
]);
} catch (e) {
// Surfaced, not silent: a load failure must be distinguishable from a feed
@@ -83,9 +121,16 @@ async function build(sb, opts = {}) {
const defBy = latestBy(defense, (r) => r.team, (r) => r.as_of_date);
const platBy = latestBy(platoon, (r) => r.player_key, (r) => r.as_of_date);
// On the dated path several as-of rows per player are in scope, so collapse to
// the latest WITHIN the bound before indexing. On the live path there is
// exactly one row per player and this is a no-op.
const statcastLatest = asOf
? [...latestBy(statcast, (r) => `${r.source_id}|${r.role}`, (r) => r.as_of_date).values()]
: statcast;
const batBy = new Map();
const pitBy = new Map();
for (const r of statcast) {
for (const r of statcastLatest) {
if (!r.player_key) continue;
if (r.role === 'pitcher') pitBy.set(r.player_key, r);
else batBy.set(r.player_key, r);
@@ -98,17 +143,28 @@ async function build(sb, opts = {}) {
return n > 1 ? n / 100 : n;
};
const resolver = (prop) => {
/**
* @param {object} prop
* @param {string} [sport]
* @param {object} [keys] FIX A6 — externally resolved { opponent,
* opposing_pitcher }. Omitted on the LIVE path, so live behaviour is
* byte-identical to before this argument existed; supplied only by the
* SHADOW resolve, which never reaches the served grade.
*/
const resolver = (prop, sport, keys) => {
const key = nameKey(prop && prop.player);
if (!key) return null;
const bat = batBy.get(key);
const bats = bat && bat.bats ? String(bat.bats)[0] : null;
// The opposing team and its starter, from whatever the prop carries.
const oppName = prop && (prop.opponent || prop.opp_team || null);
// The opposing team and its starter. `keys` wins when supplied; otherwise
// this reads what the prop carries, which is what it has always done — and
// what nothing ever sets (A5).
const oppName = (keys && keys.opponent) || (prop && (prop.opponent || prop.opp_team)) || null;
const def = oppName ? (defBy.get(oppName) || defBy.get(String(oppName).split(' ').pop())) : null;
const pitKey = prop && prop.opposing_pitcher ? nameKey(prop.opposing_pitcher) : null;
const pitName = (keys && keys.opposing_pitcher) || (prop && prop.opposing_pitcher) || null;
const pitKey = pitName ? nameKey(pitName) : null;
const pit = pitKey ? pitBy.get(pitKey) : null;
const sp = platBy.get(key);
@@ -124,6 +180,12 @@ async function build(sb, opts = {}) {
throws: pit && pit.throws ? String(pit.throws)[0] : null,
pitcherHardHit: pit ? asFraction(pit.hard_hit_pct) : null,
platoonSplits: splits,
// Carried, never invented. The archetype lives on the graded row
// (`model_snapshots.archetype`), not in any context table, so an audit
// caller supplies it and this passes it through for per-archetype
// conditioning. Absent stays absent.
archetype: (prop && prop.archetype) || null,
as_of: asOf,
};
// Nothing readable at all -> null, so the engine skips the factor block
// entirely rather than walking an empty context.
@@ -131,7 +193,11 @@ async function build(sb, opts = {}) {
return anything ? ctx : null;
};
resolver.asOf = asOf;
resolver.__stats = {
as_of: asOf,
mode: asOf ? 'as-of (audit)' : 'live',
statcast_source: asOf ? 'statcast_history' : 'statcast_aggregates',
spray_players: sprayBy.size,
defense_teams: defBy.size,
platoon_players: platBy.size,
+37 -16
View File
@@ -21,6 +21,7 @@
const lp = require('./lowParamCalibrator');
const cal = require('./calibration');
const { paginate } = require('../../utils/safePaginate');
const { knownNumber } = require('../../utils/known');
const MIN_FIT = 200;
@@ -74,24 +75,44 @@ function build(rows, opts = {}) {
};
}
function todayEt() {
return new Intl.DateTimeFormat('en-CA', {
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
}).format(new Date());
}
/**
* The ROW LOAD — exported so the read-integrity harness measures THE REAL
* FUNCTION rather than a restatement of its query.
*
* ── FIX A2 (2026-08-09) — THIS READ WAS 24.8% CORRUPT ────────────────────
* This is the PRIMARY calibrator (calibrationService is only its shadow), and it
* carried the byte-identical defect A1 fixed there: an unordered `.range()` walk,
* plus `if (error || !data) break` swallowing a failed read as end-of-data.
* Measured on production: 617 of 2,490 rows returned twice, an equal number never
* returned, with `rows.length` matching the server count exactly.
*
* Both are now `safePaginate` on the unique `id`. A throw means the read failed;
* `null` from `fromLedger` still means "not enough settled history to fit". Those
* are different states and collapsing them is what hid the defect.
*/
async function loadSettledRows(sb, { sport = 'mlb', stat = 'hits', before = null } = {}) {
const cutoff = before || todayEt();
return paginate(
() => sb.from('ledger_entries')
.select('id, p_win, outcome, game_date, quarantine_reason')
.eq('sport', sport).is('user_id', null).eq('stat', stat)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null)
.lt('game_date', cutoff),
{ key: 'id', pageSize: 1000, label: `lowParamService.fromLedger(${sport}/${stat})` },
);
}
/** Load settled history and build, POINT-IN-TIME (strictly before today). */
async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) {
if (!sb) return null;
const cutoff = before || new Intl.DateTimeFormat('en-CA', {
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
}).format(new Date());
const rows = [];
for (let from = 0; ; from += 1000) {
const { data, error } = await sb.from('ledger_entries')
.select('p_win, outcome, game_date, quarantine_reason')
.eq('sport', sport).is('user_id', null).eq('stat', stat)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null)
.lt('game_date', cutoff)
.range(from, from + 999);
if (error || !data || data.length === 0) break;
rows.push(...data);
if (data.length < 1000) break;
}
const cutoff = before || todayEt();
const rows = await loadSettledRows(sb, { sport, stat, before: cutoff });
const clean = rows
.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'))
.map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, date: String(r.game_date) }));
@@ -99,4 +120,4 @@ async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...
return built ? { ...built, cutoff } : null;
}
module.exports = { build, fromLedger, MIN_FIT, HOLDOUT_FRACTION };
module.exports = { build, fromLedger, loadSettledRows, MIN_FIT, HOLDOUT_FRACTION };
+144
View File
@@ -0,0 +1,144 @@
'use strict';
/**
* matchupKeys — resolve the two join keys the hits factors have always needed.
*
* Spec: specs/read-integrity-harness.md §11
*
* ── WHY THIS EXISTS ──────────────────────────────────────────────────────
* A5 measured that all three hits factors are skipped on 596/596 real props,
* because `hitsFactorContext` reads `prop.opponent` and `prop.opposing_pitcher`
* and NOTHING in the pipeline ever sets them. The matchup model is built and
* disconnected; these are the two wires.
*
* ── AT GRADE TIME, THE GAME HAS NOT HAPPENED ─────────────────────────────
* `prove-hit-factors` resolves the opponent from each hitter's own GAME LOG,
* which is correct for auditing a settled row and useless before first pitch —
* tonight's game is not in the log yet. So the grade-time source is the pair the
* schedule already publishes:
*
* player -> team `lineup_context` (as-of dated, one row per posted lineup)
* team -> matchup `getScheduleWithPitchers(gameDate)` (probable pitchers)
*
* Both are as-of-correct by construction: the lineup read is bounded by
* `as_of_date <= asOf` exactly like A4's context reads, and a probable pitcher is
* a pre-game fact.
*
* ── REFUSAL OVER GUESSING ────────────────────────────────────────────────
* A prop whose player has no posted lineup row, or whose game has no probable
* pitcher, resolves to NULL. It does NOT fall back to "the other team in the
* prop's game_id" — a prop carries `home_team` and `away_team`, so guessing
* which side a hitter bats for would be right about half the time and wrong
* invisibly. A guessed opponent would feed the defence factor a real team's
* fielders against the wrong hitter, which is worse than not firing.
*/
const { paginate } = require('../../utils/safePaginate');
const { uniqueKeyFor } = require('../../utils/tableKeys');
const { nameKey } = require('../../utils/playerName');
/** Keep the latest dated row per key within an as-of bound. */
function latestBy(rows, keyFn, dateFn) {
const m = new Map();
for (const r of rows) {
const k = keyFn(r);
if (!k) continue;
const prev = m.get(k);
if (!prev || String(dateFn(r)) > String(dateFn(prev))) m.set(k, r);
}
return m;
}
/** Normalise a team name for matching across feeds ("Chicago Cubs" vs "Cubs"). */
function teamForms(name) {
const s = String(name || '').trim();
if (!s) return [];
const last = s.split(' ').pop();
return [...new Set([s, last])];
}
/**
* Build the per-slate index.
*
* @param {object} deps
* - sb supabase client (for lineup_context)
* - getSchedule async (date) => [{ home:{team,probablePitcher}, away:{...} }]
* - gameDate 'YYYY-MM-DD'
* - asOf as-of bound for the lineup read (default: gameDate)
* @returns {object|null} { resolve(prop), stats } — null when nothing loaded
*/
async function build(deps = {}) {
const { sb, getSchedule, gameDate } = deps;
if (!sb || !getSchedule || !gameDate) return null;
const asOf = deps.asOf || gameDate;
let lineups = [];
let games = [];
try {
[lineups, games] = await Promise.all([
paginate(() => sb.from('lineup_context')
.select('as_of_date, game_date, sport, game_pk, team, side, player_key')
.eq('sport', 'mlb').eq('game_date', gameDate).lte('as_of_date', asOf),
{ key: uniqueKeyFor('lineup_context'), pageSize: 1000, label: 'matchupKeys:lineup_context' }),
getSchedule(gameDate),
]);
} catch (e) {
// Surfaced, never swallowed as an empty feed (the recurring costume).
console.warn('[matchupKeys] load FAILED (not an empty feed):', e.message);
return null;
}
// player -> the team he is posted to bat for, as of the cutoff
const teamByPlayer = latestBy(lineups, (r) => r.player_key, (r) => r.as_of_date);
// team -> { opponent, opposing_pitcher }
const matchup = new Map();
for (const g of games || []) {
const h = g && g.home; const a = g && g.away;
if (!h || !a || !h.team || !a.team) continue;
const put = (side, other) => {
const sp = other.probablePitcher && other.probablePitcher.name ? other.probablePitcher.name : null;
for (const form of teamForms(side.team)) {
if (!matchup.has(form)) matchup.set(form, { opponent: other.team, opposing_pitcher: sp });
}
};
put(h, a);
put(a, h);
}
const stats = {
game_date: gameDate,
as_of: asOf,
lineup_players: teamByPlayer.size,
scheduled_games: (games || []).length,
teams_with_matchup: matchup.size,
games_with_both_probables: (games || []).filter((g) => g && g.home && g.away
&& g.home.probablePitcher && g.away.probablePitcher).length,
};
/**
* @returns {object} { opponent, opposing_pitcher, source, refused }
* Nulls are refusals — never a guess from the prop's two teams.
*/
function resolve(prop) {
const key = nameKey(prop && (prop.player || prop.player_name));
const empty = { opponent: null, opposing_pitcher: null, source: null, refused: 'no_lineup_row' };
if (!key) return { ...empty, refused: 'no_player' };
const lu = teamByPlayer.get(key);
if (!lu || !lu.team) return empty;
const m = matchup.get(lu.team) || matchup.get(String(lu.team).split(' ').pop());
if (!m) return { opponent: null, opposing_pitcher: null, source: null, refused: 'team_not_in_schedule' };
return {
opponent: m.opponent || null,
opposing_pitcher: m.opposing_pitcher || null,
team: lu.team,
source: 'lineup_context+schedule',
refused: m.opposing_pitcher ? null : 'no_probable_pitcher',
};
}
resolve.stats = stats;
return resolve;
}
module.exports = { build, teamForms, latestBy };
+122
View File
@@ -0,0 +1,122 @@
'use strict';
/**
* WITHDRAWN VERDICTS — the append-only record of findings we can no longer stand behind.
*
* Spec: specs/read-integrity-harness.md §6
*
* WHY THIS FILE EXISTS AS A FILE: there is no verdict table.
* `featureRegistry.statVerdicts` is an in-memory Map that resets every process
* (featureRegistry.js:145), and `mc_test_ledger` stores hypothesis COUNTS for the
* Bonferroni denominator, not verdicts. The standing record has always been the
* specs plus CLAUDE.md. So a withdrawal is APPENDED here, in the repo, with its
* cause — the original findings are never edited or deleted.
*
* THE LEDGER TIGHTENING ITS OWN STANDARD IS THE RECORD WE WANT. A programme that
* only ever accumulates positive findings is not measuring; it is collecting. A
* withdrawal with a measured cause is worth more than the verdict it retracts.
*
* WHAT A WITHDRAWAL IS NOT: this module records what we may CLAIM. It does not
* change what the engine SERVES. `hitsFactors.js` hardcodes its factor list and
* consults no registry, so both withdrawn factors still move the served forecast.
* Disarming or re-proving them is a separate order with its own blast radius.
*
* REINSTATEMENT: a withdrawn verdict returns only by re-running its gate through
* readers that `src/utils/readIntegrity.js` reports PASS for ON THE DAY OF THE
* RUN. Never by argument, and never by re-reading the old numbers.
*/
const STATUS = Object.freeze({
WITHDRAWN_PENDING_REAUDIT: 'WITHDRAWN_PENDING_REAUDIT',
});
/**
* The measured cause, stated once so every entry cites the identical evidence
* rather than a paraphrase that can drift.
*/
const UNORDERED_PAGINATION_CAUSE = 'drawn through an unordered-pagination reader; '
+ '16.5–24.8% set corruption measured on the identical query 2026-08-09 '
+ '(412–617 of 2,490 rows duplicated, an equal number never returned); '
+ 'not retrospectively recoverable — the query plan varied per run and was never logged';
/**
* APPEND ONLY. Never edit an entry; never remove one. A superseded withdrawal
* gets a new entry that references the old one.
*/
const WITHDRAWALS = Object.freeze([
Object.freeze({
id: 'hits/defense_by_direction/2026-08-09',
factor: 'defense_by_direction',
sport: 'mlb',
stat: 'hits',
withdrawn_at: '2026-08-09',
status: STATUS.WITHDRAWN_PENDING_REAUDIT,
original_verdict: 'PROVES',
original_evidence: Object.freeze({
n: 528, brier_delta: -0.0034, ci: Object.freeze([-0.0059, -0.0009]), ci_level: 0.999,
source: 'Session 93',
}),
cause: UNORDERED_PAGINATION_CAUSE,
// ARITHMETIC ON THE RECORDED n, not a re-measurement. Stated as such so it is
// never quoted back as a new result.
eligibility_note: 'at the measured ~20% duplication rate the distinct n is ≈422 against '
+ 'factorGate MIN_N = 500 — this verdict was likely never eligible for adjudication at all. '
+ 'This is arithmetic on the recorded n, not a re-measurement.',
readers_implicated: Object.freeze([
'scripts/prove-hit-factors.js:136 (ledger_entries walk)',
'scripts/prove-hit-factors.js:131 (model_snapshots archetype join)',
]),
still_served: true,
still_served_note: 'hitsFactors.js:63-78 applies it unconditionally; it consults no registry',
reinstatement: 're-run the two-part gate through readers measured PASS by src/utils/readIntegrity.js',
}),
Object.freeze({
id: 'hits/pitcher_contact_profile/2026-08-09',
factor: 'pitcher_contact_profile',
sport: 'mlb',
stat: 'hits',
withdrawn_at: '2026-08-09',
status: STATUS.WITHDRAWN_PENDING_REAUDIT,
original_verdict: 'PROVES',
original_evidence: Object.freeze({
n: 741, brier_delta: -0.0066, ci: Object.freeze([-0.0114, -0.0016]), ci_level: 0.999,
source: 'Session 92',
}),
cause: UNORDERED_PAGINATION_CAUSE,
eligibility_note: 'n=741 clears MIN_N even after ~20% duplication (≈593 distinct), so unlike '
+ 'defense_by_direction it was plausibly eligible — but the CI that decided PROVES was '
+ 'computed on a corrupted sample against a corrupted leave-one-out baseline, so the '
+ 'interval is not recoverable either.',
readers_implicated: Object.freeze([
'scripts/prove-hit-factors.js:136 (ledger_entries walk)',
'scripts/prove-hit-factors.js:131 (model_snapshots archetype join)',
]),
still_served: true,
still_served_note: 'hitsFactors.js:80-87 applies it unconditionally; it consults no registry',
reinstatement: 're-run the two-part gate through readers measured PASS by src/utils/readIntegrity.js',
}),
]);
/** Is this factor's verdict withdrawn for this stat? */
function isWithdrawn(sport, stat, factor) {
return WITHDRAWALS.some((w) => w.sport === String(sport || '').toLowerCase()
&& w.stat === String(stat || '').toLowerCase()
&& w.factor === factor);
}
/** The withdrawal entries for a sport+stat (newest last — the list is append-only). */
function withdrawalsFor(sport, stat) {
const sp = String(sport || '').toLowerCase();
const st = String(stat || '').toLowerCase();
return WITHDRAWALS.filter((w) => w.sport === sp && (!st || w.stat === st));
}
/** Factors still SERVED despite a withdrawn verdict — the honesty gap, listed. */
function servedButWithdrawn() {
return WITHDRAWALS.filter((w) => w.still_served);
}
module.exports = {
WITHDRAWALS, STATUS, UNORDERED_PAGINATION_CAUSE,
isWithdrawn, withdrawalsFor, servedButWithdrawn,
};
+10 -3
View File
@@ -56,9 +56,10 @@ function etDateOf(iso) {
* forecast, and never on a mixture of the two, which is the trap that would
* otherwise be invisible once both generations sit in the same table.
*/
const MODEL_VERSION = process.env.MODEL_VERSION || 'engine1@2026-08-07-fullwindow';
/** Rows at or after this marker are eligible for forward re-audit. */
const REPAIRED_CHAMPION_VERSION = 'engine1@2026-08-07-fullwindow';
// FIX A3 — ONE SOURCE (src/config/modelVersion.js). Re-exported so existing
// importers keep working, but never re-declared: the twin default in
// ledgerService is exactly how these two drifted apart.
const { MODEL_VERSION, REPAIRED_CHAMPION_VERSION } = require('../config/modelVersion');
function codeSha() {
return process.env.SOURCE_COMMIT || process.env.GIT_SHA || process.env.COOLIFY_GIT_COMMIT_SHA || null;
@@ -149,6 +150,12 @@ function rowsFromSides(base, sides, ctx = {}) {
// The counterfactual enabler. Absent on pre-feature refusals (the juice
// gate runs before features are computed) — honestly null, never faked.
features: s._features && Object.keys(s._features).length ? s._features : null,
// FIX A5 — the RAW inputs the hits factors read, frozen at grade time so a
// re-audit reads evidence instead of reconstructing context. Inputs only:
// the multiplier is recomputable from them and is deliberately not stored.
// Null on every non-hits row and on any row with no factor context.
factor_inputs: s.factor_inputs || null,
});
}
return rows;
+29
View File
@@ -483,8 +483,37 @@ async function runSnapshot(sport, opts = {}) {
}
}
// ── FIX A6 — THE JOIN KEYS, SHADOW ONLY ─────────────────────────────────
// A5 measured all three hits factors skipped on 596/596 because nothing sets
// `prop.opponent` / `prop.opposing_pitcher`. These resolve them from the
// posted lineup + the schedule's probable pitchers, as-of-correct. They feed a
// SHADOW resolve that is FROZEN onto the row — the served forecast is NOT
// adjusted by them. Turning them on live is a separate, gated decision.
let matchupKeys = null;
if (sp === 'mlb') {
try {
const mk = deps.matchupKeys || require('./model/matchupKeys');
const sbm = require('../utils/supabase').getSupabaseServiceClient();
const mlbAdapter = deps.mlbAdapter || require('./adapters/mlbStatsAdapter');
const resolver = sbm ? await mk.build({
sb: sbm,
getSchedule: (d) => mlbAdapter.getScheduleWithPitchers(d),
gameDate: retentionGameDate,
}) : null;
if (resolver) {
matchupKeys = (prop) => resolver(prop);
console.log(`[matchup] ${sp} keys loaded — ${JSON.stringify(resolver.stats)}`);
} else {
console.log(`[matchup] ${sp} — no key index; shadow factors will not fire`);
}
} catch (e) {
console.warn('[matchup] key resolve skipped:', e.message);
}
}
await deps.gradeAndCacheSlate(sp, props, {
factorContext,
matchupKeys,
// Bisect hook (2026-08-01): lets the internal trigger run a bounded slate
// without a prod env change, so a cap regression can be isolated by
// measurement instead of guessed at. Omitted => gradeSlateService's own
+311
View File
@@ -0,0 +1,311 @@
'use strict';
/**
* snapshotSettlementService — settle `model_snapshots` ON A SCHEDULE.
*
* Spec: specs/read-integrity-harness.md §9
*
* ── WHY THIS EXISTS ──────────────────────────────────────────────────────
* `model_snapshots` is the retention table built for replay — it holds the
* model's INPUTS (the frozen feature vector) alongside its prediction, which
* `ledger_entries` does not. It was settled exactly once, by hand, via
* `scripts/settle-model-snapshots.js`. Nothing ever settled it on a cron.
*
* Measured 2026-08-09: 104,834 rows were past-dated and settleable; 89,350 of
* them had no outcome. Worse, of the 34,128 rows carrying the repaired
* champion's marker — the only rows a forward re-audit may be measured on —
* ZERO were settled. The accrual clock could never advance.
*
* ── OUTCOMES ONLY. NO CONTEXT RECONSTRUCTION. ────────────────────────────
* This joins each row to the realized stat from a box score, using the row's OWN
* keys (game_date + player_key + stat + line + side). It does not fetch, infer,
* or rebuild park, weather, defence, platoon, archetype or any other context.
* A settle pass that reconstructed context would be re-deciding what the model
* saw, which is precisely what the retention table exists to prevent.
*
* ── OUTCOME IS SIDE-ALIGNED ──────────────────────────────────────────────
* `p_win` is expressed for the GRADED SIDE, so `outcome` must be too. A raw
* `realized > line` indicator is the OVER perspective and would silently invert
* the target on every under row, making calibration measure the wrong thing.
* `actual_value` stores the raw realized stat; `outcome` stores whether the
* graded side won.
*
* ── A ROW LOGGED AFTER FIRST PITCH IS NOT A PREDICTION ───────────────────
* The pipeline runs on UTC cron hours, so a 01:00-UTC cycle is 21:00 the
* previous evening ET — same game date, three hours into the slate. Those rows
* are refused rather than settled. Preserved verbatim from the hand-run script,
* because dropping it would quietly admit post-hoc rows into every measurement.
*
* Every dependency is injectable, so the unit tests never touch a network.
*/
const { paginate } = require('../utils/safePaginate');
const { uniqueKeyFor } = require('../utils/tableKeys');
const { nameKey } = require('../utils/playerName');
const { knownNumber } = require('../utils/known');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** Eastern first pitch, conservatively. At or after this is in-game. */
const FIRST_PITCH_ET_HOUR = 19;
/**
* Dates fetched per run — bounded, because a cron that re-fetches 26 dates of
* box scores every five hours is a quota problem pretending to be thoroughness.
*
* NEWEST FIRST, and that ordering is a correctness property, not a preference.
* The first implementation drained oldest-first and DEADLOCKED: measured on
* production, 308 rows across the eight oldest dates are structurally
* unsettleable (184 logged after first pitch, 124 with no box-score line), so
* the window re-processed the same dead dates on every run and `dates_remaining`
* never moved. Newest-first also reaches the rows that matter — the repaired
* champion's, which are the only ones a forward re-audit may be measured on.
*/
const MAX_DATES_PER_RUN = Number(process.env.SNAPSHOT_SETTLE_MAX_DATES || 8);
/** Realized value per stat, from the box-score batting line. */
const FIELD = Object.freeze({
hits: (b) => knownNumber(b.hits),
total_bases: (b) => knownNumber(b.totalBases),
rbi: (b) => knownNumber(b.rbi),
runs: (b) => knownNumber(b.runs),
});
function isPreGame(capturedAt, gameDate) {
if (!capturedAt || !gameDate) return false;
const cap = new Date(capturedAt);
if (Number.isNaN(cap.getTime())) return false;
const et = new Date(cap.getTime() - 4 * 3600 * 1000); // EDT
const etDate = et.toISOString().slice(0, 10);
if (etDate < String(gameDate)) return true; // day before, fine
if (etDate > String(gameDate)) return false; // day after, post-game
return et.getUTCHours() < FIRST_PITCH_ET_HOUR;
}
function todayEt(now) {
return new Intl.DateTimeFormat('en-CA', {
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
}).format(now || new Date());
}
async function pool(items, fn, n = 6) {
const out = []; let i = 0;
await Promise.all(Array.from({ length: n }, async () => {
while (i < items.length) {
const idx = i; i += 1;
try { out[idx] = await fn(items[idx]); } catch { out[idx] = null; }
}
}));
return out.filter(Boolean);
}
/**
* Box-score batting lines for a set of dates, keyed `${date}|${nameKey}`.
* A doubleheader gives two lines; they are SUMMED, because the prop covers the
* day rather than a game.
*/
async function battingLines(dates, deps) {
const getJson = deps.getJson;
const games = [];
for (const d of dates) {
try {
const s = await getJson(`https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}`);
for (const day of (s && s.dates) || []) {
for (const g of day.games || []) {
if (String(g.status && g.status.detailedState) === 'Final') {
games.push({ pk: g.gamePk, date: g.officialDate || d });
}
}
}
} catch { /* absent day — stays absent */ }
}
const lines = {};
const loaded = await pool(games, async (g) => {
const box = await getJson(`https://statsapi.mlb.com/api/v1/game/${g.pk}/boxscore`);
const out = [];
for (const side of ['home', 'away']) {
const t = box && box.teams && box.teams[side];
if (!t) continue;
for (const id of t.batters || []) {
const pl = t.players[`ID${id}`];
const b = pl && pl.stats && pl.stats.batting;
if (!b || b.atBats == null) continue; // did not bat → absent, never zero
out.push({
date: g.date,
key: nameKey(pl.person && pl.person.fullName),
hits: b.hits, totalBases: b.totalBases, rbi: b.rbi, runs: b.runs,
});
}
}
return out;
}, deps.concurrency || 6);
for (const arr of loaded) {
for (const r of arr) {
const k = `${r.date}|${r.key}`;
if (!lines[k]) lines[k] = { ...r };
else {
lines[k].hits += r.hits; lines[k].totalBases += r.totalBases;
lines[k].rbi += r.rbi; lines[k].runs += r.runs;
}
}
}
return lines;
}
/**
* PURE — decide each row's outcome from the box-score index.
* Separated so the whole decision rule is unit-testable with no I/O.
*/
function decide(snaps, lines) {
const counts = { candidates: snaps.length, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 };
const updates = [];
const seen = new Set();
for (const s of snaps) {
if (seen.has(s.id)) {
const e = new Error(`INTEGRITY: duplicate snapshot id ${s.id}`);
e.code = 'DUPLICATE_ROW';
throw e;
}
seen.add(s.id);
if (!isPreGame(s.captured_at, s.game_date)) {
counts.post_hoc_logged += 1; counts.unresolvable += 1; continue;
}
const line = knownNumber(s.line);
if (line === null || !s.side) { counts.unresolvable += 1; continue; }
const b = lines[`${s.game_date}|${s.player_key}`];
if (!b) { counts.orphaned += 1; continue; }
const realized = FIELD[s.stat] ? FIELD[s.stat](b) : null;
if (realized === null) { counts.unresolvable += 1; continue; }
const over = realized > line;
const won = String(s.side).toLowerCase() === 'under' ? !over : over;
updates.push({ id: s.id, outcome: won ? 'hit' : 'miss', actual_value: realized });
counts.settled += 1;
}
// CONSERVATION — hard fail. Every candidate lands in exactly one bucket.
const acc = counts.settled + counts.unresolvable + counts.orphaned;
if (acc !== counts.candidates) {
const e = new Error(`INTEGRITY: conservation violated ${acc} != ${counts.candidates}`);
e.code = 'CONSERVATION';
throw e;
}
return { counts, updates };
}
/**
* Settle one pass.
*
* @param {object} deps
* - sb supabase service client (required to do anything)
* - getJson async (url) => json
* - now () => Date
* - write default true; false = dry run
* - maxDates dates fetched this run (oldest first)
* @returns {object} { skipped?, counts, dates, written, repaired_champion_settled }
*/
async function settleSnapshots(deps = {}) {
const sb = deps.sb;
if (!sb) return { skipped: 'supabase not configured', counts: null, written: 0 };
const now = (deps.now || (() => new Date()))();
const cutoff = todayEt(now);
const write = deps.write !== false;
// Unsettled, PAST-DATED rows only — a game that has not finished cannot be
// settled, and asking would produce an orphan rather than an absence.
const snaps = await paginate(
() => sb.from('model_snapshots')
.select('id, game_date, captured_at, stat, player_key, line, side, model_version')
.eq('sport', 'mlb').in('stat', STATS).is('outcome', null)
.lt('game_date', cutoff),
{ key: uniqueKeyFor('model_snapshots'), pageSize: PAGE, label: 'snapshotSettlement' },
);
if (!snaps.length) return { counts: { candidates: 0, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 }, dates: [], written: 0, repaired_champion_settled: 0 };
// CHOOSE DATES FROM ROWS THAT COULD ACTUALLY SETTLE.
//
// `isPreGame` is pure, so a row logged after first pitch is known-unsettleable
// WITHOUT fetching anything. Measured on production, 19,074 such rows sit in
// the recent dates; letting them pick the window meant the same dead dates
// were re-fetched on every run and the backlog never converged. Filtering
// first is what makes the drain terminate.
const settleable = snaps.filter((r) => isPreGame(r.captured_at, r.game_date));
// NEWEST FIRST: the eligible rows — the repaired champion's — are the newest,
// and an older date whose remainder cannot settle must not block them.
const allDates = [...new Set(settleable.map((r) => r.game_date))].sort().reverse();
const perWindow = deps.maxDates || MAX_DATES_PER_RUN;
const maxWindows = deps.maxWindows || Number(process.env.SNAPSHOT_SETTLE_MAX_WINDOWS || 4);
// ADVANCE PAST A DRAINED WINDOW. The newest dates keep rows that can never
// settle (a player with no box-score line never gets one), so a fixed window
// would sit on them forever while older settleable dates were never reached.
// Move to the next window when this one produces nothing, bounded so a run
// still costs a predictable number of box-score fetches.
let dates = []; let counts = null; let updates = []; let windows = 0;
for (let w = 0; w < maxWindows; w += 1) {
const slice = allDates.slice(w * perWindow, (w + 1) * perWindow);
if (!slice.length) break;
windows = w + 1;
/* eslint-disable no-await-in-loop */
const lines = await battingLines(slice, { getJson: deps.getJson, concurrency: deps.concurrency });
const res = decide(snaps.filter((r) => slice.includes(r.game_date)), lines);
/* eslint-enable no-await-in-loop */
dates = slice; counts = res.counts; updates = res.updates;
if (updates.length) break; // progress — stop here
}
if (!counts) return { counts: { candidates: 0, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 }, dates: [], written: 0, repaired_champion_settled: 0 };
const inWindow = snaps.filter((r) => dates.includes(r.game_date));
let written = 0;
let repairedSettled = 0;
const byId = new Map(inWindow.map((r) => [r.id, r]));
if (write && updates.length) {
const settledAt = now.toISOString();
for (let i = 0; i < updates.length; i += 500) {
const batch = updates.slice(i, i + 500);
/* eslint-disable no-await-in-loop */
const results = await Promise.all(batch.map((u) => sb.from('model_snapshots')
.update({
outcome: u.outcome,
actual_value: u.actual_value,
settled_at: settledAt,
settlement_source: 'statsapi_boxscore',
})
// IDEMPOTENT: only an unsettled row is written, so a re-run can never
// overwrite an outcome that is already on the record.
.eq('id', u.id).is('outcome', null)));
/* eslint-enable no-await-in-loop */
results.forEach((r, k) => {
if (r.error) return;
written += 1;
const row = byId.get(batch[k].id);
if (row && row.model_version === require('../config/modelVersion').REPAIRED_CHAMPION_VERSION) {
repairedSettled += 1;
}
});
}
}
return {
counts,
dates,
dates_remaining: Math.max(0, allDates.length - (windows * perWindow)),
windows_scanned: windows,
// Rows that can never settle, counted rather than hidden: a row logged after
// first pitch is not a prediction and will never become one.
permanently_unsettleable: snaps.length - settleable.length,
written,
repaired_champion_settled: repairedSettled,
mode: write ? 'write' : 'dry-run',
};
}
module.exports = {
settleSnapshots, decide, isPreGame, battingLines,
STATS, FIELD, MAX_DATES_PER_RUN, FIRST_PITCH_ET_HOUR,
};
+27
View File
@@ -65,6 +65,11 @@ function startSnapshotScheduler(opts = {}) {
// Session 58 — Phase 1 truth infrastructure: settle the persistent ledger
// (outcome + actual + CLV) in the same pre-grade settle pass.
const settleLedgers = opts.settleAllLedgers || require('./services/ledgerService').settleAllLedgers;
// FIX A3 — model_snapshots was settled ONCE, by hand, and never on a cron.
// 89,350 settleable rows sat unsettled and ZERO of the 34,128 repaired-champion
// rows had an outcome, so the re-audit accrual clock could not advance.
const settleSnaps = opts.settleSnapshots
|| require('./services/snapshotSettlementService').settleSnapshots;
const notify = opts.notify || require('./utils/opsNotify').notify;
const cacheGet = opts.cacheGet || require('./utils/redis').cacheGet;
const cacheSet = opts.cacheSet || require('./utils/redis').cacheSet;
@@ -196,6 +201,28 @@ function startSnapshotScheduler(opts = {}) {
title: 'VYNDR settlement', priority: 'high', tags: ['rotating_light'],
});
}
// FIX A3 — settle the RETENTION table too, on the same pass that settles the
// ledger. Outcomes only: it joins each row to a box score by the row's own
// keys and never reconstructs context. Best-effort by design — the retention
// table is a measurement asset, not the public record, so a failure here must
// never take the ledger settle or the grade with it.
try {
const sbc = require('./utils/supabase').getSupabaseServiceClient();
const axios = require('axios');
const snapRes = await settleSnaps({
sb: sbc,
getJson: async (url) => (await axios.get(url, { timeout: 45_000 })).data,
now: () => d,
});
if (snapRes && snapRes.counts) {
console.log(`[snapshots] settle pass — ${snapRes.written} settled `
+ `(${snapRes.repaired_champion_settled} repaired-champion), `
+ `${snapRes.counts.orphaned} orphaned, ${snapRes.counts.post_hoc_logged} post-hoc, `
+ `${snapRes.dates_remaining} date(s) of backlog remaining`);
}
} catch (e) {
console.warn('[snapshots] settle run failed:', e.message);
}
// Session 8 — zero-settle alarm, MORNING slot only (the book-closing pass).
// Signal = the ledger settle's own Postgres-backed return values (see
// opsWatch.zeroSettleAlarm for why not snapshot:{sport}:previous). Deduped
Binary file not shown.
+124
View File
@@ -0,0 +1,124 @@
'use strict';
/**
* safePaginate — the ONE way this codebase walks a paginated PostgREST read.
*
* Spec: specs/read-integrity-harness.md §7
*
* It exists because the ad-hoc walk it replaces failed in two independent ways,
* both measured on production, both of which returned a plausible-looking result:
*
* 1. NO STABLE ORDER. `.range(from, from+PAGE-1)` without an ORDER BY returned
* the correct row COUNT and the wrong ROWS — 410-617 of 2,490 duplicated on
* the hits read, with an equal number never returned at all. Postgres
* guarantees no ordering without ORDER BY, and the planner can order
* differently between successive LIMIT/OFFSET queries.
*
* 2. ERROR SWALLOWED AS END-OF-DATA. `if (error || !data || !data.length) break`
* makes a failed read indistinguishable from a finished one, so a partial
* fit looks like a complete fit. This is the shape that silently killed
* settlement on 2026-08-01.
*
* ── THE KEY MUST BE UNIQUE ────────────────────────────────────────────────
* An ORDER BY on a non-unique column is NOT a fix: ties may be returned in any
* order, so pages still overlap. The default is `id`. Because "unique" is a
* promise the caller makes, this module VERIFIES it at runtime — a repeated key
* during the walk means either a non-unique key or a scrambled read, and either
* way the row set is not trustworthy, so it THROWS rather than returning it.
*
* ── COMPOSITE KEYS (Fix A2b) ──────────────────────────────────────────────
* `key` accepts a string OR an ordered list of columns forming a unique TUPLE.
* The context tables (`batter_spray`, `team_defense`, `platoon_splits`, …) are
* dated-composite by design — one as-of snapshot per entity per day — and have
* no single unique column, so before this they could not be made safe at all.
* Every column is ordered, in the declared order, and the uniqueness guard
* checks the FULL tuple. `src/utils/tableKeys.js` holds the real constraints,
* pulled from the live schema rather than assumed.
*
* Ordering on a PREFIX of a unique key is not enough — the trailing columns are
* exactly where the ties live — so the caller passes the whole tuple or the
* guard will catch it.
*
* ── REFUSAL IS NOT A SWALLOW ──────────────────────────────────────────────
* Throwing here lets the caller decide, explicitly, between "no data" and "the
* read failed" — the distinction the old code destroyed. A caller that responds
* by producing NOTHING (no calibrator, no map) is refusing honestly. A caller
* that responds by continuing with partial rows is repeating the defect.
*
* The paging mechanics are `readIntegrity.walk`, which is already unit-tested
* for short-page, empty-page, error-propagation and runaway behaviour. This
* module adds the stable order and the uniqueness guard on top; it does not
* hand-roll a second walk.
*/
const { walk, keyOf, DEFAULT_PAGE } = require('./readIntegrity');
/**
* Walk a filtered PostgREST query to exhaustion, safely.
*
* @param {function} makeQuery () => a FRESH filtered PostgrestFilterBuilder.
* A factory, not a builder: a builder is single-use, so each page needs
* its own.
* @param {object} opts
* - key {string|string[]} unique column, or an ordered list of columns
* forming a unique tuple (default 'id')
* - pageSize {number} default 1000
* - ascending {boolean} default true
* - maxPages {number} runaway guard
* - label {string} included in thrown messages so a failure names itself
* @returns {Promise<Array>} every row, exactly once
* @throws on a query error, a duplicate key tuple, or a runaway walk — never silently
*/
async function paginate(makeQuery, opts = {}) {
const keyCols = normalizeKey(opts.key);
const pageSize = opts.pageSize || DEFAULT_PAGE;
const ascending = opts.ascending !== false;
const label = opts.label || 'safePaginate';
const seen = new Set();
return walk(async (from, to) => {
// EVERY column is ordered, in the declared order. Ordering on a prefix
// leaves the trailing columns tied, which is precisely where pages overlap.
let q = makeQuery();
for (const col of keyCols) q = q.order(col, { ascending });
const { data, error } = await q.range(from, to);
// AN ERROR IS AN ERROR. Never end-of-data.
if (error) {
const e = new Error(`${label}: read failed at range ${from}-${to} — ${error.message}`);
e.code = 'READ_FAILED';
e.cause = error;
throw e;
}
const rows = data || [];
// UNIQUENESS, VERIFIED RATHER THAN ASSUMED. A repeat means the ordering key
// is not unique or the walk scrambled; both make the row set unusable.
for (const r of rows) {
const k = keyOf(r, keyCols);
if (seen.has(k)) {
const e = new Error(`${label}: duplicate key (${keyCols.join(',')})=${k} at range ${from}-${to} — `
+ 'the ordering key is not unique or the read scrambled; refusing to return a corrupted row set');
e.code = 'DUPLICATE_KEY';
throw e;
}
seen.add(k);
}
return rows;
}, pageSize, opts.maxPages);
}
/** A string key and a composite key are the same thing, one column long. */
function normalizeKey(key) {
if (key == null) return ['id'];
const cols = Array.isArray(key) ? key : [key];
const clean = cols.filter((c) => typeof c === 'string' && c.length);
if (!clean.length) {
const e = new Error('safePaginate: key must be a column name or a non-empty list of column names');
e.code = 'BAD_KEY';
throw e;
}
return clean;
}
module.exports = { paginate, normalizeKey, DEFAULT_PAGE };
+66
View File
@@ -0,0 +1,66 @@
'use strict';
/**
* tableKeys — WHAT MAKES A ROW UNIQUE, per table.
*
* Spec: specs/read-integrity-harness.md §8
*
* `safePaginate` needs a key whose tuple is unique, because an ORDER BY that
* leaves ties lets pages overlap and the whole fix evaporates. Guessing that key
* is exactly the kind of assumption this programme keeps getting burned by, so
* every entry below was PULLED FROM THE LIVE SCHEMA on 2026-08-09:
*
* SELECT c.relname, array_agg(a.attname ORDER BY k.ord)
* FROM pg_index ix ... WHERE ix.indisunique
*
* Note that only `ledger_entries` and `model_snapshots` have a single-column
* key. Every context table is dated-composite — which is the point of those
* tables (an as-of snapshot per entity per day), and the reason the A2 rollout
* could not reach them until `safePaginate` learned composite keys.
*
* IF A MIGRATION CHANGES A CONSTRAINT, CHANGE IT HERE. A stale entry does not
* fail loudly at the database — it fails as a duplicate tuple at runtime, which
* `safePaginate` throws on. That is the intended failure: loud, not silent.
*/
const UNIQUE_KEY = Object.freeze({
// single-column primary keys
ledger_entries: Object.freeze(['id']),
model_snapshots: Object.freeze(['id']),
game_context: Object.freeze(['game_id']),
// dated composite keys — one as-of snapshot per entity per day
statcast_aggregates: Object.freeze(['sport', 'season', 'source_id', 'role']),
statcast_history: Object.freeze(['as_of_date', 'sport', 'season', 'source_id', 'role']),
batter_spray: Object.freeze(['as_of_date', 'sport', 'season', 'source_id']),
team_defense: Object.freeze(['as_of_date', 'sport', 'season', 'team']),
platoon_splits: Object.freeze(['as_of_date', 'sport', 'season', 'player_key']),
park_dimensions: Object.freeze(['as_of_date', 'sport', 'venue_id']),
hitter_opportunity: Object.freeze(['as_of_date', 'sport', 'season', 'player_key']),
lineup_context: Object.freeze(['as_of_date', 'sport', 'game_pk', 'player_key']),
});
/**
* The unique key for a table.
*
* THROWS for an unknown table rather than defaulting to `id`. A silent default
* would order by a column that may not exist or may not be unique, which is the
* failure this module exists to prevent — and it would do so while looking fixed.
*/
function uniqueKeyFor(table) {
const k = UNIQUE_KEY[table];
if (!k) {
const e = new Error(`tableKeys: no unique key recorded for '${table}' — `
+ 'add it from the live schema (pg_index WHERE indisunique) rather than assuming id');
e.code = 'UNKNOWN_TABLE';
throw e;
}
return k;
}
/** Is this table's key a single column? (Informational; both shapes work.) */
function isSingleKey(table) {
return uniqueKeyFor(table).length === 1;
}
module.exports = { UNIQUE_KEY, uniqueKeyFor, isSingleKey };