Raise the grade cap 25 -> 500 on measured cost; refusals are correct

PART 1 (read-only, measured on a live prod slate, n=80) OVERTURNS THE
PREMISE. The refusal rate is not a data problem -- it is 98% correct
behaviour. The cap is the entire problem, and it is worse than "25 of 546".

Composition: GRADED 44 (55.0%) | POLICY-SUPPRESSION 35 (43.8%) |
FETCHABLE-GAP 1 (1.3%) | FALSE-THRESHOLD 0 | ARCHETYPE-GAP 0 |
GENUINE-ABSENCE 0.

THE FIFTH BUCKET the order did not anticipate: all 35 "refusals" are
rare_event_over_below_line -- the 2026-07-19 betting-logic audit
deliberately refusing 0.5-line rare events, setting the SAME
insufficient_data flag as a real data gap, which is why they read as one.
They are entirely doubles (18) and stolen_bases (17), while hits (19/19),
rbi (19/19) and total_bases (5/5) grade at ~100%. Had we "fixed" this we
would have re-introduced exactly the bets a previous audit removed, and the
count would have looked like progress.

THE CAP: 585 unique gradeable props, cap 25 -> 560 discarded (95.7%).
Traced to Session 32 (f0c8b4f), commented "bound the herd" -- a guard
written before anyone measured what a grade costs. So I measured it:
721ms mean / 666ms median / 1024ms p90 per grade => ~72s for 500 props at
concurrency 5. Both callers tolerate that: the cron runs 5x/day and
recordDownstream is fire-and-forget.

PART 2 -- item 3 ONLY, because that is what the diagnosis supports.
DEFAULT_LIMIT 25 -> 500, env-tunable via GRADE_SLATE_LIMIT. Concurrency
stays 5 deliberately: the cap raise already multiplies load ~20x, and
concurrency decides how hard we hit statsapi at once. One variable at a
time.

Items 4/5/6 have nothing to act on and I am not manufacturing work for
them: 0 false thresholds to loosen (loosening would be manufacturing
grades); /context wiring is worth doing for grade QUALITY but would not
have graded one extra prop here, so it is not claimed as a coverage win;
archetypes are display-side and do not gate grading at all.

THE REFUSAL RATE DOES NOT DROP, AND THAT IS CORRECT. No threshold lowered,
no grade forced. The board grows because the cap stops discarding 95.7% of
the slate.

Flagged in advance rather than discovered later: snapshot payload and
ledger volume both scale with the same multiple. If the response gets
unwieldy the fix is a response-side cap on what the BOARD returns, never a
re-cap on what gets graded -- grading everything and serving a slice is
honest; grading a slice and calling it the slate is what this fixes.

Gates: 4,052 tests / 324 suites green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
Kev
2026-08-01 02:06:25 -04:00
parent d18a19f6aa
commit a7d6cf8e36
3 changed files with 293 additions and 1 deletions
+120
View File
@@ -0,0 +1,120 @@
'use strict';
/**
* refusalDiagnostics + the grade-slate cap (2026-08-01).
*
* Locks the two things the Part-1 diagnosis established:
* 1. POLICY-SUPPRESSION is not a data gap and must never be counted as one.
* 2. A no-projection refusal is only a GENUINE-ABSENCE when the player truly
* has no history — otherwise it is a wiring gap with something to fix.
*/
const { diagnose, __internals } = require('../../src/services/refusalDiagnostics');
const { uniqueGradeable } = __internals;
const prop = (player, stat, line, book = 'draftkings') => ({
player, stat_type: stat, line, book, over_odds: -110, under_odds: -110,
});
const isModelBook = (b) => ['draftkings', 'fanduel', 'betmgm', 'betrivers'].includes(b);
describe('uniqueGradeable — mirrors gradeSlateService.dedupeProps', () => {
it('keeps model books only, first row wins, one row per player+stat+line', () => {
const out = uniqueGradeable([
prop('A', 'hits', 1.5, 'prizepicks'), // DFS — never reaches the model
prop('A', 'hits', 1.5, 'betmgm'),
prop('A', 'hits', 1.5, 'draftkings'), // duplicate prop, different book
prop('B', 'rbi', 0.5, 'novig'), // exchange — display only
], isModelBook);
expect(out).toHaveLength(1);
expect(out[0]).toMatchObject({ player: 'A', book: 'betmgm' });
});
it('drops rows missing a player, stat or line rather than guessing', () => {
const out = uniqueGradeable([
{ player: null, stat_type: 'hits', line: 1.5, book: 'draftkings' },
{ player: 'A', stat_type: null, line: 1.5, book: 'draftkings' },
{ player: 'A', stat_type: 'hits', line: null, book: 'draftkings' },
], isModelBook);
expect(out).toEqual([]);
});
});
describe('diagnose — bucket separation', () => {
const baseDeps = (analyze, getStatRows) => ({
sport: 'mlb',
sample: 10,
concurrency: 2,
isModelBook,
getStatRows,
analyze,
getOdds: async () => ({
props: [
prop('Graded One', 'hits', 1.5),
prop('Suppressed One', 'doubles', 0.5),
prop('NoProj Has History', 'stolen_bases', 0.5),
prop('NoProj No History', 'triples', 0.5),
],
}),
});
const analyze = async (p) => {
if (p.stat_type === 'hits') return { grade: 'B', insufficient_data: false };
if (p.stat_type === 'doubles') {
return {
grade: null, insufficient_data: true, suppressed: true,
suppressed_reason: 'rare_event_over_below_line',
reasoning: { summary: 'no read — juiced rare-event market' },
};
}
return {
grade: null, insufficient_data: true,
reasoning: { summary: 'INSUFFICIENT DATA — no read.' },
};
};
// Only the player WITH history should count as a fixable gap.
const getStatRows = async (player, _sport, stat) => (
player === 'NoProj Has History' ? [{ date: '2026-07-30', [stat]: 1 }] : []
);
it('counts a deliberate suppression as POLICY, never as a data gap', async () => {
const r = await diagnose(baseDeps(analyze, getStatRows));
expect(r.outcome.e_POLICY_SUPPRESSION).toBe(1);
expect(r.policy_suppression_reasons).toEqual({ rare_event_over_below_line: 1 });
// A suppression must NOT be counted in the no-projection split at all —
// treating it as missing data would send us hunting for data that exists.
expect(r.no_projection_split.probed).toBe(2);
});
it('splits no-projection into FETCHABLE vs GENUINE by real history, not by assumption', async () => {
const r = await diagnose(baseDeps(analyze, getStatRows));
expect(r.no_projection_split.b_fetchable_gap).toBe(1);
expect(r.no_projection_split.d_genuine_absence).toBe(1);
expect(r.no_projection_split.fetchable_examples[0]).toMatchObject({
player: 'NoProj Has History', rows_with_stat: 1,
});
});
it('reports what the cap discards, which is the point of the exercise', async () => {
const r = await diagnose(baseDeps(analyze, getStatRows));
expect(r.slate.unique_gradeable_props).toBe(4);
expect(r.slate).toHaveProperty('capped_out');
expect(r.read_only).toBe(true);
});
it('a thrown grade is bucketed, never silently counted as graded', async () => {
const boom = async () => { throw new Error('kaboom'); };
const r = await diagnose(baseDeps(boom, getStatRows));
expect(r.outcome.graded).toBe(0);
expect(r.outcome.x_THREW).toBe(4);
});
});
describe('gradeSlateService cap', () => {
it('is raised well past the live slate size and is env-tunable', () => {
const { DEFAULT_LIMIT } = require('../../src/services/gradeSlateService').__internals;
// The measured live MLB slate carries ~585 unique gradeable props; a cap of
// 25 discarded 95.7% of the product.
expect(DEFAULT_LIMIT).toBeGreaterThanOrEqual(500);
});
});