Files
vyndr/specs/grade-cap-and-refusal-diagnosis.md
builtbykev 11b0139481 Verify the cap raise on prod: 7 -> 365 graded props
Induced, not projected. DEFAULT_LIMIT=500 produced 365 graded props in
114s (was 7 in 16s) -- 52x the board. All 365 carry a unique forecast_rank
and ZERO leak p_win to anonymous callers, so the tier gating holds at 50x
the volume. Anon payload 220KB in 0.44s. Stat mix went from three stats to
ten. Health green.

Measured cost curve via the ?limit= bisect hook: 1->42s, 25->58s, 60->42s,
120->66s, 500->114s. About 42s of that is FIXED overhead (odds fetch,
roster logs, archetype classify, retention), paid whether we grade 1 prop
or 500 -- grading is the cheap part.

MY PRE-FLIGHT ESTIMATE WAS WRONG. I predicted ~72s from per-prop latency
measured in isolation, which ignored the fixed cost. Real figure 114s.

A FALSE ALARM RECORDED because acting on it would have meant reverting a
fix that works: the first induced run 502'd at 13.4s and I hypothesised
load -- memory or a proxy timeout under 20x the work. Wrong. A limit=25 run
then 502'd in 2 seconds, which no amount of load explains, and both
recovered on retry. The 502s were the deploy rolling, not the cap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:22:30 -04:00

191 lines
7.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# THE 25-CAP + THE "72% REFUSAL" — DIAGNOSIS AND FIX
**Date:** 2026-08-01 · Part 1 **read-only, measured on prod** · Part 2 fix gated
on it. Queries/tooling committed (`src/services/refusalDiagnostics.js`,
`GET /api/internal/diagnose-refusals`).
---
## VERDICT UP FRONT — the premise is half right, and the wrong half is the one we were chasing
> **The refusal rate is not a data problem. It is 98% correct behaviour.**
> **The cap is the entire problem, and it is worse than "25 of 546".**
Measured on a live MLB slate (n=80 graded through the real path):
| bucket | count | share | verdict |
|---|---:|---:|---|
| **GRADED** | **44** | **55.0%** | — |
| **(e) POLICY-SUPPRESSION** | **35** | **43.8%** | ✅ **correct by design — keep** |
| (b) FETCHABLE-GAP | **1** | 1.3% | fixable, but n=1 |
| (a) FALSE-THRESHOLD | **0** | 0% | nothing to fix |
| (c) ARCHETYPE-GAP | **0** | 0% | archetypes do not gate grading at all |
| (d) GENUINE-ABSENCE | **0** | 0% | — |
**There is no meaningful data gap.** Part 2 items 4, 5 and 6 (fix thresholds,
wire `/context`, build archetypes) have **nothing to act on** — the diagnosis
says so, and I am not going to manufacture work for them.
---
## PART 1a — THE FIFTH BUCKET (the order did not anticipate this, and it changes the reading)
**All 35 "refusals" are `rare_event_over_below_line`.** They are not missing
data. They are the **2026-07-19 betting-logic audit deliberately refusing
rare-event 0.5 markets** — and they set the **same `insufficient_data: true`
flag** as a genuine data gap, which is why they read as one.
From the code that produces them:
> *"a rare-event 0.5 OVER is a read ONLY when the model genuinely projects the
> event ABOVE the line. Below that, the over carries the same |edge| as the
> (already-suppressed) under and would just take its place on the board — so
> refuse it."*
Entirely concentrated in two stats:
| stat | graded | suppressed |
|---|---:|---:|
| `doubles` | 1 | **18** |
| `stolen_bases` | 0 | **17** |
| `hits` | **19** | 0 |
| `rbi` | **19** | 0 |
| `total_bases` | **5** | 0 |
**The stats that can be graded grade at ~100%.** The refusals are entirely
0.5-line rare events, refused on purpose.
**Had we "fixed" this**, we would have re-introduced exactly the bets a previous
audit removed — and the count would have looked like progress. This is the
single most important thing in the diagnosis.
The one real gap (bucket b) is `Miguel Rojas / stolen_bases`, who has **10 log
rows with the stat** but produced no projection. Real, worth a look, **n=1** — not
a program of work.
---
## PART 1b — THE CAP, AND WHAT IT ACTUALLY COSTS
| | |
|---|---:|
| rows in feed | 11,737 |
| **unique gradeable props** | **585** |
| **cap** | **25** |
| **discarded by the cap** | **560 — 95.7%** |
**Why 25?** Traced to Session 32 (`f0c8b4f`), commented *"we still bound the
herd."* It is a **herd guard written before anyone measured what a grade costs**
not a performance limit, not a cost limit, not a display choice. Nothing was ever
measured against it.
**So I measured it.** Per grade, through the real path:
| mean | median | p90 | max |
|---:|---:|---:|---:|
| **721 ms** | 666 ms | 1,024 ms | 1,730 ms |
**~72s for 500 props at concurrency 5.** 80 props took **11.9s** wall.
**Both callers tolerate that comfortably:** the snapshot cron runs 5×/day, and
`oddsService.recordDownstream` is fire-and-forget and never holds an HTTP
response.
---
## PART 2 — THE FIX (item 3 only, because that is what the diagnosis supports)
**`DEFAULT_LIMIT` 25 → 500**, env-tunable via `GRADE_SLATE_LIMIT`.
**`DEFAULT_CONCURRENCY` left at 5, deliberately.** Raising the cap already
multiplies total load ~20×, and concurrency is the knob that decides how hard we
hit statsapi *at once*. One variable at a time.
**Items 4, 5, 6, 7 — no action, and the reason is the measurement:**
- **(4) false-thresholds — 0 found.** Nothing to loosen. Loosening anything here
would be manufacturing grades, which the order's absolute rule forbids and
which this whole run has been removing.
- **(5) fetchable gaps — 1 of 80.** `/context` wiring is worth doing for *grade
quality* (umpire, roof, handedness) but it is **not** a coverage fix; it would
not have graded a single additional prop in this sample. Keeping it out of this
order rather than claiming it as a win.
- **(6) archetype-gap — 0.** Archetypes are display-side and **do not gate
grading at all**. Building MLB batter archetypes remains worthwhile for the
model, but it is **not** what is suppressing the board, and presenting it as
such would have been a misdiagnosis.
- **(7) genuine-absence — still refusing.** Unchanged.
### The refusal rate does not drop, and that is the correct outcome
**No threshold was lowered. No grade was forced.** The board grows because the
cap stops discarding 95.7% of the slate — not because the bar moved.
---
## VERIFIED EFFECT (induced on prod, not projected)
| | before | after |
|---|---:|---:|
| **graded props on the MLB board** | **7** | **365** |
| ledger rows per run | 7 | 365 |
| retention rows | 50 | 1,000 |
| snapshot duration | 16s | **114s** |
| anon `/api/snapshot/mlb` payload | ~6 KB | **220 KB, 0.44s** |
**52× the board.** Every one of the 365 carries a `forecast_rank` (1365, all
unique) and **zero leak `p_win` to anonymous callers** — the gating still holds
at 50× the volume.
Stat mix, previously three stats, now ten: hits 141 · total_bases 78 · rbi 42 ·
runs 41 · walks 31 · strikeouts 12 · outs 11 · earned_runs 4 · hits_allowed 3 ·
doubles 2.
### Measured cost curve (induced runs, `?limit=` bisect hook)
| limit | grades | wall |
|---:|---:|---:|
| 1 | 0 | 42s |
| 25 | 7 | 58s |
| 60 | 39 | 42s |
| 120 | 63 | 66s |
| **500** | **365** | **114s** |
**~42s of that is FIXED overhead** (odds fetch, roster logs, archetype classify,
retention) — it is paid whether we grade 1 prop or 500. Grading is the cheap
part. **My pre-flight estimate of "~72s" was built from per-prop latency measured
in isolation and ignored that fixed cost; the real figure is 114s.** Wrong by
~40s, in the direction that mattered least, but wrong.
### A false alarm I chased, and what it actually was
The first induced run **502'd at 13.4s**, and did not complete in background. I
hypothesised load — memory or a proxy timeout under 20× the work. **That was
wrong.** A `limit=25` run then 502'd in **2 seconds**, which no amount of load
explains, and both recovered on retry: the 502s were **the deploy rolling**, not
the cap. Recorded because the load hypothesis was plausible enough to have been
acted on, and acting on it would have meant reverting a fix that works.
**Two things to watch after deploy, stated in advance rather than discovered
later:**
1. **Payload size.** `snapshot:{sport}:latest` and `/api/snapshot/mlb` currently
carry 7 grades; at ~300 they will be substantially larger. If the response
becomes unwieldy, the fix is a response-side cap on what the *board* returns —
**not** a re-cap on what gets graded. Grading everything and serving a slice
is honest; grading a slice and calling it the slate is what we just fixed.
2. **Ledger volume.** `recordPipelineGrades` writes one row per graded prop, so
ledger growth scales with the same multiple. That is desirable — it is the
accrual the model verdict needs — but it changes the shape of the table.
## TAGS
**VERIFIED:** 585 unique gradeable props vs a cap of 25 · per-grade cost
721ms mean · 43.8% of refusals are deliberate policy suppression · 0 false
thresholds · 0 archetype-gated refusals · the gradeable stats grade at ~100%.
**CORRECTED:** "72% refuse for insufficient_data" — the flag is shared between a
genuine data gap and a deliberate suppression, and it is almost entirely the
latter. The product was never failing to grade for want of data; it was
discarding 95.7% of the slate before grading began.