Work-order #4 closed. First post-deploy snapshot (2026-07-17 14:00:53 UTC) re-graded with the fix. Before → after: - projection==0: 9/25 → 0/25 (the nine now refuse) - grade<->confidence: mismatch → 25/25 agree - edge_pct: {20,60,100,140} cluster → 7 continuous values, all-positive projections The lone remaining edge=100 is a REAL projection (Abreu hits, line 0.5, proj 1.0 over = 100% by (model-line)/line), not the old proj=0 degeneracy. Verified. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
4.7 KiB
MLB Grade Pipeline Degradation — FIXED (2026-07-17)
Source: phone-audit P1-7 (broken edge board) + P2-9 (B grades at 45% confidence).
Diagnosed against LIVE GET /api/snapshot/mlb on 2026-07-17. Backend grading
bug, fixed at the source in the generic grade path (engine1 +
analyzeViaEngine1), which grades EVERY sport.
Before (25 live MLB grades, degraded)
- projection == 0 for 9/25 — graded on a zero projection.
- edge_pct quantized to {20, 60, 100, 140} — the 100s were the proj=0
degeneracy
(line - 0)/line = 100%. - grade ↔ confidence mismatch — 10/25 disagreed even at the 4-letter level (25/25 vs the stricter 11-step bands).
Root causes + fixes (commit 888d103)
- projection=0 bypassed the refusal gate.
projectionForreturnedl5_avgeven when 0 (finite → the== nullgate passed it). FIX: a non-positive reference is not a projection —projectionForskips it and falls through to the next POSITIVE reference (l5 → l20 → per_90 → xg); when none is positive it returns null and the read REFUSES (insufficient_data). The gate also gained an explicit> 0guard — the invariant is structural. - edge_pct. Formula was already
(model - line)/line(the intended semantics); the {100} cluster was purely the proj=0 degeneracy. With fix 1 those refuse. Main-line edge now reuses the VALIDATED projection so edge and the persistedprojectioncan't diverge. - confidence/letter split.
engine1.GRADE_TO_CONFIDENCEwas hand-rolled and drifted a full sub-tier low (B → 0.55, whichgrade_thresholds.jsoncalls B-). FIX: confidence is now DERIVED from each grade's band MIDPOINT ingrade_thresholds.json— one source of truth. Applying the threshold table to any grade's displayed confidence resolves back to the same letter (proven for all 11 grades intests/unit/mlbGradeDegradation.test.js).
Blast radius (commit 9fc4edf) — work-order #6
The degraded grades (projection=0 → model_value = 0) are already settled in
the append-only ledger_entries and are NOT deleted. Functional marking:
getModelAggregate now filters .gt('model_value', 0) on the settled AND
pending queries — the rows stay in the ledger but leave the public model record
(their hit/miss is noise, not skill). Post-fix no new such row can be written.
Exact count NOT queryable from the dev box (*.supabase.co is unreachable
here — curl 000; only vyndr.app/api.vyndr.app resolve; no VYNDR_INTERNAL_KEY
locally). Proxy signal: 9/25 (36%) of the current live slate. For the precise
figure, run in Supabase SQL:
SELECT count(*) FILTER (WHERE outcome IS NOT NULL) AS settled_degraded,
count(*) AS total_degraded
FROM ledger_entries
WHERE user_id IS NULL AND model_value = 0;
Other sports — work-order #5
NBA/WNBA/soccer grade through the SAME analyzeViaEngine1 → engine1 path
(gradeSlateService does not branch by sport; mlbGrader.js is dead code). So
they SHARE the disease and are fixed by the same commit. They rarely grade in
prod today (stats service offline off-season → refuse anyway). No separate fix.
Live validation — work-order #4
The fix deploys immediately, but the SNAPSHOT only re-grades on the full cron (UTC hours 14,19,22,1,3). Run after the next 14:00 UTC snapshot post-deploy:
node scripts/validate-grade-fix.js
PASS criteria: projection==0 count → 0; edge_pct no longer contains 100 and is not the four-value cluster; grade↔confidence agreement 25/25. The before-state (this file's "Before") is the diff baseline; the script's output is the fingerprint.
AFTER — VALIDATED LIVE (2026-07-17 14:00:53 UTC regrade, fingerprint)
Ran scripts/validate-grade-fix.js against the first post-deploy snapshot.
| Signal | Before (03:00 UTC) | After (14:00 UTC) |
|---|---|---|
| projection == 0 | 9 / 25 | 0 / 25 ✅ |
| grade ↔ confidence agree | 15 / 25 (4-letter); 0/25 vs 11-step | 25 / 25 ✅ |
| edge_pct distinct values | 4 — {20, 60, 100, 140} | 7 — {20, 60, 70, 76, 78, 82, 100} ✅ |
| projection distribution | contained 0 | all positive (min 0.09, no zeros) ✅ |
The nine projection-0 grades vanished (those props now refuse). Grade and
confidence agree on every row. Edges are continuous, not the degenerate cluster.
The single remaining 100 is NOT the old bug: Wilyer Abreu · hits · line 0.5 ·
projection 1.0 · over → (1.0 − 0.5)/0.5 = 100%, a real model call of double
a small line. The (model − line)/line metric inherently produces large % on
0.5-step lines — the frontend |edge| > 40 guard is the intended safety net for
exactly that, and it stays. All three bugs resolved; validated in production.