Completes the diagnosis. Report only; no grade logic or thresholds changed.
The grade is an integer index (GRADE_SCALE, NEUTRAL_INDEX 3) moved by a
flat sum of +/-1.0 and +/-0.5 factor deltas, then clamped and rounded.
grade_thresholds.json is NOT an input mapper in the JS path — engine1
reads it BACKWARDS, taking the letter the index already produced and
looking up that band's midpoint to manufacture `confidence`. So
confidence is a cosmetic re-encoding of the letter: zero information
beyond it, and it can never disagree with it. There is no
data-sufficiency penalty in the live path (the one CLAUDE.md describes is
in mlbGrader.js, which is dead code).
Six of thirteen factors are wired to features nothing populates —
verified: refreshTeamStats has ZERO production callers (so opp_rank_stat
is permanently null, killing a +/-1.0), teamId/season_type/
game_count_in_7d are never passed (gameContext is built as {home_away}
and nothing else), and MLB consistency starves on the same dead
gameLogService path as Finding 2. Also verified: BOTH l20 branches are
delta +1.0 — there is no negative L20 contribution at all.
Arithmetic: an A needs sum >= +4.5; the live maximum is +3.0 (+2.0 on a
back-to-back, and MLB rest_days is 0 most days). D needs <= -1.51; the
live minimum is -1.5 and Math.round(1.5)=2, so it misses by one rounding
tick. Reachable band is index 2..6 = {C-,C,C+,B-,B}, which the adapter's
FOUR_LETTER_MAP (a 3->1 collapse) renders as exactly {C,B} — the observed
output, derived from first principles. Reachable confidences {42,47,52,
57,63} match the live values {47,52,57,63} exactly; C- is truncated by
gradeSlateService keeping the higher-confidence side.
mlb-grade-degradation.md's "25/25 grade<->confidence agreement" is a
TAUTOLOGY, not a validation — confidence is derived from the letter, so
it would report 25/25 even if every grade were wrong.
Recommends feeding the starving factors (restores A/D on merit) and
explicitly REJECTS re-scaling thresholds, which would mint A's without
adding information — every "A" would be a relabelled B.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
13 KiB
GRADE COLLAPSE + DEAD PROBABILITY LAYER — diagnosis
REPORT ONLY. No grade logic, thresholds, or engine code changed (Kev's instruction). Two findings. The second one is bigger than the question I was asked.
Data: live Supabase ledger_entries (604 rows, all users) + live prod API
api.vyndr.app, 2026-07-19 ~22:05 UTC.
FINDING 1 — THE COLLAPSE IS REAL, LIVE, AND STRUCTURAL
Not a thin-slate artifact. Across 604 ledger rows and both sports:
- 2 distinct grades ever emitted: B and C. Zero A+, A, A−, B+, B−, C−, D, F.
- 9 distinct confidence values ever emitted: 63, 57, 55, 52, 47, 45, 35, 25, 20.
- Confidence ceiling = 63. It has never exceeded 63 in the recorded era.
Still true today (Jul 18 + 19, post every fix, both sports): 4 confidence values (63/57/52/47), 2 grades.
| Sport | Grade | n | conf min | conf max |
|---|---|---|---|---|
| mlb | B | 276 | 45 | 63 |
| mlb | C | 107 | 20 | 52 |
| wnba | B | 130 | 45 | 63 |
| wnba | C | 91 | 35 | 52 |
Confidence does NOT determine the letter
| conf | grade | n |
|---|---|---|
| 63 | B | 54 |
| 57 | B | 175 |
| 55 | B | 29 |
| 52 | C | 100 |
| 47 | C | 14 |
| 45 | B | 148 |
| 35 | C | 80 |
conf 45 → B, but conf 47 and 52 → C. The mapping is non-monotonic, so the surfaced
confidence is not the quantity the letter was derived from. This confirms the
mlb-grade-degradation.md "grade↔confidence mismatch" as a display artifact: two
different quantities are being shown as if one explains the other.
(At conf 45→B the avg edge is 103; at conf 52→C it is 49 — so the letter tracks the engine composite/edge, not the displayed confidence.)
Collateral: the edge scale is still broken and still live
- 311 of 604 rows (51.5 %) have |edge| > 40 — the frontend's
EDGE_BOARD_SANE_MAX, i.e. over half the board's edge is nulled at render. - 39 rows have |edge| > 100 — impossible as a percentage. Worst: 620.
- Live today: edges of 140, 180, 220 on Jul 18–19 rows.
U-deg's projection == 0 leak IS closed (0 since 07-18). The edge_pct scale is
not — it remains open and is now quantified.
FINDING 2 — 🔴 THE ENTIRE PROBABILITY LAYER IS DEAD IN PRODUCTION
Found while fingerprinting Arc 1 (U-fp). This is the headline.
Live fingerprint, GET /api/snapshot/mlb, 8 graded props
| Field | Present |
|---|---|
projection, confidence, book_odds, fair_odds, takeable, devig_method, alt_lines |
8 / 8 |
p_win |
0 / 8 |
kelly |
0 / 8 |
ev_pct |
0 / 8 |
model_odds |
0 / 8 |
value |
0 / 8 |
Control: alt_lines is present 8/8 and is Desk-gated in tierGating.js:55, which
proves the payload is not being tier-stripped. These fields are genuinely never
computed — not hidden.
Root cause — a one-line sport gate, and an S46 fix that was only half-applied
analyzeViaEngine1.js:509 feeds the estimator from meta.gameLogs:
const est = estimateProbability({ gameLogs: meta.gameLogs, line: prop.line, ... });
meta.gameLogs comes from computeFeatures.js:173-181 → gameLogService.getGameLogs.
And gameLogService.js:21-26:
function pythonPath(sport) {
switch (sport) {
case 'nba': return '/stats/last-n';
case 'wnba': return '/wnba/stats/last-n';
default: return null; // ← MLB exits here
}
}
with getGameLogs line 31: if (!path) return null;
So:
- MLB — returns
nullby construction. Never had game logs on this path. - NBA/WNBA — hits the Python stats service, which is offline in prod (documented in CLAUDE.md; degrades to null).
⇒ meta.gameLogs is [] for every sport in production ⇒
estimateProbability returns {p_over: null, reason:'insufficient_data'}
(probabilityEstimator.js:55-57) ⇒ pWin is null ⇒ every field guarded by
if (pWin != null) is skipped: p_win, kelly, model_odds, ev_pct, value.
This is the S46 bug, second location, never fixed. CLAUDE.md records that
gameLogService.getGameLogs being NBA/WNBA-only starved MLB, and that the fix was an
MLB branch in featureCache.gameLogFeatures. That fixed the feature path — which
is why projection, confidence, and grades still work. The estimator path was never
given the same branch, so it has been silently dead the whole time.
What this actually breaks
- EV — the Model Train's entire ranking signal — does not exist in production.
Arc 1 shipped
ev_pctand it has never once been computed on a live prop. - Hero v2 is non-functional.
pickHeroProprequires a finiteev_pct(heroPropService.js:84), so the EV loop matches nothing and always falls through to the "most recent graded read" fallback. Live proof:/api/hero-propreturns"is_recent": true— the fallback path, every time. The hero has not been an EV pick since the day it shipped. - Quarter-Kelly is dead — same
pWindependency (analyzeViaEngine1.js:516-520). This is a promise-audit issue: Kelly sizing is sold on the pricing page andPROMISE-AUDIT.mdlists it as BUILT. It is built and never runs. - The "value triplet" is a duet live —
book_odds+fair_oddsrender;model_oddsis always absent. valueis never true, so the VALUE marker can never light up.
Why this reframes the whole train
- C-led would persist a column of nulls. Do not build EV persistence until EV exists.
- G-a's
EV_FLEX_THRESHOLDwould gate on a permanently-null value. WithEV_FLEX_ENFORCE=0(Kev's ruling) this is harmless today — but had we enforced it, the flex band would have been cut to zero, becauseev_pct >= 4can never be true. The ruling to ship it disabled accidentally prevented an outage. - S-b (rank board on EV) would rank on nulls.
RELATIONSHIP BETWEEN THE TWO FINDINGS
They are adjacent, not identical, and both trace to the same missing input:
- The dead estimator explains why no probability-derived output exists (EV, Kelly, model_odds, p_win).
- It does not by itself explain the B/C letter collapse, because the letter comes from engine1's rule-based composite over the feature vector, which is alive.
- But they share a root: the model is running on a partial input set. One of its two probability inputs (the empirical quantile distribution over real game logs) is absent for 100 % of props, so whatever spread the composite was designed to produce is being generated from the surviving features only.
The 9-discrete-confidence-values pattern is a small set of additive rule hits — a scorer landing on a lattice rather than a continuum. Mechanism now traced in full below.
FINDING 3 — THE MECHANISM: A IS MATHEMATICALLY UNREACHABLE
Every claim here was verified directly against the source.
The grade is an integer index, not a score
engine1.js:16-17, 158-163:
const GRADE_SCALE = ['F','D','C-','C','C+','B-','B','B+','A-','A','A+'];
const NEUTRAL_INDEX = 3; // 'C'
...
let idx = NEUTRAL_INDEX;
for (const f of factors) idx += f.delta; // flat sum of ±1.0 / ±0.5
idx = clampIndex(Math.round(idx));
grade_thresholds.json is not an input mapper in the JS path. Nothing compares a
probability to those cutoffs. engine1.js:29-36 reads the table backwards — it takes
the letter the index already produced and looks up that band's midpoint to
manufacture a confidence number.
So confidence is a cosmetic re-encoding of the letter. It carries zero information
beyond the letter and by construction can never disagree with it. There is no
data-sufficiency penalty in the live path — the one CLAUDE.md describes lives in
mlbGrader.js:50-69, which is DEAD CODE. (computeFeatures.js:21 still carries a stale
comment claiming the adapter downgrades confidence; it does not.)
Six of thirteen factors are wired to features nothing populates
| Dead factor | Δ | Why it never fires |
|---|---|---|
weak/top_opponent_defense |
±1.0 | needs opp_rank_stat ← team_stats:{sport}:{abbr} ← refreshTeamStats has ZERO production callers (verified: only its own export + tests) |
consistency_elite/boom_bust |
±1.0 | MLB consistency logs come from the same dead gameLogService path as Finding 2 |
opp_starters_out |
+1.0/+0.5 | featureCache.js:280: if (!teamId) return out; — computeFeatures never passes teamId |
| playoff factors | ±0.5 | season_type never set |
heavy_workload_7d |
−0.5 | game_count_in_7d never set |
ref_* / coach_* |
±0.5 | NBA-flavored caches, absent for MLB |
computeFeatures.js:234-236 builds gameContext as { home_away } and nothing else.
The arithmetic
idx = clamp(round(3 + Σδ)). Live-firing factors for MLB reduce to: l5_* (±1.0),
l20_* (+1.0 only — verified, BOTH branches are delta: 1.0, there is no negative L20
contribution), home_game (+0.5), rest (±0.5), trap_composite_high (−1.0).
| 4-letter | needs Σδ | live reachable? |
|---|---|---|
| A (A−/A/A+) | ≥ +4.5 | NO — live max is +3.0 (+2.0 on a back-to-back, and MLB rest_days is 0 most days) |
| B | +1.5 … +4.49 | yes |
| C | −1.5 … +1.49 | yes |
| D | ≤ −1.51 | NO — live min is −1.5, and Math.round(1.5) = 2 → C−. Misses by one rounding tick. |
| F | ≤ −2.51 | NO |
An A is short by at least 1.5 index steps — and the ≥1.5 of deltas that would close the
gap (opp_rank_stat ±1.0, consistency ±1.0, injury +1.0) are exactly the permanently-
null features. The reachable index band is 2…6 = {C−, C, C+, B−, B}, which
gradeAdapter.FOUR_LETTER_MAP (gradeAdapter.js:31-37, a 3→1 collapse) renders as
exactly {C, B}. That is the observed output, derived from first principles.
Confidence corroborates exactly: reachable letters carry {42, 47, 52, 57, 63}. Live
today we observe precisely {47, 52, 57, 63} — C− (42) is absent because
gradeSlateService.js:76 keeps the higher-confidence side of each prop, truncating the
bottom. The older values in the ledger (55, 45, 35, 25, 20) are from the pre-888d103
hand-rolled table {10,15,20,25,35,45,55,65,80,90,100} — the ledger is append-only, so
it contains both eras.
Relationship to mlb-grade-degradation.md
Shared table, different bug — and its "fix" made this collapse invisible. That audit
redefined confidence as the band midpoint so the letter round-trips through the table.
The resulting "25/25 agreement" is a tautology, not a validation: confidence is
derived from the letter, so it would report 25/25 even if every grade were wrong. That
audit only examined the output encoding. This collapse is one layer upstream, on the
input side — whether computeFactors has enough live features to move the index at all.
RECOMMENDATION (no code changed pending Kev's call)
Re-sequence: fix the dead estimator FIRST — before G-a, before C-led.
Rationale: it is the cheapest fix on the board (an MLB branch in the estimator's log
source, mirroring the one already written for featureCache), and it simultaneously
restores EV, Kelly, model_odds, the VALUE flag, and hero v2. Every other Arc 2-5 item
is downstream of it. Building the gate, the persistence layer, or the board ranking on a
null signal is building on nothing.
Suggested order:
- Revive the probability layer (MLB branch + a real NBA/WNBA fallback, since Python
is offline). Fingerprint that
p_win/ev_pctappear live. - Then C-led — persist EV that now has values.
- Then G-a — with the flex band still disabled per the standing ruling.
- Then the grade range — now diagnosed (Finding 3), and it is NOT primarily a
consequence of step 1. It needs its own decision, because there are two very different
fixes and picking wrong bakes in a lie:
- (a) Feed the starving factors. Call
refreshTeamStats(nothing does), passteamId/season_type/game_count_in_7dthroughgameContext, givesafeGetConsistencythe same MLB branch as step 1. This restores ±3.0 of range and makes A/D reachable on merit. - (b) Re-scale the index/thresholds so the current narrow spread spans more letters. This is the tempting one and it is the wrong one — it would mint A's without adding a single bit of information, and every "A" would be a relabelled B. It converts a visible limitation into an invisible lie. Recommend (a), explicitly reject (b). If (a) proves infeasible, the honest fallback is to keep the two-letter output and stop advertising a scale we don't produce — not to stretch the scale.
- (a) Feed the starving factors. Call
Copy consequence, either way: "A-RATED" appears on public surfaces and AccuracyBadge
for a grade the engine has never emitted. Until (a) lands, that copy is unsupported.
Open question for Kev: NBA/WNBA have no free game-log source on this path with Python
down. espnStatsAdapter.getPlayerGameLog (Wave 0) already solves exactly this for
featureCache — reusing it here is the obvious candidate, and costs no quota.
Diagnosis 2026-07-19. Data + live API. No engine code changed.