Files
vyndr/specs/model-train.md
T
builtbykev 1a94ef5fcf Revive the dead probability layer + restore grade range ON MERIT
Folds re-sequenced steps 1+2 into one change (Kev's call): same bug
family — features wired to sources that return null.

THE PROBABILITY LAYER WAS DEAD IN PRODUCTION. p_win/ev_pct/kelly/
model_odds/value were absent on 0/8 live grades because
gameLogService.getGameLogs returns null for MLB by construction and
depends on the offline Python service for NBA/WNBA, so meta.gameLogs was
[] for every sport. This was the S46 bug in a second location — that fix
gave featureCache an MLB branch (why grades still worked) but never the
estimator. featureCache.getStatRows now supplies normalized rows
([{date,[statType]:v}], most-recent-first) for every sport, feeding the
estimator AND consistency AND game_count_in_7d from one fetch.
VERIFIED on real props: p_win 25/25 WNBA, 8/8 MLB (was 0).

GRADE RANGE, ON MERIT — never by rescaling (permanent founder ruling:
minting A's without new information is a relabelled B sold as an A and
corrupts an append-only ledger).
- refreshTeamStats wired into runSnapshot — it had ZERO production
  callers, so opp_rank_stat was permanently null and a +/-1.0 factor
  could never fire. Test-env no-op (opsNotify precedent).
- L20 made SYMMETRIC: both branches were delta +1.0, so the season
  baseline could only ever ADD. No negative path was a structural reason
  D was unreachable. New l20_contradicts_* carries -1.0.
- game_count_in_7d derived from real logged dates (heavy_workload_7d).
- NOT wired, deliberately, with reasons inline: teamId (no team_id
  column; getFeatures reads it top-level; factor also needs a starter-id
  list) and season_type (ESPN 2 = REGULAR season; threading it raw would
  fire veteran_in_playoffs in July). Dead code dressed as a fix is the
  thing we are removing, not adding.

CALIBRATION GUARD (found by verifying, not assuming): consistency CV is
NBA-tuned; for a Poisson-ish stat cv ~ 1/sqrt(mean), so any stat with
mean < 4 auto-classifies boom_bust. First verification run showed 8/8 MLB
props boom_bust — a blanket -1.0 that dropped the board to all-C. Floored
at CONSISTENCY_MIN_MEAN=4 -> 'unknown' below. Absent beats wrong. MLB
low-count stats therefore still get no consistency factor: honest, not
fixed. Scale-free index-of-dispersion classifier is the open follow-up.

CONFIDENCE IS NOT A PROBABILITY: payloads carry confidence_basis:
'grade_band'. Corrected mlb-grade-degradation.md — its "25/25
grade<->confidence agreement" is a TAUTOLOGY (confidence is derived FROM
the letter, so it would report 25/25 even if every grade were wrong), not
a validation. Removed dead mlbGrader.js (referenced only by its own test)
and the stale computeFeatures comment claiming a penalty that never ran.

VERIFICATION (scripts/verify-grade-range.js, real props/logs/engine):
WNBA 25 props B 68%->32%, C 32%->64%, D 0->1 (4%); 11-step spread went
from 2 steps to 5 (C/C+/B-/D). The D is earned: Angel Reese assists o2.5,
p_win 0.365. Nothing flooded — grades got HARDER. A did not emit locally
because opp_rank_stat needs the Redis cache only prod populates (local
ceiling +3.0 vs the +4.5 A needs); reachability is proven arithmetically
and locked in tests. Prod A-emission is the outstanding fingerprint.

MARKETING HOLD: "A-RATED" (AccuracyBadge, TopSignals) is unsupported
until that fingerprint. Confirmed honest fallbacks render today —
/api/ledger/accuracy has B and C buckets only, so the badge shows
"MODEL · 63% HIT" and TopSignals self-hides. Nothing fabricated ships.

Suite 276/3286 green, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 18:54:51 -04:00

17 KiB
Raw Blame History

MODEL TRAIN — VALUE ENGINE

Status: Arc 1 (steps 16) SHIPPED on main 7a925f4 (2026-07-19). Arcs 2+ OPEN. Written retroactively (2026-07-19) per CLAUDE.md rule #1 — Arc 1 was built from a plan that lived only in a session context that was lost. Everything in the "SHIPPED" sections below was read off the code ON DISK, not from that plan or from memory.

Doctrine (Kev): VYNDR promotes bets people actually take — roughly the 160 to +200 band — that ALSO carry a genuine vig-free edge. Not plus-money-only, not heavy chalk. Grade answers "is this a good read"; value answers "does the price pay you." They are separate fields and must stay separate.


1. WHAT ARC 1 SHIPPED

1.1 De-vig — src/utils/devig.js (new)

Two-way multiplicative (proportional / normalized-implied-probability) de-vig. Each side's implied probability is divided by the sum of both; the sum's excess over 1 is the overround. Method is recorded on the payload as devig_method:'multiplicative'.

Export Behavior
americanToImpliedProb(a) +: 100/(a+100); : a/(a+100). Null on non-finite or 0.
americanToDecimal(a) Total-return multiple incl. stake. Null on non-finite or 0.
impliedProbToAmerican(p) Null outside (0,1). p>0.5 → negative (favorite).
devigTwoWay(overOdds, underOdds) {method, overround, over:{fair_prob,fair_odds}, under:{…}}. Returns null if EITHER side is missing/invalid — fair values are never synthesized from one side.
evPct(modelProb, american) modelProb × decimal 1, as a percentage rounded to one decimal. Null on bad input.

Rounding: probabilities/overround to 3 dp (round3); ev_pct to 1 dp.

1.2 Config knobs — the ACTUAL values on disk

src/config/valueEngine.js (new — the promotion gates):

Knob Env var Default in code Meaning
TAKEABLE_ODDS_CEILING TAKEABLE_ODDS_CEILING 160 Most-juiced favorite we will PROMOTE.
TAKEABLE_ODDS_MAX TAKEABLE_ODDS_MAX +200 Longest dog we will PROMOTE.
VALUE_EV_THRESHOLD VALUE_EV_THRESHOLD 2 (i.e. 2 % EV) Minimum EV for the value flag.
  • isTakeable(american)a >= 160 && a <= +200. Strict: null, '', undefined, and non-finite are false (the Number(null)===0 fabrication trap).
  • isValue(american, evPct)isTakeable(american) && Number.isFinite(evPct) && evPct >= 2.

src/config/rareEventMarkets.js (pre-existing, from the same-day rare-event work — the refusal layer that sits UNDERNEATH the promotion gates):

Knob Env var Default in code Meaning
JUICE_ODDS_FLOOR JUICE_ODDS_FLOOR 400 Graded side at or past this price → refused, not graded.
RARE_EVENT_LINE_MAX — (constant) 0.5 Line at/below which the rare-event rules apply.
RARE_EVENT_STATS — (constant) doubles, triples, home_runs, stolen_bases, steals, blocks Backup list for props with NO odds.

Knobs that DO NOT EXIST (verified by grep across src/, web/src/, tests/): TAKEABLE_CEILING (real name is TAKEABLE_ODDS_CEILING), EDGE_FLEX_WALL, HARD_JUICE_WALL, LADDER_ODDS_MAX, MIN_RUNG_PROBABILITY. If those were in the original plan, they are unbuilt — see §3.

1.3 What the gate ACTUALLY does today

Reported as-built, not as-designed. There are two independent layers, and only the first one can refuse a grade:

Layer A — the refusal wall (JUICE_ODDS_FLOOR, 400). In analyzeViaEngine1, the FIRST thing that happens — before computeFeaturesForProp, before any projection:

if (isTooJuiced(rawProp)) → suppressedRareResult('juiced_no_edge')

Characteristics of the shipped behavior:

  • A flat price wall, NOT edge-aware. It looks only at the graded side's American price. No EV, no projection, no edge, no flex band participates in the decision.
  • The wall is 400, not 250. There is no 250 constant anywhere in the codebase.
  • It runs pre-feature, so a refused prop costs zero feature computation.
  • Odds absent → isTooJuiced is false (can't judge from a price that isn't there); the structural RARE_EVENT_STATS rule is the backup for that case.
  • Output is a refusal, not a bad grade: grade:null, insufficient_data:true, suppressed:true, suppressed_reason:'juiced_no_edge', plus branded "No read —…" copy in reasoning.summary.

The other two refusals in the same family: rare_event_under (rare stat, under, line ≤ 0.5 — pre-feature, needs no projection) and rare_event_over_below_line (rare stat, over, line ≤ 0.5, and the model does NOT project above the line — checked after the projection exists).

Layer B — the takeable band (160 … +200). This gate never refuses a grade. It is a promotion filter, applied only where a read is featured:

  • Today its ONLY enforcement point is heroPropService.pickHeroProp (if (!isTakeable(g.book_odds)) continue).
  • Every graded read still carries takeable / value as flags for consumers.
  • The full board shows every graded read regardless of band. Parlay Lab is exempt by design (juiced legs combine into takeable payouts).

So: edge-aware promotion, price-only refusal. The -400 wall decides what gets graded; EV + the 160/+200 band decide what gets promoted.

1.4 The value triplet

Attached in analyzeViaEngine1's existing p_win/Kelly try block (so it inherits the "real quantile probability × real book odds, or nothing" rule, and the whole block is additive — a throw never breaks the read):

Field Source Absent when
book_odds the graded side's actual price (under_odds if under, else over_odds) no side odds
fair_odds, fair_prob devigTwoWay(over, under) for the graded side either side unpriced
model_odds impliedProbToAmerican(pWin) no quantile probability
overround, devig_method from the de-vig either side unpriced
ev_pct evPct(pWin, sideOdds) no pWin or no side odds
takeable isTakeable(sideOdds) no side odds
value isValue(sideOdds, ev) no ev

The intended render (Design's, not built): "book 145 · vig-free 132 · model 110."

1.5 Hero v2 — heroPropService

Old rule: largest |projection line| among A/B candidates. New rule: among A/B candidates, keep only those with a finite ev_pct and a takeable book_odds, then take the highest ev_pct. The gap is still computed and returned for display, but it no longer selects. Empty-slate fallback (most recent real graded read, any grade, by timestamp) is unchanged. toHero now also passes through ev_pct / value / takeable / book_odds / fair_odds / model_odds.

1.6 Distribution + tests

  • Fields ride the existing ...result / ...data spreads into grades:{sport}, /api/snapshot/:sport, /api/hero-prop, /api/scan. Documented in BACKEND_HANDOFF.md ("Value Engine fields").
  • src/utils/tierGating.js deletes alt_lines (below Desk) and kelly (below the Kelly tier). It does not touch the value fields — ev_pct, value, takeable, and the triplet are currently visible to every tier including free. That is an unmade product decision, not a verified choice (see §3).
  • Suite: 276 suites / 3306 tests, green, next build exit 0. New: tests/unit/devig.test.js, tests/unit/valueEngine.test.js; tests/unit/heroPropService.test.js rewritten for the EV rule.

2. LAWS THIS TRAIN INHERITS (do not relax)

  • Absent beats wrong. One side unpriced → fair_odds is absent, never extrapolated. No price → takeable is absent, not false-by-coercion.
  • Number(null) === 0 is the standing fabrication bug of this codebase. Every new numeric path gets a strict guard (valueEngine.isTakeable is the model).
  • Grade ≠ value. An A read with value:false is honest output: right read, price gone. Never let the value flag re-letter a grade.
  • Refusal is a product surface, not an error. Refused reads carry branded copy.
  • Market values are REAL book numbers at a timestamp; only projection / grade / edge / EV / fair price are model output and must be labeled MODEL.

2A. ARC 2+ BOARD — STATUS (Kev's arc list, 2026-07-19)

Full arc definitions live in the Session-63 order. Status only here; update as each ships.

Item Status Note
G-a the real gate HELD for ruling Built nothing yet — the G-b report changes the recommended dials. See specs/audit-data/gate-simulation.md §2.4.
G-b gate simulation REPORTED specs/audit-data/gate-simulation.md. Replayed over the ledger, NOT snapshots (no 30d snapshot store exists).
G-c never-empty honesty open Rare but real on thin MLB nights.
S-a line-moved truth open Shares plumbing with C-clv — build together.
S-b board ranks on EV open Blocked-ish: EV not persisted; live payload has it.
L-a per-rung odds? REPORTED — NO Rungs are synthetic, carry no price. Feed offers no alternate markets. §3 of the report.
L-b per-rung EV ladder BLOCKED Needs a real price per rung. Probe PropLine first.
C-cal calibration REPORTED Confidence monotonic but ~20-25pts miscalibrated; only B/C grades ever emitted.
C-led ledger speaks value open — recommended next locked_odds already 99.1% populated: units/ROI need no backfill. EV columns are net-new, forward-only.
C-clv fix C4 open Confirmed broken in data (359/376 MLB closes == lock). Keep suppressed.
D-ref visible refusals partial Copy already ships on refusals; new gate_* reasons land with G-a.
D-par parlay lab open Gate must NOT apply inside the Lab.
D-tier tier gating DECIDED, not enforced Free/Analyst: value marker + grade + triplet. Desk: ladder + per-rung EV + Kelly. Today ladder/Kelly are already Desk-gated; triplet ungated = correct per this ruling.
D-ev consolidate EV open devig.evPct vs processing/EVCalculator.js (used only by UnifiedOddsProvider).
U-deg MLB degradation STATUS REPORTED projection==0 leak already closed (0 occurrences since 07-18). edge_pct scale still broken.
U-fp Arc 1 fingerprint open Do it on the first deploy this train ships.

SESSION 63 — PROBABILITY LAYER + GRADE RANGE RESTORED (shipped)

The re-sequenced step 1+2, folded into one change. Full write-up: specs/audit-data/grade-collapse.md.

  • The probability layer was DEAD in productionp_win/ev_pct/kelly/ model_odds/value were absent on 0/8 live grades because gameLogService returns null for MLB by construction and the Python service is offline for NBA/WNBA. featureCache.getStatRows now supplies normalized rows for every sport. Verified: p_win 25/25 on real WNBA props, 8/8 MLB (was 0).
  • Hero v2 had never once selected on EV (it requires a finite ev_pct) and silently fell through to the recent-read fallback every time.
  • Grade range: refreshTeamStats wired into runSnapshot (it had ZERO callers, so opp_rank_stat was permanently null), game_count_in_7d derived from real logs, and L20 made symmetric (there was no negative branch at all). D now emits on merit (WNBA 1/25, an earned p_win 0.365); A is proven reachable arithmetically but has not yet emitted in production — that is the outstanding fingerprint.
  • Calibration guard: consistency CV was NBA-tuned; for any stat with mean < 4, cv ≈ 1/√mean forces boom_bust. It would have stamped a blanket 1.0 on nearly every MLB prop. Floored at CONSISTENCY_MIN_MEAN=4unknown below it.
  • Confidence is NOT a probability — payloads now carry confidence_basis: 'grade_band'. The real signal is p_win.
  • mlbGrader.js REMOVED (dead; referenced only by its own test).
  • 🔴 MARKETING HOLD: "A-RATED" copy (AccuracyBadge, TopSignals) is unsupported until a production fingerprint shows real A grades. Honest fallbacks confirmed rendering ("MODEL · 63% HIT"); nothing fabricated ships.

🔶 OPEN DECISION — FLEX BAND ENFORCEMENT (Kev, 2026-07-19)

Ruling: build EDGE_FLEX_WALL (250) + EV_FLEX_THRESHOLD (default 4 %, = 2× VALUE_EV_THRESHOLD) but ship with EV_FLEX_ENFORCE=0. Flex-band props (161…−250) grade exactly as they do today — the band is NOT cut. The ledger shows it is our most profitable segment (+2.2 % ROI, n=70) and we will not restrict a proven-profitable band on an unvalidated threshold. Record ev_pct on every flex prop now so real in-production data accumulates.

TRIGGER TO REVISIT: once ~2 weeks of production ev_pct data exists on the 161…−250 band, report the ev_pct distribution vs settled outcomes for that band, then Kev decides the threshold — or whether to enforce at all. Do NOT flip EV_FLEX_ENFORCE=1 without that report and explicit sign-off. It comes back as a data-backed decision, never a silent flip.

Unaffected — these ship and enforce normally: HARD_JUICE_WALL (250, never grades), LADDER_ODDS_MAX (+400), MIN_RUNG_PROBABILITY (0.25), the no-odds refusal, and the folded projection > 0 check.

Backtest harness

None exists (grep-verified). migrations/006 defines grade_outcomes + player_calibrated_weights and no code reads or writes them. The G-b/C-cal replay was done in SQL against ledger_entries; a real harness is still owed before any weight change.

3. NOT YET BUILT — checklist to reconcile against the full arc list

Verified absent from the codebase as of 7a925f4. Kev supplies the complete arc list; mark each done/open against this.

Engine / gates

  • Edge-aware juice wall — a soft/flex band where a big enough edge can survive a juiced price (EDGE_FLEX_WALL, HARD_JUICE_WALL — neither exists; today's wall is the single flat JUICE_ODDS_FLOOR 400).
  • TAKEABLE_ODDS_* enforced on any surface beyond the daily hero (featured rows, top-of-board, alerts). Today: hero only.
  • Alt-line ladder value gates (LADDER_ODDS_MAX, MIN_RUNG_PROBABILITY — neither exists). The ladder in analyzeViaEngine1 is still the S62 build: fixed shifts [1, 0.5, 0, +0.5, +1], filtered to line > 0, sorted by edge_pct. It has no odds ceiling and no per-rung probability floor, and its rungs carry no ev_pct / fair_odds / price at all.
  • EV-rank the board itself. slateAdapter.flattenToEdgeBoard still ranks on edge (guarded by EDGE_BOARD_SANE_MAX = 40), not ev_pct.
  • Reconcile the orphan src/services/processing/EVCalculator.js (a separate, older EV path used only by UnifiedOddsProvider) against devig.evPct — two EV implementations currently coexist.
  • Multi-way / three-way de-vig, and alternatives to multiplicative (Shin, power/log). Two-way multiplicative is the only method implemented.
  • Best-price de-vig across books (today the triplet de-vigs whichever single book's two-sided price came through on the prop).

Surfaces (nothing on the frontend reads these fields yet — grep-verified zero hits for ev_pct / fair_odds / model_odds / book_odds / suppressed_reason under web/src/)

  • Value triplet rendered on the reveal card / board row / hero.
  • VALUE marker (and the honest "A but no value" state).
  • Visible-refusal moment — the "No read —…" copy is generated and shipped on the payload but is not displayed anywhere.
  • EV shown as the headline number instead of / alongside edge %.

Product / policy

  • Tier policy for the value fields — currently ungated to free (§1.6).
  • Alerts on takeable +EV reads (the config comment anticipates "future alerts").
  • Ledger/record split by value (does the model's +EV subset beat its overall record?). ledger_entries stores no EV or fair price today.
  • Settlement/CLV interaction with fair price — CLV is still locked-vs-closing LINE, not vs closing FAIR price. (Note: C4 — CLV capture is broken and suppressed on public surfaces; see specs/audit-data/clv-capture-broken.md. Any fair-price CLV work is blocked behind C4.)

Verification debt

  • Arc 1 is not deploy-fingerprinted. Post-deploy, confirm /api/hero-prop returns a hero with a takeable book_odds and a finite ev_pct, and that /api/snapshot/mlb grades carry the triplet.
  • Known upstream contamination: specs/audit-data/mlb-grade-degradation.md — projection == 0 on ~9/25 MLB grades and a broken edge_pct scale. EV is computed from the quantile probability, not from edge_pct, so it is not directly poisoned — but hero ranking and any EV-vs-edge comparison sit on top of that pipeline. Fix the degradation before trusting arc-1 numbers in aggregate.

Arc 1 recorded 2026-07-19 from the code at 7a925f4. Update this file in the same commit as any arc that ships.