NORTH STAR (design philosophy, not built): VYNDR measures players by MODERN FUNCTION, not legacy label — the principle already under the archetype system, from Rashad Phillips' Basketball Position Metric. The rule: every proprietary metric is baselined against the player's functional ARCHETYPE's CURRENT-SEASON behavior, never the position's inherited standard. The edge is that the market often prices today's players against yesterday's baselines, so archetype-vs-position baseline disagreement is a repeatable mispricing. Generalizes across sports. Moat = proprietary metrics x current-game calibration x our private outcome data. Metrics ship as VALIDATED FAMILIES: hypothesis, flagged build, backtest, ship-or-delete with the negative result written down. Nothing is real until the harness proves it predicts better. SOURCING SCOPE (report, no code): MLB opponent strength IS derivable from statsapi, verified live — one free call returns all 30 teams' pitching splits (era/whip/avg/slg/ops/homeRuns/strikeOuts/HR9), which beats the ESPN field we were reaching for because it is STAT-SPECIFIC, exactly what opp_rank_stat wants. NBA/WNBA cannot use ESPN (its team endpoint carries only a team's own stats, no defensive rating or pace); options are stats.nba.com dashboards, deriving allowed-points from scoreboard finals we already fetch, or API-Sports. API-Sports is a fallback tier at best — 100/day will not survive per-team-per-day. ESPN stays last, always behind an adapter. Proposed the SOURCE-ADAPTER pattern: one interface per feed, config-driven primary+fallback per (sport x capability), normalized output so vendor quirks stay in adapters, fallback announced rather than silent, sources with zero callers deleted rather than left as corpses, and a health check that PAGES when a source returns empty or broken — where EMPTY IS A FAILURE. Tonight's crash (captured 0 / errored 15) and the months-null opp_rank_stat are both exactly what that check exists to catch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
29 KiB
MODEL TRAIN — VALUE ENGINE
Status: Arc 1 (steps 1–6) SHIPPED on main 7a925f4 (2026-07-19). Arcs 2+ OPEN.
Written retroactively (2026-07-19) per CLAUDE.md rule #1 — Arc 1 was built from a
plan that lived only in a session context that was lost. Everything in the "SHIPPED"
sections below was read off the code ON DISK, not from that plan or from memory.
Doctrine (Kev): VYNDR promotes bets people actually take — roughly the −160 to +200 band — that ALSO carry a genuine vig-free edge. Not plus-money-only, not heavy chalk. Grade answers "is this a good read"; value answers "does the price pay you." They are separate fields and must stay separate.
1. WHAT ARC 1 SHIPPED
1.1 De-vig — src/utils/devig.js (new)
Two-way multiplicative (proportional / normalized-implied-probability) de-vig.
Each side's implied probability is divided by the sum of both; the sum's excess over 1
is the overround. Method is recorded on the payload as devig_method:'multiplicative'.
| Export | Behavior |
|---|---|
americanToImpliedProb(a) |
+: 100/(a+100); −: −a/(−a+100). Null on non-finite or 0. |
americanToDecimal(a) |
Total-return multiple incl. stake. Null on non-finite or 0. |
impliedProbToAmerican(p) |
Null outside (0,1). p>0.5 → negative (favorite). |
devigTwoWay(overOdds, underOdds) |
{method, overround, over:{fair_prob,fair_odds}, under:{…}}. Returns null if EITHER side is missing/invalid — fair values are never synthesized from one side. |
evPct(modelProb, american) |
modelProb × decimal − 1, as a percentage rounded to one decimal. Null on bad input. |
Rounding: probabilities/overround to 3 dp (round3); ev_pct to 1 dp.
1.2 Config knobs — the ACTUAL values on disk
src/config/valueEngine.js (new — the promotion gates):
| Knob | Env var | Default in code | Meaning |
|---|---|---|---|
TAKEABLE_ODDS_CEILING |
TAKEABLE_ODDS_CEILING |
−160 | Most-juiced favorite we will PROMOTE. |
TAKEABLE_ODDS_MAX |
TAKEABLE_ODDS_MAX |
+200 | Longest dog we will PROMOTE. |
VALUE_EV_THRESHOLD |
VALUE_EV_THRESHOLD |
2 (i.e. 2 % EV) | Minimum EV for the value flag. |
isTakeable(american)→a >= −160 && a <= +200. Strict:null,'',undefined, and non-finite arefalse(theNumber(null)===0fabrication trap).isValue(american, evPct)→isTakeable(american) && Number.isFinite(evPct) && evPct >= 2.
src/config/rareEventMarkets.js (pre-existing, from the same-day rare-event work —
the refusal layer that sits UNDERNEATH the promotion gates):
| Knob | Env var | Default in code | Meaning |
|---|---|---|---|
JUICE_ODDS_FLOOR |
JUICE_ODDS_FLOOR |
−400 | Graded side at or past this price → refused, not graded. |
RARE_EVENT_LINE_MAX |
— (constant) | 0.5 | Line at/below which the rare-event rules apply. |
RARE_EVENT_STATS |
— (constant) | doubles, triples, home_runs, stolen_bases, steals, blocks |
Backup list for props with NO odds. |
Knobs that DO NOT EXIST (verified by grep across src/, web/src/, tests/):
TAKEABLE_CEILING (real name is TAKEABLE_ODDS_CEILING), EDGE_FLEX_WALL,
HARD_JUICE_WALL, LADDER_ODDS_MAX, MIN_RUNG_PROBABILITY. If those were in the
original plan, they are unbuilt — see §3.
1.3 What the gate ACTUALLY does today
Reported as-built, not as-designed. There are two independent layers, and only the first one can refuse a grade:
Layer A — the refusal wall (JUICE_ODDS_FLOOR, −400). In analyzeViaEngine1, the
FIRST thing that happens — before computeFeaturesForProp, before any projection:
if (isTooJuiced(rawProp)) → suppressedRareResult('juiced_no_edge')
Characteristics of the shipped behavior:
- A flat price wall, NOT edge-aware. It looks only at the graded side's American price. No EV, no projection, no edge, no flex band participates in the decision.
- The wall is −400, not −250. There is no −250 constant anywhere in the codebase.
- It runs pre-feature, so a refused prop costs zero feature computation.
- Odds absent →
isTooJuicedisfalse(can't judge from a price that isn't there); the structuralRARE_EVENT_STATSrule is the backup for that case. - Output is a refusal, not a bad grade:
grade:null,insufficient_data:true,suppressed:true,suppressed_reason:'juiced_no_edge', plus branded "No read —…" copy inreasoning.summary.
The other two refusals in the same family: rare_event_under (rare stat, under, line
≤ 0.5 — pre-feature, needs no projection) and rare_event_over_below_line (rare stat,
over, line ≤ 0.5, and the model does NOT project above the line — checked after the
projection exists).
Layer B — the takeable band (−160 … +200). This gate never refuses a grade. It is a promotion filter, applied only where a read is featured:
- Today its ONLY enforcement point is
heroPropService.pickHeroProp(if (!isTakeable(g.book_odds)) continue). - Every graded read still carries
takeable/valueas flags for consumers. - The full board shows every graded read regardless of band. Parlay Lab is exempt by design (juiced legs combine into takeable payouts).
So: edge-aware promotion, price-only refusal. The -400 wall decides what gets
graded; EV + the −160/+200 band decide what gets promoted.
1.4 The value triplet
Attached in analyzeViaEngine1's existing p_win/Kelly try block (so it inherits the
"real quantile probability × real book odds, or nothing" rule, and the whole block is
additive — a throw never breaks the read):
| Field | Source | Absent when |
|---|---|---|
book_odds |
the graded side's actual price (under_odds if under, else over_odds) |
no side odds |
fair_odds, fair_prob |
devigTwoWay(over, under) for the graded side |
either side unpriced |
model_odds |
impliedProbToAmerican(pWin) |
no quantile probability |
overround, devig_method |
from the de-vig | either side unpriced |
ev_pct |
evPct(pWin, sideOdds) |
no pWin or no side odds |
takeable |
isTakeable(sideOdds) |
no side odds |
value |
isValue(sideOdds, ev) |
no ev |
The intended render (Design's, not built): "book −145 · vig-free −132 · model −110."
1.5 Hero v2 — heroPropService
Old rule: largest |projection − line| among A/B candidates. New rule: among A/B
candidates, keep only those with a finite ev_pct and a takeable book_odds, then
take the highest ev_pct. The gap is still computed and returned for display, but
it no longer selects. Empty-slate fallback (most recent real graded read, any grade, by
timestamp) is unchanged. toHero now also passes through
ev_pct / value / takeable / book_odds / fair_odds / model_odds.
1.6 Distribution + tests
- Fields ride the existing
...result/...dataspreads intogrades:{sport},/api/snapshot/:sport,/api/hero-prop,/api/scan. Documented inBACKEND_HANDOFF.md("Value Engine fields"). src/utils/tierGating.jsdeletesalt_lines(below Desk) andkelly(below the Kelly tier). It does not touch the value fields —ev_pct,value,takeable, and the triplet are currently visible to every tier including free. That is an unmade product decision, not a verified choice (see §3).- Suite: 276 suites / 3306 tests, green,
next buildexit 0. New:tests/unit/devig.test.js,tests/unit/valueEngine.test.js;tests/unit/heroPropService.test.jsrewritten for the EV rule.
2. LAWS THIS TRAIN INHERITS (do not relax)
- Absent beats wrong. One side unpriced →
fair_oddsis absent, never extrapolated. No price →takeableis absent, notfalse-by-coercion. Number(null) === 0is the standing fabrication bug of this codebase. Every new numeric path gets a strict guard (valueEngine.isTakeableis the model).- Grade ≠ value. An A read with
value:falseis honest output: right read, price gone. Never let the value flag re-letter a grade. - Refusal is a product surface, not an error. Refused reads carry branded copy.
- Market values are REAL book numbers at a timestamp; only projection / grade / edge / EV / fair price are model output and must be labeled MODEL.
2C. METRICS-ENGINE NORTH STAR (design philosophy — NOT yet built)
VYNDR classifies and measures players by MODERN FUNCTION, not legacy label. This is the principle already underneath the archetype system, drawn from Rashad Phillips' Basketball Position Metric — the work VYNDR's classification started from. Writing it down so every future metric inherits it.
The principle
A "center" who shoots 7 threes a game is not doing a center's job. A contact hitter and a launch-angle hitter share the label "outfielder" and share almost nothing else. Positions are inherited labels; function is what a player actually does this season. VYNDR models function.
The rule every metric must follow
Every proprietary metric is baselined against the player's functional ARCHETYPE's CURRENT-SEASON behavior — never against the position's inherited or historical standard.
A STRETCH BIG's rebounding is judged against how stretch bigs rebound this season, not against what centers rebounded in 2015.
Why it is an edge, not just a nicety
The market frequently prices today's players against yesterday's baselines. Books and public models lean on positional priors that lag the way the game is actually played. Where the archetype baseline and the positional baseline disagree, that gap is a real, repeatable mispricing — and it is ours to measure because we already classify by function.
It generalizes across sports
Launch-angle vs contact hitters · mobile vs pocket QBs · position-less wings · bullpen-game openers vs traditional starters. The archetype registry is already cross-sport; the baseline rule is the same everywhere.
The moat
Proprietary metrics, calibrated to the CURRENT game, validated on our private outcome data (the ledger). None of those three is individually rare. Together, and compounding as the ledger grows, they are hard to copy.
How metrics get built — VALIDATED FAMILIES, never a big-bang dump
- Propose ONE metric family with an explicit hypothesis about what it predicts.
- Implement behind a flag, computed but not surfaced.
- Run the backtest harness: does it predict better WITH the metric than without? Hit rate, ROI, Brier, CLV where capture exists.
- Ships only if it earns its place. Fails → deleted, and the negative result is written down so it is not re-proposed.
- Then, and only then, it becomes a surface.
Nothing in this vision is real until the harness can prove a metric predicts better than without it. The harness (Phase 2) is the gate for all of it.
2D. SOURCING SCOPE — REPORT (Session 64, no code written)
Triggered by tonight's failure: we wired refreshTeamStats into production and
it crashed on every team for a shape change nobody noticed, because it had no
caller, no health check, and no alarm. The feed was dead and invisible.
Q: can opponent strength be DERIVED from data we already ingest?
MLB — YES, and better than what we were reaching for. ✅ VERIFIED LIVE.
GET https://statsapi.mlb.com/api/v1/teams/stats?season=YYYY&group=pitching&stats=season&sportIds=1
returns all 30 teams in ONE call, free, official, no quota, no auth. Verified
keys include era, whip, avg (opponent batting average against), obp,
slg, ops, homeRuns, strikeOuts, runsScoredPer9, homeRunsPer9.
That is strictly better than the ESPN field we tried to fetch, because it is
stat-specific — which is what opp_rank_stat actually wants:
| Prop stat | Opponent-strength input |
|---|---|
| hits | opponent pitching avg (BAA) |
| total_bases / home_runs | slg, homeRunsPer9 |
| strikeouts (batter) | team strikeOuts per BF |
| earned_runs / pitcher props | opposing lineup's ops (group=hitting) |
Ranking is a normalize-across-30-teams pass we already have in teamStatsCache.
NBA/WNBA — ESPN cannot do it. Verified: /teams/{id}/statistics carries only
a team's OWN stats (rebounds, FG%, blocks) — no defensive rating, no pace, no
opponent-allowed anything. Options, in order: (a) stats.nba.com team
dashboards, which DO carry real DEF_RATING/PACE (needs careful headers, is
rate-limited, and is the source the offline Python service used); (b) derive
allowed-points ourselves from scoreboard finals we already fetch — free, slower
to build, fully under our control; (c) API-Sports basketball. WNBA is the
in-season priority.
API-Sports family — we hold API_FOOTBALL_KEY (validated, free tier 100/day,
active to 2027, currently dormant). Same vendor covers basketball/baseball on
separate keys. Worth evaluating as the fallback tier, not primary: the free
quota is thin (100/day) and would not survive a per-team-per-day pattern.
ESPN stays LAST and always behind an adapter. It is free and broad, but it is an undocumented site API that changes shape without notice — which is precisely what bit us tonight.
PROPOSAL: the SOURCE-ADAPTER pattern (design only — not built)
Every feed behind one interface, so a source can be swapped without touching feature code, and so a dead source is loud instead of invisible.
interface StatSource {
id: 'mlb-statsapi' | 'espn' | 'nba-stats' | 'api-sports'
supports(sport, capability): boolean // 'team_defense' | 'game_logs' | ...
fetch(sport, capability, params): Promise<Normalized|null>
health(): Promise<{ ok, checked_at, sample_nonempty, note }>
}
- Registry, config-driven: per (sport × capability), an ordered
[primary, ...fallbacks]. MLB team_defense →[mlb-statsapi, espn]. WNBA team_defense →[nba-stats, derived-from-scoreboard, espn]. - Normalized output only. Adapters own every vendor quirk; nothing upstream
learns a vendor's shape. (Tonight's
results.statsobject-vs-array would have been one adapter's problem, not a pipeline crash.) - Health check that PAGES. Each adapter self-tests on a known entity and
asserts a NON-EMPTY, plausible result. Run on a schedule + before the snapshot.
Empty is a FAILURE, not a pass —
captured: 0, errored: 15must page, and so mustcaptured: 30where every value is null. - Fallback is announced, never silent. Dropping to a fallback logs and pages once per period; a silent degrade is how we ended up with a feature that had been null for months.
- A source with zero callers is deleted or health-checked — never left as a
corpse to be wired up later, which is exactly what
refreshTeamStatswas. - Capability coverage is a first-class report: which (sport × capability) pairs have a live primary today. This is also the NFL/NBA/soccer scaling gate.
2B. FOUNDATION-FIRST RE-ORDER (Kev, 2026-07-19) — supersedes the arc order
New features are paused until the foundation is real. Phases run in order: 1 (tooling+safety) → 2 (the instrument) → 3 (corruption fixes) → 4 (depth).
| # | Item | Phase | Status |
|---|---|---|---|
| 1 | Manual regrade trigger | 1 | ✅ scripts/run-snapshot.js shipped — but see ACCESS BLOCKER |
| 2 | Backup cron INSTALLED | 1 | ✅ shipped as CODE (src/backupScheduler.js) — deploy == installed |
| 3 | Backtest harness | 2 | open — REPORT-FIRST on replayable history |
| 4 | Settlement-correctness audit | 2 | open — REPORT |
| 5 | Sample-size discipline | 2 | open — report the floor |
| 6 | edge_pct → ev_pct migration | 3 | open — REPORT-FIRST (ledger bleed) |
| 7 | Grade-lock + directional CLV + C4 | 3 | open — REPORT-FIRST (does a snapshot overwrite a prior grade?) |
| 8 | Model uncertainty | 4 | open — REPORT-FIRST (design proposal) |
| 9 | Calibration by odds band | 4 | open — runs ON the harness |
🔴 ACCESS BLOCKER (Session 64) — I cannot reach the box
Verified this session: no VYNDR_INTERNAL_KEY in the local .env (it holds
only Supabase + ODDS_API_KEY), and SSH to git.builtbykev.com /
api.vyndr.app / vyndr.app times out from WSL2. So I can neither call the
internal endpoints that already exist (S45 shipped
POST /api/internal/snapshot/:sport|/all) nor docker exec anything.
Consequences, stated plainly:
- The manual trigger is BUILT and correct, but only Kev can run it until one
of these exists:
VYNDR_INTERNAL_KEYshared with the agent env, or box SSH. - Verifying a grading fix still depends on the cron or on Kev running one command.
- This is the single highest-leverage unblock for every future phase — Phase 2 and 3 both need on-demand regrade + settle runs to verify anything.
Standing cautions (Kev, logged 2026-07-19)
- CLV ledger stays PRIVATE until the model is backtest-proven. Publishing CLV before then broadcasts our weaknesses to sharps. (Also currently broken — C4.)
- "Self-improving model" is UNSUPPORTED marketing until the loop actually
closes: backtest → calibration → weight correction. No such loop exists today
(no harness, no calibration gate,
weightAdjusteris not in the grade path). Do not claim it. - The engine is MLB/WNBA-calibrated. NFL/NBA/soccer are NOT. "Unified engine" is currently "MLB engine, others guessing." Each sport needs its own calibration before the offseason hub grades it — this is a scaling gate, not a nice-to-have.
2A. ARC 2+ BOARD — STATUS (Kev's arc list, 2026-07-19)
Full arc definitions live in the Session-63 order. Status only here; update as each ships.
| Item | Status | Note |
|---|---|---|
| G-a the real gate | HELD for ruling | Built nothing yet — the G-b report changes the recommended dials. See specs/audit-data/gate-simulation.md §2.4. |
| G-b gate simulation | ✅ REPORTED | specs/audit-data/gate-simulation.md. Replayed over the ledger, NOT snapshots (no 30d snapshot store exists). |
| G-c never-empty honesty | open | Rare but real on thin MLB nights. |
| S-a line-moved truth | open | Shares plumbing with C-clv — build together. |
| S-b board ranks on EV | open | Blocked-ish: EV not persisted; live payload has it. |
| L-a per-rung odds? | ✅ REPORTED — NO | Rungs are synthetic, carry no price. Feed offers no alternate markets. §3 of the report. |
| L-b per-rung EV ladder | BLOCKED | Needs a real price per rung. Probe PropLine first. |
| C-cal calibration | ✅ REPORTED | Confidence monotonic but ~20-25pts miscalibrated; only B/C grades ever emitted. |
| C-led ledger speaks value | open — recommended next | locked_odds already 99.1% populated: units/ROI need no backfill. EV columns are net-new, forward-only. |
| C-clv fix C4 | open | Confirmed broken in data (359/376 MLB closes == lock). Keep suppressed. |
| D-ref visible refusals | partial | Copy already ships on refusals; new gate_* reasons land with G-a. |
| D-par parlay lab | open | Gate must NOT apply inside the Lab. |
| D-tier tier gating | DECIDED, not enforced | Free/Analyst: value marker + grade + triplet. Desk: ladder + per-rung EV + Kelly. Today ladder/Kelly are already Desk-gated; triplet ungated = correct per this ruling. |
| D-ev consolidate EV | open | devig.evPct vs processing/EVCalculator.js (used only by UnifiedOddsProvider). |
| U-deg MLB degradation | ✅ STATUS REPORTED | projection==0 leak already closed (0 occurrences since 07-18). edge_pct scale still broken. |
| U-fp Arc 1 fingerprint | open | Do it on the first deploy this train ships. |
✅ SESSION 63 — PROBABILITY LAYER + GRADE RANGE RESTORED (shipped)
The re-sequenced step 1+2, folded into one change. Full write-up:
specs/audit-data/grade-collapse.md.
- The probability layer was DEAD in production —
p_win/ev_pct/kelly/model_odds/valuewere absent on 0/8 live grades becausegameLogServicereturns null for MLB by construction and the Python service is offline for NBA/WNBA.featureCache.getStatRowsnow supplies normalized rows for every sport. Verified:p_win25/25 on real WNBA props, 8/8 MLB (was 0). - Hero v2 had never once selected on EV (it requires a finite
ev_pct) and silently fell through to the recent-read fallback every time. - Grade range:
refreshTeamStatswired intorunSnapshot(it had ZERO callers, soopp_rank_statwas permanently null),game_count_in_7dderived from real logs, and L20 made symmetric (there was no negative branch at all). D now emits on merit (WNBA 1/25, an earnedp_win0.365); A is proven reachable arithmetically but has not yet emitted in production — that is the outstanding fingerprint. - Calibration guard: consistency CV was NBA-tuned; for any stat with mean < 4,
cv ≈ 1/√meanforcesboom_bust. It would have stamped a blanket −1.0 on nearly every MLB prop. Floored atCONSISTENCY_MIN_MEAN=4→unknownbelow it. - Confidence is NOT a probability — payloads now carry
confidence_basis: 'grade_band'. The real signal isp_win. mlbGrader.jsREMOVED (dead; referenced only by its own test).- 🔴 MARKETING HOLD: "A-RATED" copy (AccuracyBadge, TopSignals) is unsupported until a production fingerprint shows real A grades. Honest fallbacks confirmed rendering ("MODEL · 63% HIT"); nothing fabricated ships.
🔶 OPEN DECISION — FLEX BAND ENFORCEMENT (Kev, 2026-07-19)
Ruling: build EDGE_FLEX_WALL (−250) + EV_FLEX_THRESHOLD (default 4 %, = 2×
VALUE_EV_THRESHOLD) but ship with EV_FLEX_ENFORCE=0. Flex-band props
(−161…−250) grade exactly as they do today — the band is NOT cut. The ledger shows
it is our most profitable segment (+2.2 % ROI, n=70) and we will not restrict a
proven-profitable band on an unvalidated threshold. Record ev_pct on every flex prop
now so real in-production data accumulates.
TRIGGER TO REVISIT: once ~2 weeks of production ev_pct data exists on the
−161…−250 band, report the ev_pct distribution vs settled outcomes for that band,
then Kev decides the threshold — or whether to enforce at all.
Do NOT flip EV_FLEX_ENFORCE=1 without that report and explicit sign-off. It comes
back as a data-backed decision, never a silent flip.
Unaffected — these ship and enforce normally: HARD_JUICE_WALL (−250, never grades),
LADDER_ODDS_MAX (+400), MIN_RUNG_PROBABILITY (0.25), the no-odds refusal, and the
folded projection > 0 check.
🔴 OPEN — U-deg PART 2: edge_pct IS ON A BROKEN SCALE (next U-deg item)
projection == 0 is closed. edge_pct is not, and it is the one still being
read by users. Live proof from the S63 fingerprint response: edge_pct: 100 on a
prop whose real edge is single-digit. Ledger-wide: 311/604 rows (51.5 %) exceed
EDGE_BOARD_SANE_MAX (40), 39 rows exceed 100 (impossible as a percentage), worst
620.
Who reads what — this is the problem in one line: every user-facing surface reads the broken number, and nothing reads the good one.
| Consumers | |
|---|---|
edge_pct (BROKEN) — frontend |
Slate.tsx, PropRow.tsx, GradeCard.tsx, GradeResultCard.tsx, MobileEdgeBoard.tsx, TierRecord.tsx, DemoScan.tsx, SoccerGradeResult.tsx, slateAdapter.js, gradeAdapter.js, scan/page.tsx, soccer/page.tsx, api/scan/route.ts |
edge_pct (BROKEN) — backend |
deskShowcaseService, contentTemplateService, parlayScanService, tierGating (free-tier "hook"), ledgerService:199 → persists it to the edge column |
ev_pct (GOOD) — frontend |
NONE |
ev_pct (GOOD) — backend |
heroPropService only |
Consequences to fix together, not piecemeal:
- The frontend's
EDGE_BOARD_SANE_MAX = 40guard is damage control that nulls half the board — it hides the bug rather than fixing it. ledger_entries.edgeis being written with garbage right now, permanently, into an append-only table. Any future edge-based analysis inherits it.- S-b (rank board on EV) is the real remedy — it moves ranking off
edge_pctontoev_pct, which is now live and correct. Do S-b and the edge-scale fix as one piece of work. - Free tier is sold
edge_pctas "the hook" (tierGating:62). That hook is currently a wrong number.
🔶 OPEN — INDEX-OF-DISPERSION CONSISTENCY CLASSIFIER (modelling change)
CONSISTENCY_MIN_MEAN = 4 is an honest stopgap, not the answer. It stops the
NBA-tuned CV thresholds from stamping boom_bust on every low-count MLB stat
(cv ≈ 1/√mean, so mean < 4 always trips boom_bust) — but the cost is that MLB
low-count stats get no consistency factor at all, leaving a ±1.0 dead for the
sport that carries most of our volume.
The real fix: classify on the index of dispersion (variance / mean) against the Poisson baseline of 1.0 — under-dispersed (< 1) = genuinely consistent, over-dispersed (> 1) = genuinely boom/bust. It is scale-free, so one set of thresholds works for MLB hits (mean 0.6) and NBA points (mean 20) alike.
VALIDATION REQUIRED BEFORE IT SHIPS — this changes grades, so per the standing rule no weight change ships without a backtest:
- Build the backtest harness (still does not exist — see below).
- Replay settled outcomes with dispersion-based consistency vs the current floor.
- Show hit-rate/ROI by grade tier does not degrade, and that the restored ±1.0 moves grades on merit rather than flooding one letter.
- Report before flipping. Env-gate the switchover.
Backtest harness
None exists (grep-verified). migrations/006 defines grade_outcomes +
player_calibrated_weights and no code reads or writes them. The G-b/C-cal replay was
done in SQL against ledger_entries; a real harness is still owed before any weight change.
3. NOT YET BUILT — checklist to reconcile against the full arc list
Verified absent from the codebase as of 7a925f4. Kev supplies the complete arc list;
mark each done/open against this.
Engine / gates
- Edge-aware juice wall — a soft/flex band where a big enough edge can survive a
juiced price (
EDGE_FLEX_WALL,HARD_JUICE_WALL— neither exists; today's wall is the single flatJUICE_ODDS_FLOOR−400). TAKEABLE_ODDS_*enforced on any surface beyond the daily hero (featured rows, top-of-board, alerts). Today: hero only.- Alt-line ladder value gates (
LADDER_ODDS_MAX,MIN_RUNG_PROBABILITY— neither exists). The ladder inanalyzeViaEngine1is still the S62 build: fixed shifts[−1, −0.5, 0, +0.5, +1], filtered toline > 0, sorted byedge_pct. It has no odds ceiling and no per-rung probability floor, and its rungs carry noev_pct/fair_odds/ price at all. - EV-rank the board itself.
slateAdapter.flattenToEdgeBoardstill ranks onedge(guarded byEDGE_BOARD_SANE_MAX = 40), notev_pct. - Reconcile the orphan
src/services/processing/EVCalculator.js(a separate, older EV path used only byUnifiedOddsProvider) againstdevig.evPct— two EV implementations currently coexist. - Multi-way / three-way de-vig, and alternatives to multiplicative (Shin, power/log). Two-way multiplicative is the only method implemented.
- Best-price de-vig across books (today the triplet de-vigs whichever single book's two-sided price came through on the prop).
Surfaces (nothing on the frontend reads these fields yet — grep-verified zero hits
for ev_pct / fair_odds / model_odds / book_odds / suppressed_reason under
web/src/)
- Value triplet rendered on the reveal card / board row / hero.
- VALUE marker (and the honest "A but no value" state).
- Visible-refusal moment — the "No read —…" copy is generated and shipped on the payload but is not displayed anywhere.
- EV shown as the headline number instead of / alongside edge %.
Product / policy
- Tier policy for the value fields — currently ungated to free (§1.6).
- Alerts on takeable +EV reads (the config comment anticipates "future alerts").
- Ledger/record split by
value(does the model's +EV subset beat its overall record?).ledger_entriesstores no EV or fair price today. - Settlement/CLV interaction with fair price — CLV is still locked-vs-closing
LINE, not vs closing FAIR price. (Note: C4 — CLV capture is broken and
suppressed on public surfaces; see
specs/audit-data/clv-capture-broken.md. Any fair-price CLV work is blocked behind C4.)
Verification debt
- Arc 1 is not deploy-fingerprinted. Post-deploy, confirm
/api/hero-propreturns a hero with a takeablebook_oddsand a finiteev_pct, and that/api/snapshot/mlbgrades carry the triplet. - Known upstream contamination:
specs/audit-data/mlb-grade-degradation.md— projection == 0 on ~9/25 MLB grades and a brokenedge_pctscale. EV is computed from the quantile probability, not fromedge_pct, so it is not directly poisoned — but hero ranking and any EV-vs-edge comparison sit on top of that pipeline. Fix the degradation before trusting arc-1 numbers in aggregate.
Arc 1 recorded 2026-07-19 from the code at 7a925f4. Update this file in the same
commit as any arc that ships.