Files
vyndr/specs/edge-shading-challenger.md
T
builtbykev c2c7abbb65 Edge-shading challenger: built + measured. Flooding NOT fixed — input scale is the bug
Challenger only. Champion grade byte-identical (verified by diff). Nothing
promoted, no live grade re-lettered, no ledger row deleted or re-settled.

BUILT src/services/challengers/efficiencyShading.js (measured-never-served):
  adjusted_edge = raw_edge * f(efficiency); grade = band(adjusted_edge) against
  ONE fixed bar (A+>=10, A>=5, B>=3, C>=1, D>=0, F<0) that never moves.
  f(e) = E_SOFTEST/e bounded to (0,1] — soft markets intact (never amplified),
  sharp shaded toward but not past zero, unscored -> f=1 and FLAGGED.
  A fence test asserts no production grade path imports it.
  Cross-market behaviour is unit-proven: the same raw 6% edge grades A in soft
  mlb:total_bases and B in sharp nba:points.

MEASURED on 1250 live ledger rows — Phase 2.5's answer is NO, the flooding is
not gone: challenger 79.0% A and 80.9% A/B (MLB 93.4% A) vs champion 0.2% A.

TWO findings explain why, and they are the point of the order:

1. The shading is a NO-OP on the live board: rows_actually_shaded = 0 of 1250.
   96.5% of rows are UNSCORED (f=1), and the one scored market present
   (mlb:total_bases) is the anchor so its f is 1.0 by construction.
   mlb:strikeouts and nba:points do not appear in the ledger at all (our
   basketball is wnba, not nba). Challenger vs baseline: 0 rows changed.

2. Placement was never the bug — the INPUT SCALE is. Against a fixed 5% bar the
   RAW edge already clears A on 100% of MLB doubles, 89.6% of hits, before any
   shading. MLB median raw edge is 60%, twelve times the bar. Decisive test:
   apply the sharpest score in the spec (f=0.647) to EVERY row — the maximum
   the design permits — and 75.8% still clear A (MLB 91.7%). Since f is bounded
   <= 1, no achievable shading can close a 12x overshoot. Moving the multiply
   from the threshold to the edge does not change the outcome.

This is edge_pct behaving as the 2026-07-29 diagnosis described: a price-free
(proj-line)/line gap whose scale is a function of line size. It is not a
betting edge, so no fixed betting-edge bar is meaningful against it.

2.6 efficient-market over-suppression: CANNOT DETERMINE — zero live rows are
shaded, so there is no efficient market in the data to over-suppress.

Phase 3: takeable tagging was completed in the previous order (migration 034,
1246/1254 rows) and is not repeated. The model-version boundary is again NOT
applied: nothing promoted, so no boundary exists.

Unblocking needs the input replaced, not the multiply moved: p_win vs
fair_prob (both already computed) instead of edge_pct, plus scores FIT from our
own record for the markets we actually grade.

Floor: 313 suites / 3899 tests green (9 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:39:13 -04:00

6.4 KiB
Raw Blame History

SPEC — EDGE-SIDE EFFICIENCY SHADING CHALLENGER (built + measured; NOT promoted)

2026-07-31. Champion grade UNCHANGED (byte-identical, verified by diff). Push scoring untouched.


REVIEW ZERO

  • 0.1 raw edge — EXISTS. analyzeViaEngine1.edgePctFor = ((proj line) / line) × 100, signed by direction; persisted as ledger_entries.edge.
  • 0.1 threshold application — DOES NOT EXIST. The champion grade is engine1's additive factor index; there is no edge-vs-threshold step on the live path.
  • 0.1 the flat additive ±0.05 efficiency nudge — DOES NOT EXIST (third re-check). No efficiency reference anywhere in src/services/intelligence/ or gradeAdapter. There is nothing to remove from the challenger path.
  • 0.2 fixed base thresholds — DO NOT EXIST live. So the challenger DEFINES one and it never moves: A+ ≥ 10, A ≥ 5, B ≥ 3, C ≥ 1, D ≥ 0, F < 0, identical for every market. (The spec's per-sport MLB-5%/NBA-7% variant is deliberately NOT used — a bar that differs by sport is a moving bar, which this design forbids.)
  • 0.3 efficiency scores — 5 exist in spec, and they cover almost nothing we grade. nfl:passing_yards .85 · nba:points .80 · mlb:strikeouts .60 · mlb:total_bases .55 · nba:role_player .55. Of the 11 live ledger markets, only mlb:total_bases matches. Unscored markets default to f = 1 and are FLAGGED.

THE BUILD (src/services/challengers/efficiencyShading.js, measured-never-served)

adjusted_edge = raw_edge × f(efficiency)
grade         = band(adjusted_edge)   ← ONE fixed bar, never moved

f(e) = E_SOFTEST / e, bounded to (0, 1]. At the softest scored market (0.55) f = 1.0 — the edge is left intact, never amplified (the guardrail). Sharper markets shade down monotonically: f(0.60)=0.917, f(0.80)=0.688, f(0.85)=0.647. Unscored → f = 1, flagged. Absent edge → null (absent, never graded F). A fence test asserts no production grade path imports this module.

Cross-market behaviour works as designed (unit-proven): the same raw 6% edge grades A in soft mlb:total_bases and B in sharp nba:points.


PHASE 2 — THE MEASUREMENT, AND IT IS A NO

2.5 "Confirm the naive-formula pathology is GONE" → IT IS NOT GONE.

population n challenger %A challenger %A/B champion %A champion %A/B
MLB 696 93.4 93.7 0.0 57.9
WNBA 554 61.0 64.8 0.5 58.3
TOTAL 1,250 79.0 80.9 0.2 58.1

80.9% A/B — the flooding the fix was meant to remove is still there (the naive formula's reported 86% A/B is not meaningfully improved on).

TWO FINDINGS THAT EXPLAIN WHY, AND THEY ARE THE POINT OF THIS ORDER

(1) The shading is a NO-OP on the live board. rows_actually_shaded = 0 of 1,250. 96.5% of rows (1,206) are UNSCOREDf = 1. And the one scored market present (mlb:total_bases) is the anchor, so its f is 1.0 by construction. mlb:strikeouts and nba:points do not appear in the ledger at all (our basketball is wnba, not nba). Challenger vs baseline: 0 rows changed. The design cannot act on a board it has no scores for.

(2) Placement was never the bug — the INPUT SCALE is. Against a fixed 5% bar, the raw edge already clears A on 100% of MLB doubles, 100% stolen_bases, 89.6% hits, 88.6% total_bases — before any shading at all. MLB's median raw edge is 60%, twelve times the bar.

The decisive test: apply the SHARPEST score in the entire spec (f = 0.647) to EVERY row — the maximum shading the design permits — and 75.8% still clear the A bar (MLB 91.7%). Since f is bounded ≤ 1, no achievable shading can fix this. Moving the multiply from the threshold to the edge does not change the outcome, because a bounded multiplier cannot close a 12× overshoot.

This is edge_pct doing exactly what the 2026-07-29 diagnosis said it does: it is a price-free (proj line)/line gap whose scale is a function of line size, so on 0.5-lines it explodes (median 60, max 620). It is not a betting edge, so no fixed betting-edge bar is meaningful against it.

2.6 efficient-market over-suppression → CANNOT DETERMINE. Zero live rows are shaded, so

there is no efficient market in the data to over-suppress. The failure mode is untestable here.

2.7 cross-market behaviour → works in unit test, ZERO live effect (see finding 1).


PHASE 3 — TAGGING

  • Takeable tagging: ALREADY DONE in the previous order (migration 034; 1,246 of 1,254 rows tagged, 781 takeable / 465 below floor, null_despite_price = 0, floor 160 recorded per row and labelled policy-not-derived). Not repeated — re-running would be a no-op.
  • Model-version boundary: DELIBERATELY NOT APPLIED, again. It marks pre/post a scaling change. Nothing promoted, so no boundary exists — stamping one would record a model transition that never happened. config/modelEras.js is its home the day one promotes.

WHAT WOULD ACTUALLY UNBLOCK THIS

  1. Replace the input, not the multiply. A bar-based grade needs a quantity that is a real edge — model probability vs the de-vigged market price. That is p_win vs fair_prob, both already computed and stored. edge_pct cannot carry a fixed bar at any shading.
  2. Fit the scores from our own record, per market we actually grade — 9 of 11 live markets have no score, and inventing 9 numbers would be fabricated calibration.
  3. Then re-run this challenger; the mechanic itself (bounded, edge-side, fixed bar) is sound and already built.

VERIFY

Champion byte-identical (diff empty across intelligence/, gradeSlateService, gradeAdapter, snapshotService, grade_thresholds.json). Fixed bar never moved — efficiency multiplies the edge only, and a test asserts shading can only lower a grade. 313 suites / 3,899 tests green (9 new), web build exit 0. Nothing promoted, no live grade re-lettered, no ledger row deleted or re-settled. Auth-gated visuals → Chrome audit; nothing visual changed.

TAGS

VERIFIED: the no-op (0 rows shaded, 96.5% unscored); the flooding (80.9% A/B) and that max shading still leaves 75.8% clearing A; champion untouched. CANNOT DETERMINE: efficient-market over-suppression (no shaded rows exist). BLOCKED: promotion — the input scale must change first.