diff --git a/specs/audit-data/gate-simulation.md b/specs/audit-data/gate-simulation.md new file mode 100644 index 0000000..f2e3170 --- /dev/null +++ b/specs/audit-data/gate-simulation.md @@ -0,0 +1,283 @@ +# GATE SIMULATION + CALIBRATION — Model Train G-b / C-cal + +**REPORT-FIRST deliverable. No engine code was written for this. G-a is HELD pending Kev's ruling.** +Dataset snapshot: `2026-07-19 22:01:58 UTC`, live Supabase `ledger_entries`, `user_id IS NULL` +(the public model record). **576 rows · 6 game days (2026-07-11 → 07-19) · 470 settled · +571 with locked odds.** Numbers move: the pipeline inserted 8 rows mid-census, so every +table here is pinned to that timestamp. + +--- + +## 0. THE BLOCKER: "last 30 days of stored snapshots" DOES NOT EXIST + +There is no 30-day snapshot store to replay. Verified on disk: + +- `snapshot:{sport}:latest` and `:previous` — **two generations, 24h TTL** + (`snapshotService.js:404,412`). No `snapshot:{sport}:{date}` archive exists anywhere. +- The intraday line `history:[{t,line}]` rides **inside** the same 24h blob + (`intradayRefreshService.js:68-84,144`) — it dies with it. +- `outcomes:{sport}:log` (30d TTL, cap 1000) carries **no odds, no confidence, no + projection** (`outcomeService.js:233-237`) — unusable for an odds-aware table. +- **No backtest/replay harness exists** — grep for `backtest|replay` across `src/`, + `scripts/`, `tests/` returns zero. `migrations/006` defines `grade_outcomes` + + `player_calibrated_weights`; **no code reads or writes either table.** + +**So the replay ran against `ledger_entries`, the only permanent store.** Two consequences: + +1. **6 days, not 30.** Not a choice — that is the entire history. +2. **The denominator is the GRADED board, not the raw slate.** `ledgerService.js:173-181` + skips no-grade / `insufficient_data` / no-captured-line props. Reads the current gate + already refused were never written. So "cut %" below means *cut from what we publish + today*, which is the right question for G-b, but it cannot tell us about props that + never got that far. + +--- + +## 1. G-b — WHAT THE NEW GATE DOES TO THE BOARD + +### 1.1 Historical board by odds band + +| Band | Props | % board | Settled | Hit % | **Units** | **ROI** | +|---|---|---|---|---|---|---| +| takeable −160…+200 | 257 | 44.6 % | 209 | 52.6 % | −0.72 | **−0.3 %** | +| flex −161…−250 | 77 | 13.4 % | 70 | 67.1 % | +1.56 | **+2.2 %** | +| juiced −251…−400 | 4 | 0.7 % | 4 | 75.0 % | −0.12 | −3.0 % | +| **past −400** | **221** | **38.4 %** | 173 | 80.3 % | **−13.29** | **−7.7 %** | +| dog +201…+400 | 10 | 1.7 % | 8 | 37.5 % | +4.42 | +55.3 % | +| longshot > +400 | 2 | 0.3 % | 1 | 0.0 % | −1.00 | −100 % | +| no odds | 5 | 0.9 % | 5 | — | — | — | + +Flat 1u staking; ROI = units / settled. n≥20 holds only for **takeable (209)** and +**flex (70)** — every other row is anecdote and is reported for shape, not for truth. + +### 1.2 The three findings that matter + +**FINDING 1 — the −400 floor you already shipped was the whole win.** +Past −400 hit **80.3 %** but needed **86.9 %** to break even: **−13.29 units, −7.7 % ROI** +on 173 settled bets. That single band is essentially the entire historical loss. The +guard shipped this morning (`f72f063`/`348a82b`) already kills it. The board proves it: +MLB Jul 18 = 103 graded props, 71 past −400; MLB Jul 19 (post-guard) = 15 props, **0** +past −400. + +**FINDING 2 — Arc 2's incremental bite is small, and it lands almost entirely on the +flex band.** Against the *already-live* gate, the new knobs cut: `−251…−400` **4 props +total across 6 days**, no-odds **5**, `> +400` **2**. That is 11 props in 6 days. The +only material change is `EDGE_FLEX_WALL` gating **77 props (13.4 %)** behind a 2× EV test. + +**FINDING 3 — the flex band is the most profitable band we have, and the takeable band +is flat.** Flex −161…−250: **+2.2 % ROI** (67.1 % actual vs 65.8 % break-even, n=70). +Takeable −160…+200: **−0.3 % ROI** (52.6 % vs 53.1 % break-even, n=209). The band +doctrine wants to promote is break-even; the band Arc 2 proposes to restrict is the one +carrying positive ROI. Neither result is statistically strong at these n's, but the +direction is the opposite of the assumption behind `EDGE_FLEX_WALL`. + +### 1.3 Survival, per night (the "is it livable" question) + +| Date | Sport | Graded board | Survives outright | Flex (needs EV) | Cut: wall | no-odds | longshot | +|---|---|---|---|---|---|---|---| +| 07-11 | mlb | 74 | 21 | 15 | 34 | 4 | 0 | +| 07-12 | mlb | 54 | 8 | 6 | 40 | 0 | 0 | +| 07-16 | mlb | 48 | 14 | 15 | 18 | 0 | 1 | +| 07-16 | wnba | 32 | 26 | 6 | 0 | 0 | 0 | +| 07-17 | mlb | 86 | 14 | 10 | 61 | 1 | 0 | +| 07-17 | wnba | 70 | 66 | 4 | 0 | 0 | 0 | +| 07-18 | mlb | 103 | 21 | 9 | 72 | 0 | 1 | +| 07-18 | wnba | 62 | 54 | 8 | 0 | 0 | 0 | +| 07-19 | mlb | 15 | 13 | 2 | 0 | 0 | 0 | +| 07-19 | wnba | 32 | 30 | 2 | 0 | 0 | 0 | + +**Verdict: the board stays livable.** WNBA is already almost entirely takeable (54/62, +66/70, 30/32) and is barely touched. MLB thins to **~15-25 survivors a night** on the +old data — but nearly all of that thinning is the −400 mass *already removed today*. +A typical post-Arc-2 night looks like **~15-25 MLB + ~30-60 WNBA = 45-85 reads**. That +is not a starved board. + +Caveat: 07-19 is a partial day (census taken 22:01 UTC, before the 01/03 UTC slots). + +### 1.4 What could NOT be simulated — and it's the important half + +`ev_pct` is **not stored on any ledger row** (no column; Arc 1 computes it in-request and +it rides only the live payload). Neither is `p_win`. **Therefore the EV half of the +proposed gate — `EDGE_FLEX_WALL` + `EV_FLEX_THRESHOLD` — cannot be replayed against +history at all.** Everything in §1.3 is the *price-only* portion of the gate. The flex +column says "needs EV", not "survives". + +To ever answer this we must start persisting it — see §4 (C-led). + +--- + +## 2. C-cal — CALIBRATION + +### 2.1 Confidence is monotonic but badly miscalibrated + +| Bucket | n | Claimed conf | **Actual hit %** | Units | ROI | +|---|---|---|---|---|---| +| conf < 45 | 81 | 34.6 | 59.3 % | −6.47 | −8.0 % | +| conf 45-54 | 214 | 47.3 | 64.0 % | −6.61 | −3.1 % | +| conf 55-64 | 170 | 57.5 | 68.8 % | +3.94 | **+2.3 %** | + +**Ordering is real** — hit rate rises monotonically with confidence (59.3 → 64.0 → 68.8), +and so does ROI. **Absolute calibration is not** — confidence understates the hit rate by +~20-25 points at every level. This is the audit's known grade↔confidence mismatch +(`mlb-grade-degradation.md`), now quantified against settled outcomes. + +**Consequence for the engine:** `confidence` must NEVER be fed into an EV or probability +calculation. Arc 1 is already correct here — `evPct` uses `p_win` from the quantile +estimator, not confidence. Do not "fix" confidence by rescaling it into a probability; +it is a display ordering, and the two must stay separate. + +### 2.2 The grade distribution is degenerate + +| Grade | Rows | Takeable | Avg conf | Settled hit % | Units | ROI | +|---|---|---|---|---|---|---| +| B | 384 | 149 | 53.0 | 67.5 % | −7.82 | −2.5 % | +| C | 184 | 102 | 43.9 | 59.9 % | −1.33 | −0.8 % | + +**The entire public ledger contains two letters: B and C. Zero A, zero A+, zero D/F.** +Confidence spans only ~34-64. The model is not using its own scale. + +This directly breaks hero v2: `pickHeroProp` filters `isAB(g.grade)`, so the hero can +only ever be a B. It also means "A-RATED" copy on public surfaces describes a grade the +engine has never emitted in the recorded era. + +**This is the single biggest finding in the report** and it is upstream of the whole value +engine — a gate that ranks EV within an undifferentiated B pool is tuning the wrong knob. + +### 2.3 EV_FLEX_THRESHOLD — proposal with reasoning + +Kev asked me to confirm against calibration if available, else propose. Calibration exists +but **cannot validate an EV threshold** (§1.4: no stored EV). So this is a proposal, and I +want it labelled as such rather than dressed up as data-driven. + +**Proposal: keep `EV_FLEX_THRESHOLD` = 2× `VALUE_EV_THRESHOLD` (= 4 %) as the default — +but ship it OFF-by-default in the flex band until EV is persisted and re-checked.** + +Reasoning: +- The flex band is the only clearly-positive band we have (+2.2 %, n=70). Restricting it + on an unvalidated threshold risks cutting the profitable third of the board to enforce + a rule we cannot yet measure. +- A 4 % EV bar is defensible on first principles: at −200 you need 66.7 % to break even, so + 4 % EV ≈ 69 % model probability — a real, not rounding, edge. I'm comfortable with the + *number*; I'm not comfortable *enforcing* it blind. +- Concretely: implement the knob, wire the band, default `EV_FLEX_ENFORCE=0` so flex props + grade as they do today while `ev_pct` gets recorded. Flip to `1` after ~2 weeks of stored + EV shows what the threshold actually cuts. Zero board impact now, full data later. + +### 2.4 Recommended dial changes vs your spec + +| Knob | Your spec | My recommendation | Why | +|---|---|---|---| +| `TAKEABLE_ODDS_CEILING` | −160 | **−160, unchanged** | promotion band, works | +| `HARD_JUICE_WALL` | −250 | **−250, ship it** | costs 4 props/6 days; band is −3 % ROI; the −400→−250 tightening is nearly free | +| `EDGE_FLEX_WALL` | −250 | **−250 band, enforcement OFF by default** | §2.3 — the band is our best performer, threshold unvalidated | +| `EV_FLEX_THRESHOLD` | 2× (=4 %) | **4 %, dormant until measured** | §2.3 | +| `LADDER_ODDS_MAX` | +400 | **+400, ship it** | costs 2 props/6 days; the one settled longshot lost | +| `MIN_RUNG_PROBABILITY` | 0.25 | **ship the knob, but it binds on nothing today** | §3 — no per-rung odds exist, so there are no rungs to price | +| no-odds → refuse | yes | **ship it** | 5 props; and all 5 had `model_value = 0` — pure garbage | +| `JUICE_ODDS_FLOOR` | keep as backstop | **keep at −400** | dead code once the −250 wall lands, but harmless and legacy-safe | + +--- + +## 3. L-a — DO `alt_lines` CARRY PER-RUNG ODDS? **NO.** + +- A rung is exactly `{ line, grade, edge_pct, base }` (`analyzeViaEngine1.js:492`). No + price, no probability. +- The rungs are **synthetic** — fixed `[−1,−0.5,0,+0.5,+1]` shifts off the main line, + re-graded on the same feature vector. They are not market lines. +- **The feed offers nothing better today:** grep for `alternate|alt_line` across + `proplineAdapter.js`, `oddsNormalizer.js`, `oddsService.js` returns **zero hits**. No + alternate-line market is requested, normalized, or stored. + +**L-b (per-rung EV) is BLOCKED on a data source, not on engine work.** Per-rung EV needs a +real price per rung. Options, in cost order: (a) check whether PropLine exposes alternate +markets on the existing 3 free keys — costs nothing but a probe; (b) odds-api alternate +markets — blocked, quota is 0/500 and Kev's ruling is hold the line; (c) derive rung prices +from a distribution around the main line — **rejected, that fabricates market data and +violates the Data Semantics Rule.** + +Recommend: probe PropLine (free) before any L-b work is scheduled. + +--- + +## 4. C-led — HOW MUCH HISTORY HAS ODDS ATTACHED? + +**Good news: 571 of 576 rows (99.1 %) already carry `locked_odds`, and 465+ settled rows +have both odds and an outcome.** `locked_odds` was in migration 019 from the start +(`019:17-43`) and `rowsFromSnapshot` has always populated it (`ledgerService.js:186-205`). + +**So C-led needs no backfill for odds — the ledger can speak units TODAY.** That is why +§1.1's ROI table exists at all. + +What is genuinely missing and needs new columns (not backfill — these were never computed +historically): `ev_pct`, `p_win`, `fair_odds`, `takeable`, `value`. Honest handling: add +the columns, populate going forward, and label the start date. **"Tracking began +2026-07-XX" is the honest answer — do not backfill EV by recomputing it from today's +model against yesterday's lines.** That would be a fabricated record of a model that +didn't exist yet. + +`getModelAggregate` currently reads only `outcome, clv_result, clv, player_key, grade, +model_value` (`ledgerService.js:454`) — it never reads `locked_odds`, so units/ROI/ +record-by-odds-band are all new aggregate work on data that already exists. + +--- + +## 5. C-clv / S-a — CLV IS CONFIRMED BROKEN IN THE DATA + +The C4 write-up is right, and the ledger proves it: + +| Sport | rows | closing_odds present | **closing_line == locked line** | clv = 0 | clv ≠ 0 | +|---|---|---|---|---|---| +| mlb | 380 | 376 | **359** | 285 | 21 | +| wnba | 196 | 196 | **164** | 138 | 26 | + +`captureClosing` re-records the lock, so CLV is structurally ~0. `clvCaptureReliable()` +correctly suppresses `beat_close_pct` and `clv_distribution` on every public surface +(`ledgerService.js:45,532-547`). **Keep it suppressed.** The fix shares plumbing with S-a +(line-moved truth) exactly as you scoped — both need a real "fresher odds at serve time" +read, which the 24h `snapshot:latest` + 20-min intraday refresh can supply without new +quota. + +--- + +## 6. U-deg — STATUS: THE `projection == 0` LEAK APPEARS ALREADY CLOSED + +`model_value = 0` by day: 07-11 **20** · 07-12 **17** · 07-16 **10** · 07-17 **8** · +07-18 **0** · 07-19 **0**. + +It stopped. 55 historical rows carry `model_value = 0`; **all 55 are grade B**, and 50 of +them sit past −400 (so the shipped juice floor would have refused them anyway). The +`projection > 0` refusal at `analyzeViaEngine1.js:422` is doing its job. + +Two things still true and worth carrying into G-a: +1. `getModelAggregate` already excludes them via `.gt('model_value', 0)` + (`ledgerService.js:464`) — public accuracy was never polluted by these rows. +2. Folding the `projection > 0` check into G-a's gate (as you asked) is still correct — + it's currently a separate refusal *after* feature computation, and moving it into the + gate chokepoint makes the refusal reasons uniform. It is a **tidy-up, not a leak fix.** + +The `edge_pct` scale problem is NOT resolved and is untouched by this report. + +--- + +## 7. WHAT I RECOMMEND HAPPENS NEXT + +1. **Kev rules on §2.4** (dial table) — specifically flex-band enforcement on/off. +2. **G-a ships** with the agreed dials + `gate_*` refusal reasons + no-odds refusal + + the folded `projection > 0` check. Low risk: ~11 props in 6 days of incremental cuts. +3. **C-led first, not later** — add `ev_pct`/`p_win`/`fair_odds` columns and start + recording. Nothing else in this train can be validated until EV is on disk. This is the + cheapest highest-leverage item on the board. +4. **Escalate §2.2 (only B and C grades ever emitted)** to its own investigation. It + undermines hero v2, "A-RATED" copy, and any EV ranking within a flat grade pool. +5. Probe PropLine for alternate markets before scheduling L-b. + +**HANDOFF — Design (Session 2):** nothing in this report is renderable yet. When the gate +ships, the refusal reasons `gate_too_juiced` / `gate_thin_edge` / `gate_longshot` / +`no_odds` will each carry user copy in the payload (D-ref), and §1.3 says the empty-board +state (G-c) will be rare but real on a thin MLB night — it needs a designed empty state, +not a blank grid. + +--- + +*Report generated 2026-07-19 against live prod data. No engine code changed. G-a held for +ruling.* diff --git a/specs/model-train.md b/specs/model-train.md index 56f84f5..caf3b10 100644 --- a/specs/model-train.md +++ b/specs/model-train.md @@ -153,6 +153,34 @@ timestamp) is unchanged. `toHero` now also passes through --- +## 2A. ARC 2+ BOARD — STATUS (Kev's arc list, 2026-07-19) + +Full arc definitions live in the Session-63 order. Status only here; update as each ships. + +| Item | Status | Note | +|---|---|---| +| **G-a** the real gate | **HELD for ruling** | Built nothing yet — the G-b report changes the recommended dials. See `specs/audit-data/gate-simulation.md` §2.4. | +| **G-b** gate simulation | ✅ **REPORTED** | `specs/audit-data/gate-simulation.md`. Replayed over the ledger, NOT snapshots (no 30d snapshot store exists). | +| **G-c** never-empty honesty | open | Rare but real on thin MLB nights. | +| **S-a** line-moved truth | open | Shares plumbing with C-clv — build together. | +| **S-b** board ranks on EV | open | Blocked-ish: EV not persisted; live payload has it. | +| **L-a** per-rung odds? | ✅ **REPORTED — NO** | Rungs are synthetic, carry no price. Feed offers no alternate markets. §3 of the report. | +| **L-b** per-rung EV ladder | **BLOCKED** | Needs a real price per rung. Probe PropLine first. | +| **C-cal** calibration | ✅ **REPORTED** | Confidence monotonic but ~20-25pts miscalibrated; **only B/C grades ever emitted**. | +| **C-led** ledger speaks value | open — **recommended next** | `locked_odds` already 99.1% populated: units/ROI need no backfill. EV columns are net-new, forward-only. | +| **C-clv** fix C4 | open | Confirmed broken in data (359/376 MLB closes == lock). Keep suppressed. | +| **D-ref** visible refusals | partial | Copy already ships on refusals; new `gate_*` reasons land with G-a. | +| **D-par** parlay lab | open | Gate must NOT apply inside the Lab. | +| **D-tier** tier gating | **DECIDED, not enforced** | Free/Analyst: value marker + grade + triplet. Desk: ladder + per-rung EV + Kelly. Today ladder/Kelly are already Desk-gated; triplet ungated = correct per this ruling. | +| **D-ev** consolidate EV | open | `devig.evPct` vs `processing/EVCalculator.js` (used only by `UnifiedOddsProvider`). | +| **U-deg** MLB degradation | ✅ **STATUS REPORTED** | `projection==0` leak **already closed** (0 occurrences since 07-18). `edge_pct` scale still broken. | +| **U-fp** Arc 1 fingerprint | open | Do it on the first deploy this train ships. | + +### Backtest harness +**None exists** (grep-verified). `migrations/006` defines `grade_outcomes` + +`player_calibrated_weights` and **no code reads or writes them**. The G-b/C-cal replay was +done in SQL against `ledger_entries`; a real harness is still owed before any weight change. + ## 3. NOT YET BUILT — checklist to reconcile against the full arc list Verified absent from the codebase as of `7a925f4`. Kev supplies the complete arc list;