Model Train G-b/C-cal: gate simulation + calibration report (docs only)

REPORT-FIRST per the arc order. G-a is HELD — the data changes the
recommended dials. No engine code touched.

Replayed against live ledger_entries (576 rows, 6 game days, 470 settled)
because the "30 days of stored snapshots" does not exist: snapshot Redis
keys are latest/previous only at 24h TTL, and no backtest harness exists
anywhere in the repo.

Findings that change the plan:
- The -400 floor shipped this morning was the whole win: past -400 hit
  80.3% against an 86.9% breakeven = -13.29u / -7.7% ROI on 173 settled.
- Arc 2's incremental cut over the live gate is ~11 props in 6 days. The
  only material change is gating the flex band behind 2x EV.
- The flex band (-161..-250) is our BEST band (+2.2% ROI, n=70) and the
  takeable band is flat (-0.3%, n=209) — the opposite of the assumption
  behind EDGE_FLEX_WALL. Recommend shipping the knob with enforcement
  OFF until EV is persisted and measured.
- ev_pct/p_win are on NO ledger row, so the EV half of the gate cannot be
  replayed at all. C-led (persist EV) is now the highest-leverage item.
- Confidence is monotonic but understates hit rate by ~20-25 points, and
  the entire public ledger contains only B and C grades — zero A/A+.
  That breaks hero v2 (isAB) and undermines "A-RATED" copy. Escalated.
- L-a answered: alt_lines carry NO odds and the feed has no alternate
  markets. L-b is blocked on a data source, not engine work.
- C-led needs no odds backfill (locked_odds 99.1% populated).
- U-deg: the projection==0 leak is already closed (0 since 07-18).
- C4 confirmed in data (359/376 MLB closes == the lock). Stays suppressed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
This commit is contained in:
Kev
2026-07-19 18:04:29 -04:00
parent 18bf3ecb51
commit 669479097c
2 changed files with 311 additions and 0 deletions
+283
View File
@@ -0,0 +1,283 @@
# GATE SIMULATION + CALIBRATION — Model Train G-b / C-cal
**REPORT-FIRST deliverable. No engine code was written for this. G-a is HELD pending Kev's ruling.**
Dataset snapshot: `2026-07-19 22:01:58 UTC`, live Supabase `ledger_entries`, `user_id IS NULL`
(the public model record). **576 rows · 6 game days (2026-07-11 → 07-19) · 470 settled ·
571 with locked odds.** Numbers move: the pipeline inserted 8 rows mid-census, so every
table here is pinned to that timestamp.
---
## 0. THE BLOCKER: "last 30 days of stored snapshots" DOES NOT EXIST
There is no 30-day snapshot store to replay. Verified on disk:
- `snapshot:{sport}:latest` and `:previous`**two generations, 24h TTL**
(`snapshotService.js:404,412`). No `snapshot:{sport}:{date}` archive exists anywhere.
- The intraday line `history:[{t,line}]` rides **inside** the same 24h blob
(`intradayRefreshService.js:68-84,144`) — it dies with it.
- `outcomes:{sport}:log` (30d TTL, cap 1000) carries **no odds, no confidence, no
projection** (`outcomeService.js:233-237`) — unusable for an odds-aware table.
- **No backtest/replay harness exists** — grep for `backtest|replay` across `src/`,
`scripts/`, `tests/` returns zero. `migrations/006` defines `grade_outcomes` +
`player_calibrated_weights`; **no code reads or writes either table.**
**So the replay ran against `ledger_entries`, the only permanent store.** Two consequences:
1. **6 days, not 30.** Not a choice — that is the entire history.
2. **The denominator is the GRADED board, not the raw slate.** `ledgerService.js:173-181`
skips no-grade / `insufficient_data` / no-captured-line props. Reads the current gate
already refused were never written. So "cut %" below means *cut from what we publish
today*, which is the right question for G-b, but it cannot tell us about props that
never got that far.
---
## 1. G-b — WHAT THE NEW GATE DOES TO THE BOARD
### 1.1 Historical board by odds band
| Band | Props | % board | Settled | Hit % | **Units** | **ROI** |
|---|---|---|---|---|---|---|
| takeable 160…+200 | 257 | 44.6 % | 209 | 52.6 % | 0.72 | **0.3 %** |
| flex 161…−250 | 77 | 13.4 % | 70 | 67.1 % | +1.56 | **+2.2 %** |
| juiced 251…−400 | 4 | 0.7 % | 4 | 75.0 % | 0.12 | 3.0 % |
| **past 400** | **221** | **38.4 %** | 173 | 80.3 % | **13.29** | **7.7 %** |
| dog +201…+400 | 10 | 1.7 % | 8 | 37.5 % | +4.42 | +55.3 % |
| longshot > +400 | 2 | 0.3 % | 1 | 0.0 % | 1.00 | 100 % |
| no odds | 5 | 0.9 % | 5 | — | — | — |
Flat 1u staking; ROI = units / settled. n≥20 holds only for **takeable (209)** and
**flex (70)** — every other row is anecdote and is reported for shape, not for truth.
### 1.2 The three findings that matter
**FINDING 1 — the 400 floor you already shipped was the whole win.**
Past 400 hit **80.3 %** but needed **86.9 %** to break even: **13.29 units, 7.7 % ROI**
on 173 settled bets. That single band is essentially the entire historical loss. The
guard shipped this morning (`f72f063`/`348a82b`) already kills it. The board proves it:
MLB Jul 18 = 103 graded props, 71 past 400; MLB Jul 19 (post-guard) = 15 props, **0**
past 400.
**FINDING 2 — Arc 2's incremental bite is small, and it lands almost entirely on the
flex band.** Against the *already-live* gate, the new knobs cut: `251…−400` **4 props
total across 6 days**, no-odds **5**, `> +400` **2**. That is 11 props in 6 days. The
only material change is `EDGE_FLEX_WALL` gating **77 props (13.4 %)** behind a 2× EV test.
**FINDING 3 — the flex band is the most profitable band we have, and the takeable band
is flat.** Flex 161…−250: **+2.2 % ROI** (67.1 % actual vs 65.8 % break-even, n=70).
Takeable 160…+200: **0.3 % ROI** (52.6 % vs 53.1 % break-even, n=209). The band
doctrine wants to promote is break-even; the band Arc 2 proposes to restrict is the one
carrying positive ROI. Neither result is statistically strong at these n's, but the
direction is the opposite of the assumption behind `EDGE_FLEX_WALL`.
### 1.3 Survival, per night (the "is it livable" question)
| Date | Sport | Graded board | Survives outright | Flex (needs EV) | Cut: wall | no-odds | longshot |
|---|---|---|---|---|---|---|---|
| 07-11 | mlb | 74 | 21 | 15 | 34 | 4 | 0 |
| 07-12 | mlb | 54 | 8 | 6 | 40 | 0 | 0 |
| 07-16 | mlb | 48 | 14 | 15 | 18 | 0 | 1 |
| 07-16 | wnba | 32 | 26 | 6 | 0 | 0 | 0 |
| 07-17 | mlb | 86 | 14 | 10 | 61 | 1 | 0 |
| 07-17 | wnba | 70 | 66 | 4 | 0 | 0 | 0 |
| 07-18 | mlb | 103 | 21 | 9 | 72 | 0 | 1 |
| 07-18 | wnba | 62 | 54 | 8 | 0 | 0 | 0 |
| 07-19 | mlb | 15 | 13 | 2 | 0 | 0 | 0 |
| 07-19 | wnba | 32 | 30 | 2 | 0 | 0 | 0 |
**Verdict: the board stays livable.** WNBA is already almost entirely takeable (54/62,
66/70, 30/32) and is barely touched. MLB thins to **~15-25 survivors a night** on the
old data — but nearly all of that thinning is the 400 mass *already removed today*.
A typical post-Arc-2 night looks like **~15-25 MLB + ~30-60 WNBA = 45-85 reads**. That
is not a starved board.
Caveat: 07-19 is a partial day (census taken 22:01 UTC, before the 01/03 UTC slots).
### 1.4 What could NOT be simulated — and it's the important half
`ev_pct` is **not stored on any ledger row** (no column; Arc 1 computes it in-request and
it rides only the live payload). Neither is `p_win`. **Therefore the EV half of the
proposed gate — `EDGE_FLEX_WALL` + `EV_FLEX_THRESHOLD` — cannot be replayed against
history at all.** Everything in §1.3 is the *price-only* portion of the gate. The flex
column says "needs EV", not "survives".
To ever answer this we must start persisting it — see §4 (C-led).
---
## 2. C-cal — CALIBRATION
### 2.1 Confidence is monotonic but badly miscalibrated
| Bucket | n | Claimed conf | **Actual hit %** | Units | ROI |
|---|---|---|---|---|---|
| conf < 45 | 81 | 34.6 | 59.3 % | 6.47 | 8.0 % |
| conf 45-54 | 214 | 47.3 | 64.0 % | 6.61 | 3.1 % |
| conf 55-64 | 170 | 57.5 | 68.8 % | +3.94 | **+2.3 %** |
**Ordering is real** — hit rate rises monotonically with confidence (59.3 → 64.0 → 68.8),
and so does ROI. **Absolute calibration is not** — confidence understates the hit rate by
~20-25 points at every level. This is the audit's known grade↔confidence mismatch
(`mlb-grade-degradation.md`), now quantified against settled outcomes.
**Consequence for the engine:** `confidence` must NEVER be fed into an EV or probability
calculation. Arc 1 is already correct here — `evPct` uses `p_win` from the quantile
estimator, not confidence. Do not "fix" confidence by rescaling it into a probability;
it is a display ordering, and the two must stay separate.
### 2.2 The grade distribution is degenerate
| Grade | Rows | Takeable | Avg conf | Settled hit % | Units | ROI |
|---|---|---|---|---|---|---|
| B | 384 | 149 | 53.0 | 67.5 % | 7.82 | 2.5 % |
| C | 184 | 102 | 43.9 | 59.9 % | 1.33 | 0.8 % |
**The entire public ledger contains two letters: B and C. Zero A, zero A+, zero D/F.**
Confidence spans only ~34-64. The model is not using its own scale.
This directly breaks hero v2: `pickHeroProp` filters `isAB(g.grade)`, so the hero can
only ever be a B. It also means "A-RATED" copy on public surfaces describes a grade the
engine has never emitted in the recorded era.
**This is the single biggest finding in the report** and it is upstream of the whole value
engine — a gate that ranks EV within an undifferentiated B pool is tuning the wrong knob.
### 2.3 EV_FLEX_THRESHOLD — proposal with reasoning
Kev asked me to confirm against calibration if available, else propose. Calibration exists
but **cannot validate an EV threshold** (§1.4: no stored EV). So this is a proposal, and I
want it labelled as such rather than dressed up as data-driven.
**Proposal: keep `EV_FLEX_THRESHOLD` = 2× `VALUE_EV_THRESHOLD` (= 4 %) as the default —
but ship it OFF-by-default in the flex band until EV is persisted and re-checked.**
Reasoning:
- The flex band is the only clearly-positive band we have (+2.2 %, n=70). Restricting it
on an unvalidated threshold risks cutting the profitable third of the board to enforce
a rule we cannot yet measure.
- A 4 % EV bar is defensible on first principles: at 200 you need 66.7 % to break even, so
4 % EV ≈ 69 % model probability — a real, not rounding, edge. I'm comfortable with the
*number*; I'm not comfortable *enforcing* it blind.
- Concretely: implement the knob, wire the band, default `EV_FLEX_ENFORCE=0` so flex props
grade as they do today while `ev_pct` gets recorded. Flip to `1` after ~2 weeks of stored
EV shows what the threshold actually cuts. Zero board impact now, full data later.
### 2.4 Recommended dial changes vs your spec
| Knob | Your spec | My recommendation | Why |
|---|---|---|---|
| `TAKEABLE_ODDS_CEILING` | 160 | **160, unchanged** | promotion band, works |
| `HARD_JUICE_WALL` | 250 | **250, ship it** | costs 4 props/6 days; band is 3 % ROI; the 400→−250 tightening is nearly free |
| `EDGE_FLEX_WALL` | 250 | **250 band, enforcement OFF by default** | §2.3 — the band is our best performer, threshold unvalidated |
| `EV_FLEX_THRESHOLD` | 2× (=4 %) | **4 %, dormant until measured** | §2.3 |
| `LADDER_ODDS_MAX` | +400 | **+400, ship it** | costs 2 props/6 days; the one settled longshot lost |
| `MIN_RUNG_PROBABILITY` | 0.25 | **ship the knob, but it binds on nothing today** | §3 — no per-rung odds exist, so there are no rungs to price |
| no-odds → refuse | yes | **ship it** | 5 props; and all 5 had `model_value = 0` — pure garbage |
| `JUICE_ODDS_FLOOR` | keep as backstop | **keep at 400** | dead code once the 250 wall lands, but harmless and legacy-safe |
---
## 3. L-a — DO `alt_lines` CARRY PER-RUNG ODDS? **NO.**
- A rung is exactly `{ line, grade, edge_pct, base }` (`analyzeViaEngine1.js:492`). No
price, no probability.
- The rungs are **synthetic** — fixed `[1,0.5,0,+0.5,+1]` shifts off the main line,
re-graded on the same feature vector. They are not market lines.
- **The feed offers nothing better today:** grep for `alternate|alt_line` across
`proplineAdapter.js`, `oddsNormalizer.js`, `oddsService.js` returns **zero hits**. No
alternate-line market is requested, normalized, or stored.
**L-b (per-rung EV) is BLOCKED on a data source, not on engine work.** Per-rung EV needs a
real price per rung. Options, in cost order: (a) check whether PropLine exposes alternate
markets on the existing 3 free keys — costs nothing but a probe; (b) odds-api alternate
markets — blocked, quota is 0/500 and Kev's ruling is hold the line; (c) derive rung prices
from a distribution around the main line — **rejected, that fabricates market data and
violates the Data Semantics Rule.**
Recommend: probe PropLine (free) before any L-b work is scheduled.
---
## 4. C-led — HOW MUCH HISTORY HAS ODDS ATTACHED?
**Good news: 571 of 576 rows (99.1 %) already carry `locked_odds`, and 465+ settled rows
have both odds and an outcome.** `locked_odds` was in migration 019 from the start
(`019:17-43`) and `rowsFromSnapshot` has always populated it (`ledgerService.js:186-205`).
**So C-led needs no backfill for odds — the ledger can speak units TODAY.** That is why
§1.1's ROI table exists at all.
What is genuinely missing and needs new columns (not backfill — these were never computed
historically): `ev_pct`, `p_win`, `fair_odds`, `takeable`, `value`. Honest handling: add
the columns, populate going forward, and label the start date. **"Tracking began
2026-07-XX" is the honest answer — do not backfill EV by recomputing it from today's
model against yesterday's lines.** That would be a fabricated record of a model that
didn't exist yet.
`getModelAggregate` currently reads only `outcome, clv_result, clv, player_key, grade,
model_value` (`ledgerService.js:454`) — it never reads `locked_odds`, so units/ROI/
record-by-odds-band are all new aggregate work on data that already exists.
---
## 5. C-clv / S-a — CLV IS CONFIRMED BROKEN IN THE DATA
The C4 write-up is right, and the ledger proves it:
| Sport | rows | closing_odds present | **closing_line == locked line** | clv = 0 | clv ≠ 0 |
|---|---|---|---|---|---|
| mlb | 380 | 376 | **359** | 285 | 21 |
| wnba | 196 | 196 | **164** | 138 | 26 |
`captureClosing` re-records the lock, so CLV is structurally ~0. `clvCaptureReliable()`
correctly suppresses `beat_close_pct` and `clv_distribution` on every public surface
(`ledgerService.js:45,532-547`). **Keep it suppressed.** The fix shares plumbing with S-a
(line-moved truth) exactly as you scoped — both need a real "fresher odds at serve time"
read, which the 24h `snapshot:latest` + 20-min intraday refresh can supply without new
quota.
---
## 6. U-deg — STATUS: THE `projection == 0` LEAK APPEARS ALREADY CLOSED
`model_value = 0` by day: 07-11 **20** · 07-12 **17** · 07-16 **10** · 07-17 **8** ·
07-18 **0** · 07-19 **0**.
It stopped. 55 historical rows carry `model_value = 0`; **all 55 are grade B**, and 50 of
them sit past 400 (so the shipped juice floor would have refused them anyway). The
`projection > 0` refusal at `analyzeViaEngine1.js:422` is doing its job.
Two things still true and worth carrying into G-a:
1. `getModelAggregate` already excludes them via `.gt('model_value', 0)`
(`ledgerService.js:464`) — public accuracy was never polluted by these rows.
2. Folding the `projection > 0` check into G-a's gate (as you asked) is still correct —
it's currently a separate refusal *after* feature computation, and moving it into the
gate chokepoint makes the refusal reasons uniform. It is a **tidy-up, not a leak fix.**
The `edge_pct` scale problem is NOT resolved and is untouched by this report.
---
## 7. WHAT I RECOMMEND HAPPENS NEXT
1. **Kev rules on §2.4** (dial table) — specifically flex-band enforcement on/off.
2. **G-a ships** with the agreed dials + `gate_*` refusal reasons + no-odds refusal +
the folded `projection > 0` check. Low risk: ~11 props in 6 days of incremental cuts.
3. **C-led first, not later** — add `ev_pct`/`p_win`/`fair_odds` columns and start
recording. Nothing else in this train can be validated until EV is on disk. This is the
cheapest highest-leverage item on the board.
4. **Escalate §2.2 (only B and C grades ever emitted)** to its own investigation. It
undermines hero v2, "A-RATED" copy, and any EV ranking within a flat grade pool.
5. Probe PropLine for alternate markets before scheduling L-b.
**HANDOFF — Design (Session 2):** nothing in this report is renderable yet. When the gate
ships, the refusal reasons `gate_too_juiced` / `gate_thin_edge` / `gate_longshot` /
`no_odds` will each carry user copy in the payload (D-ref), and §1.3 says the empty-board
state (G-c) will be rare but real on a thin MLB night — it needs a designed empty state,
not a blank grid.
---
*Report generated 2026-07-19 against live prod data. No engine code changed. G-a held for
ruling.*
+28
View File
@@ -153,6 +153,34 @@ timestamp) is unchanged. `toHero` now also passes through
---
## 2A. ARC 2+ BOARD — STATUS (Kev's arc list, 2026-07-19)
Full arc definitions live in the Session-63 order. Status only here; update as each ships.
| Item | Status | Note |
|---|---|---|
| **G-a** the real gate | **HELD for ruling** | Built nothing yet — the G-b report changes the recommended dials. See `specs/audit-data/gate-simulation.md` §2.4. |
| **G-b** gate simulation | ✅ **REPORTED** | `specs/audit-data/gate-simulation.md`. Replayed over the ledger, NOT snapshots (no 30d snapshot store exists). |
| **G-c** never-empty honesty | open | Rare but real on thin MLB nights. |
| **S-a** line-moved truth | open | Shares plumbing with C-clv — build together. |
| **S-b** board ranks on EV | open | Blocked-ish: EV not persisted; live payload has it. |
| **L-a** per-rung odds? | ✅ **REPORTED — NO** | Rungs are synthetic, carry no price. Feed offers no alternate markets. §3 of the report. |
| **L-b** per-rung EV ladder | **BLOCKED** | Needs a real price per rung. Probe PropLine first. |
| **C-cal** calibration | ✅ **REPORTED** | Confidence monotonic but ~20-25pts miscalibrated; **only B/C grades ever emitted**. |
| **C-led** ledger speaks value | open — **recommended next** | `locked_odds` already 99.1% populated: units/ROI need no backfill. EV columns are net-new, forward-only. |
| **C-clv** fix C4 | open | Confirmed broken in data (359/376 MLB closes == lock). Keep suppressed. |
| **D-ref** visible refusals | partial | Copy already ships on refusals; new `gate_*` reasons land with G-a. |
| **D-par** parlay lab | open | Gate must NOT apply inside the Lab. |
| **D-tier** tier gating | **DECIDED, not enforced** | Free/Analyst: value marker + grade + triplet. Desk: ladder + per-rung EV + Kelly. Today ladder/Kelly are already Desk-gated; triplet ungated = correct per this ruling. |
| **D-ev** consolidate EV | open | `devig.evPct` vs `processing/EVCalculator.js` (used only by `UnifiedOddsProvider`). |
| **U-deg** MLB degradation | ✅ **STATUS REPORTED** | `projection==0` leak **already closed** (0 occurrences since 07-18). `edge_pct` scale still broken. |
| **U-fp** Arc 1 fingerprint | open | Do it on the first deploy this train ships. |
### Backtest harness
**None exists** (grep-verified). `migrations/006` defines `grade_outcomes` +
`player_calibrated_weights` and **no code reads or writes them**. The G-b/C-cal replay was
done in SQL against `ledger_entries`; a real harness is still owed before any weight change.
## 3. NOT YET BUILT — checklist to reconcile against the full arc list
Verified absent from the codebase as of `7a925f4`. Kev supplies the complete arc list;