REPORT-FIRST per the arc order. G-a is HELD — the data changes the recommended dials. No engine code touched. Replayed against live ledger_entries (576 rows, 6 game days, 470 settled) because the "30 days of stored snapshots" does not exist: snapshot Redis keys are latest/previous only at 24h TTL, and no backtest harness exists anywhere in the repo. Findings that change the plan: - The -400 floor shipped this morning was the whole win: past -400 hit 80.3% against an 86.9% breakeven = -13.29u / -7.7% ROI on 173 settled. - Arc 2's incremental cut over the live gate is ~11 props in 6 days. The only material change is gating the flex band behind 2x EV. - The flex band (-161..-250) is our BEST band (+2.2% ROI, n=70) and the takeable band is flat (-0.3%, n=209) — the opposite of the assumption behind EDGE_FLEX_WALL. Recommend shipping the knob with enforcement OFF until EV is persisted and measured. - ev_pct/p_win are on NO ledger row, so the EV half of the gate cannot be replayed at all. C-led (persist EV) is now the highest-leverage item. - Confidence is monotonic but understates hit rate by ~20-25 points, and the entire public ledger contains only B and C grades — zero A/A+. That breaks hero v2 (isAB) and undermines "A-RATED" copy. Escalated. - L-a answered: alt_lines carry NO odds and the feed has no alternate markets. L-b is blocked on a data source, not engine work. - C-led needs no odds backfill (locked_odds 99.1% populated). - U-deg: the projection==0 leak is already closed (0 since 07-18). - C4 confirmed in data (359/376 MLB closes == the lock). Stays suppressed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
15 KiB
GATE SIMULATION + CALIBRATION — Model Train G-b / C-cal
REPORT-FIRST deliverable. No engine code was written for this. G-a is HELD pending Kev's ruling.
Dataset snapshot: 2026-07-19 22:01:58 UTC, live Supabase ledger_entries, user_id IS NULL
(the public model record). 576 rows · 6 game days (2026-07-11 → 07-19) · 470 settled ·
571 with locked odds. Numbers move: the pipeline inserted 8 rows mid-census, so every
table here is pinned to that timestamp.
0. THE BLOCKER: "last 30 days of stored snapshots" DOES NOT EXIST
There is no 30-day snapshot store to replay. Verified on disk:
snapshot:{sport}:latestand:previous— two generations, 24h TTL (snapshotService.js:404,412). Nosnapshot:{sport}:{date}archive exists anywhere.- The intraday line
history:[{t,line}]rides inside the same 24h blob (intradayRefreshService.js:68-84,144) — it dies with it. outcomes:{sport}:log(30d TTL, cap 1000) carries no odds, no confidence, no projection (outcomeService.js:233-237) — unusable for an odds-aware table.- No backtest/replay harness exists — grep for
backtest|replayacrosssrc/,scripts/,tests/returns zero.migrations/006definesgrade_outcomes+player_calibrated_weights; no code reads or writes either table.
So the replay ran against ledger_entries, the only permanent store. Two consequences:
- 6 days, not 30. Not a choice — that is the entire history.
- The denominator is the GRADED board, not the raw slate.
ledgerService.js:173-181skips no-grade /insufficient_data/ no-captured-line props. Reads the current gate already refused were never written. So "cut %" below means cut from what we publish today, which is the right question for G-b, but it cannot tell us about props that never got that far.
1. G-b — WHAT THE NEW GATE DOES TO THE BOARD
1.1 Historical board by odds band
| Band | Props | % board | Settled | Hit % | Units | ROI |
|---|---|---|---|---|---|---|
| takeable −160…+200 | 257 | 44.6 % | 209 | 52.6 % | −0.72 | −0.3 % |
| flex −161…−250 | 77 | 13.4 % | 70 | 67.1 % | +1.56 | +2.2 % |
| juiced −251…−400 | 4 | 0.7 % | 4 | 75.0 % | −0.12 | −3.0 % |
| past −400 | 221 | 38.4 % | 173 | 80.3 % | −13.29 | −7.7 % |
| dog +201…+400 | 10 | 1.7 % | 8 | 37.5 % | +4.42 | +55.3 % |
| longshot > +400 | 2 | 0.3 % | 1 | 0.0 % | −1.00 | −100 % |
| no odds | 5 | 0.9 % | 5 | — | — | — |
Flat 1u staking; ROI = units / settled. n≥20 holds only for takeable (209) and flex (70) — every other row is anecdote and is reported for shape, not for truth.
1.2 The three findings that matter
FINDING 1 — the −400 floor you already shipped was the whole win.
Past −400 hit 80.3 % but needed 86.9 % to break even: −13.29 units, −7.7 % ROI
on 173 settled bets. That single band is essentially the entire historical loss. The
guard shipped this morning (f72f063/348a82b) already kills it. The board proves it:
MLB Jul 18 = 103 graded props, 71 past −400; MLB Jul 19 (post-guard) = 15 props, 0
past −400.
FINDING 2 — Arc 2's incremental bite is small, and it lands almost entirely on the
flex band. Against the already-live gate, the new knobs cut: −251…−400 4 props
total across 6 days, no-odds 5, > +400 2. That is 11 props in 6 days. The
only material change is EDGE_FLEX_WALL gating 77 props (13.4 %) behind a 2× EV test.
FINDING 3 — the flex band is the most profitable band we have, and the takeable band
is flat. Flex −161…−250: +2.2 % ROI (67.1 % actual vs 65.8 % break-even, n=70).
Takeable −160…+200: −0.3 % ROI (52.6 % vs 53.1 % break-even, n=209). The band
doctrine wants to promote is break-even; the band Arc 2 proposes to restrict is the one
carrying positive ROI. Neither result is statistically strong at these n's, but the
direction is the opposite of the assumption behind EDGE_FLEX_WALL.
1.3 Survival, per night (the "is it livable" question)
| Date | Sport | Graded board | Survives outright | Flex (needs EV) | Cut: wall | no-odds | longshot |
|---|---|---|---|---|---|---|---|
| 07-11 | mlb | 74 | 21 | 15 | 34 | 4 | 0 |
| 07-12 | mlb | 54 | 8 | 6 | 40 | 0 | 0 |
| 07-16 | mlb | 48 | 14 | 15 | 18 | 0 | 1 |
| 07-16 | wnba | 32 | 26 | 6 | 0 | 0 | 0 |
| 07-17 | mlb | 86 | 14 | 10 | 61 | 1 | 0 |
| 07-17 | wnba | 70 | 66 | 4 | 0 | 0 | 0 |
| 07-18 | mlb | 103 | 21 | 9 | 72 | 0 | 1 |
| 07-18 | wnba | 62 | 54 | 8 | 0 | 0 | 0 |
| 07-19 | mlb | 15 | 13 | 2 | 0 | 0 | 0 |
| 07-19 | wnba | 32 | 30 | 2 | 0 | 0 | 0 |
Verdict: the board stays livable. WNBA is already almost entirely takeable (54/62, 66/70, 30/32) and is barely touched. MLB thins to ~15-25 survivors a night on the old data — but nearly all of that thinning is the −400 mass already removed today. A typical post-Arc-2 night looks like ~15-25 MLB + ~30-60 WNBA = 45-85 reads. That is not a starved board.
Caveat: 07-19 is a partial day (census taken 22:01 UTC, before the 01/03 UTC slots).
1.4 What could NOT be simulated — and it's the important half
ev_pct is not stored on any ledger row (no column; Arc 1 computes it in-request and
it rides only the live payload). Neither is p_win. Therefore the EV half of the
proposed gate — EDGE_FLEX_WALL + EV_FLEX_THRESHOLD — cannot be replayed against
history at all. Everything in §1.3 is the price-only portion of the gate. The flex
column says "needs EV", not "survives".
To ever answer this we must start persisting it — see §4 (C-led).
2. C-cal — CALIBRATION
2.1 Confidence is monotonic but badly miscalibrated
| Bucket | n | Claimed conf | Actual hit % | Units | ROI |
|---|---|---|---|---|---|
| conf < 45 | 81 | 34.6 | 59.3 % | −6.47 | −8.0 % |
| conf 45-54 | 214 | 47.3 | 64.0 % | −6.61 | −3.1 % |
| conf 55-64 | 170 | 57.5 | 68.8 % | +3.94 | +2.3 % |
Ordering is real — hit rate rises monotonically with confidence (59.3 → 64.0 → 68.8),
and so does ROI. Absolute calibration is not — confidence understates the hit rate by
~20-25 points at every level. This is the audit's known grade↔confidence mismatch
(mlb-grade-degradation.md), now quantified against settled outcomes.
Consequence for the engine: confidence must NEVER be fed into an EV or probability
calculation. Arc 1 is already correct here — evPct uses p_win from the quantile
estimator, not confidence. Do not "fix" confidence by rescaling it into a probability;
it is a display ordering, and the two must stay separate.
2.2 The grade distribution is degenerate
| Grade | Rows | Takeable | Avg conf | Settled hit % | Units | ROI |
|---|---|---|---|---|---|---|
| B | 384 | 149 | 53.0 | 67.5 % | −7.82 | −2.5 % |
| C | 184 | 102 | 43.9 | 59.9 % | −1.33 | −0.8 % |
The entire public ledger contains two letters: B and C. Zero A, zero A+, zero D/F. Confidence spans only ~34-64. The model is not using its own scale.
This directly breaks hero v2: pickHeroProp filters isAB(g.grade), so the hero can
only ever be a B. It also means "A-RATED" copy on public surfaces describes a grade the
engine has never emitted in the recorded era.
This is the single biggest finding in the report and it is upstream of the whole value engine — a gate that ranks EV within an undifferentiated B pool is tuning the wrong knob.
2.3 EV_FLEX_THRESHOLD — proposal with reasoning
Kev asked me to confirm against calibration if available, else propose. Calibration exists but cannot validate an EV threshold (§1.4: no stored EV). So this is a proposal, and I want it labelled as such rather than dressed up as data-driven.
Proposal: keep EV_FLEX_THRESHOLD = 2× VALUE_EV_THRESHOLD (= 4 %) as the default —
but ship it OFF-by-default in the flex band until EV is persisted and re-checked.
Reasoning:
- The flex band is the only clearly-positive band we have (+2.2 %, n=70). Restricting it on an unvalidated threshold risks cutting the profitable third of the board to enforce a rule we cannot yet measure.
- A 4 % EV bar is defensible on first principles: at −200 you need 66.7 % to break even, so 4 % EV ≈ 69 % model probability — a real, not rounding, edge. I'm comfortable with the number; I'm not comfortable enforcing it blind.
- Concretely: implement the knob, wire the band, default
EV_FLEX_ENFORCE=0so flex props grade as they do today whileev_pctgets recorded. Flip to1after ~2 weeks of stored EV shows what the threshold actually cuts. Zero board impact now, full data later.
2.4 Recommended dial changes vs your spec
| Knob | Your spec | My recommendation | Why |
|---|---|---|---|
TAKEABLE_ODDS_CEILING |
−160 | −160, unchanged | promotion band, works |
HARD_JUICE_WALL |
−250 | −250, ship it | costs 4 props/6 days; band is −3 % ROI; the −400→−250 tightening is nearly free |
EDGE_FLEX_WALL |
−250 | −250 band, enforcement OFF by default | §2.3 — the band is our best performer, threshold unvalidated |
EV_FLEX_THRESHOLD |
2× (=4 %) | 4 %, dormant until measured | §2.3 |
LADDER_ODDS_MAX |
+400 | +400, ship it | costs 2 props/6 days; the one settled longshot lost |
MIN_RUNG_PROBABILITY |
0.25 | ship the knob, but it binds on nothing today | §3 — no per-rung odds exist, so there are no rungs to price |
| no-odds → refuse | yes | ship it | 5 props; and all 5 had model_value = 0 — pure garbage |
JUICE_ODDS_FLOOR |
keep as backstop | keep at −400 | dead code once the −250 wall lands, but harmless and legacy-safe |
3. L-a — DO alt_lines CARRY PER-RUNG ODDS? NO.
- A rung is exactly
{ line, grade, edge_pct, base }(analyzeViaEngine1.js:492). No price, no probability. - The rungs are synthetic — fixed
[−1,−0.5,0,+0.5,+1]shifts off the main line, re-graded on the same feature vector. They are not market lines. - The feed offers nothing better today: grep for
alternate|alt_lineacrossproplineAdapter.js,oddsNormalizer.js,oddsService.jsreturns zero hits. No alternate-line market is requested, normalized, or stored.
L-b (per-rung EV) is BLOCKED on a data source, not on engine work. Per-rung EV needs a real price per rung. Options, in cost order: (a) check whether PropLine exposes alternate markets on the existing 3 free keys — costs nothing but a probe; (b) odds-api alternate markets — blocked, quota is 0/500 and Kev's ruling is hold the line; (c) derive rung prices from a distribution around the main line — rejected, that fabricates market data and violates the Data Semantics Rule.
Recommend: probe PropLine (free) before any L-b work is scheduled.
4. C-led — HOW MUCH HISTORY HAS ODDS ATTACHED?
Good news: 571 of 576 rows (99.1 %) already carry locked_odds, and 465+ settled rows
have both odds and an outcome. locked_odds was in migration 019 from the start
(019:17-43) and rowsFromSnapshot has always populated it (ledgerService.js:186-205).
So C-led needs no backfill for odds — the ledger can speak units TODAY. That is why §1.1's ROI table exists at all.
What is genuinely missing and needs new columns (not backfill — these were never computed
historically): ev_pct, p_win, fair_odds, takeable, value. Honest handling: add
the columns, populate going forward, and label the start date. "Tracking began
2026-07-XX" is the honest answer — do not backfill EV by recomputing it from today's
model against yesterday's lines. That would be a fabricated record of a model that
didn't exist yet.
getModelAggregate currently reads only outcome, clv_result, clv, player_key, grade, model_value (ledgerService.js:454) — it never reads locked_odds, so units/ROI/
record-by-odds-band are all new aggregate work on data that already exists.
5. C-clv / S-a — CLV IS CONFIRMED BROKEN IN THE DATA
The C4 write-up is right, and the ledger proves it:
| Sport | rows | closing_odds present | closing_line == locked line | clv = 0 | clv ≠ 0 |
|---|---|---|---|---|---|
| mlb | 380 | 376 | 359 | 285 | 21 |
| wnba | 196 | 196 | 164 | 138 | 26 |
captureClosing re-records the lock, so CLV is structurally ~0. clvCaptureReliable()
correctly suppresses beat_close_pct and clv_distribution on every public surface
(ledgerService.js:45,532-547). Keep it suppressed. The fix shares plumbing with S-a
(line-moved truth) exactly as you scoped — both need a real "fresher odds at serve time"
read, which the 24h snapshot:latest + 20-min intraday refresh can supply without new
quota.
6. U-deg — STATUS: THE projection == 0 LEAK APPEARS ALREADY CLOSED
model_value = 0 by day: 07-11 20 · 07-12 17 · 07-16 10 · 07-17 8 ·
07-18 0 · 07-19 0.
It stopped. 55 historical rows carry model_value = 0; all 55 are grade B, and 50 of
them sit past −400 (so the shipped juice floor would have refused them anyway). The
projection > 0 refusal at analyzeViaEngine1.js:422 is doing its job.
Two things still true and worth carrying into G-a:
getModelAggregatealready excludes them via.gt('model_value', 0)(ledgerService.js:464) — public accuracy was never polluted by these rows.- Folding the
projection > 0check into G-a's gate (as you asked) is still correct — it's currently a separate refusal after feature computation, and moving it into the gate chokepoint makes the refusal reasons uniform. It is a tidy-up, not a leak fix.
The edge_pct scale problem is NOT resolved and is untouched by this report.
7. WHAT I RECOMMEND HAPPENS NEXT
- Kev rules on §2.4 (dial table) — specifically flex-band enforcement on/off.
- G-a ships with the agreed dials +
gate_*refusal reasons + no-odds refusal + the foldedprojection > 0check. Low risk: ~11 props in 6 days of incremental cuts. - C-led first, not later — add
ev_pct/p_win/fair_oddscolumns and start recording. Nothing else in this train can be validated until EV is on disk. This is the cheapest highest-leverage item on the board. - Escalate §2.2 (only B and C grades ever emitted) to its own investigation. It undermines hero v2, "A-RATED" copy, and any EV ranking within a flat grade pool.
- Probe PropLine for alternate markets before scheduling L-b.
HANDOFF — Design (Session 2): nothing in this report is renderable yet. When the gate
ships, the refusal reasons gate_too_juiced / gate_thin_edge / gate_longshot /
no_odds will each carry user copy in the payload (D-ref), and §1.3 says the empty-board
state (G-c) will be rare but real on a thin MLB night — it needs a designed empty state,
not a blank grid.
Report generated 2026-07-19 against live prod data. No engine code changed. G-a held for ruling.