Takeable audit: the ledger is contaminated, and I caused it
READ-ONLY. Nothing enforced or fixed; the five challengers untouched. VERDICT: gaps exist, and one is LIVE CONTAMINATION of the ledger -- the exact table every accruing holdout resolves against. book, locked_odds and the takeable flag ITSELF are being stamped from books you cannot bet: DFS dabble (707 rows, 24% of all rows), offshore bovada (214), onexbet (42), exchange kalshi (7, mean |odds| 1120). 0% before 2026-08-01. 47.9% on 08-01. 42.5% on 08-02. It began the day I widened the books for display. LEAK LOCATED, not inferred: recordPipelineGrades indexes byKey over the FULL display-widened props list, then prefers that prop -- book: (prop && prop.book) || g.book, and locked_odds/takeable both fall back to oddsForSide(prop). The grade is computed on a MODEL book and the ledger row is then re-stamped from whatever book indexed first. The takeable flag is therefore not merely mislabelled: it is computed FROM the contaminated price, so it is wrong on its own terms. The served grade path is clean TODAY (428 grades, 100% MODEL books), so dedupeProps' gate works. But MODEL_BOOKS is NOT a subset of TAKEABLE_BOOKS -- pinnacle is model-eligible and correctly not takeable -- so the projection may anchor to a reference line by design. Harmless while pinnacle returns nothing; live again when it recovers. BLAST RADIUS bounded but growing: 47 contaminated rows have already settled (21% of settled rows since 08-01) and ~700 are still pending and will settle into the holdouts. The damage is mostly ahead of us, which is what makes this urgent rather than historical. NOT VERIFIED and not claimed either way: whether the stored `line` is also contaminated. It traces to the graded prop, but I did not check it end-to-end; the enforcement order should. The prediction-vs-reference distinction HOLDS and must not be collapsed: the prediction target must be takeable, while fair_prob / consensus / edge stay reference. The bug is not the three-way split -- it is that one write path ignores it. Stack sequenced in the plan: (a) takeable enforcement, (b) structural Number(null)===0 guard (hits will re-trigger it -- its 0.5 lines make P(0) the whole game), (c) hits. Carry-forward: tb-v1 verdict, the third pre-registered branch, and the 100s Cloudflare timeout vs a ~115s snapshot. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
+24
-15
@@ -29,25 +29,34 @@ stay provisional until re-run** · documented ≠ verified.
|
|||||||
|
|
||||||
## ▶ NEXT EXECUTABLE ORDER
|
## ▶ NEXT EXECUTABLE ORDER
|
||||||
|
|
||||||
**AWAIT the tb-v1 verdict, then `hits`.** `specs/tb-v1-compound.md`.
|
**🔴 (a) TAKEABLE ENFORCEMENT — urgent.** `specs/takeable-enforcement-audit.md`.
|
||||||
|
|
||||||
**tb-v1 is LIVE as a challenger** — total_bases modelled as the compound outcome
|
**The ledger — the table every accruing holdout resolves against — is stamping
|
||||||
it is (per-component Poisson rates convolved exactly) instead of one negative
|
`book`, `locked_odds` and the `takeable` flag itself from books you cannot bet**
|
||||||
binomial that treats a home run as four events. Firing on **10/10 TB props**;
|
(DFS `dabble` 707 rows, offshore `bovada` 214, `onexbet` 42, exchange `kalshi` 7).
|
||||||
means agree with the ladder, **only the shape differs** — the intended
|
**0% before 2026-08-01, 47.9% on 08-01, 42.5% on 08-02** — it began with the
|
||||||
intervention.
|
display widening. `recordPipelineGrades` indexes over the FULL widened props list
|
||||||
|
and prefers that prop's book/odds over the graded one.
|
||||||
|
|
||||||
**The diagnosis is on trial, and the branch is pre-registered:**
|
**47 contaminated rows have already settled; ~700 are pending and will settle into
|
||||||
- **tb-v1 improves TB resolution** → family-mismatch **HOLDS**, similarity stays
|
the holdouts.** The damage is mostly ahead of us.
|
||||||
**off** the critical path, **`hits` is next** (AB-bounded, mostly 0.5 lines, so
|
|
||||||
it rides on P(0)).
|
|
||||||
- **tb-v1 does not improve** → hypothesis **WRONG**, the **mean-weakness /
|
|
||||||
MLB-similarity branch REOPENS**.
|
|
||||||
|
|
||||||
**Verdict n-blocked:** 0 settled TB rows carry tb-v1 yet; first settle is
|
Also latent: **`MODEL_BOOKS` ⊄ `TAKEABLE_BOOKS`** (`pinnacle` is model-eligible but
|
||||||
tomorrow. `scripts/tb-compound-holdout.sql` is TB-rows-only and direction-aligned.
|
not takeable) — harmless while pinnacle returns nothing, live again if it recovers.
|
||||||
|
|
||||||
**Then:** `archetype_x_archetype` → retire/rewrite `bayesianEngine`.
|
**Keep the split:** the *prediction target* must be takeable; `fair_prob` /
|
||||||
|
consensus / `edge` **stay reference** and must not be over-enforced.
|
||||||
|
|
||||||
|
**Then (b) a structural `Number(null) === 0` guard** — six occurrences, most
|
||||||
|
recently in my own new module. **`hits` will re-trigger it**: its 0.5 lines make
|
||||||
|
P(0) the whole game, where a null rate read as zero is maximally wrong.
|
||||||
|
|
||||||
|
**Then (c) the `hits` fix**, built on both invariants — not before them.
|
||||||
|
|
||||||
|
**Carry-forward:** tb-v1 verdict (accruing) · the third pre-registered branch
|
||||||
|
(improves-ranking-but-low → promote **and** recalibrate) · the **100s Cloudflare
|
||||||
|
origin timeout** vs a ~115s snapshot (work completes; the trigger returns 524 —
|
||||||
|
flag it before it reads as a failure).
|
||||||
|
|
||||||
### FIVE challengers accruing in parallel — do NOT re-run early
|
### FIVE challengers accruing in parallel — do NOT re-run early
|
||||||
Verified firing on a real prod snapshot (293 grades), not inferred:
|
Verified firing on a real prod snapshot (293 grades), not inferred:
|
||||||
|
|||||||
@@ -0,0 +1,147 @@
|
|||||||
|
# TAKEABLE-ENFORCEMENT AUDIT
|
||||||
|
|
||||||
|
**Date:** 2026-08-02 · **READ-ONLY** — nothing enforced, built or fixed. The five
|
||||||
|
accruing challengers were not touched.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## VERDICT: 🔴 GAPS EXIST, AND ONE IS LIVE CONTAMINATION OF THE LEDGER
|
||||||
|
|
||||||
|
> **The ledger — the exact table every accruing holdout resolves against — is
|
||||||
|
> stamping `book`, `locked_odds` and the `takeable` flag itself from books you
|
||||||
|
> cannot bet: DFS pick'em, offshore, and exchanges.**
|
||||||
|
>
|
||||||
|
> **0% before 2026-08-01. 47.9% on 08-01. 42.5% on 08-02.**
|
||||||
|
>
|
||||||
|
> **I caused this.** It started the day I widened the books for display.
|
||||||
|
|
||||||
|
**Enforcement is the next order, ahead of `hits`.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## PER-PATH ANCHOR REPORT
|
||||||
|
|
||||||
|
| path | anchors to | takeable? |
|
||||||
|
|---|---|---|
|
||||||
|
| **PROJECTION** | `g.line` via `dedupeProps` → `MODEL_BOOKS` | ⚠️ **model-gated, not takeable-gated** |
|
||||||
|
| **GRADE** | same | ⚠️ same — clean in practice today |
|
||||||
|
| **CHALLENGERS** (arch-v1, tb-v1, proj-v1.1) | read off the grade | ✅ inherit the grade's anchor |
|
||||||
|
| **LEDGER** | `prop.book` / `oddsForSide(prop)` over the **full widened props list** | 🔴 **CONTAMINATED** |
|
||||||
|
| **DISPLAY / RANKING** | `takeablePWin` → `isTakeable(price)` | ⚠️ gate is correct, **but applied to a contaminated price** |
|
||||||
|
|
||||||
|
### The grade path is clean *today*, but not by construction
|
||||||
|
|
||||||
|
Served snapshot, 428 grades: **draftkings 422, fanduel 6 — 100% MODEL books.**
|
||||||
|
`dedupeProps`' `MODEL_BOOKS` filter works.
|
||||||
|
|
||||||
|
**But `MODEL_BOOKS` ⊄ `TAKEABLE_BOOKS`.** `MODEL_BOOKS` contains **`pinnacle`**,
|
||||||
|
which `isTakeableBook` correctly excludes — it is a reference book, not a US
|
||||||
|
retail one. So the projection *may* anchor to a reference line by design.
|
||||||
|
|
||||||
|
**Currently harmless** (pinnacle has returned 0 props since 07-31) and therefore
|
||||||
|
**exactly the kind of latent gap that bites when the feed recovers.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## THE CONTAMINATION, LOCATED
|
||||||
|
|
||||||
|
`ledgerService.recordPipelineGrades`:
|
||||||
|
|
||||||
|
```js
|
||||||
|
const byKey = indexProps(oddsProps); // ← the FULL, display-widened list
|
||||||
|
const prop = byKey[`${nameKey(player)}|${stat}`]; // ← first match wins
|
||||||
|
...
|
||||||
|
locked_odds: locked.odds != null ? String(locked.odds) : oddsForSide(prop, side),
|
||||||
|
takeable: takeableFor(locked.odds != null ? locked.odds : oddsForSide(prop, side)),
|
||||||
|
book: (prop && prop.book) || g.book || null,
|
||||||
|
```
|
||||||
|
|
||||||
|
The grade is computed on a MODEL book, then the ledger row is **re-stamped** from
|
||||||
|
whatever book happened to index first in the widened list.
|
||||||
|
|
||||||
|
### Books actually stamped on ledger rows since 2026-08-01
|
||||||
|
|
||||||
|
| book | rows | role | settled |
|
||||||
|
|---|---:|---|---:|
|
||||||
|
| draftkings | 518 | ✅ takeable | 49 |
|
||||||
|
| **dabble** | **707** | 🔴 **DFS pick'em** | **44** |
|
||||||
|
| betmgm | 348 | ✅ takeable | 126 |
|
||||||
|
| hardrockbet | 298 | ✅ takeable | 0 |
|
||||||
|
| **bovada** | **214** | 🔴 reference/offshore | 3 |
|
||||||
|
| **onexbet** | **42** | 🔴 excluded | 0 |
|
||||||
|
| **pinnacle** | **36** | 🔴 reference | 0 |
|
||||||
|
| **kalshi** | **7** | 🔴 exchange — mean \|odds\| **1120** | 0 |
|
||||||
|
|
||||||
|
**`dabble` alone is 24% of all rows.** `kalshi`'s mean absolute price of **1120**
|
||||||
|
is a self-evident tell: no US bettor is taking that as the anchor of a graded prop.
|
||||||
|
|
||||||
|
### What is confirmed contaminated, and what is not
|
||||||
|
|
||||||
|
**Confirmed:** `book`, `locked_odds`, and **`takeable`** — the flag is computed
|
||||||
|
*from* the contaminated price, so it is not merely mislabelled, it is **wrong on
|
||||||
|
its own terms**.
|
||||||
|
|
||||||
|
**Not confirmed:** `line`. It comes from `locked.line`/`g.line`, which trace to the
|
||||||
|
graded (model-book) prop. **I did not verify it end-to-end and am not claiming it
|
||||||
|
is clean.** The enforcement order should check it.
|
||||||
|
|
||||||
|
### Blast radius on the accruing verdicts — bounded, but growing
|
||||||
|
|
||||||
|
| | settled rows since 08-01 |
|
||||||
|
|---|---:|
|
||||||
|
| clean-book | 178 |
|
||||||
|
| **contaminated-book** | **47 (21%)** |
|
||||||
|
|
||||||
|
**Most contaminated rows have not settled yet** — 707 `dabble` rows are pending and
|
||||||
|
**will settle into the holdouts tonight and tomorrow.** The damage is mostly
|
||||||
|
*ahead* of us, which is what makes this urgent rather than merely historical.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## PREDICTION-TARGET vs REFERENCE-RULER — the distinction holds
|
||||||
|
|
||||||
|
**Do NOT over-enforce takeable onto the reference ruler.** They answer different
|
||||||
|
questions and collapsing them would break market measurement:
|
||||||
|
|
||||||
|
| quantity | must be | current | correct? |
|
||||||
|
|---|---|---|---|
|
||||||
|
| the line the model **projects against** | **TAKEABLE** | model-gated | ⚠️ gap |
|
||||||
|
| the price the row is **scored on** (`locked_odds`) | **TAKEABLE** | contaminated | 🔴 gap |
|
||||||
|
| `fair_prob` / consensus / `edge` diagnostics | **REFERENCE** (incl. exchanges, pinnacle) | reference | ✅ **correct — leave alone** |
|
||||||
|
|
||||||
|
`bookRoles`' three-way split is right and must not be collapsed. **The bug is not
|
||||||
|
the split — it is that one write path ignores it.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## SEQUENCED STACK
|
||||||
|
|
||||||
|
**(a) 🔴 TAKEABLE ENFORCEMENT — next order.** Index the ledger's prop lookup over
|
||||||
|
**takeable books only**, verify `line` end-to-end, decide whether `pinnacle` may
|
||||||
|
anchor a projection (it may price the *ruler*, but should it price the *target*?),
|
||||||
|
and decide what to do with the ~1,031 already-contaminated rows: **re-stamp, or
|
||||||
|
tag and exclude from holdouts.** Do not silently pool them.
|
||||||
|
|
||||||
|
**(b) Structural `Number(null) === 0` guard.** Six occurrences, the most recent in
|
||||||
|
my own new module yesterday. A shared unknown-vs-zero helper. **`hits` will
|
||||||
|
re-trigger it** — its 0.5 lines make P(0) the whole game, and a null rate read as
|
||||||
|
zero is maximally wrong exactly there.
|
||||||
|
|
||||||
|
**(c) `hits` fix** — built on both invariants above, not before them.
|
||||||
|
|
||||||
|
**Carry-forward:** the tb-v1 verdict (accruing, slower clock) · the third
|
||||||
|
pre-registered branch (*improves ranking but stays low → promote **and**
|
||||||
|
recalibrate*) · the **100s Cloudflare origin timeout** on a ~115s snapshot (the
|
||||||
|
work completes server-side, but the trigger returns 524 — worth a flag before it
|
||||||
|
is mistaken for a failure).
|
||||||
|
|
||||||
|
## TAGS
|
||||||
|
|
||||||
|
**VERIFIED:** ledger `book`/`locked_odds`/`takeable` contaminated on 42–48% of rows
|
||||||
|
since 2026-08-01, 0% before · leak located at `recordPipelineGrades`' index over
|
||||||
|
the widened props list · served grade path clean today (100% MODEL books) ·
|
||||||
|
`MODEL_BOOKS` ⊄ `TAKEABLE_BOOKS` via `pinnacle` · 47 contaminated rows already
|
||||||
|
settled, ~700 pending.
|
||||||
|
|
||||||
|
**NOT VERIFIED:** whether the stored `line` is also contaminated. Check it in the
|
||||||
|
enforcement order rather than assuming either way.
|
||||||
Reference in New Issue
Block a user