3ba3dd28f3
PROMOTE-THE-EARNED. Nothing was promoted, because nothing earned it -- not because the bar was held high. Measured on the same bar that refuted hits-v1: own rows only, direction-aligned, paired bootstrap, promote only on a CI excluding zero. arch-v1 n=1741 delta 0.0000 CI[-0.0050,+0.0054] inconclusive contact-v1 n=1055 delta +0.0008 CI[-0.0052,+0.0069] inconclusive proj-v1.1 n=1664 delta -0.0301 CI[-0.0543,-0.0060] reliably WORSE matchup/tb-v1/hits-v1 n=0 genuinely pending (rows dated 08-02+) arch-v1 is the interesting one: it MOVED 76% of rows by 2.5 points on average and resolution is identical to the champion to four decimals, on the moved rows too. That is active movement carrying no information -- a finding, not a pending verdict. These are true prospective holdouts: arch-v1 and contact-v1 wrote p_win at grade time into their own columns before the game. Nothing recomputed. THE 429, read-only. The premise was that we re-pull the full picture every slot and blow the quota. Measured: PropLine is at 5 calls of 3,000/day -- 0.17%. One snapshot is ONE PropLine call per sport, all markets comma-joined. There is no request-pattern problem, so a change-based pull cannot fix it and no tier upgrade is needed. The 429 is odds-api: 478/500 MONTHLY, blocked at 95%. oddsService falls through silently when PropLine returns empty, and the backup's quota gate throws the error -- so an empty slate is indistinguishable from an outage and the message names the wrong provider. Flagged for its own order. Could NOT verify PropLine movement endpoints: docs are auth-gated and the keys are production-only. Not asserted either way. The movement-as-data argument stands on its own merits and should be justified that way, not as a quota fix it isn't. Book-breadth invariant written down: we never discard books. All are kept and shown (DISPLAY_BOOKS = MODEL + REFERENCE + DFS); DFS pick'em is excluded from PRICING only, because a fixed-payout shaded number is not a market price. Verified this is already what bookRoles.js does. Champion byte-identical; every challenger stays wired. 4,159 tests green (332 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
110 lines
6.1 KiB
Markdown
110 lines
6.1 KiB
Markdown
# CHALLENGER SCOREBOARD — 2026-08-03
|
||
|
||
> **Nothing was promoted. Nothing earned it yet.** Not because the bar was held
|
||
> too high, but because no challenger's confidence interval excludes zero on the
|
||
> good side. The champion is byte-identical; every challenger stays wired.
|
||
|
||
The headline of this session is not the scoreboard. It is that **the scoreboard
|
||
was unmeasurable until a two-day-old settlement outage was found and fixed** —
|
||
see §1. Settled sample went **493 → 1,741** the moment it was repaired.
|
||
|
||
---
|
||
|
||
## 1. Why "n-blocked" was the wrong diagnosis
|
||
|
||
The order said: don't repeat "n-blocked" without counting. Counting is what found
|
||
the real problem.
|
||
|
||
Three of the four axes read **exactly zero** settled rows — not low, *zero*:
|
||
environment 1,496 rows / 0 settled, opportunity 922 / 0, matchup 603 / 0. Rows
|
||
whose games had been **played days earlier** and never settled, with
|
||
`settle_attempts = 0` — never even attempted.
|
||
|
||
**Root cause:** `settleLedger` fetched open ids, then refetched full rows via
|
||
`.in('id', ids)`. PostgREST puts filters in the URL, so 500 UUIDs became an
|
||
**18,499-character request** that the fetch layer rejects with `TypeError: fetch
|
||
failed`. The result was destructured as `const { data: rows } = ...` with **no
|
||
error binding**, so `rows` came back null, the loop never ran, and the function
|
||
returned `{settled:0, voided:0, unrecoverable:0, pending:0}` — byte-identical to
|
||
a healthy "nothing to settle."
|
||
|
||
It hid for two days because it is **volume-triggered**: daily volume ran 20–260
|
||
rows and settled perfectly for weeks. **2026-08-01 was the first day past the
|
||
500-row fetch limit** and settlement died that night. Worse, the zero-settle ops
|
||
alarm reads these same return values, so `pending: 0` told the watchdog the
|
||
backlog was empty — *the alarm built to catch exactly this could not see it.*
|
||
|
||
Fixed, deployed, and drained: **1,444 rows from 2026-08-01 settled (1,376
|
||
hit/miss + 68 void, 0 remaining).** `captureClosing` carried the same shape one
|
||
level down and is now chunked at 100 ids.
|
||
|
||
## 2. THE SCOREBOARD
|
||
|
||
Bar: the challenger's **own rows only**, direction-aligned, **paired bootstrap**
|
||
(4,000 resamples, deterministic seed) on the difference in resolution, because
|
||
both models score the same rows and independent standard errors would overstate
|
||
certainty. **PROMOTE requires the CI to exclude zero on the good side.** Same bar
|
||
that refuted hits-v1 — no lighter test for a would-be winner.
|
||
|
||
| challenger | settled n | resolution (chal / champ) | Δ vs champion | CI95 | verdict |
|
||
|---|---|---|---|---|---|
|
||
| arch-v1 (market-relative nudge) | **1,741** | 0.4599 / 0.4599 | 0.0000 | [−0.0050, +0.0054] | **STAY WIRED** (inconclusive) |
|
||
| arch-v1 · rows it MOVED only | 1,325 | 0.4886 / 0.4886 | −0.0001 | [−0.0064, +0.0061] | **STAY WIRED** (inconclusive) |
|
||
| contact-v1 (season contact quality) | **1,055** | 0.4070 / 0.4061 | +0.0008 | [−0.0052, +0.0069] | **STAY WIRED** (inconclusive) |
|
||
| contact-v1 · rows it MOVED only | 511 | 0.3966 / 0.3948 | +0.0018 | [−0.0107, +0.0144] | **STAY WIRED** (inconclusive) |
|
||
| proj-v1.1 ladder (all stats) | **1,664** | 0.4339 / 0.4640 | **−0.0301** | **[−0.0543, −0.0060]** | **STAY WIRED** (measured WORSE) |
|
||
| arch-v1 · environment axis rows | 871 | 0.5651 / 0.5679 | −0.0028 | [−0.0089, +0.0036] | **STAY WIRED** (inconclusive) |
|
||
| arch-v1 · opportunity axis rows | 539 | 0.5311 / 0.5310 | +0.0001 | [−0.0091, +0.0090] | **STAY WIRED** (inconclusive) |
|
||
| arch-v1 · matchup axis rows | **0** | — | — | — | **STILL PENDING** |
|
||
| tb-v1 (total_bases only) | **0** | — | — | — | **STILL PENDING** |
|
||
| hits-v1 (hits only) | **0** | — | — | — | **STILL PENDING** (refuted by replay, `specs/hits-v1-binomial.md`) |
|
||
|
||
**These are true prospective holdouts, not backtests.** arch-v1 and contact-v1
|
||
wrote their probability at grade time, into their own columns, before the game
|
||
was played. Nothing was recomputed. That is the strongest evidence available and
|
||
it is why no replay was needed here.
|
||
|
||
## 3. What the numbers actually say
|
||
|
||
**arch-v1 moves a lot and changes nothing.** It moved **1,325 of 1,741 rows
|
||
(76%)**, mean absolute move **2.5 points**, max 10.9 — and resolution is
|
||
identical to the champion to four decimal places, on the moved rows too. This is
|
||
not "too small to detect." It is movement that carries **no information about the
|
||
outcome**. A nudge this active with an effect this precisely zero is a finding,
|
||
not a pending verdict.
|
||
|
||
**The projection ladder is reliably worse than the champion.** −0.0301 with a CI
|
||
excluding zero, across 1,664 rows and all stats. Combined with hits-v1's refutation
|
||
(`specs/hits-v1-binomial.md`), the projection family now has two independent
|
||
measurements pointing the same way: it is not the champion's equal on any stat
|
||
measured so far. That is an argument for diagnosing its *inputs*, not for shipping
|
||
another variant of it.
|
||
|
||
**Three are genuinely pending, for a legitimate reason now.** matchup, tb-v1 and
|
||
hits-v1 all have rows written only on 2026-08-02/03, which settle after ET
|
||
midnight. matchup has ~496 rows queued, tb-v1 65, hits-v1 pending its first
|
||
snapshot write. They will read within a day or two — and now that settlement
|
||
works, they actually will.
|
||
|
||
## 4. Provenance
|
||
|
||
All 1,741 arch-v1 rows carry a single `model_version` (`engine1@2026-07-20`). The
|
||
older `pre-retention-unknown` rows (320 settled) carry no challenger values at
|
||
all, so they cannot influence any verdict. **No verdict here depends on
|
||
mixed-provenance rows** — the split was checked, not assumed.
|
||
|
||
Contamination excluded throughout: `quarantine_reason LIKE 'nontakeable_book%'`.
|
||
|
||
## 5. Promotion mechanics — specified, deliberately unused
|
||
|
||
No flip was performed because nothing qualified. When one does, the shape is:
|
||
challenger-first (write the promoted value into the served path while the
|
||
champion column keeps recording), version-tagged, atomic, with the previous
|
||
version one env flag away. Recorded here so a future promotion is a decision,
|
||
not an improvisation.
|
||
|
||
## 6. Reproduce
|
||
|
||
`SUPABASE_URL=... node scripts/challenger-scoreboard.js` — prints the full board,
|
||
the moved-rows-only slice, the per-axis slice and the provenance split.
|