f897c7ec06
The prod-write fingerprint that was blocked by the odds outage has landed on the first snapshot after deploy. hits-v1 records exactly as the live-board verification predicted, and the takeable axis behaves as specified -- scope is book identity, never price shape. The verdict is unchanged: hits-v1 is REFUTED and stays unpromoted. This confirms only that it is recording, so the forward accrual can judge the backtest. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
361 lines
17 KiB
Markdown
361 lines
17 KiB
Markdown
# hits-v1 — HITS AS A BINOMIAL OVER AT-BATS
|
||
|
||
> **VERDICT: REFUTED. DO NOT PROMOTE.** The hypothesis was tested honestly and
|
||
> lost. §7 has the numbers, and §8 records where the pre-registration in §6 was
|
||
> itself wrong. hits-v1 stays wired as a challenger only, writing its own ledger
|
||
> columns, so the forward accrual can confirm or overturn the backtest.
|
||
|
||
**Status:** built, firing (99.4% on the live board), holdout PRE-REGISTERED
|
||
(§1–§6 were written *before* the holdout was run) and RESOLVED (§7–§8).
|
||
**Scope:** `stat = 'hits'`, MLB, challenger only. Champion, current ladder,
|
||
ranking, calibration, the reference ruler and the four accruing challenger
|
||
verdicts are untouched and byte-identical.
|
||
|
||
---
|
||
|
||
## 1. The diagnosis this challenger is built on
|
||
|
||
Measured 2026-08-02 on the public ledger, hits rows only, direction-aligned,
|
||
contamination excluded:
|
||
|
||
```
|
||
n = 245 matched settled rows base rate 0.588
|
||
resolution champion 0.2044
|
||
current ladder (proj-v1.1) 0.0595
|
||
```
|
||
|
||
The current ladder is not merely worse than the champion on hits — it is close to
|
||
uninformative. And the line distribution says why it matters so much:
|
||
|
||
```
|
||
line 0.5 842 of 1,006 hits rows (84%)
|
||
line 1.5 144
|
||
line 2.5 20
|
||
```
|
||
|
||
**84% of hits props are a single question: does he get one?** That makes the
|
||
whole stat ride on `P(0 hits)` — precisely where the choice of count family does
|
||
the most damage, because P(0) is the one point of a count distribution that is
|
||
most sensitive to its assumed shape.
|
||
|
||
**This is the HYPOTHESIS, and it is on trial:** hits is being modelled by the
|
||
wrong family. It is not a low-rate event count. It is a *bounded conversion*.
|
||
|
||
## 2. The structure hits actually has
|
||
|
||
A hitter gets `N` official at-bats and converts each at some rate `q`:
|
||
|
||
```
|
||
N ~ the player's own EMPIRICAL at-bat distribution
|
||
hits | N ~ Binomial(N, q)
|
||
P(hits >= k) = Σ_n P(N=n) · P(Bin(n,q) >= k)
|
||
```
|
||
|
||
At the traded line of 0.5 this reduces to the axis that matters, stated directly
|
||
rather than inferred:
|
||
|
||
```
|
||
P(>=1 hit) = 1 − E_N[(1 − q)^N]
|
||
```
|
||
|
||
A negative binomial has **unbounded support and no notion of opportunity at
|
||
all**. It must infer from the count alone what are actually two separate things:
|
||
how many chances a hitter got, and how good he is. Two hitters with the same
|
||
mean hits and very different at-bat profiles are the same object to it.
|
||
|
||
This is the same doctrine as tb-v1, one level deeper: model the stat by its
|
||
generative structure, not by a family that fits its name. Total bases is a
|
||
weighted **sum**; hits is a bounded **conversion**. Different wrongness, different
|
||
fix.
|
||
|
||
### Two deliberate modelling choices
|
||
|
||
1. **The at-bat distribution is EMPIRICAL, not a fitted family.** At-bats per
|
||
game are driven by lineup slot and how the game goes, and are strongly
|
||
*under*-dispersed (3, 4 or 5; almost never 0 or 9). Poisson would overstate
|
||
that spread badly. His own log *is* the distribution.
|
||
|
||
2. **Zero-at-bat games are KEPT** — a conditioning argument. A statsapi game log
|
||
contains only games the player appeared in, so it is already conditioned on
|
||
appearing, which matches the settled population exactly: a prop on a player
|
||
who never appears produces no game-log row and never settles. An appearance
|
||
with 0 official at-bats (walked twice, pinch-ran) is a real outcome that
|
||
settles as 0 hits, and carrying it is a structural advantage over a count
|
||
model, which must infer that mass.
|
||
|
||
**The multiplier moves `q`, not `N`.** Park × weather × platoon × matchup are all
|
||
effects on whether a batted ball becomes a hit; none of them changes how often a
|
||
hitter comes to the plate. Scaling the count mean, as the negative-binomial path
|
||
does, silently mixes opportunity with conversion.
|
||
|
||
## 3. Inputs — confirmed present before the model was built (STEP 0)
|
||
|
||
`scripts/hits-input-coverage.js`, run against the 30 real players carrying the
|
||
most hits props in the public ledger, pulling **real statsapi game logs**:
|
||
|
||
```
|
||
probed 30 · resolved 30 · usable combined inputs 30 · coverage 100%
|
||
mean at-bats/game 3.518
|
||
mean per-AB hit rate 0.248 range 0.131 – 0.312
|
||
```
|
||
|
||
Both binomial inputs exist at full coverage on real rows. **Unknown is not
|
||
zero:** every read goes through `knownRate`. A game-log row with no `atBats`
|
||
field is *dropped*, never counted as a 0-at-bat game — reading it as a measured
|
||
zero would assert "no opportunity", the strongest statement available, out of an
|
||
absence of data. That defect has shipped seven times in this codebase; the guard
|
||
is unit-tested here so it does not ship an eighth.
|
||
|
||
## 4. Which takeable axis this reads, and why
|
||
|
||
Three questions once shared one word (`src/config/takeability.js`). hits-v1 reads
|
||
two of them, for two different purposes, and never the deprecated `takeable`
|
||
mirror:
|
||
|
||
| axis | question | used for |
|
||
|---|---|---|
|
||
| `isTakeableMarket(book)` | can a bettor place this? **book identity** | what is a real market to **MODEL** |
|
||
| `isWithinPromotionBand(odds)` | is it worth crowning? **price band** | recorded only; **never gates the model** |
|
||
|
||
**No price-shape rule excludes a hits prop from being modelled.** Baseball hits
|
||
markets are genuinely thin, genuinely juiced and genuinely one-sided, and all
|
||
three are normal structure rather than a bad quote. A −300 hits-over from
|
||
BetRivers is `market_takeable: true, within_promotion_band: false` — both true at
|
||
once, and it gets modelled. Each row carries `market.price_filtered: false` as a
|
||
stated invariant, and the unit suite asserts the −300 one-sided case and the
|
||
+1300 longshot case are both modelled.
|
||
|
||
## 5. Where it lives — additive, never substitutive
|
||
|
||
- `src/services/projection/binomialHits.js` — the pure model.
|
||
- `src/services/projectionChallenger.js` — emits `proj_hits_p_over` +
|
||
`proj_hits_meta` beside `proj_p_over_line`, which is **untouched**.
|
||
- `ledger_entries.proj_hits_p_over` / `proj_hits_meta` — own columns, so it joins
|
||
to the same outcome and the same close as the champion and the ladder.
|
||
- NULL on non-hits props and on underivable inputs. Never a fabricated 0.
|
||
|
||
## 6. THE HOLDOUT — pre-registered
|
||
|
||
Two measurements, kept separate because they answer different questions:
|
||
|
||
**(a) Point-in-time replay (available now, n≈245).** For each settled hits row,
|
||
rebuild the player's game log **strictly before that row's `game_date`** and
|
||
re-run hits-v1 with the row's **stored grade-time `combined_multiplier`**. No
|
||
lookahead: the log is truncated at the game being predicted, and the multiplier
|
||
is the real one recorded at lock. Hits rows only, direction-aligned to the over
|
||
basis. Reported: resolution of hits-v1 vs the current ladder vs the champion, on
|
||
matched rows, with n.
|
||
|
||
*Stated limits of the replay:* game-log reconstruction uses statsapi as it reads
|
||
today (retroactive stat corrections would not be visible); `LEAGUE_HIT_RATE`
|
||
(0.248) and `PRIOR_AB` (20) are constants chosen today, though at 20 at-bats
|
||
against a regular's 200–400 the prior moves a settled hitter by thousandths. It
|
||
is a backtest, and it is labelled one.
|
||
|
||
**(b) Forward ledger accrual.** hits-v1 writes at grade time from the next
|
||
snapshot onward and is judged by the same instrument as every other challenger.
|
||
This is the verdict of record. Hits is high-volume (1,006 rows, 405 settled), so
|
||
it accrues faster than tb-v1.
|
||
|
||
### The branch, recorded before the result is known
|
||
|
||
- **If hits resolution improves materially** → the family-mismatch mechanism
|
||
holds for hits as it did for total bases, and the per-stat structural doctrine
|
||
has a second confirmation. Promote hits-v1 for `stat='hits'` only.
|
||
- **If it does not improve** → the P(0)/family hypothesis is **wrong for hits**,
|
||
and the honest reading is that hits at a 0.5 line may be genuinely
|
||
low-resolution *for anyone*: 84% of the rows ask one near-coin-flip question at
|
||
a base rate of 0.588, and there may be little separable signal in it. That is an
|
||
**honest ceiling**, to be measured against what is achievable rather than
|
||
forced. It would NOT license rescaling, re-tuning to fit, or re-labelling.
|
||
Record it and say so.
|
||
|
||
Promotion requires a **material** improvement in hits resolution, not a
|
||
directional one. No edge or CLV term is in this gate.
|
||
|
||
---
|
||
|
||
## 7. THE RESULT (measured 2026-08-02, after §1–§6 were fixed)
|
||
|
||
### Firing — `scripts/verify-hits-v1.js`, live production snapshot
|
||
|
||
Real board (`api.vyndr.app/api/snapshot/mlb`, 428 grades, 159 hits props), run
|
||
through `attachProjection` — the exact function the snapshot calls — with the
|
||
real statsapi adapter behind it:
|
||
|
||
```
|
||
fired 158 / 159 firing coverage 99.4%
|
||
abstained 1 reason: inputs_underivable (an abstention, not a zero)
|
||
hits-v1 mean p 0.545 sd 0.124 range 0.210 – 0.728
|
||
ladder mean p 0.523 sd 0.112
|
||
```
|
||
|
||
**The takeable axis earned its keep.** Of 159 real hits props, **94 sit OUTSIDE
|
||
the promotion band** and **93 of those were modelled anyway**; `price_filtered`
|
||
is false on every row. A rule that inferred un-takeability from price shape would
|
||
have thrown away 59% of the live hits board. Every row was takeable by book
|
||
identity, which is the axis that decides what is a real market to model.
|
||
|
||
### Holdout — `scripts/hits-v1-holdout.js`, point-in-time replay
|
||
|
||
242 of 245 matched settled rows (3 dropped: thin prior log), 184/184 players
|
||
resolved, hits rows only, direction-aligned, game logs truncated strictly before
|
||
each row's game date, real grade-time multipliers.
|
||
|
||
```
|
||
RESOLUTION (corr with outcome)
|
||
overall n=242 line 0.5 n=189
|
||
champion 0.1951 0.2760
|
||
current ladder 0.0483 0.1439
|
||
hits-v1 0.0259 0.0899
|
||
base rate 0.5868 0.5979
|
||
```
|
||
|
||
Paired bootstrap on the differences (4,000 resamples, same rows, deterministic
|
||
seed — the models are scored on identical rows, so treating their errors as
|
||
independent would overstate the uncertainty):
|
||
|
||
```
|
||
hits-v1 − ladder overall −0.0224 CI95 [−0.0460, −0.0003] P(improves) 0.023
|
||
hits-v1 − ladder line 0.5 −0.0540 CI95 [−0.0933, −0.0197] P(improves) 0.001
|
||
champion − ladder overall +0.1468 CI95 [+0.0423, +0.2495] P(improves) 0.997
|
||
champion − hits-v1 overall +0.1692 CI95 [+0.0635, +0.2749] P(improves) 0.999
|
||
```
|
||
|
||
**hits-v1 does not improve hits resolution. It is slightly and reliably worse**
|
||
— the confidence interval excludes zero at the traded line, where 84% of the
|
||
rows live. The promotion gate required a material improvement. It got a small
|
||
regression. **Not promoted.**
|
||
|
||
## 8. WHAT THIS ELIMINATES — including one thing I pre-registered wrongly
|
||
|
||
**The family-mismatch hypothesis is REFUTED for hits.** Modelling hits by its
|
||
true generative structure — a bounded conversion over at-bats, with P(0) stated
|
||
directly instead of inferred from a count family — moved resolution by nothing,
|
||
in the wrong direction. The mechanism that worked for total bases does not
|
||
transfer. The per-stat structural doctrine is not universal, and this is the case
|
||
that shows its limit.
|
||
|
||
**The mean was not the binding constraint either, and that is now measurable.**
|
||
The working diagnosis was "it's the MEAN not the shape." hits-v1 corrected the
|
||
mean — at the 0.5 line the ladder averages 0.554 against a 0.598 base rate, and
|
||
hits-v1 averages 0.581, materially closer — **while resolution got worse**. So
|
||
both candidate explanations, the mean and the shape, are now eliminated. What is
|
||
left is per-prop DISCRIMINATION: the projection ladder's inputs do not separate
|
||
one hitter's night from another's, and no redistribution of the same inputs will
|
||
fix that. Both models are near-uninformative because they are reading nearly the
|
||
same thing.
|
||
|
||
### The pre-registered fallback in §6 was WRONG, and I am recording it as wrong
|
||
|
||
§6 said: if hits-v1 does not improve, "hits may be genuinely low-resolution at
|
||
0.5 — an honest ceiling for anyone." **The data refutes that fallback.** On the
|
||
identical 189 rows at the identical line, the champion resolves **0.276**, and
|
||
the paired bootstrap puts champion-over-ladder at +0.147 with a CI excluding
|
||
zero. There is real, separable signal in hits, and something in this codebase
|
||
already extracts it. Calling this an honest ceiling would have been the
|
||
comfortable reading, not the honest one.
|
||
|
||
The error was mine and it was avoidable: the champion's number was already in
|
||
hand when I wrote the fallback, and I wrote a ceiling hypothesis that its own
|
||
control refuted. A pre-registration is only worth something if it is written
|
||
against the evidence already available. Recorded here rather than quietly
|
||
dropped, because a wrong branch that gets deleted teaches nothing.
|
||
|
||
**The honest ceiling for hits is therefore not 0.06. It is at least 0.276 at the
|
||
traded line** — that is the achievable bar, measured, on a real instrument. The
|
||
projection ladder is not near it.
|
||
|
||
## 9. WHAT STAYS, AND WHAT IS UNTOUCHED
|
||
|
||
- hits-v1 **remains wired as a challenger**, writing `proj_hits_p_over` /
|
||
`proj_hits_meta` on its own columns. The replay is a backtest with stated
|
||
limits; the verdict of record is the forward accrual, and the cheapest way to
|
||
confirm a negative is to let it accrue. It is never served and never promoted.
|
||
- **Byte-identical, verified:** the diff against `projectionChallenger.js` and
|
||
`ledgerService.js` contains **zero deleted lines** — it is purely additive. The
|
||
champion (`p_win`), the current ladder (`proj_p_over_line`), the ranking, the
|
||
calibration, the reference ruler and the four accruing challenger verdicts
|
||
(arch-v1, contact-v1, proj-v1.1, tb-v1) are all untouched.
|
||
- Gates: full suite **332 suites / 4,156 tests green**; `web` build **exit 0**.
|
||
|
||
## 10. OPEN — the next real question, NOT started here
|
||
|
||
If the deficit is discrimination rather than distribution, the next order is
|
||
about INPUTS, and it should begin by asking what the champion's `p_win` reads on
|
||
a hits prop that the projection ladder does not. That is a diagnosis to run, not
|
||
a model to build — and building another projection variant before running it
|
||
would repeat exactly the mistake this document just recorded.
|
||
|
||
**Unrelated finding, logged not fixed:** the local `.env` carries a transposed
|
||
Supabase project ref (`zmdnczhtdxcddszxttub`; the real ref is
|
||
`zmdnczhtdxcddsxzttub`), so any local script reading it fails DNS. Production is
|
||
unaffected — the deployed API and the MCP connection both use the correct ref.
|
||
The holdout was run with an explicit `SUPABASE_URL` override rather than editing
|
||
your credentials file.
|
||
|
||
## 11. POST-DEPLOY FINGERPRINT — **PASSED** 2026-08-03
|
||
|
||
Confirmed on the first snapshot after deploy (game_date 2026-08-03):
|
||
|
||
```
|
||
hits rows 72 · proj_hits_p_over written 72 · 100%
|
||
mean p 0.538
|
||
takeable by identity 72/72
|
||
outside promotion band 45 — MODELLED anyway
|
||
```
|
||
|
||
hits-v1 writes in production exactly as the live-board verification predicted,
|
||
and the takeable axis behaves as specified: 45 of 72 real hits props sit outside
|
||
the promotion band and were modelled regardless, because scope is book identity
|
||
and never price shape.
|
||
|
||
The verdict is unchanged — **hits-v1 is REFUTED and stays unpromoted** (§7–§8).
|
||
This fingerprint confirms only that the challenger is recording, so the forward
|
||
accrual can confirm or overturn the backtest.
|
||
|
||
### Original PENDING note, kept for the record
|
||
|
||
Deployed 2026-08-02 ~23:05 UTC (the API 502'd and returned at 23:05:35, right
|
||
after the push). The **prod-write fingerprint has NOT landed yet**, and the
|
||
reason is upstream of this work:
|
||
|
||
```
|
||
POST /api/internal/snapshot/mlb -> {"status":"error","gradeCount":0,
|
||
"reason":"Odds data temporarily unavailable"}
|
||
GET /api/odds/mlb -> HTTP 429
|
||
GET /api/internal/snapshot/status -> overdue: true, mlb last 2026-08-02T07:30:46Z
|
||
```
|
||
|
||
**The odds provider is quota-exhausted (429), so the pipeline cannot grade a
|
||
slate at all.** The MLB board has been frozen since 07:30 UTC and the 14/19/22
|
||
UTC cron slots all failed the same way — every one of them before this change was
|
||
deployed. hits-v1 is inside the snapshot's existing try/catch and is purely
|
||
additive; it cannot be the cause, and there were zero grades for it to attach to.
|
||
|
||
If PropLine's quota resets at 00:00 UTC, the 01:00 UTC cron slot should be the
|
||
first successful run. The fingerprint to confirm hits-v1 is writing in prod:
|
||
|
||
```sql
|
||
select count(*) rows_today,
|
||
count(proj_hits_p_over) hits_v1_written,
|
||
round(avg(proj_hits_p_over)::numeric, 3) mean_p,
|
||
count(*) filter (where proj_hits_meta->'market'->>'market_takeable' = 'true') takeable,
|
||
count(*) filter (where proj_hits_meta->'market'->>'within_promotion_band' = 'false') outside_band
|
||
from public.ledger_entries
|
||
where sport = 'mlb' and user_id is null and stat = 'hits'
|
||
and graded_at > now() - interval '6 hours';
|
||
```
|
||
|
||
Expected on a healthy run, from the live-board verification in §7: `hits_v1_written`
|
||
at roughly 99% of `rows_today`, and a substantial `outside_band` count that was
|
||
modelled anyway.
|
||
|
||
**Firing itself is already verified** — §7 ran the production `attachProjection`
|
||
over the real production snapshot payload and got 158/159. What is pending is
|
||
only the confirmation that the deployed process writes the columns, which needs a
|
||
slate the pipeline can actually fetch.
|
||
|
||
**Flagged for its own order, out of scope here:** the odds quota exhaustion is a
|
||
live production outage of the whole grading pipeline, not a hits problem.
|