Files
vyndr/specs/hits-v1-binomial.md
T
builtbykev 2394fb04a1 Record the hits-v1 fingerprint as PENDING, and why
The prod-write fingerprint did not land: the odds provider is returning 429
(quota exhausted), so the snapshot refuses with gradeCount 0 and the MLB board
has been frozen since 07:30 UTC. The 14/19/22 UTC cron slots failed the same
way, all before this change deployed -- hits-v1 sits inside the snapshot's
existing try/catch, is purely additive, and had zero grades to attach to.

Firing is already verified against the real production snapshot through the
real attachProjection path (158/159). What is pending is only confirmation
that the deployed process writes the columns, which needs a slate the pipeline
can fetch. The exact fingerprint query is recorded in the spec.

The odds quota exhaustion is a live outage of the whole grading pipeline and
is flagged for its own order, not folded into this one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-02 19:07:58 -04:00

16 KiB
Raw Blame History

hits-v1 — HITS AS A BINOMIAL OVER AT-BATS

VERDICT: REFUTED. DO NOT PROMOTE. The hypothesis was tested honestly and lost. §7 has the numbers, and §8 records where the pre-registration in §6 was itself wrong. hits-v1 stays wired as a challenger only, writing its own ledger columns, so the forward accrual can confirm or overturn the backtest.

Status: built, firing (99.4% on the live board), holdout PRE-REGISTERED (§1–§6 were written before the holdout was run) and RESOLVED (§7–§8). Scope: stat = 'hits', MLB, challenger only. Champion, current ladder, ranking, calibration, the reference ruler and the four accruing challenger verdicts are untouched and byte-identical.


1. The diagnosis this challenger is built on

Measured 2026-08-02 on the public ledger, hits rows only, direction-aligned, contamination excluded:

n = 245 matched settled rows       base rate 0.588
resolution   champion  0.2044
             current ladder (proj-v1.1)  0.0595

The current ladder is not merely worse than the champion on hits — it is close to uninformative. And the line distribution says why it matters so much:

line 0.5   842 of 1,006 hits rows   (84%)
line 1.5   144
line 2.5    20

84% of hits props are a single question: does he get one? That makes the whole stat ride on P(0 hits) — precisely where the choice of count family does the most damage, because P(0) is the one point of a count distribution that is most sensitive to its assumed shape.

This is the HYPOTHESIS, and it is on trial: hits is being modelled by the wrong family. It is not a low-rate event count. It is a bounded conversion.

2. The structure hits actually has

A hitter gets N official at-bats and converts each at some rate q:

N        ~  the player's own EMPIRICAL at-bat distribution
hits | N ~  Binomial(N, q)
P(hits >= k) = Σ_n P(N=n) · P(Bin(n,q) >= k)

At the traded line of 0.5 this reduces to the axis that matters, stated directly rather than inferred:

P(>=1 hit) = 1  E_N[(1  q)^N]

A negative binomial has unbounded support and no notion of opportunity at all. It must infer from the count alone what are actually two separate things: how many chances a hitter got, and how good he is. Two hitters with the same mean hits and very different at-bat profiles are the same object to it.

This is the same doctrine as tb-v1, one level deeper: model the stat by its generative structure, not by a family that fits its name. Total bases is a weighted sum; hits is a bounded conversion. Different wrongness, different fix.

Two deliberate modelling choices

  1. The at-bat distribution is EMPIRICAL, not a fitted family. At-bats per game are driven by lineup slot and how the game goes, and are strongly under-dispersed (3, 4 or 5; almost never 0 or 9). Poisson would overstate that spread badly. His own log is the distribution.

  2. Zero-at-bat games are KEPT — a conditioning argument. A statsapi game log contains only games the player appeared in, so it is already conditioned on appearing, which matches the settled population exactly: a prop on a player who never appears produces no game-log row and never settles. An appearance with 0 official at-bats (walked twice, pinch-ran) is a real outcome that settles as 0 hits, and carrying it is a structural advantage over a count model, which must infer that mass.

The multiplier moves q, not N. Park × weather × platoon × matchup are all effects on whether a batted ball becomes a hit; none of them changes how often a hitter comes to the plate. Scaling the count mean, as the negative-binomial path does, silently mixes opportunity with conversion.

3. Inputs — confirmed present before the model was built (STEP 0)

scripts/hits-input-coverage.js, run against the 30 real players carrying the most hits props in the public ledger, pulling real statsapi game logs:

probed 30 · resolved 30 · usable combined inputs 30 · coverage 100%
mean at-bats/game 3.518
mean per-AB hit rate 0.248   range 0.131  0.312

Both binomial inputs exist at full coverage on real rows. Unknown is not zero: every read goes through knownRate. A game-log row with no atBats field is dropped, never counted as a 0-at-bat game — reading it as a measured zero would assert "no opportunity", the strongest statement available, out of an absence of data. That defect has shipped seven times in this codebase; the guard is unit-tested here so it does not ship an eighth.

4. Which takeable axis this reads, and why

Three questions once shared one word (src/config/takeability.js). hits-v1 reads two of them, for two different purposes, and never the deprecated takeable mirror:

axis question used for
isTakeableMarket(book) can a bettor place this? book identity what is a real market to MODEL
isWithinPromotionBand(odds) is it worth crowning? price band recorded only; never gates the model

No price-shape rule excludes a hits prop from being modelled. Baseball hits markets are genuinely thin, genuinely juiced and genuinely one-sided, and all three are normal structure rather than a bad quote. A 300 hits-over from BetRivers is market_takeable: true, within_promotion_band: false — both true at once, and it gets modelled. Each row carries market.price_filtered: false as a stated invariant, and the unit suite asserts the 300 one-sided case and the +1300 longshot case are both modelled.

5. Where it lives — additive, never substitutive

  • src/services/projection/binomialHits.js — the pure model.
  • src/services/projectionChallenger.js — emits proj_hits_p_over + proj_hits_meta beside proj_p_over_line, which is untouched.
  • ledger_entries.proj_hits_p_over / proj_hits_meta — own columns, so it joins to the same outcome and the same close as the champion and the ladder.
  • NULL on non-hits props and on underivable inputs. Never a fabricated 0.

6. THE HOLDOUT — pre-registered

Two measurements, kept separate because they answer different questions:

(a) Point-in-time replay (available now, n≈245). For each settled hits row, rebuild the player's game log strictly before that row's game_date and re-run hits-v1 with the row's stored grade-time combined_multiplier. No lookahead: the log is truncated at the game being predicted, and the multiplier is the real one recorded at lock. Hits rows only, direction-aligned to the over basis. Reported: resolution of hits-v1 vs the current ladder vs the champion, on matched rows, with n.

Stated limits of the replay: game-log reconstruction uses statsapi as it reads today (retroactive stat corrections would not be visible); LEAGUE_HIT_RATE (0.248) and PRIOR_AB (20) are constants chosen today, though at 20 at-bats against a regular's 200400 the prior moves a settled hitter by thousandths. It is a backtest, and it is labelled one.

(b) Forward ledger accrual. hits-v1 writes at grade time from the next snapshot onward and is judged by the same instrument as every other challenger. This is the verdict of record. Hits is high-volume (1,006 rows, 405 settled), so it accrues faster than tb-v1.

The branch, recorded before the result is known

  • If hits resolution improves materially → the family-mismatch mechanism holds for hits as it did for total bases, and the per-stat structural doctrine has a second confirmation. Promote hits-v1 for stat='hits' only.
  • If it does not improve → the P(0)/family hypothesis is wrong for hits, and the honest reading is that hits at a 0.5 line may be genuinely low-resolution for anyone: 84% of the rows ask one near-coin-flip question at a base rate of 0.588, and there may be little separable signal in it. That is an honest ceiling, to be measured against what is achievable rather than forced. It would NOT license rescaling, re-tuning to fit, or re-labelling. Record it and say so.

Promotion requires a material improvement in hits resolution, not a directional one. No edge or CLV term is in this gate.


7. THE RESULT (measured 2026-08-02, after §1–§6 were fixed)

Firing — scripts/verify-hits-v1.js, live production snapshot

Real board (api.vyndr.app/api/snapshot/mlb, 428 grades, 159 hits props), run through attachProjection — the exact function the snapshot calls — with the real statsapi adapter behind it:

fired 158 / 159   firing coverage 99.4%
abstained 1       reason: inputs_underivable  (an abstention, not a zero)
hits-v1  mean p 0.545  sd 0.124  range 0.210  0.728
ladder   mean p 0.523  sd 0.112

The takeable axis earned its keep. Of 159 real hits props, 94 sit OUTSIDE the promotion band and 93 of those were modelled anyway; price_filtered is false on every row. A rule that inferred un-takeability from price shape would have thrown away 59% of the live hits board. Every row was takeable by book identity, which is the axis that decides what is a real market to model.

Holdout — scripts/hits-v1-holdout.js, point-in-time replay

242 of 245 matched settled rows (3 dropped: thin prior log), 184/184 players resolved, hits rows only, direction-aligned, game logs truncated strictly before each row's game date, real grade-time multipliers.

                        RESOLUTION (corr with outcome)
                   overall n=242      line 0.5 n=189
champion               0.1951             0.2760
current ladder         0.0483             0.1439
hits-v1                0.0259             0.0899
                        base rate 0.5868   0.5979

Paired bootstrap on the differences (4,000 resamples, same rows, deterministic seed — the models are scored on identical rows, so treating their errors as independent would overstate the uncertainty):

hits-v1  ladder      overall  0.0224  CI95 [0.0460, 0.0003]  P(improves) 0.023
hits-v1  ladder      line 0.5 0.0540  CI95 [0.0933, 0.0197]  P(improves) 0.001
champion  ladder     overall  +0.1468  CI95 [+0.0423, +0.2495]  P(improves) 0.997
champion  hits-v1    overall  +0.1692  CI95 [+0.0635, +0.2749]  P(improves) 0.999

hits-v1 does not improve hits resolution. It is slightly and reliably worse — the confidence interval excludes zero at the traded line, where 84% of the rows live. The promotion gate required a material improvement. It got a small regression. Not promoted.

8. WHAT THIS ELIMINATES — including one thing I pre-registered wrongly

The family-mismatch hypothesis is REFUTED for hits. Modelling hits by its true generative structure — a bounded conversion over at-bats, with P(0) stated directly instead of inferred from a count family — moved resolution by nothing, in the wrong direction. The mechanism that worked for total bases does not transfer. The per-stat structural doctrine is not universal, and this is the case that shows its limit.

The mean was not the binding constraint either, and that is now measurable. The working diagnosis was "it's the MEAN not the shape." hits-v1 corrected the mean — at the 0.5 line the ladder averages 0.554 against a 0.598 base rate, and hits-v1 averages 0.581, materially closer — while resolution got worse. So both candidate explanations, the mean and the shape, are now eliminated. What is left is per-prop DISCRIMINATION: the projection ladder's inputs do not separate one hitter's night from another's, and no redistribution of the same inputs will fix that. Both models are near-uninformative because they are reading nearly the same thing.

The pre-registered fallback in §6 was WRONG, and I am recording it as wrong

§6 said: if hits-v1 does not improve, "hits may be genuinely low-resolution at 0.5 — an honest ceiling for anyone." The data refutes that fallback. On the identical 189 rows at the identical line, the champion resolves 0.276, and the paired bootstrap puts champion-over-ladder at +0.147 with a CI excluding zero. There is real, separable signal in hits, and something in this codebase already extracts it. Calling this an honest ceiling would have been the comfortable reading, not the honest one.

The error was mine and it was avoidable: the champion's number was already in hand when I wrote the fallback, and I wrote a ceiling hypothesis that its own control refuted. A pre-registration is only worth something if it is written against the evidence already available. Recorded here rather than quietly dropped, because a wrong branch that gets deleted teaches nothing.

The honest ceiling for hits is therefore not 0.06. It is at least 0.276 at the traded line — that is the achievable bar, measured, on a real instrument. The projection ladder is not near it.

9. WHAT STAYS, AND WHAT IS UNTOUCHED

  • hits-v1 remains wired as a challenger, writing proj_hits_p_over / proj_hits_meta on its own columns. The replay is a backtest with stated limits; the verdict of record is the forward accrual, and the cheapest way to confirm a negative is to let it accrue. It is never served and never promoted.
  • Byte-identical, verified: the diff against projectionChallenger.js and ledgerService.js contains zero deleted lines — it is purely additive. The champion (p_win), the current ladder (proj_p_over_line), the ranking, the calibration, the reference ruler and the four accruing challenger verdicts (arch-v1, contact-v1, proj-v1.1, tb-v1) are all untouched.
  • Gates: full suite 332 suites / 4,156 tests green; web build exit 0.

10. OPEN — the next real question, NOT started here

If the deficit is discrimination rather than distribution, the next order is about INPUTS, and it should begin by asking what the champion's p_win reads on a hits prop that the projection ladder does not. That is a diagnosis to run, not a model to build — and building another projection variant before running it would repeat exactly the mistake this document just recorded.

Unrelated finding, logged not fixed: the local .env carries a transposed Supabase project ref (zmdnczhtdxcddszxttub; the real ref is zmdnczhtdxcddsxzttub), so any local script reading it fails DNS. Production is unaffected — the deployed API and the MCP connection both use the correct ref. The holdout was run with an explicit SUPABASE_URL override rather than editing your credentials file.

11. POST-DEPLOY FINGERPRINT — armed, PENDING (blocked upstream, not by this change)

Deployed 2026-08-02 ~23:05 UTC (the API 502'd and returned at 23:05:35, right after the push). The prod-write fingerprint has NOT landed yet, and the reason is upstream of this work:

POST /api/internal/snapshot/mlb  ->  {"status":"error","gradeCount":0,
                                      "reason":"Odds data temporarily unavailable"}
GET  /api/odds/mlb               ->  HTTP 429
GET  /api/internal/snapshot/status -> overdue: true, mlb last 2026-08-02T07:30:46Z

The odds provider is quota-exhausted (429), so the pipeline cannot grade a slate at all. The MLB board has been frozen since 07:30 UTC and the 14/19/22 UTC cron slots all failed the same way — every one of them before this change was deployed. hits-v1 is inside the snapshot's existing try/catch and is purely additive; it cannot be the cause, and there were zero grades for it to attach to.

If PropLine's quota resets at 00:00 UTC, the 01:00 UTC cron slot should be the first successful run. The fingerprint to confirm hits-v1 is writing in prod:

select count(*)                                        rows_today,
       count(proj_hits_p_over)                         hits_v1_written,
       round(avg(proj_hits_p_over)::numeric, 3)        mean_p,
       count(*) filter (where proj_hits_meta->'market'->>'market_takeable' = 'true')  takeable,
       count(*) filter (where proj_hits_meta->'market'->>'within_promotion_band' = 'false') outside_band
from public.ledger_entries
where sport = 'mlb' and user_id is null and stat = 'hits'
  and graded_at > now() - interval '6 hours';

Expected on a healthy run, from the live-board verification in §7: hits_v1_written at roughly 99% of rows_today, and a substantial outside_band count that was modelled anyway.

Firing itself is already verified — §7 ran the production attachProjection over the real production snapshot payload and got 158/159. What is pending is only the confirmation that the deployed process writes the columns, which needs a slate the pipeline can actually fetch.

Flagged for its own order, out of scope here: the odds quota exhaustion is a live production outage of the whole grading pipeline, not a hits problem.