tb-v1 report + plan: diagnosis on trial, branch pre-registered
Firing verified on a real prod snapshot: 10/10 total_bases props carry proj_tb_p_over. The snapshot HTTP call returned 524 (Cloudflare's 100s origin timeout vs a ~115s snapshot) but the work completed server-side -- confirmed from the ledger rather than assumed. Face validity is good and diagnostic: means agree almost exactly with the ladder (1.813 vs 1.833), so this is a SHAPE-ONLY intervention, which is what was intended. Component rates are plausible, and Carroll's triples rate (0.112, far above his peers) is a clean check -- he is a speed player and the model sees it. AN OBSERVATION I AM NOT RESOLVING BY EYE: tb-v1 reads systematically LOWER than the ladder (0.424 vs 0.540 at the same mean). That is the expected DIRECTION, since the NB overstates P(>=2) by treating a home run as four accumulating events -- but whether 0.424 is right or an overcorrection is not knowable from face validity. A ~1.8-TB hitter clearing 1.5 empirically sits nearer 45-50%, between the two. I am not claiming tb-v1 is better; the holdout decides. BRANCH PRE-REGISTERED, before the result, so the verdict cannot be reinterpreted afterward: improves -> family-mismatch HOLDS, similarity stays off the critical path, hits is next; does not improve -> hypothesis WRONG and the mean-weakness/similarity branch REOPENS. Also recorded: I hit Number(null)===0 in my own new module -- a null component rate treated as a measured zero, the difference between "never triples" and "we don't know his triple rate". A test caught it. Sixth appearance of this trap in this codebase, and it caught the person writing the warnings about it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
@@ -0,0 +1,139 @@
|
||||
# tb-v1 — TOTAL BASES AS A COMPOUND OUTCOME (challenger)
|
||||
|
||||
**Date:** 2026-08-02 · current ladder + champion `p_win` **byte-identical** ·
|
||||
ledger-only · **firing verified on a real prod snapshot**.
|
||||
|
||||
**Gates:** 4,104 tests / 329 suites green · `next build` exit 0 · tb-v1 on
|
||||
**10/10 total_bases props** · holdout committed, verdict n-blocked.
|
||||
|
||||
---
|
||||
|
||||
## STEP 0 — COMPONENTS CONFIRMED ON REAL DATA
|
||||
|
||||
statsapi has **no `singles` field**. But on a real 10-game log:
|
||||
|
||||
```
|
||||
singles = hits − doubles − triples − homeRuns
|
||||
singles + 2·doubles + 3·triples + 4·homeRuns == stored totalBases ✓ exact
|
||||
```
|
||||
|
||||
**The decomposition is exact, not an approximation.** Verified before any code
|
||||
was written — this fix depended on it.
|
||||
|
||||
---
|
||||
|
||||
## THE MODEL
|
||||
|
||||
Each component gets its own per-game Poisson rate; TB is their weighted sum, and
|
||||
the PMF is built by **exact convolution** rather than simulated (TB support is
|
||||
small). It inherits the **same combined multiplier** proj-v1.1 computes, so the
|
||||
two models differ **only in structure**.
|
||||
|
||||
### Why this is the fix, in one number
|
||||
|
||||
With **identical mean TB of 1.0**:
|
||||
|
||||
| hitter | P(TB≥2) | **P(TB≥4)** |
|
||||
|---|---:|---:|
|
||||
| pure home-run hitter | 0.221 | **0.221** |
|
||||
| pure singles hitter | 0.264 | **0.019** |
|
||||
|
||||
**A 12× difference at the 4-base rung that an NB on TB alone cannot express**,
|
||||
because it treats one home run as four independent events. A test asserts this
|
||||
separation, and asserts that P(TB≥4) for a pure-HR hitter equals P(at least one
|
||||
HR) **exactly**.
|
||||
|
||||
### The approximation, stated
|
||||
|
||||
Components are modelled as **independent** Poissons. They are not: a plate
|
||||
appearance that becomes a double cannot also become a single, so they are weakly
|
||||
negatively correlated and **independence slightly overstates the tail**. Closer
|
||||
to the truth than what it replaces; **not a solved problem**, and labelled
|
||||
`independence_caveat: true` on every record.
|
||||
|
||||
---
|
||||
|
||||
## FIRING ON PROD — 10/10 TB PROPS
|
||||
|
||||
| | |
|
||||
|---|---:|
|
||||
| rows written | 241 |
|
||||
| total_bases props | 10 |
|
||||
| **with `proj_tb_p_over`** | **10 (100%)** |
|
||||
| range / sd | 0.295–0.424 / 0.045 |
|
||||
|
||||
Real records:
|
||||
|
||||
| player | line | ladder mean | **ladder P** | **tb-v1 P** | tb-v1 mean | rates (1B/2B/3B/HR) | games |
|
||||
|---|---:|---:|---:|---:|---:|---|---:|
|
||||
| Ketel Marte | 1.5 | 1.813 | **0.540** | **0.424** | 1.833 | .614/.208/.028/.180 | 106 |
|
||||
| Corbin Carroll | 1.5 | 1.814 | 0.540 | 0.424 | 1.815 | .468/.206/**.112**/.150 | 107 |
|
||||
| Gabriel Moreno | 1.5 | 1.590 | 0.471 | 0.405 | 1.608 | .773/.221/.000/.098 | 84 |
|
||||
|
||||
**The means agree almost exactly** (1.813 vs 1.833) — both use the same rate
|
||||
machinery. **Only the shape differs**, which is precisely the intended
|
||||
intervention. Carroll's triples rate (0.112, far above the others) is a good
|
||||
face-validity check: he is a speed player, and the model sees it.
|
||||
|
||||
### An observation the holdout must settle
|
||||
|
||||
**tb-v1 is systematically LOWER than the ladder** (0.424 vs 0.540 at the same
|
||||
mean). That is the expected *direction* — the NB overstates P(≥2) by treating a
|
||||
home run as four accumulating events. But **whether 0.424 is right, or an
|
||||
overcorrection, is not knowable from face validity.** A ~1.8-TB hitter clearing
|
||||
1.5 empirically sits nearer 45–50%, which is between the two.
|
||||
|
||||
**I am not claiming tb-v1 is better. The holdout decides.**
|
||||
|
||||
---
|
||||
|
||||
## HOLDOUT — COMMITTED, VERDICT n-BLOCKED
|
||||
|
||||
`scripts/tb-compound-holdout.sql`, with two guards baked in:
|
||||
|
||||
1. **TB rows ONLY.** total_bases is 49 of 437 settled rows; averaging into other
|
||||
stats would hide the effect entirely.
|
||||
2. **DIRECTION-ALIGNED.** `p_win` is P(graded side); both projection values are
|
||||
P(over). 31.4% of rows are under-graded, and the unaligned comparison is the
|
||||
artifact that accounted for **41%** of the ladder's apparent loss.
|
||||
|
||||
**Settled TB rows carrying tb-v1: 0** — it went live today, on games not yet
|
||||
played. First settle pass is tomorrow.
|
||||
|
||||
---
|
||||
|
||||
## THE DIAGNOSIS IS ON TRIAL — recorded in advance
|
||||
|
||||
> **If tb-v1 materially improves TB resolution** → the family-mismatch hypothesis
|
||||
> **HOLDS**, similarity stays **off** the critical path, and **`hits` is next**
|
||||
> (its own order: AB-bounded, mostly 0.5 lines, so it rides on P(0)).
|
||||
>
|
||||
> **If it does NOT improve** → the hypothesis is **WRONG**, and the
|
||||
> **mean-weakness / MLB-similarity branch REOPENS**.
|
||||
|
||||
Written down now, before the result, so the verdict cannot be reinterpreted after
|
||||
the fact.
|
||||
|
||||
---
|
||||
|
||||
## A TRAP I HIT IN MY OWN CODE
|
||||
|
||||
`Number(null) === 0` — again, and this time in the new module. A null component
|
||||
rate passed a naive finite check and was treated as a **measured zero**. That is
|
||||
the difference between *"this player never triples"* and *"we do not know his
|
||||
triple rate"*, and it would have silently narrowed the distribution.
|
||||
|
||||
**A test caught it.** Both `tbPmf` and `tbMean` now reject `null`/`''`/boolean
|
||||
strictly. Recording it because this is the sixth time this specific trap has
|
||||
appeared in this codebase, and it caught the person who has been writing the
|
||||
warnings about it.
|
||||
|
||||
## TAGS
|
||||
|
||||
**VERIFIED:** the TB decomposition is exact on real logs · tb-v1 fires on 10/10 TB
|
||||
props · means agree with the ladder (shape-only intervention) · component rates
|
||||
face-valid (Carroll's triples) · 12× separation at the 4-base rung.
|
||||
|
||||
**NOT CLAIMED:** that tb-v1 is better. n-blocked until TB rows settle.
|
||||
|
||||
**STATED LIMITATION:** component independence overstates the tail.
|
||||
Reference in New Issue
Block a user