The causally-correct defence atom proves where the crude one did not

Kev's insight holds, and the data says so cleanly. defense_by_direction PROVES
on hits -- n=528, Brier -0.0034, interval [-0.0059, -0.0009] at the 99.9% level
the cumulative correction now demands -- while team-average defence remains not
proven, its interval still spanning zero. Same signal, same rows, different
unit.

The detail worth keeping is that the causally-correct atom moves the number
LESS THAN HALF as much as the crude one, 0.013 against 0.030, and is the one
that is reliably right. The team average was moving more and knowing less. Big
movement is not evidence of a good factor; it is frequently the tell.

Both halves turned out to be free, as the order expected. Savant's batted-ball
leaderboard carries pull/straight/oppo crossed with ground/air for 609 hitters
-- the statcast leaderboard we already pull does not, it has nineteen columns
and no direction at all -- and the OAA feed already carries each fielder's
position, so per-position defence is a regrouping of last week's ingest rather
than a new source. Verified in production: 609 spray profiles, 31 teams.

Handedness is what joins them and getting it backwards would have been
invisible. Pull for a right-handed hitter is the left side; for a left-handed
hitter it is the right side. A model that ignored `bats` would send half the
league's grounders to the wrong infielders and still look like it was reading
defence, and nothing downstream would have caught it. Switch hitters bat
opposite the pitcher, which this does not resolve, so they are unreadable
rather than guessed.

Unmeasured zones are renormalised away rather than contributing a zero, since a
zero asserts an exactly-average fielder standing there, and coverage states
honestly what share of a hitter's contact we could actually read.

ATOM 2 is input-blocked rather than sample-blocked, and the distinction matters
because waiting will not fix it. The weather free-source check passes --
Open-Meteo is already wired and exposes temperature, wind speed, wind direction
and precipitation -- but those raw fields are collapsed into a single scalar
modifier and wx_forecast is empty on all 1,119 settled rows. Park DIMENSIONS
are not ingested at all; parkFactors holds coefficients, not wall heights or
fence distances. A park-and-weather-to-hit-type conversion needs both, so it is
scoped rather than half-built: retaining the raw weather fields is the cheap
half, dimensions are the missing one.

Proven factors for hits are now pitcher_contact_profile and
defense_by_direction, both pooled; every per-archetype slot remains
sample-blocked.

4,297 tests green (342 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
This commit is contained in:
Kev
2026-08-04 20:01:09 -04:00
parent 405180e791
commit 20c45cbcd1
3 changed files with 73 additions and 0 deletions
+30
View File
@@ -1689,6 +1689,36 @@ phased plan in the Session-57 conversation / BUILD-STATE Next section).
`parkFactors.STAT_BASE` maps `hits → run_base`, so there is **no hits-specific
park factor**: a park that turns outs into hits without scoring is invisible.
## Causally-correct atoms (Session 93 — non-obvious)
- **KEV'S INSIGHT IS CONFIRMED ON DATA: crude operationalizations under-prove.**
`defense_by_direction` **PROVES** (n=528, Brier 0.0034, CI [0.0059,0.0009]
at 99.9%) where team-average `defense` does NOT (CI spans zero). And the
causally-correct atom **MOVES THE NUMBER LESS THAN HALF AS MUCH** (0.013 vs
0.030) while being reliably right — the crude version was moving more and
knowing less. Bigger movement is not better; it is often the tell.
- **`src/services/model/sprayDefense.js`** = spray×trajectory × positional OAA,
joined by HANDEDNESS. Pull for a RHB is the LEFT side (3B/SS/LF); for a LHB the
RIGHT side (1B/2B/RF). Getting that backwards sends half the league's grounders
to the wrong infielders and **still looks like it's reading defence** — nothing
downstream would catch it. Switch hitters are UNREADABLE (they bat opposite the
pitcher, unresolved here), not guessed.
- **Both halves were already free.** Savant's `leaderboard/batted-ball` carries
pull/straight/oppo × ground/air (609 hitters) — the `statcast` leaderboard does
NOT (19 cols, no direction). And the OAA feed already carries each fielder's
position, so per-position defence is a REGROUPING of last week's ingest.
`team_defense.position_oaa` + `batter_spray` (dated). Prod-verified 609/31.
- **Unmeasured zones are RENORMALISED AWAY, never zero** — a zero asserts an
exactly-average fielder standing there. `coverage` states what share of a
hitter's contact we could actually read; nothing readable → null.
- **ATOM 2 (park+weather→hit-type) is INPUT-BLOCKED, not sample-blocked.**
Weather's free source check PASSES (Open-Meteo already wired via
`weatherService`, exposing temp_f/wind_mph/wind_dir/precip_mm) — but those raw
fields are collapsed into a scalar `env_weather_mod` and `wx_forecast` is
**0/1119** on settled rows. And **park DIMENSIONS are not ingested at all**
(parkFactors holds coefficients, not wall heights or fence distances). A
hit-type conversion needs both; retaining the raw weather fields is the cheap
half, dimensions are the missing one.
## Active Skills
- vyndr-voice (all user-facing output)
- prop-analysis (grading methodology)