Compare commits

...

35 Commits

Author SHA1 Message Date
builtbykev 387ae4d54e Handoff doc: ground-truth build state at 71d3b7b
Written from the repo, not summary -- every claim grep- or run-verified,
UNKNOWN where the repo cannot say.

State: HEAD and gitea main in sync at 71d3b7b, tree clean, 4,539 tests
passing, web build exit 0. Push remote is gitea, never origin.

DONE this arc, existence-verified: E1 movement strip, F9-F11 offseason hub
shell, E10 Report template, E12 /report archive, content engine + studio
API + preview, Wave-D1 motion primitives, honest served grade, archetype
doctrine. Also verified done and previously mis-boarded: D1 glyphs (39 of
83 -- every one that maps to a real archetype), A1 card token, B1 boundary
channel, book comparison, THE WIRE.

GATED with each gate named and each absence verified by 0-file grep: F5
article media and E16/F8 on the card-system reconciliation, in-season hub
IA on the content formula, E9/E15 on model accrual, E2/E6 on licensing,
E3 crown on measurement, E13 on another order. NO UNGATED WAVE-2 TARGETS
REMAIN -- the next move is a decision, not a build.

The two open decisions are stated with what each unblocks. Card-system
reconciliation is the cheapest: it frees F5 and E16/F8 together, and it
exists because I built a 1080x1350 renderer without checking whether a
designed card system existed. It did.

Accrual clock recorded at 0 eligible dates on all four items, with the
first trigger (10 calibration dates) and the attempt-floor-not-trust-floor
caveat carried forward.

Standing doctrine carried: Truth Law including designer samples and the
inverse Number(null) breach (zero is a real fact); the 83-glyph taxonomy
with 44 dormant slots; the repaired champion and the standing flag that
every prior factor verdict was measured against the broken baseline;
calibration withdrawn; unissuable A with a B+ 0.663 ceiling; both CI
guards; and compose-don't-fork, which this arc failed three times and
caught twice.

Both parallel workstreams written as seed prompts.

specs/HANDOFF-2026-08-08.md

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-08 00:10:34 -04:00
builtbykev 71d3b7b786 E10 Report issue template + E12 /report archive, to spec
PHASE 0 — the spec, read not recalled. E10: "Hybrid: dark billboard header
that survives every client, light paper body Gmail can't wreck. 600px,
stacked, no webfont dependence." Content law: "One email per slate day.
Top read, what changed, the record. Nothing else." E12: "EVERY ISSUE SHOWS
ITS OWN DAY RECORD -- THE ARCHIVE IS A LEDGER TOO."

COMPOSED, NOT FORKED. The audit had E10 as PARTIAL, not absent:
newsletterService already builds the daily report's CONTENT and lints its
voice. What was missing is the designed hybrid SHELL, so reportTemplate.js
is a template over that builder rather than a second report -- the same
call made for the movement strip, and for the same reason.

PHASE 1 — the hybrid shell is an ENGINEERING constraint, not a look, and
the tests say so: Gmail strips style blocks, Outlook ignores flexbox, and
a dark body renders as a black rectangle in several clients. Hence tables,
inline styles, 600px fixed, system fonts, no image required to read, and
the green SHIFTS from #00D4A0 to #00A57D on paper because the dark-mode
green is unreadable there.

FACT-CONTRACTED: a section whose data is absent is OMITTED and NAMED in
`omitted`, never filled. There is no code path producing a placeholder
figure. The honesty block carries the real numbers -- graded count,
cleared-ceiling count, the realized rate against baseline, and that we do
not issue A grades.

E1'S LAW TRAVELS EVEN THOUGH ITS RENDERING CANNOT. An SVG strip is not
reliable in email, so movementText carries the RULE: green only when the
move favours the read, and a flat market says FLAT · [N]D rather than
showing nothing.

NO DESIGNER SAMPLE DATA. Nabers 1,120.5, No 128, DAY RECORD 9-4 are a spec
for what a live issue renders; pasting them in would be fabrication
carrying a designer's authority and would look entirely correct. Tested.

PHASE 2 — /report is now the real archive, REPLACING the S41 redirect to
/blog. That redirect existed because the surface did not; E12 built it, so
the placeholder is correctly gone and the S41 test is updated rather than
worked around. Every row carries its own day record, and an unknown record
says UNSETTLED -- never a dash that reads as zero. Empty archive is an
honest state.

Backend: public read-only /api/report over Redis issues, plus the Next
proxy. Both surfaces registered under the reachability guard.

A test bug I made twice now: my check for forbidden sample values matched
the template's own doc block, which NAMES those values as things never to
paste. Documentation worth keeping, so both suites strip comments before
matching -- a guard that reads its own warning is not reading the code.

WAVE-2 STATUS: E1, F9-F11, E10, E12 done. Still gated -- F5 article media
and E16/F8 on the card-system reconciliation; the in-season hub IA on the
social chat's formula; E9/E15 on model; E2/E6 on licensing.

Read-only throughout; serving fingerprint unchanged including
newsletterService; accrual clock unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 22:17:59 -04:00
builtbykev 49e76068da Doctrine + E1 movement strip + F9-F11 offseason hub shell, to spec
PHASE 0 — specs/ARCHETYPE-TAXONOMY-DOCTRINE.md records the ruling as
shared law: 83 designed glyphs are the full four-sport taxonomy; a glyph
renders ONLY where its archetype is modeled and proven. 39 of 83 map to a
real archetype and are wired; the 44 unmapped are DORMANT SLOTS for
WNBA/NBA/Soccer, not a wiring gap. Wiring them would mean inventing 44
archetypes to consume artwork -- decoration presented as classification,
which is forbidden. DUAL THREAT and PAINT BOSS are modeled archetypes with
no mark: the mirror gap, flagged to the design side. When a sport's
archetypes ship, activation is a MANIFEST lookup, not new art.

PHASE 1 — E1 movement strip. The spec's own line is "the movement strip is
defined once here and reused everywhere a line has a past", so it is a
primitive, not a fourth chart.

RECONCILED RATHER THAN FORKED: lib/gradeShift.js ALREADY implements E1's
colour law -- toward/against/flat, including the direction flip that makes
an UNDER's favourable move the opposite sign of an OVER's. MovementStrip
CONSUMES buildGradeTimeline instead of reimplementing it, and a test
asserts it never redefines isUnder. GradeShift stays the grade-history
view; this is the reusable strip. That is the card-fork lesson applied
before it could happen again.

Spec laws honoured: STEPS NOT CURVES (H then V, no smoothing -- a curve
invents prices that never traded, and a test rejects any C/S/Q/T command);
green only when the move FAVOURS the read; FLAT renders as a hairline plus
FLAT · [N]D because a flat market is a finding; and too little history
says NO MOVEMENT HISTORY rather than rendering blank.

PHASE 2 — F9-F11 offseason hub shell, built from Vyndr Offseason.dc.html.
The spec's load-bearing words are used verbatim: "OUTLOOKS REPRICE ON NEWS
· NOT GAME ODDS" (an offseason number is not a game line), the QUIET WIRE
empty state ("No outlook-moving news since X. We don't manufacture
movement."), WHAT CHANGED TODAY as the hero with the countdown ambient and
top-right, the tag-colour-is-meaning row anatomy, the open -> NOW -> FAIR
triplet with the movement strip embedded, and the OUTLOOK ONLY block where
every row carries NOT GRADED.

THE DESIGN FILE'S SAMPLE DATA IS NOT IN THE COMPONENT. Wembanyama +420 ->
+330, Nabers cleared 11:42 AM, the Summer League names -- all of it is a
SPEC for what a live feed renders, and copying it in would be fabrication
carrying a designer's authority. A test asserts none of those strings
appear.

The IN-SEASON information architecture is NOT invented here. The spec
covers an offseason hub; nothing specifies how content, articles, wire and
the live slate share year-round navigation. That remains the open design
gap, and the route notes it.

Two test bugs caught and fixed: my first assertions matched my own doc
comments -- the ordering check found "WHAT CHANGED TODAY" in the header
block and the no-curves check caught the word "curve" in the sentence
explaining why curves are wrong. A guard that reads its own explanation is
not reading the render; both now strip comments first.

PHASE 3 — both surfaces registered under the reachability guard. Read-only
throughout, serving fingerprint unchanged, accrual clock unchanged at 0
eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 21:24:36 -04:00
builtbykev 60469422af Wave D1 primitives — and the audit says E1/E9 were not Wave 1
PHASE 0 — the order proposed E1 + E9 as Wave-1 and told me to follow the
audit if it disagreed. It disagrees: E9 calibration curve is WAVE D3,
gated on MODEL work ("resolve n>=20 vs N30, accrue buckets"), and E1
movement strip is WAVE D6, a large surface build. E9's gate is live right
now -- calibration is WITHDRAWN at 0 eligible dates, so the curve could
only render its empty state today. Building it would ship a component
whose entire purpose is unavailable.

PHASE 1 — three of the five real Wave D1 items were ALREADY DONE, and the
2026-07-31 audit has aged:

  D1 glyph library  audit: 38/83 wired (46%)
                    now:   COMPLETE for everything wireable -- 39 of 83
                           designed glyphs map to a real archetype, all 39
                           are wired, colours match the registry exactly
                           (0 disagreements).
  A1 card token     audit: BUILT-BUT-DRIFTED, "in only 1 file"
                    now:   BUILT-TO-SPEC -- it IS the --bg-1 token,
                           consumed by 32 files. The audit counted literal
                           hex, which is what a correctly tokenised value
                           looks like.
  B1 boundary blue  audit: PARTIAL, hex in 2 files
                    now:   BUILT -- --priced-out/#8fb2de is a token with a
                           documented colour law, 4 consumers.

The 44 unwired glyphs are NOT a wiring gap: they have no backend
archetype, so wiring them means inventing 44 archetypes to consume
artwork -- the fabrication this programme refuses. That is the 41-vs-74
scope question and it is Kev's call. Separately, 2 registry archetypes
have NO designed glyph (DUAL THREAT, PAINT BOSS) -- a design gap.

PHASE 2 — what was genuinely absent is now built. web/src/lib/motion.js:
nudge() capped at 180ms so it reads as acknowledgement rather than
latency; bootStagger capped at 240ms because uncapped, row 40 waits 1.1s
and the stagger BECOMES the latency it exists to disguise; rowHover
returns handlers not CSS so touch cannot stick a hover state; and
revealOnIntersect returns an unobserve in every path and reveals
IMMEDIATELY when there is no IntersectionObserver or motion is reduced --
content is never hidden behind a capability check.

Reduced motion is honoured, not softened. The sharpest of the 10 tests:
bootStagger under reduced motion returns opacity 1, not merely delay 0 --
if the CSS animation supplies the opacity, skipping it leaves the row
invisible forever.

PHASE 3 — the reachability guard gains a PRIMITIVES section: a module
built to be embedded must declare its exports AND name its intended
consumers, because a primitive imported by nothing is the same
built-but-unread class as an unmounted component.

WAVE-2 UNGATED: F9-F11 offseason hub, F5 article media, E10/E12 Report,
E1 movement strip. GATED: E9 + E15 on model, E16/F8 on the resolution tail
and the card-system reconciliation, E2/E6 on licensing, E13 on another
order.

Read-only throughout; serving fingerprint unchanged; accrual clock
unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 20:31:55 -04:00
builtbykev 7a318ccce0 Media/hub design inventory: the specs already exist, and so does a gap audit
READ-ONLY. Nothing designed, built or mounted.

THE HEADLINE: this inventory was largely already done. specs/
design-vs-build-gap-audit.md is a 61-item design-vs-build audit from
2026-07-31, every claim grep-verified, with a wave-ordered build plan. The
useful work is reconciling it, not redoing it.

WHAT ALREADY HAS A FULL DESIGN SPEC (and is NOT built):
- Offseason hub -- Vyndr Offseason.dc.html, 119KB: hub home desktop+390, an
  NBA Summer-League variant with an OUTLOOK ONLY / NOT GRADED honesty
  block, season-long board, season-read reveal with WHAT WOULD CHANGE THIS
  READ, news/outlook feed row anatomy, quiet-wire empty state.
- Article media (S3) -- hero template plus four hero graphic archetypes,
  all GENERATED data visuals never stock, inline figures with a one-pull-
  quote max and a caption law that every figure names its data, article
  card, OG 1200x630 with the master already rasterised.
- Movement strip (E1) -- steps not curves, green only when the move favours
  the read, FLAT as hairline. THIS CORRECTS MY 2026-08-07 BOARD, which
  listed it UNKNOWN and possibly satisfied by GradeShift. It is a specified
  primitive that does not exist.

THE WIRE is BUILT (vyndr/Ticker on real ticker exhaust) and is a DIFFERENT
thing from NewsWire.tsx, which is the offseason news/outlook feed.

THE ONE REAL DESIGN GAP: the hub's information architecture across seasons.
The Offseason file specifies an OFFSEASON hub; nothing specifies what that
surface is IN-season, or how content, articles, wire and the live slate
share one navigation. Every component exists on paper; their composition
into a year-round media surface does not.

A FINDING AGAINST MY OWN RECENT WORK: I built a card renderer at 1080x1350
without checking whether a designed card system existed. It does -- five
master sizes with defined layouts and rasterised PNGs in exports/. The
content engine's cards are a parallel invention. Not wrong, but they should
conform to E16/F8 rather than diverge, and that is a design decision.

CORRECTION: the order lists content API, preview page, book-comparison
mount and guard widening as still pending. All four shipped at c575a70,
and book comparison was never unmounted -- my earlier board grepped only
web/src/app and missed component-level mounting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 19:18:36 -04:00
builtbykev c575a708c7 Content studio API + preview page; widen the reachability guard; correct
two inventory errors

INVENTORY CORRECTION, and it was mine. Phase 2's two "orphans" are NOT
orphans -- my board grepped only web/src/app and missed component-level
mounting. The transitive check says both are already mounted:

  BookComparisonPanel -> GradeResultCard -> app/scan/page.tsx
  NewsWire            -> ExploreHub      -> app/explore/page.tsx

So book comparison is DONE (wired to /api/books, rendering on the grade
card) and THE WIRE is DONE-BY-DESIGN, mounted in ExploreHub. Its header
names an "Offseason Hub" as its home, and that hub genuinely does not
exist -- but that is board item #8, not a mounting bug, and inventing a
surface to satisfy a comment would be the wrong fix.

The lesson is the same one this session keeps teaching: I checked one
directory and reported a conclusion the check could not support.

ALSO CAUGHT: I overwrote src/routes/content.js, which was the Session-29
content-templates route, by picking a filename without looking. Restored
from git with no work lost; the new surface lives at
/api/content-studio and both now coexist.

PHASE 0/1 — /api/content-studio serves finished posts (copy, branded card,
card_svg, the fact_contract each was REQUIRED to have, and the facts that
actually backed it) plus a POST for editorial status in Redis. Private via
internal key; the Next proxy holds the key server-side so the browser
never does. /studio renders it as a thin client -- copy and card side by
side with the fact contract visible, because reviewing copy by reading it
is exactly how a wrong number ships. Never-blank: a night with nothing
generated says so.

API-FIRST is the point: the endpoint an autonomous poster will call is the
one the page already renders, so the agent handoff is a pointer change,
not a rebuild. Contract documented at docs/CONTENT-STUDIO-API.md.

EXPRESS 5 BROKE 23 SUITES at first: `router.get('/:date?')` throws at
mount time in Express 5, taking down everything that imports app.js. Two
explicit routes instead.

PHASE 3 — the reachability guard is widened from grade-fields-only to a
general built-but-unread check. Book comparison, THE WIRE and the content
studio are now registered surfaces; a page counts as its own entry point
(Next mounts it by convention) while everything else must trace to one.
22 checks green; a registered-but-unimported surface still goes red.

FULLY ISOLATED: read-only on model/slate/ledger, serving fingerprint
verified unchanged, accrual clock unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 18:51:52 -04:00
builtbykev 74aa75945e Content engine: posts that structurally cannot lie
PHASE 0 — contentEngine makes Truth Law structural, not careful. Copy is
token-substituted and an unbacked {token} REFUSES to render -- there is no
code path that produces a plausible default. The fact contract is asserted
before any string is built. Card and copy render from ONE fact object, so
a caption and a card cannot disagree. No live model writes factual claims:
the voice is in the template, the facts are pulled, and the voice-polish
port is deliberately unwired, because an LLM that can rewrite a sentence
can rewrite a number.

18 tests carry the proof. The one that matters most: ZERO IS PRESENT.
"0 cleared B+" is our most honest possible post, and treating 0 as missing
would be the Number(null)===0 breach wearing its opposite coat -- it would
silently delete exactly the post the brand is built on.

PHASE 1 — three templates, generating real posts from tonight's data:
hot hitters off the repaired full-season log, the honesty flex off the
real servedGrade distribution (2,140 graded / 70 cleared B+ / 42% not
separable / A unissuable), and streaks verified from settled outcomes only.

THE ENGINE CAUGHT A BUG IN ITSELF, and it is the sharpest lesson here. The
first run published "No hitter is meaningfully hot tonight -- we could
dress up a middling week as a streak. We don't." That was FALSE: the
box-score cache spans only the settled window, every player had under 20
games, and the pool was empty. A broken pull was publishing as considered
editorial judgement -- the fourth appearance of this class tonight and the
first where our OWN HONESTY COPY was the disguise.

Fixed structurally rather than by patching the number: an absent() variant
may now DECLINE to speak, and the template separates "no candidates at
all" (SKIP with a reason) from "candidates judged, none hot" (honest
absence). Both locked by test. Source corrected to mlbStatsAdapter.fullLog,
the same log the repaired champion reads.

PHASE 2 — cardRenderer emits SVG rather than canvas: it is text, so it
diffs in review and its numbers are greppable, which matters when the
whole claim is that the numbers are real. VYND white + R green, slashed-Y,
scanlines, mono. The card never formats its own facts -- every string
arrives pre-rendered and gate-checked.

PHASE 3 — scripts/generate-content.js writes copy + card per template to
.content-out/<date>/. Template N+1 is a registry entry: requires, pull,
copy, card, absent. Queued as stubs, not built: hot takes, daily reads,
"grades we DIDN'T give", cross-sport streak variants (the streak template
is already sport-agnostic -- settled outcomes and a noun).

FULLY ISOLATED: read-only on every source, zero writes to serving, model
or ledger tables. Serving fingerprint verified unchanged. The accrual clock
is untouched at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 16:34:57 -04:00
builtbykev 08791520fc Board: mechanical inventory of every outstanding item, read from the repo
Inventory only, no build. Every line cites a file or a query; unverifiable
items are marked UNKNOWN rather than asserted.

Corrections to the assumed board:
- Scanner is DONE on the blue-boundary format (scan/page.tsx:715-717), not
  still amber.
- Screen conversion is DONE except ONE file -- RouteStub survives only in
  app/notifications/page.tsx across 47 route dirs.
- Terminal is deliberately retired (S57 redirect), not debt.
- ev_pct DOES render, in 5 frontend files. That flag was stale.
- edge_pct is NOT a scale bug: (model-line)/line is correct arithmetic that
  explodes on small lines (a 5.5 projection on a 0.5 line is a legitimate
  1000%). Live top values 900/860/700 on 68,364 rows. It is a display
  defect, not a broken computation, and it reaches a user only via
  alt_lines typing -- the grade card computes edge independently.
- All FOUR sports have real archetype registries (nba 15, mlb 15, soccer 6,
  wnba 5) and ACTIVE_SPORTS runs all four. Only MLB BATTERS is MODELED;
  everything else is scaffolding with a sound base rate and no proven
  factors. MLB pitchers: pitcherEngine.js exists at 261 lines and is read
  by ZERO serving code.
- Book comparison is built and mounted NOWHERE -- the same built-but-unread
  class the reachability guard exists for, and NOT covered by it, since
  that contract is grade fields only.

UNKNOWN and marked as such: opp_rank_stat prod population, internal-key
rotation, the ~9% void/DNP rate, and whether MovementStrip is a distinct
spec item or satisfied by GradeShift/MarketBreadth.

The ordered board puts five PARTIAL items first and flags accrual
sensitivity: seven items are isolated design/guard work buildable today
with zero clock impact; five touch the serving path and would muddy the
accruing repaired-champion dates. edge_pct is the sharpest tension -- a
real honesty defect a user can see, whose fix touches serving mid-accrual.

Clock today: 0 eligible dates, all four re-audit items WAITING.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 15:13:02 -04:00
builtbykev 55b210cb95 Fix the dormant basketball window-bug before it ships; guard the class
PHASE 0 — audit. espnStatsAdapter's slice(0,20) was already fixed at
494c83c, so the ESPN branch of getStatRows inherits a full log and is
CORRECT. MLB's l20_avg is a real seasonTotal/games aggregate, verified,
also CORRECT. TWO basketball defects remained:

  DEFECTIVE  nbaGameLogFeatures   m20 = avg(vals.slice(0, 20))
             -- l20_avg is the season reference projectionFor reads, and
             slicing to 20 made it a twenty-game average wearing a season
             label. The MLB l20 bug, unfixed for basketball.
  DEFECTIVE  getStatRows python   getGameLogs(playerName, sp, 20)

PHASE 1 — both fixed via a named SEASON_LOG_DEPTH = 100, past an 82-game
season so a request can never truncate one. API COST: ZERO. The count is a
request parameter, so asking for a season is the same single call. No
extra request, no extra quota.

MEASUREMENT DEFERRED, EXPLICITLY: basketball is offline, there are no
settled basketball rows, and resolution before/after CANNOT be measured
now. This is a code fix, not a measured improvement -- exactly like the
deferred MLB sibling paths.

PHASE 2 — the window guard asserts the property on source: no fixed N may
stand in for a season. It immediately caught TWO MORE instances I had
missed in Phase 1 -- a hardcoded 20 at featureCache:377 and
gameLogService's own `count = 20` DEFAULT, which would have handed a
twenty-game window to any caller that omitted the argument. That is a
seventh path, found by the guard rather than by me.

Retro-confirmed: run unchanged against 981a05c it goes 3 failed / 6
passed, flagging the basketball slice, the game-log request and the
season-reference check.

The class in full, now six paths across two sports, every one of which
looked like ordinary code. `slice(0, 20)` is unremarkable; what made it a
defect was the QUESTION it answered -- "what is this player's season
rate?" -- and no test could see that mismatch because the value produced
was always a plausible number.

PHASE 3 — THIS CLOSES THE NON-ACCRUAL ARC. Everything buildable without
settled rows is built: push unblocked (it was the wrong remote, not the
firewall), grade surface honest and rendering, reachability guarded,
champion repaired across every path including dormant basketball, and the
window class guarded so it cannot return.

The program is now correctly IDLE on modeling. Today's count: 0 eligible
dates, all four re-audit items WAITING. FIRST TRIGGER: 10 eligible
calibration dates on the repaired champion, at which point the resumption
order is calibration re-fit, hits factor lift, prior verdict re-audit,
rbi lineup-slot gate.

No basketball measurement claimed. No NBA chain/archetype build. MLB
serving verified byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 15:00:24 -04:00
builtbykev 981a05cbd6 Render-reachability guard: make built-but-unread a CI failure
Three consecutive orders shipped a backend-correct field that never
reached a screen, all on a green suite: gradeBands (required by no
serving code), served_grade (dropped at the adapter boundary),
GradeScaleLegend (imported by nothing). Each was caught by luck on a later
re-check, and in two of the three I had already reported the wiring done.

WHY GREEN TESTS COULD NOT SEE IT: backend tests stop at the API payload.
They prove a field is PRODUCED and say nothing about whether it is
CONSUMED. Invisible by construction, not an oversight in any one test.

THE TRAP, NAMED: the difficulty pools in the backend, so by the time a
field exists on the payload it feels finished. What remains is a
three-line adapter change nobody considers worth verifying, so it gets
claimed rather than traced. The last inch is the one with no friction,
which is exactly why it gets skipped. "I added the field" and "a user can
see it" are different claims and only the first is fun.

THE GUARD traces each promised field the whole way: payload -> adapter
consumes -> component renders -> component is MOUNTED. Mounted is
transitive to a Next entry point (page/layout/template), the only thing
that puts a pixel on screen, depth-limited so an import cycle cannot hang
the suite.

Container rows are exempted EXPLICITLY, not silently: served_grade carries
container:true plus a rendersVia list, and a separate assertion checks
every named part actually renders. The exemption is auditable and cannot
hide an unrendered field.

The guard tests itself -- an orphan component must report unmounted, and
the contract must be non-empty, since an empty contract passing vacuously
is how this would most plausibly rot.

RETRO-PROOF: run unchanged against 3591c76, before the wiring, it goes
11 failed / 8 passed and names the exact bugs -- "the ceiling stance /
grade scale legend - its component is MOUNTED, not merely written", "the
served grade object - the adapter consumes it", "whether the band
separates from the baseline - a component actually renders it". Green on
the current tree.

HONEST SCOPE LIMIT: gradeBands is NOT in the contract and would not be
caught. It is a backend module, not a promised user-facing field, and it
is correctly unwired -- every band collapses to base-rate at current
resolution. Out of scope by design, not oversight.

Now in the standing suite, so the three-gate floor is tests green
(including reachability) + build exit 0 + fingerprint. No serving or model
change. p_win never mutated. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:44:17 -04:00
builtbykev 09186ea609 Wire the composed surface: the honest fields now reach a user
PHASE 0 re-check caught the same failure a THIRD time, mine again. Last
turn I created GradeScaleLegend.tsx and never mounted it, and I reported
that the separation flag "renders per band" -- it did not. grep: zero
frontend references to separates_from_base_rate, served_grade or
factor_adjustment. The honest grade was reaching the API payload and dying
at the adapter boundary.

That is three occurrences in three orders of the same shape: built,
correct, unread. gradeBands, then served_grade, then the legend.

PHASE 2 — the composed surface is wired end to end:
  analyzeViaEngine1 -> served_grade + factor_adjustment on the payload
  scan/page.tsx     -> ScanResponse types them and forwards them
  gradeAdapter      -> gradeMeaning, separatesFromBaseRate,
                       bandRealizedRate, factorsApplied, refusalReason
  GradeResultCard   -> renders "WHAT THIS GRADE MEANS"

What a user now sees that they could not before: what the band has
actually realized, an amber note when the read CANNOT be separated from
the baseline, and -- only where a factor actually fired, with its proven
sign -- what moved the read, in plain language rather than feature names
("where he hits it vs who is standing there").

No narrative on props where nothing fired: factorsApplied is empty and the
block self-hides. Only the three PROVEN hits factors have labels, so an
unproven factor cannot acquire prose by being added to the map.

PHASE 1 — GradeScaleLegend is now MOUNTED on the grade card (compact). The
ceiling is a stated position where the grade is, not a page a user would
have to find.

PHASE 3 hand-verified through the real adapter across twelve states: B+
with 3/0 factors, B with 1, C+/C/C- all flagged not-separable, D, F, a
switch-hitter case where 2 of 3 factors fire, two refusals and a
no-forecast. never-blank PASS, no-manufactured-A PASS, no-narrative-when-
nothing-fired PASS, separation-flag-reaches-card PASS.

No A-threshold loosening. No calibrated number leaks. p_win never mutated.
engine_grade still read by zero serving code. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:36:03 -04:00
builtbykev 3591c7626e Total grade cutover + the ceiling stated as a position
PHASE 0 caught my own repeat of the failure I diagnosed one order ago.
91927a4 attached `served_grade` BESIDE the old letter and left `grade`
alone -- so the honest grade reached nobody, exactly as gradeBands had
been built-correct-and-unread. grep showed served_grade appearing in one
file (where I set it) and all 14+ consumers -- scan route, dashboard,
parlay, newsletter, desk, content templates, retention -- still reading
`.grade`, i.e. still the dishonest letter.

CUTOVER IS NOW TOTAL: legacy.grade IS the honest letter. Overwriting the
one field every consumer already reads cuts every surface over at once
instead of editing fourteen call sites and missing one. engine1's index is
preserved as `engine_grade` and verified read by ZERO serving code.

Confidence follows the letter: it came from a grade-band midpoint of the
OLD letter, so leaving it would have paired a served B+ with a C's
confidence. Both now derive from p_win, kept on the existing 0-100 scale.

MEASURED BLAST RADIUS before shipping: 303 of 47,991 non-refused
snapshots (0.6%) have a grade but no p_win, and now render NO READ instead
of a letter. That is correct -- their old letter came from the retired
index carrying 0.48% resolution, i.e. noise -- and NO READ is a rendered
state with a reason, so never-blank holds.

PHASE 1 — the ceiling is now a STATED POSITION, not a confusing absence.
servedGrade.SCALE_LEGEND plus web GradeScaleLegend.tsx say it plainly: we
do not issue A grades, no band has hit at a rate that would justify one,
our honest ceiling is a strong B+ (~66% realized vs ~60% baseline), and if
the model earns an A the legend changes and we say why. The
separates_from_base_rate flag renders per band -- C+/C/C- are labelled
"we cannot separate this from the baseline", which is most of any slate.

PHASE 3 hand-verified across every state: B+ with 3 factors (basis
forecast_plus_matchup_factors), B+ with none (forecast_only), C flagged
not-separable, F, and three refusal states rendering NO READ with reasons.
never-blank PASS, no-manufactured-A PASS.

Test fallout was real and is documented rather than papered over: engine
BEHAVIOUR assertions moved to engine_grade, suppression assertions stayed
on grade (a suppressed prop has no letter either way), and the confidence
78 -> 95 change is the grade-band midpoint being replaced by p_win.

No A-threshold loosening. No calibrated number leaks (deployed set empty).
p_win never mutated. Ten frozen modules verified unchanged including
engine1 and probabilityEstimator. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:07:02 -04:00
builtbykev 91927a4a8a Serve an honest grade: the letter was carrying 1/6 the information of the
number beside it

PHASE 0 corrects the order's premise. A grade letter has been served all
along -- engine1.gradeProp builds it from an additive factor index,
computed INDEPENDENTLY of p_win. gradeBands is orphaned for a different
reason than assumed: it defines what a letter MEANS from realized
outcomes, and every band collapses to base-rate at current resolution.

The measurement that changed this order, on 3,417 settled props:

  grade  n      realized  mean p_win
  A         8    0.500      0.647     <- the TOP grade did WORST
  B       985    0.640      0.700
  C     1,695    0.602      0.676
  D       303    0.558      0.604
  F       426    0.535      0.588

  letter resolution 0.00116 (0.48% of variance)
  p_win  resolution 0.00715 (2.98%)
  -> the letter carried 0.16x the information of the number beside it

Concretely, from the hand-verify: Christian Encarnacion's 0.95 over
graded C and his 0.05 under ALSO graded C -- same hitter, opposite
forecasts, same letter. The gap was never that grades don't ship; it is
that the weaker of two available signals shipped as the headline.

PHASE 1 — model/servedGrade.js derives the letter from p_win with bands
anchored on MEASURED realized rates (B+ 0.663 / B 0.646 / C+ 0.615 /
C 0.589 / C- 0.548 / D 0.512 / F 0.447, base 0.6005).

NO MANUFACTURED A, structurally: A+/A/A- are UNISSUABLE, not rare. The
realized rate plateaus at 0.65-0.68 above p_win 0.70, so no band has
earned a top letter; a test sweeps every p_win 0..1 and asserts none
produces one. Even 0.99 tops out at B+ with its realized 0.663 attached.
Raising that ceiling later is a deliberate, visible act.

Bands that cannot separate SAY so -- C+/C/C- carry
separates_from_base_rate false and copy naming it, which is the honest
description of a forecast explaining 3% of variance. Every grade states
its basis (forecast_only vs forecast_plus_matchup_factors, naming which
factors fired) and calibrated:false. engine1.grade is preserved as
engine_grade so nothing downstream breaks.

PHASE 2 — refusals render real states: insufficient_data -> "not enough
history to call this one"; juiced_no_edge -> "the book has priced the vig
past any edge on this side". 1,870 refused snapshots carry exactly those
two reasons and both now surface.

PHASE 3 — hand-verified on 12 real served props. Freeman/Rice/Encarnacion
0.95 overs now B+ (was B, C, B); the 0.05 unders now F (was C). Refused
doubles render NO READ with their reason. never-blank PASS,
no-manufactured-A PASS.

Serving change; nine frozen model modules unchanged including engine1;
p_win never mutated; no calibrated number leaks (deployed set empty); no
Bonferroni slot.

STILL TRUE: the forecast explains ~3% of outcome variance. This order did
not make the model better. It made the letter stop overstating it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 04:53:02 -04:00
builtbykev 9159b7e1b9 Diagnose the push block: it was the wrong remote, not the firewall
The GATE-0 theory is REJECTED for VYNDR, tested rather than assumed. Both
remotes are ALREADY HTTPS -- there is no git@ SSH URL in this repo to be
blocked -- and both hosts answer on 443: git.builtbykev.com HTTP 200 in
0.64s, github.com HTTP 200 in 0.17s, Gitea's git endpoint 200, GitHub's
401 (auth required, reachable). No connectivity failure of any kind.
COLYRA's HTTPS-remote fix was right for COLYRA; VYNDR was already in the
state that fix produces, so applying it would have minted a new token to
solve a problem that did not exist -- and the pre-existing credential
would have made the new token look like the cure.

The real cause is in the error text, which named it exactly and was
misread all session: "could not read Username for 'https://github.com'".
Every attempt used `git push origin main`, and origin is GitHub
(kev3109/betonblk) with no stored credential. The WORKING remote is
`gitea` (git.builtbykev.com/builtbykev/vyndr.git), and a valid credential
for it sat in ~/.git-credentials the entire time. The habit of typing
`origin` is what kept twenty commits local. Nothing was blocked.

FIX: git push gitea main. Pushed 6452926..ecf78b9, 21 commits, verified by
git ls-remote matching local HEAD exactly.

No firewall rule touched, no token created, no code or model change --
infra only. The bundle and patch series stay as belt-and-braces; they are
no longer the only copy.

docs/GIT-PUSH-DIAGNOSIS.md records it so no future session re-reaches for
an infrastructure theory when the error message names a credential and a
specific host.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 04:24:21 -04:00
builtbykev ecf78b911c Accrual watch + pre-registered resumption; the program is idle on modeling
PHASE 1 — the existence risk is resolved. Nineteen commits were local-only
with no push credentials. Three artifacts now exist off the working tree:

  ~/vyndr-full-history-2026-08-07.bundle   6.8M  SELF-CONTAINED, clone it
  ~/vyndr-session-2026-08-07.bundle        244K  needs existing history
  ~/vyndr-session-patches/  (20 patches)   1.5M

Use the full-history bundle -- `git bundle verify` shows the session bundle
requires ref 6452926, so it only applies onto a repo that already has this
history. The full one clones standalone. Copying one off the machine is
the top non-accrual action item.

PHASE 2 — the watch, today's starting line: 0 eligible rows, 0 eligible
dates, everything WAITING. Zero is correct, since the repair ships in this
session's commits and no settled row can carry the marker yet.

THRESHOLDS ARE ATTEMPT FLOORS, NOT TRUST FLOORS, and this is written into
the ledger so a future session cannot misread it. Reaching 10 dates means
the calibration re-fit CAN BE MEASURED, not that it is trustworthy -- we
lived that distinction tonight with 2-4 cluster CIs, a LODO gate at
1.4-9.3% power, and a >=40 bar that was right for promotion and wrong for
deploy. A 10-date map is thin, deploys PROVISIONAL with auto-demotion, and
its interval will still be wide.

Real-time estimates stated so the wait reads as designed: ~2 weeks for
calibration and hits lift, 3+ weeks for the verdict re-audit and the rbi
gate, since rbi accrues slowest and a date only counts once its props
settle.

PHASE 3 — resumption order fixed, triggered purely by date thresholds:
calibration re-fit @10, hits composed-lift re-measure @10 (the 1.39%
figure is VOID, direction unknown), prior verdict re-audit @14 (not
pre-priced), rbi lineup-slot two-part gate @14. FIRST TRIGGER: when
reAuditEligibility reports 10 eligible calibration dates.

Until then the program is honestly IDLE on modeling. That is the correct
state, not a gap to fill -- the alternative is measuring on reconstructions
of a retired forecast, which this program has now refused by name three
times.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:50:53 -04:00
builtbykev 2391574f00 Live-surface integrity on the repaired champion; fix the label my own
repair falsified

PHASE 0 — four checks PASS, one defect found and fixed.

PASS  gradeBands is required by NO serving code -- built across several
      orders, never wired. No stale band derived from the retired
      ten-game champion can reach a user, because none reaches a user
      at all.
PASS  CALIBRATION_DEPLOYED is [] and the calibrate loop iterates it, so
      calibrate() is never called and p_win_calibrated is never set. The
      only assignment site sits inside that empty loop. No withdrawn map
      can leak.
PASS  projectionFor reads l20_avg, which mlbGameLogFeatures now builds
      from fullLog -- so refusals are computed on the repaired
      full-window reference, not the retired ten-game one.
PASS  factors still fire with correct sign across the repaired base
      range (0.35/0.50/0.65/0.80): defense lowers, pitcher-contact
      raises, platoon raises at every point. Mechanical firing check
      only -- NOT a lift re-measurement, which waits for accrual.

DEFECT FIXED — my own repair falsified a user-facing sentence. The grade
card rendered "Last 20 games average: X" from l20_avg, and l20_avg is now
a FULL SEASON average. The number changed and the label did not, so the
surface was stating something the data no longer supported. Copy now reads
"Season average"; trapDetection's L20 explanations likewise. The field
name is kept -- it is read in many places -- but no rendered sentence
claims a window that isn't there.

That is the same class as everything else tonight, one layer out: a
correct-looking string describing data that moved underneath it.

Serving change; frozen model modules unchanged; p_win never mutated; no
Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:49:54 -04:00
builtbykev 494c83cf76 Hunt the window-bug class: three more paths, and the forward re-audit rule
in code

PHASE 0 — getStatRows is the single base-rate path, so every branch is
audited, plus the feature builders since l20_avg is the season reference
projectionFor reads:

  getStatRows MLB -> estimator base    fullLog            CORRECT (929fd81)
  mlbGameLogFeatures l5/l10/l20        last10 = 10        DEFECTIVE
  espnStatsAdapter.parseGameLog        slice(0,20)        DEFECTIVE
  getStatRows NBA/WNBA ESPN branch     inherits 20-cap    DEFECTIVE via source
  getStatRows NBA/WNBA python branch   getGameLogs(...,20) dormant (offline)
  pitcherEngine / skillProjection      statcast profiles  N/A
  pitcher props via getStatRows MLB    fullLog            CORRECT
  settleSource                         full log (S64)     CORRECT

THE PITCHER ANSWER IS GOOD NEWS: pitcher props run through the same
getStatRows MLB branch, so 929fd81 repaired them too. There is no separate
defective pitcher base-rate path.

THE ONE HIDING IN PLAIN SIGHT: mlbGameLogFeatures carries the comment
"l20 = all available (the season per-game reference projectionFor needs)"
while building from last10 -- so l20_avg was a TEN-GAME AVERAGE WEARING A
SEASON LABEL, feeding both the consistency pull inside the estimator and
projectionFor, which decides refusals. It survived the previous repair
because that fix touched only getStatRows.

PHASE 1 — mlbGameLogFeatures now reads fullLog; espnStatsAdapter drops its
slice(0,20) cap. ZERO new API calls on both: each widens data already
fetched and then discarded, the same shape as the original repair. The
python branch is left alone -- the service is offline in prod and fixing it
would be speculative.

Their before/after resolution is NOT measured, deliberately: the only way
to measure today is to reconstruct the repaired forecast over old rows,
which is the reconstruction-vs-served trap this order refuses. Code fix
now, measurement at accrual.

PHASE 2 — MODEL_VERSION bumped to engine1@2026-08-07-fullwindow, so every
forward snapshot is self-identifying (retentionService already stamps it;
no new plumbing). model/reAuditEligibility.js encodes the rule: isEligible
accepts only the repaired marker, assess counts eligible DATES not rows,
and ACCRUAL is frozen at calibration 10 / hits-lift 10 / verdict-reaudit
14 / rbi-gate 14. A test locks the invisible case -- a MIXED table of 330
rows with 30 repaired returns eligible_dates 3, not 330 rows of false
confidence. Once both generations share a table a naive count would fit a
map on a blend of two forecasters.

PHASE 3 — the board, each consequence labelled: calibration WITHDRAWN
(refits at 10 dates, never on reconstructions); factor verdicts SUSPECT
(all measured against a champion worse than a frequency table, direction
UNKNOWN, not pre-priced, 14 dates); hits factor lift UN-REMEASURABLE (10
dates, factors still wired and transmitting); rbi lineup-slot RE-QUEUED
(14 dates). Pre-registered order: calibration, hits lift, verdict
re-audit, rbi gate.

Then STOP and accrue. Nothing further can be honestly measured until the
board fills with rows the repaired champion produced.

Serving-path changes by design for the MLB feature path and NBA/WNBA logs;
eleven frozen model modules verified unchanged. p_win never mutated. No
Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:40:20 -04:00
builtbykev 929fd81940 Repair the champion: it was reading ten games, not a season
PHASE 0 — the defect is real past the peek. Against a FAIR point-in-time
baseline (each player's rate over games strictly before that date, >=10
prior games, box scores back to 05-01), the served champion LOSES on all
four stats, three of four CIs excluding zero:

  hits  0.00251 vs 0.00774  CI [-0.0074,-0.0011]
  TB    0.00393 vs 0.00619  CI [-0.0055,-0.0003]
  rbi   0.02481 vs 0.03133  CI [-0.0153,-0.0005]
  runs  0.00181 vs 0.00683  CI [-0.0114,+0.0008]

PHASE 1 — the cause is the WINDOW, not the weights. estimateProbability
builds its base rate as the frequency over every row it is handed, and
featureCache.getStatRows handed it res.last10. So the "season rate" was a
TEN-GAME rate, and 0.4 of the forecast was the last five OF THOSE TEN. The
0.40 recency weight costs resolution on all four stats (-0.00086,
-0.00107, -0.00562, -0.00365). Nudges are mixed and small -- harmful on
hits and rbi, marginally helpful on TB and runs -- so they are left alone.

PHASE 2 — two lines, no new data, no extra API call, because fullLog was
already fetched by the same adapter call that produced last10:
getStatRows now reads fullLog, and RECENCY_WEIGHT goes 0.40 -> 0.20.

  hits  0.00251 -> 0.00817  (tripled; now above the fair baseline)
  TB    0.00393 -> 0.00734  (above baseline; vs old CI [0.0020,0.0067])
  rbi   0.02481 -> 0.02727  (still below baseline, CI includes zero)
  runs  0.00181 -> 0.00436  (still below baseline, CI includes zero)

Gate stated exactly: hits and TB now exceed the fair baseline on the point
estimate; rbi and runs remain below but EVERY CI now includes zero, so no
stat reliably loses to a frequency table. That is a tie on rbi/runs, not a
win, and it is reported as one. Only TB's improvement over the old
champion is CI-confirmed; the rest are directional.

STALE-FIT GATE: CALIBRATION_DEPLOYED is now EMPTY. The low-param maps were
fitted on the retired forecast and fromLedger cannot rescue them -- settled
ledger rows still carry OLD p_win, so refitting today would refit the
retired forecast. Nothing is served calibrated until dates settle under
the repaired champion, and the favourite-longshot bias must be re-measured
rather than assumed to survive. The shadow duel is void.

PHASE 3 — the hits factor lift is NOT re-measured, and cannot be yet: it
needs settled rows produced BY the repaired champion, which ships in this
commit. Replaying would score the factors against a reconstruction rather
than the served forecast. Deferred, explicitly. The factors remain wired
and transmitting; only their lift is unquantified on the new baseline.

PHASE 4 — standing flag, and it is large: EVERY factor verdict in this
programme, every null and every THEATER, was measured against a champion
worse than a frequency table. Signal added to noise reads as noise. Prior
verdicts may deserve re-audit. Logged, not re-run.

Re-queued not built: rbi lineup-slot / RISP opportunity through the
two-part gate, now landing on a repaired champion.

Serving-path change by design; the byte-identical invariant inverted and
all four stats move. Nine frozen model modules verified unchanged. No
Bonferroni slot -- resolution accounting on the champion's own knobs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:28:33 -04:00
builtbykev 65ca6493db Decompose the rbi anomaly: it is lineup ROLE, and the counter is
out-resolved by a frequency table on three of four stats

PHASE 0 — the 14.51% is REAL. Re-derived with a paged pull asserted
against an exact count (rbi 7,930 == 7,930; hits 11,690; TB 12,086; runs
6,440), since this harness produced a false null three times tonight. rbi
resolution 0.03268 reproduces, deciles are monotone through the middle,
and 20 raw rows are in the artifact for hand audit.

CAVEAT GOVERNING EVERYTHING BELOW: the naive forecasts are leave-one-out
ON THE EVALUATION WINDOW, so they see the rows they are scored on while
the model is strictly point-in-time. They are upper bounds on available
resolution, not fair competitors, and every comparison is read that way.

PHASE 1 — the split:

  stat  MODEL    (a)player-base  (b)lineup-slot  (c)within-stratum
  rbi   0.03268     0.01167         0.03608          0.01908
  hits  0.00252     0.00446         0.00100          0.00473
  TB    0.00442     0.01331         0.03448          0.00607
  runs  0.00130     0.00262         0.01156          0.01170

FINDING 1 — rbi's resolution is LINEUP ROLE almost exactly. Batting-order
slot alone resolves 0.03608 against the model's 0.03268. A single integer
accounts for the whole anomaly and slightly more. That is opportunity, not
skill -- the cleanup hitter bats with runners on. 36% is matched by player
identity alone. Within similar-base-rate strata the model still resolves
0.01908, 58% of its total and higher than any other stat's ENTIRE model
resolution, so genuine within-role discrimination exists on top.

FINDING 2 — on three of four stats the model is beaten by "he's a .270
hitter". Base-rate-only out-resolves the model 1.8x on hits, 3.0x on TB,
2.0x on runs. Even allowing for the window-peeking advantage, a 1.8-3.0x
gap is not explained by that alone: the served counter appears to DESTROY
discrimination relative to the player's own rate. rbi is the one stat
where the model beats the naive baseline.

FINDING 3 — lineup slot out-resolves the MODEL on three stats: TB 7.8x,
runs 8.9x, rbi 1.1x. Hits is the only stat where batting order carries
less, which is mechanically right -- a hit is a hit wherever you bat, but
runs, RBI and total bases all scale with opportunity.

PHASE 2 — all three worlds are partly true, in measured proportions.
World A ~90% true (slot covers rbi's entire resolution). World B ~36% true
for rbi, but the WHOLE story for hits/TB/runs where base rate alone wins.
World C true with a low ceiling: hits' total available spread resolution
is 0.00446, i.e. 1.8% of variance from a forecast that has seen the
answers.

PHASE 3 — the next arc is NOT "strengthen hits factors". Hits has the
lowest available resolution on the board and last order's wiring already
took it to 1.39% of a ~1.8% ceiling. Named first factor order for next
session: LINEUP SLOT / RISP OPPORTUNITY on rbi through the two-part gate --
input already ingested and prod-verified (S89), resolution measured not
hypothesised, causally-correct unit is plate appearances with runners on.
Measured availability is not a pass; it still faces the gate.

And higher-value than either: the counter being out-resolved by a
frequency table on three of four stats is a defect in the CHAMPION, not a
factor problem, and it costs nothing to test -- the recency blend and the
+/-0.03 / +/-0.015 nudges are three lines in probabilityEstimator.

The hits transmission win from 43f65d3 stands: the conduit is real and
permanent. This order changes only which stat has the most worth flowing
through it.

Diagnostic only -- no factor wired, no serving path changed, p_win
untouched, all frozen modules byte-identical. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:09:44 -04:00
builtbykev 43f65d30cb Wire the three proven hits factors pre-grade: transmission proven, gain
inconclusive

THE BUG THIS NEARLY SHIPPED AS A FINDING. The first audit reported 0
factors fired on all 1,140 rows. Not a result -- my paging helper ordered
by `id`, and batter_spray, team_defense, platoon_splits and
statcast_aggregates have composite primary keys with NO id column. The
query errored, the loop broke on error, and four fully-populated tables
read as empty. hitsFactorContext.js -- the PRODUCTION loader -- had the
identical defect, so live wiring would have loaded nothing and served
unadjusted while logging success. Third occurrence of this class in one
session. Both loaders now order by a real column and THROW rather than
degrade. The Phase 2 gate is what caught it: no resolution number was
quoted until transmission was proved.

PHASE 1 — pipeline is now base -> FACTORS -> CALIBRATE -> GRADE. Context
built in snapshotService BEFORE gradeAndCacheSlate (was line 640+, grade
at 454), threaded per prop, applied to p_over before p_win is set with
p_win_prefactor and a full trace retained. Hits only. Coverage 859/1140
rows (75%): 474 with all three factors, 256 two, 129 one, 281 none.

PHASE 2 — TRANSMISSION PROVEN, 12/12 sign-correct, 4/4 per factor, each
applied IN ISOLATION. My first table compared each factor's expected sign
against the COMPOSITE change and showed 3 false failures -- with three
factors firing the net can oppose any single member; that was a flaw in
the test, not the wiring. Two under-side rows confirm the flip is handled:
a factor raising p(over) correctly lowers p_win. Switch hitters (Bailey,
Bell, Rocchio) took no spray adjustment while their other factors fired
normally -- the refusal is selective, not a blanket skip.

PHASE 3/4 — both maps refit on the factor-adjusted forecast; the
shadow-duel baseline is VOID and restarts, since it accumulated against a
different forecast. Point-in-time, 765 held-out rows:

  reliability 0.00795 -> 0.00828
  RESOLUTION  0.00229 -> 0.00345   (variance explained 0.93% -> 1.39%)
  Brier       0.25398 -> 0.25305   delta -0.00093  CI [-0.00225,+0.00002]

Resolution rose 51% relative. The CI TOUCHES ZERO on 4 eval dates, so the
composition does NOT earn a proven keep -- three isolated passes did not
grant a composed pass. INCONCLUSIVE, reported as such. The gain is far
below the sum of the isolated effects, which is expected: all three run
through the same pitcher-batter confrontation and share signal.

PHASE 5 — 1.39% of variance is still far below what band separation
needs. The pivot was correct and incomplete: the plumbing defect was real
and is fixed, three proven factors reach the served number for the first
time, and transmission alone did not buy grade separation. Next arc is
factor STRENGTH and BREADTH, not more plumbing.

PHASE 6 — rbi anomaly logged, not chased: 14.51% variance explained vs
hits 1.03%, on the stat we do not serve corrected and which has no proven
factors. Either the biggest lever on the board or a mirage; it deserves
its own order.

The byte-identical invariant INVERTED for hits by design. All 13 frozen
non-hits modules verified unchanged, probabilityEstimator included -- the
factors ride outside it. No new Bonferroni slot; the composed OOS claim is
reported with its CI and not claimed as a pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 02:53:48 -04:00
builtbykev e872eff4ce Instrument the calibration duel forward; diagnose the resolution ceiling
— the proven factors were never wired in

PHASE 0 — two truths recorded. The swap is a BET, not an OOS win:
isotonic beat low-param on identical held-out rows (hits +0.0028, rbi
+0.0042, TB tied) and we serve low-param anyway on an untestable prior
about shared daily structure. At 19 dates nothing here can test it. And
the MIN_SLOPE catch is preserved as standing rationale: a near-zero or
negative slope collapses toward base-rate-for-everything, which LOWERS
Brier while destroying all resolution -- a metric win that guts the
product.

PHASE 1 — the duel is now falsifiable. Both corrections computed on every
hits/TB prop; p_win_lowparam served, p_win_isotonic_shadow logged in its
own try so it can never break serving. calibrationDuel.adjudicate encodes
the rule IN CODE before any forward date exists: >=10 forward dates and
isotonic winning with a date-block CI excluding zero => REFUTED, revert;
otherwise UPHELD; under 10 dates PENDING regardless of the numbers. A
date counts as forward only if NEITHER map was fitted on it -- otherwise
we would be scoring which map memorised better. Nothing swaps now.

PHASE 2 — the ceiling, quantified via Murphy decomposition:

  stat   reliability  RESOLUTION  uncertainty  variance explained
  hits      0.01353     0.00252      0.24532        1.03%
  TB        0.01419     0.00442      0.24329        1.82%
  rbi       0.00654     0.03268      0.22531       14.51%
  runs      0.00788     0.00130      0.23182        0.56%

Calibration did exactly what theory says and nothing more: hits
reliability 0.01353 -> 0.00233 (-0.0112, 83% of the error removed) while
resolution moved -0.0002. Unexpected: rbi has 13x the resolution of hits
and is the one stat we do NOT serve corrected -- it needs calibration
least and discriminates most.

PHASE 2 DIAGNOSIS — NOT-TRANSMITTED, and not weak, ABSENT. Traced in code:
sprayDefense.js and platoonSeverity.js are required by NOTHING in src/,
only by analysis scripts and their own tests. The served p_win
(intelligence/probabilityEstimator.js:54) reads exactly four inputs --
game-log frequency, opp_rank_stat +/-0.03, home_away +/-0.015, and a cv
pull -- with zero occurrences of spray, platoon, hard-hit or
contact-profile. And snapshotService grades at line 454 while computing
challenger/context at 640+, so everything proven is computed DOWNSTREAM of
the grade it would inform. The three proven hits factors have never once
moved a served number.

That reframes the recent nulls: "calibrated p_win does not separate within
archetype" was never a statement about factors. The factors were not in
the forecast.

PHASE 3 — bands rebuilt on SERVED values (hits/TB low-param, rbi/runs
raw): 28 archetype slots across four stats, ZERO show lift. No longer an
open shrug -- it is the arithmetic of resolution 0.0013-0.0327 against
uncertainty ~0.23. A forecast explaining 1% of variance cannot produce
separating bands, and no correction to its numbers will change that.

HEADLINE: calibration is complete, delivered honest numbers on two stats
and zero grade separation, because the counter has no resolution -- and
the proven factors are not wired into the forecast at all. The second is
the reason for the first, and it is plumbing rather than a modelling wall.
Per-archetype grades need proven factors that actually reach p_win. Last
calibration order.

Serving unchanged from 74cf1ce. p_win never mutated. No Bonferroni slot.
Counter and frozen clusters verified file-by-file (15 modules).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 02:25:45 -04:00
builtbykev 74cf1ce974 Robust bias established; low-parameter correction replaces isotonic
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.

PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.

PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:

  hits         a=0.406 shrink 0.565  0.2626 -> 0.2540  CI [-0.0112,-0.0069]  DEPLOY
  total_bases  a=0.472 shrink 0.333  0.2490 -> 0.2429  CI [-0.0062,-0.0059]  DEPLOY
  rbi          a=0.775 shrink 0.231  0.2011 -> 0.2007  CI [-0.0007, 0]       REFUSE
  runs         a=-0.032                                                      REFUSE

A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.

STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.

PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.

PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.

p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 23:20:05 -04:00
builtbykev ced40421ed Audit the LODO instrument: it cannot evaluate any stat, and both prior
FAILs were false

PHASE 0 — the gate at 1f40014 was mine and was an incoherent pair. A 1-SE
informativeness bar with a ZERO-reversal rule: at exactly 1 SE a stable
stat's drop reverses with prob Phi(-1)=0.1587, so on four informative
drops P(>=1 reversal | perfectly stable) = 1 - 0.8413^4 = 0.50. It failed
stable stats half the time by construction. And the pooled n*=70
mis-credited EVERY stat -- too low for hits (own 77) and runs (81), too
high for total_bases (60) and rbi (54).

PHASE 1, blind. Per-stat (g, sigma_row): hits -0.01288/0.11251, TB
-0.01380/0.10680, rbi -0.00884/0.06459, runs -0.00902/0.08080. All four
clear z=1.96 at full n, so none is NO-EFFECT. Committed k=1 with per-stat
n* and a binomial cutoff holding FP at 0.004-0.031.

THE FINDING THAT DOMINATES: the test has no power. Against a strong
instability (date-to-date SD equal to the effect) it detects a failure
1.4%-9.3% of the time, and across every k from 1.0 to 2.0 the best any
stat reaches is 0.337. A gate that cannot fail cannot pass, so
LODO_POWER_FLOOR=0.50 makes UNTESTABLE structural -- "could not test" can
never read as "passed".

PHASE 2/3 cold, at each stat's OWN n*:

  hits  5 informative, 0 reversals, cutoff 2, power 0.093  UNTESTABLE
  TB    5 informative, 0 reversals, cutoff 2, power 0.093  UNTESTABLE
  rbi   4 informative, 1 reversal,  cutoff 2, power 0.045  UNTESTABLE
  runs  3 informative, 2 reversals, cutoff 2, power 0.014  UNTESTABLE

Setting the power floor aside entirely, NOT ONE STAT EXCEEDS ITS CUTOFF.

PHASE 4 — rbi's FAIL was false, as the order suspected. So was RUNS' --
which the order did not anticipate, having classified it DATE-DRIVEN on a
244-row reversal; two reversals in three drops does not clear a cutoff of
2. TB's PASS was vacuous: the test could not have failed it. hits' own n*
is LARGER than the pooled one (77 vs 70), and it remains untestable.

PHASE 5 — deploy basis is now the date-clustered CI alone:

  hits  CI [-0.0139,-0.0097], 4 date clusters   relabelled ci_only
  TB    CI [-0.0061,-0.0045], 2 date clusters   RELABELLED, kept
  rbi   CI [-0.0092,-0.0010], 2 date clusters   NEWLY DEPLOYED
  runs  no fittable map at its split            REFUSE, no CI either

Every deployed stat carries calibration_basis ci_only_lodo_untestable and
auto-demotion is the SOLE stability guard, not a backstop to a passed
test. Stated plainly: those intervals rest on 2-4 date clusters, which is
thin, and it is now the only support. rbi gains chainAcross stackability;
its bands rebuilt on p_win_calibrated (425 rows) are every-archetype
base_rate. runs is queued for the low-param calibrator for the ordinary
reason -- no fittable map -- not on the date-driven finding, which was an
artefact.

PHASE 6 — the deploy set was set by a coin-flip-power ruler; it is now set
by a per-stat power-coherent pre-committed test whose first act was to
report that it cannot evaluate anything. The audit was permitted to wound
the live deploy and did: total_bases lost its LODO claim. Standing
question unchanged -- 18 archetype slots across three deployed stats, every
one a single band indistinguishable from base rate.

Blind ordering held. p_win never mutated. No Bonferroni slot. Counter and
frozen clusters verified file-by-file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 23:01:57 -04:00
builtbykev 1f40014256 Power-derive the LODO threshold: hits restored through the gate, rbi/runs
routed as date-driven

PHASE 0 — threshold derived BLIND, before any stat was re-read. A
reversal is informative only if that date's Brier delta is
distinguishable from zero at its row count. Per-row Brier difference
d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and
n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single
stat's verdict could shape the threshold deciding it:

  pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175
  n* = (0.09816/0.01175)^2 = 69.8 -> 70

The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect
this removes, and why the previous verdict moved with the number.
Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with
LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it
equals the constant, so it cannot drift from its own justification. The
derivation script prints no stat verdict, no date and no reversal.

PHASE 1 — LODO at n*, applied cold:

  hits         5 informative drops, 0 reversals   PASS
  total_bases  4 informative drops, 0 reversals   PASS
  rbi          reverses 2026-08-01 (n=99)         FAIL
  runs         reverses 08-01 (n=86), 08-05 (244) FAIL

hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across
123-272 row dates, favourite sign holding on every testable drop. THIS IS
THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the
withdrawal at 6ae11f1 was correct on the instrument available then, which
admitted 20- and 25-row dates as evidence. Nothing about hits changed;
the threshold stopped being chosen.

PHASE 2 — both failures are DATE-DRIVEN, not underpowered. Every
reversal sits above n*=70 (99, 86, 244), so no threshold and no further
accrual rescues either: isotonic is fitting day-structure. Routed to the
low-parameter calibrator queue (Platt/beta), not built here.

PHASE 3 — CALIBRATION_DEPLOYED is now ['hits','total_bases'], frozen and
tested, both PROVISIONAL with auto-demotion armed and the >=40
date-cluster promotion bar unchanged. hits stackability for
chain.chainAcross is RESTORED, and the record shows it returned through
the powered gate rather than by fiat. hits bands rebuilt on
p_win_calibrated (765 eval rows): every archetype still one band, still
base_rate -- calibrated YES, proven-per-archetype NO.

PHASE 4 logged: the deploy set is now set by a power-derived,
pre-committed, tested constant rather than an operator-chosen number. At
6ae11f1 that rule moved the live path AGAINST the operator; it has now
moved it back on the same evidence because the instrument changed. Both
directions are the rule working. And calibrated p_win separates within
archetype no better than raw across 13 archetype slots on two deployed
stats -- per-archetype separation will come from proven factors or not at
all.

p_win never mutated; no Bonferroni slot consumed; counter and frozen
clusters verified byte-identical file by file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 20:15:32 -04:00
builtbykev 6ae11f1193 LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration
layer without challenging the binding. That floor is a cluster-robust
interval bar for a CAUSAL claim. Calibration makes no causal claim, has a
bounded failure mode (it can only over- or under-shrink) and consumes no
Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and
leave-one-date-out tests that directly -- a STRICTER bar, since a cluster
count cannot detect a single day carrying the effect. The >=40 floor is
retained, correctly scoped as the PROMOTION bar.

PHASE 1 — both guards codified, 11 tests, green before Phase 2.
Demonstrated on live data: raw population violated=true, mean_p 0.4962,
both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The
null guard's test demonstrates the trap explicitly, since (null-1)**2 is
1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate.

PHASE 2 — LODO:

  hits         n=1140 dates=17  2 reversals (07-22 n=20, 07-26 n=25)  FAIL
  total_bases  n=1050 dates=7   0 reversals, 0 sign flips             PASS
  rbi          n= 630 dates=5   1 reversal  (08-01 n=99)              FAIL
  runs         n= 597 dates=5   2 reversals (08-01 n=86, 08-05 n=244) FAIL

Threshold sensitivity reported because the verdict moves: total_bases
passes at every held-size threshold, runs fails at every one, and hits
fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20
before seeing which stats passed and did not move it afterwards to
preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a
standard error several times the effect, so the instrument is
underpowered per-drop -- an argument for pre-registering a higher
threshold, which is a Roundtable call, not one to make while holding the
results.

PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and
runs REFUSE.

HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService
hardcoded it since S91; it fails LODO, so it is out. A stat that cannot
survive dropping one day was never calibrated, it was fitted to that day.
The consequence is real -- hits props become unstackable for
chain.chainAcross -- and it errs toward withdrawing a claim rather than
preserving one on a fragile verdict. Deployment is now driven by a frozen,
tested CALIBRATION_DEPLOYED set, not a hardcoded stat name.

PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither
waivable. reverify auto-demotes on the first breach (CI stops excluding
zero, or the favourite bias flips sign) and logs the breaking date.
Promotion needs the original >=40 bar. A provisional deploy that cannot be
taken away is just a deploy.

PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The
two-bar rule still bites: calibrated YES, proven NO, so they stay a
base-rate read, now honestly numbered. Every archetype still collapses to
one band -- calibrated p_win separates within archetype no better than raw.

PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI,
and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did
not survive); the refused set is a map of missing inputs; a low-parameter
calibrator is queued unbuilt.

p_win never mutated; calibration rides as p_win_calibrated with
calibration_status provisional. No Bonferroni slot consumed. Counter and
frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 18:31:19 -04:00
builtbykev f976df47b8 Settle model_snapshots + four-stat calibration: works, deploys nowhere
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.

PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.

PHASE 1, two integrity findings both caught by the gate:

1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').

2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.

THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.

THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.

PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):

  hits  n=1140  bias +0.0868  brier 0.2626 -> 0.2511  d -0.0115  CI [-0.0139,-0.0097]
  TB    n=1050  bias +0.0834  brier 0.2490 -> 0.2438  d -0.0052  CI [-0.0061,-0.0045]
  rbi   n= 630  bias +0.0164  brier 0.2011 -> 0.1965  d -0.0046  CI [-0.0092,-0.0010]
  runs  n= 597  bias +0.0410  no map fittable (173 fit rows < 200)

ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.

A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.

PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.

PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.

PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.

p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 15:30:30 -04:00
builtbykev 23d1b13176 runs + RBI: mostly-base-rate confirmed, and one level deeper than expected
Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest
grade is the POOLED base rate.

PREMISE NOTE: the order's closing line says the batter board is
per-archetype-graded after this. Nothing has been rescaled for hits or
total_bases either -- no archetype slot has ever reached sample and
gradeBands remains built, gated and unwired. This is the fourth stat
measured, not the completion of three.

AUDIT: RBI 935 clean / 43 games; RUNS 617 clean / 33 games. Zero
quarantined. No archetype slot reaches 500 -- and the signature
archetypes the order names are the two SMALLEST slots on the board,
RBI->DRIVER at n=24 and runs->CATALYST at n=9. RUNS is refused
structurally before any factor is tested: 33 game clusters against a 40
floor.

INPUTS RECONSTRUCTED rather than declared missing. lineup_context only
covers 08-04 onward while settled rows start 07-31, so 187/617 runs rows
joined. But the play-by-play cache runs from 05-01 and the batting order
IS the order batters first appear -- slot, power-behind and reach-base
all rebuilt point-in-time, coverage 187 -> 574.

RBI, all THEATER: risp_opportunity +0.0047, extra_base_skill +0.0010,
risp x extra_base +0.0056. RUNS, all refused on clusters and all pointing
the wrong way: +0.0043 / +0.0008 / +0.0054.

THE COMPOUND IS THE WORST VERSION IN BOTH STATS. The causally-correct
compound was the most promising factor on the sheet and is the most
harmful in each. Two multipliers that individually carry nothing do not
cancel -- they compound each other's noise. Distinct from the
collapsed-sequence lesson: there the product of two REAL effects was too
small to use; here the product of two NULL effects is worse than either.

THE ARCHETYPE DOES NOT RESCUE IT, and this is where the session nearly
went wrong. The base rates look strongly differentiated (RBI DRIVER 0.609
vs BOMBER 0.413; runs GHOST 0.716 vs BOMBER 0.460). Gated directly
against the pooled base rate: RBI +0.0010 CI [-0.0034,+0.0050] THEATER;
runs -0.0028 CI [-0.0147,+0.0108] candidate at k=33. DRIVER's 0.609 is
n=23 -- small-slot noise wearing a decimal point. Read off the table
instead of gated, this would have shipped as "archetype differentiation
is real and large". It is not.

THE CROSS-STAT PATTERN THAT IS REAL -- the counter over-predicts every
batter counting stat measured:

  total_bases  p_win 0.5698 vs actual 0.5074  bias +0.0624
  rbi          p_win 0.4860 vs actual 0.4313  bias +0.0547
  runs         p_win 0.5949 vs actual 0.5749  bias +0.0200

Across four stats and three sessions, calibration is the systematic
defect and factor scarcity is not. TB's held-out isotonic fix (-0.0039)
still outperforms every factor tried on any stat, all null or theatre.

NO RESCALE. Nothing proved, nothing certified calibrated, no slot at
sample, and for RBI the archetype split is itself theatre -- so the
honest band is the pooled base rate, which gradeBands returns by
construction.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 14:42:33 -04:00
builtbykev 6a327d9114 total_bases: every power factor is THEATER, and a units bug nearly hid it
PREMISE CORRECTION: the per-archetype rescale is not "proven and live on
hits". gradeBands was built, gated and explicitly NOT wired two orders
ago -- no hits archetype slot reached sample, every band came back
base-rate, and only defense_by_direction proved pooled. This applies an
unvalidated-at-archetype-level method to a second stat.

FULL-HISTORY AUDIT: 988 clean settled TB rows (101 quarantined, 948 with
p_win), 341 players, and only 9 DISTINCT GAME DATES. No archetype slot
reaches 500 -- BOMBER 340, GHOST 147, BRUSH 55. Confirmed short on full
history, not a windowed artifact. The 9-date figure matters more than the
row count: ~49 games means any game- or venue-borne factor has almost no
replication here.

THE BASELINE HAD TO CHANGE, to a harder null. TB lines vary (1.5 on 559
rows, 0.5 on 345), so a per-line personal base rate would rest on ~2 rows
per player-line and would have to be invented. The null is the counter's
own p_win, which already prices the line -- beating the champion, not
beating "he's due".

THE UNITS BUG, caught, and it had produced the best result in the
programme. The first run reported barrel_rate at Brier -0.0095, the
largest improvement ever measured here. fromStatcastRow returns
barrel_pct as a FRACTION (0.06) while the raw table stores 0-100, so
(0.06 - 7.8) * 0.018 clamped EVERY row to the maximum negative shift.
That uniform downward push "improved" Brier purely by leaning on the
counter's over-prediction and contained no barrel information at all.
Same family as the S80 trap, inverted. exit_velo was a second bug -- the
column is avg_exit_velo, so it read null on every row and reported n=0. A
zero is a wiring bug until proven an honest absence.

GATE with units fixed, 138 cumulative tests:

  barrel_rate           n=707  shift 0.0364  brier +0.0036  THEATER
  exit_velo             n=707  shift 0.0229  brier +0.0022  THEATER
  hard_contact_allowed  n=707  shift 0.0260  brier +0.0033  THEATER
  park_weather_hit_type n=651  36 entities   PENDING (k<40)
  platoon_severity      n=481  PENDING (n<500)

THE PREDICTED INVERSION WENT THE OTHER WAY. BOMBER x barrel_rate is
+0.0114, the single most harmful cell in the table, exactly where the
strongest proof was predicted. GHOST +0.0012. All sample-blocked so not a
verdict, but recorded so it is not claimed later.

AND IT IS NOT DOUBLE-COUNTING -- tested and refuted: corr(barrel, p_win)
= -0.061, the counter is not pricing barrel at all. The duller answer is
corr(barrel, counter RESIDUAL) = -0.012. Barrel is a real skill that
carries no information about what the counter gets wrong at this line.
That also closes the S81 lead: hard_hit r=0.153 at n=295 drifted to 0.135
at n=383 and is THEATER at n=707.

THE REAL FINDING: TB is miscalibrated, not under-factored. mean p_win
0.5698 vs actual 0.5074, bias +0.0624. Held out on a strict time split
(fit < 2026-08-02, eval 651 unseen rows): raw 0.25007, constant de-bias
0.24740 (-0.00267), isotonic 0.24621 (-0.00386). Worth more than any
factor tested and the only intervention pointing the right way -- and
still refused at the corrected bar on 32 clusters. A CANDIDATE, not a
result. It also explains the units bug's fake success exactly: a blanket
downward shift is a crude de-bias.

NO RESCALE. Nothing proved, nothing certified calibrated, no slot at
sample -- every band would be the honest base-rate band gradeBands
already returns by construction.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 03:13:35 -04:00
builtbykev 8ab6557faa Collapsed sequence edge: two proven links whose product is too small to use
Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.

FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.

TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.

GATE on the concentrated subset, 114 cumulative tests:

  early-exit x WEAK pen    n=1931  brier -0.0001  CI [-0.0014,+0.0010]  NOT_PROVEN
  early-exit x STRONG pen  n=2574  brier  0.0000  CI [-0.0011,+0.0010]  THEATER
  all early-exit later ABs n=6869  brier -0.0001  CI [-0.0007,+0.0005]  NOT_PROVEN
  pooled all later ABs    n=17891  brier  0.0000  CI [-0.0004,+0.0003]  THEATER

Not pooled-diluted -- the concentrated subset was gated alone and is no
better.

THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:

  P(faces pen | early-exit flagged)  0.8075
  P(faces pen | starter goes deep)   0.7149
    exposure the flag actually buys  0.0925
  hit-rate swing across pen quality  0.0394
  MAX JUSTIFIABLE ADJUSTMENT         0.00365
  actually applied                   0.01930   -> 5.3x over-movement

A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.

A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.

NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.

THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.

Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 02:43:23 -04:00
builtbykev b2e4c6c4fb Link 2 at the coarse grain: pen QUALITY proves, archetype does not
The refinement was right. Naming the individual reliever failed; the same
question at the grain the chain needs passes, and it transmits more than
anything else measured in this chain.

WHY IT WAS WORTH RE-ASKING: last session's null (the pen is on average no
softer, +0.0010 on 35,760 PAs) does NOT rule this out, and treating it as
though it did would have been the error. An average washing out is fully
consistent with quality VARIATION mattering. It does -- actual arm quality
moves the hit rate monotonically across quartiles, 0.2244 / 0.2293 /
0.2410 / 0.2501, a 2.57pp spread, larger than the whole times-through-
the-order effect.

CLUSTER UNIT CORRECTED, THEN CHECKED RATHER THAN ARGUED. Last session
refused Link 2 partly as team-borne (30 bullpens, the park ceiling). My
first re-check was that 76% of pen-quality variance is within-team -- but
that is a statement about TREATMENT variance, not about where errors
correlate, and stopping there would have been picking the convenient
answer. Measured the actual thing: ICC of prediction error by team =
0.0261, design effect 1.41, SEs inflated ~19%. So the verdict was run
three ways:

  unclustered            CI [-0.0067,-0.0010]  excludes zero
  team-clustered (30)    CI [-0.0086,-0.0003]  excludes zero (below the
                         40-cluster floor -- indicative, not a pass)
  design-effect adjusted CI [-0.0072,-0.0005]  excludes zero

QUALITY GRAIN PROVES on the concentrated elevated-early-exit subset:
n=501 team-games, 426 clusters, MAE 0.0294 -> 0.0260, delta -0.0034, CI
[-0.0063,-0.0005] at 110 cumulative tests. Pooled also proves, so it is
not a subset artefact.

ARCHETYPE GRAIN DOES NOT: 0.5669 vs a 0.5309 modal-guess baseline,
corrected interval [-0.1073,+0.0268] spans zero. Two grains tested, one
earned a place -- penQuality.js exposes no archetype and a test asserts
it.

WHAT LINK 3 RECEIVES, which is the number that actually matters -- not
the MAE gain but realized outcome separation, prediction strictly
point-in-time:

  predicted BEST pen   167 games  2,044 PAs  hit rate 0.2231 +/-0.0180
  predicted WORST pen  167 games  1,799 PAs  hit rate 0.2501 +/-0.0200

2.70pp separated, intervals non-overlapping, capturing nearly all the
2.57pp available at the quartile grain. Caveat stated not buried: the
tercile cut is chosen in-sample; the prediction driving it is not.

BUILT: penQuality.js + 9 tests. Abstains below 5 prior club games and 40
arm appearances -- a league-average stand-in would assert "this is an
ordinary bullpen", which is a claim, and usually the wrong one for exactly
the clubs whose pens just turned over.

Link 3 is unblocked on a proven Link 2 at the quality grain only. Not run
here; this order scopes to building and gating Link 2.

Parallel track logged unchanged: TB n=948 pooled, BOMBER x TB 340, short
by 160.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 02:17:00 -04:00
builtbykev e4dae0e6b0 Reliever chain: Link 1 proves, Link 2 does not, and the premise inverts
The causal insight is right -- the game is a sequence and the matchup does
shift mid-game. The direction is backwards, measured on 93,663 plate
appearances from 1,238 games pulled free from statsapi.

LINK 1 PROVES. Starter batters-faced, point-in-time from his own prior
starts only, clustered on the pitcher: MAE 3.2226 -> 2.7990, delta
-0.4236, CI [-0.6006,-0.2731] at 0.9995 corrected for 107 tests, 1,706
starts across 204 pitchers. It finds the tail the chain needed -- early
exits are a 23.2% base rate, model-flagged starts are 34.0% early, lift
+10.8pp.

Scope correction inside Link 1: the order specifies fatigue x GAME
SCRIPT, but game script is not available at grade time -- whether he gets
hit tonight is the thing being projected, not an input to it. Only the
workload half is measured; the in-game half is recorded as a live feature,
out of scope, rather than quietly folded in.

LINK 2 DOES NOT PROVE, twice over. Model accuracy 17.2% vs an 8.6%
baseline -- doubling it sounds good and is not, since naming a specific
arm is wrong five times in six. And structurally the entity is the
BULLPEN: 39,629 post-starter plate appearances across 30 clubs is 30
readings, below the 40-cluster floor, the same permanent ceiling as park
geometry and team defence. LINK 3 NOT RUN, per the order's own rule.

THE PREMISE IS REFUTED, and this chains on nothing so it was safe to
measure:

  vs STARTER  n=48,492  hit rate 0.2444 +/-0.0038
  vs BULLPEN  n=35,760  hit rate 0.2373 +/-0.0044

The pen is 0.7pp HARDER. The specific effect the chain exists to exploit
-- early exit making later at-bats softer -- is +0.0010 on 35,760 PAs. A
well-powered null, not a sample problem.

What IS real is times through the order: TTO1 0.2351 -> TTO2 0.2515 ->
TTO3 0.2518. A starter does decay as the lineup sees him again, but that
advantage is SURRENDERED when he leaves, not extended -- the pen is
harder than his second and third time through. A modern bullpen is a
queue of fresh specialists throwing one inning each; there is no tiring
arm to punish.

So the insight survives inverted, and Link 1 stays valuable for the
opposite reason it was built: a likely early hook predicts the hitter
LOSES his third-time-through look (0.2518 -> 0.2373 on that PA). The
mispricing is on hitters who get an EXTRA look at a starter going deep.

BUILT: predictionGate.js + tests -- the two-part gate for a continuous
prediction. factorGate binarises outcomes for Brier, which would destroy
a target like batters faced. Same discipline, same THEATER verdict, real
scale.

PRE-REGISTERED NOT RUN: Link 2' using a PA-weighted bullpen AGGREGATE
rather than a named arm. Recorded rather than substituted in -- running
Link 3 on a swapped-in Link 2 is the assumed-link failure the order
forbids. Given the premise result its expected value is now low.

PARALLEL TRACK logged: total_bases n=948 pooled, BOMBER x TB 340, short
by 160. Sample-readiness only, not a verdict.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 01:58:38 -04:00
builtbykev 3081c92e00 Per-archetype grade bands: built, gated, and the rescale blocked twice
The premise does not hold. proven-status.js run fresh: PROVEN_SET is
EMPTY, no archetype x stat reaches the gate. pitcher_contact_profile has
a CI upper bound of exactly 0.0000 and platoon_severity is held on
4.5%-contaminated splits, so the proven set is one factor, pooled, not
three archetype-conditioned ones. The specific pattern the order names --
defense strong for GHOST/BRUSH, null for BOMBER -- is the one I measured
running the OTHER WAY yesterday, both noise-dominated.

But the second blocker is new and matters more, because it would stop the
rescale even if the factors had proved: the grade does not separate
within any archetype. Every archetype collapses to ONE band at the
corrected bar, because bands merge when their intervals overlap and
publishing two letters we cannot tell apart is a distinction we have not
measured.

Uncorrected, so the ranking is visible rather than hidden by the bar,
this INVERTS the order's design. The order gives contact types the
factor-rich treatment and power types honest base-rate, reasoning that
single-game hits are variance for a power profile. Measured:

  BOMBER n=466  corr(p_win,outcome) +0.207  quintiles 0.75 0.62 0.60 0.48 0.48
  GHOST  n=192  corr(p_win,outcome) -0.007  quintiles 0.47 0.63 0.74 0.58 0.45

BOMBER is the one archetype the model ranks, and it splits into a real
A 0.660 / B 0.481 at 95%. GHOST is flat, and non-monotone -- its most
confident reads hit 47% while its middle reads hit 74%. Shipping as
specified would have given the factor-rich treatment to the archetype the
model reads worst and left base-rate on the one it reads best. That is
mechanically sensible in hindsight: a power hitter's hit tracks whether
he can damage the arm, a contact hitter's depends on balls finding holes.

BOMBER's split does not survive the cumulative correction at 106 tests.
Exposing it by loosening the correction is the curve-to-make-A's the
order forbids, so it stays one band.

BUILT: gradeBands.js -- lift against the archetype's OWN base rate (the
same 62% is lift for a 45% profile and a deficit for a 68% one),
indistinguishable neighbours merged, thin bands PROVISIONAL not dropped,
Wilson intervals widened by the cumulative correction. The two-bar rule
is structural: proven-alone, calibrated-alone and neither all return
base_rate with the reason stated, so with nothing proven no
factor-informed band can be produced at all.

reasoning() is built and tested but NOT wired to the card -- there is no
per-archetype band being served, so attaching the copy now would ship
product language for a rescale that does not exist.

NOT BUILT: the specified power-type reason "the matchup edge is in
total_bases". total_bases is recorded INCONCLUSIVE (+0.0038, CI
[-0.068,+0.075]). Wiring it would assert an edge measured as
indistinguishable from zero -- the exact fabricated-reason failure this
module exists to prevent.

BOMBER x hits is 29 rows short of the gate and is the archetype the model
actually reads. That is the first slot to test, not GHOST.

Counter and frozen clusters byte-identical. No letter was moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 01:03:10 -04:00
builtbykev 6b17f79367 Per-archetype re-audit: no slot reaches 500, and the replication unit
decided everything

The premise does not hold. prove-hit-factors.js has no date filter
anywhere in it and pages the full table -- there was never a window to
widen. Full clean history is 1,266 rows, not 2,715. platoon was not
"proved" last session, it was explicitly held on 4.5%-median-contaminated
season-to-date splits, and pitcher_contact_profile was demoted. The
proven set going in was one factor, not three.

STEP 1: no archetype slot reaches n>=500 on full history. Best is BOMBER
at 408, and BOMBER is the most common archetype on the board. GHOST 173,
BRUSH 64, DRIVER 43, CATALYST 16. These are confirmed genuinely short,
not artifacts.

STEP 2 is where the real finding is. park_hits initially PROVED at 619
rows across 45 games -- but those games only ever visited 14 distinct
park values. A park effect is replicated across parks, and unmodelled
park heterogeneity is confounded with the thing being estimated. Each
factor is now clustered on the coarser of the game and the entity its
treatment rides on.

That flipped two verdicts and confirms Kev's causal-correctness thesis
from a new direction: defense_by_direction has 442 hitter-team units of
replication where crude team defense has 26. The correct atom is not just
more accurate, it is the only one measurable at all. park_hits (14) and
defense (26) can never be validated however long the ledger runs -- the
same ceiling as park dimensions, reached independently.

Also fixed a bar I got wrong last session: I transplanted the 500-row
floor onto clusters, which refused a factor with 1,059 rows over 85 games
while answering neither question. Two floors now -- rows>=500 for a stable
estimate, clusters>=40 for a trustworthy interval. Not a lowered bar:
park_hits and defense are still refused.

PROVEN: defense_by_direction only, pooled, [-0.0054,-0.0012] at 99 tests.
It stays POOLED-ONLY -- no per-archetype reasoning wired, nothing
grandfathered. The card must not say "GHOST: defence matchup strong"
because we have not earned that sentence. The predicted fingerprint did
not appear either: BOMBER -0.0036 vs GHOST -0.0024, the opposite
direction, both noise-dominated. Recorded so it is not claimed later.

RESCALE: NOT READY. One proven factor worth -0.0031 Brier. Rescaling on
that is relabelling.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-05 19:58:37 -04:00
builtbykev b818626870 BUILD-STATE: under-querying vs out-of-data session
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-05 19:30:56 -04:00
builtbykev 7b85934dc3 Under-querying vs out of data: the answer depends on the unit
The platoon test's n=452 described how much of the JOIN survived, not how
much data exists. There are 1,266 clean settled hits rows and zero
quarantined ones. platoon_splits had been ingested from tonight's lineups
only (315 players), so any hitter who settled a prop without appearing in
an ingest-day lineup was silently absent from every test.

Backfilled all 380 hitters (81 fetched, 0 unresolved). Re-ran on 1,059
rows, up from 452.

THE DEMOTION IS THE HEADLINE. pitcher_contact_profile, the strongest
proven factor in the programme (-0.0064, CI [-0.0113,-0.0014]), roughly
halved to -0.0034 on more than double the sample and its corrected
interval now spans zero. The Bonferroni denominator also rose to 55,
which widens every interval -- but a denominator cannot move a point
estimate, and that halved on its own.

platoon and platoon_severity now clear the bar and are NOT promoted.
Upper bound -0.0001, on season-to-date splits that contain the games they
predict: measured contamination is 4.5% median, 12.4% at p90, 137% worst.
I had assumed ~1%. They stay CANDIDATE pending point-in-time splits.

GAME-LEVEL IS A DIFFERENT PROBLEM. game_context held zero weather rows
ever -- not because the fetcher was wrong (it correctly targets
Open-Meteo's archive) but because ledger_entries keys a game as
mlb:2026-08-03:Away@Home and game_context keys it as mlb:823437. Every
lookup missed and NULL columns read as honest absence. Third occurrence
of that class.

Fixed the join: 96/101 settled games now carry actual archived weather,
park dimensions backfilled 15 -> 30 venues.

But 928 total_bases rows sit on 47 games at 17.6 rows per game. Park and
weather assign one value per game, so resampling rows would have
manufactured a pass. factorGate now resamples clusters when rows carry
one and judges sample against effective_n; unclustered rows keep the
original path byte-for-byte. Verdict: 47 clusters < 500, and the point
estimate is +0.0011 -- worse, not merely unproven.

Weather needs ~57 more days. Park dimensions need never: there are 30
ballparks in MLB, so a venue-constant factor can never reach 500
independent units. That bar was built for player-level factors and does
not transfer.

Wind is refused. We have speed and bearing for all 96 games; we lack park
orientation, and 220 degrees is blowing out at one park and in at
another. Using speed alone would assert an effect while discarding the
sign that decides what it is.

Counter and frozen clusters untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-05 19:30:17 -04:00
129 changed files with 14474 additions and 89 deletions
+4
View File
@@ -27,3 +27,7 @@ out/
# Vercel
.vercel/
.seq-cache/
.content-out/
+37
View File
@@ -5133,3 +5133,40 @@ Complete frontend overhaul. 18 pages, 22 API routes. `npm run build` passes with
- Timestamped records in evolution_detections table, Evolution Watch content formatter
- **Migration 008:** coaching_tendencies, player_out_history, evolution_detections, unconventional_validations (all with indexes + RLS)
- **Integration:** 3 new blueprints registered in app.py (coaching_bp, redistribution_bp, unconventional_bp), evolution + odds_scanner extended with new endpoints
---
## Session — Under-querying vs out of data (2026-08-05)
**Shipped**
- `scripts/backfill-context.js` — platoon splits backfilled to all 380 settled
hitters (was 298; ingest had only ever seen tonight's lineups).
- `scripts/reconstruct-game-environment.js` — joins the ledger's game slug to
statsapi, writes `game_context` on the LEDGER's key, pulls actual archived
Open-Meteo weather. 96/101 settled games now carry real weather; park
dimensions 15 -> 30 venues.
- `src/services/model/parkWeather.js` (+ tests) — park geometry + air read onto
HIT TYPE, not P(hit). Wind refused (no park orientation).
- `factorGate` — cluster-aware bootstrap + `effective_n`. Unclustered rows keep
the original path byte-for-byte.
- `specs/under-querying-vs-out-of-data.md` — the full record.
**Decided**
- `pitcher_contact_profile` DEMOTED: point estimate halved on 2x sample, interval
now spans zero.
- `platoon` / `platoon_severity` clear the bar but are NOT promoted — 4.5% median
contamination (season-to-date splits contain the games they predict) on a
-0.0001 bound.
- Park+weather on total_bases: 47 clusters < 500, point estimate WORSE (+0.0011).
**Next**
- Weather is reachable: ~57 days at 7 games/settled-day.
- Park geometry is NOT — 30 ballparks exist, so a venue-constant factor can never
reach 500 independent units. Needs a hierarchical fixed-effect treatment or it
stays unvalidatable.
- Point-in-time platoon splits would settle the two held passes.
- Park orientation is the one column that would unlock wind.
**Blocker**
- `git push` has no credentials in this environment; commit `7b85934` is local.
+65
View File
@@ -0,0 +1,65 @@
# `/api/content-studio` — the agent-ready contract
The endpoint Kev's `/studio` page reads today and an autonomous poster reads
later. **The page is a thin client**: no posting logic, no fact handling. Wiring
a bot means pointing it here — nothing on this side changes.
Distinct from `/api/content` (Session 29), which serves structured content
*objects* by data level. This serves finished **posts**.
## Auth
`x-internal-key: $VYNDR_INTERNAL_KEY` — private, never public. The browser never
holds the key; the Next proxy at `web/src/app/api/content-studio/[...path]`
attaches it server-side.
## `GET /api/content-studio/:date?`
`:date` optional, defaults to today ET.
```jsonc
{
"date": "2026-08-07",
"count": 3,
"posts": [{
"id": "honesty_flex",
"label": "The Honesty Flex",
"sport": "mlb",
"status": "pending", // pending | approved | skipped | regenerate_requested
"ok": true,
"skipped": false,
"reason": null, // why it was skipped, when it was
"honest_absence": false, // a real "nothing tonight" post, not a failure
"copy": "WE GRADED 2140 PROPS TONIGHT...",
"card": { "title": "...", "lines": [...] },
"card_svg": "<svg ...>", // ready to render or rasterise
"fact_contract": ["graded", "ceiling_letter", "..."], // REQUIRED fields
"facts": { "graded": 2140, "...": "..." } // what backed it
}]
}
```
**`fact_contract` + `facts` are the point.** An agent (or a reviewer) can check
what a claim rests on instead of trusting the sentence. A post whose contract
could not be met never appears with invented values — it arrives `skipped` with
a `reason`, or as an `honest_absence`.
## `POST /api/content-studio/:date/:id/status`
```jsonc
{ "status": "approved" } // approved | skipped | regenerate_requested
```
Editorial state only, stored in Redis for 14 days. **Approving a post changes
nothing about the model** — status never touches a serving, model or ledger
table.
## For the agent build
1. `GET` the date → filter `ok && !skipped`.
2. Post `copy`; rasterise or attach `card_svg`.
3. `POST` status `approved` on success.
4. **Never** synthesise a claim not present in `facts`. The engine refuses to
render an unbacked token; an agent must not reintroduce one downstream.
Adding a template changes the payload not at all — a new `id` simply appears.
+64
View File
@@ -0,0 +1,64 @@
# git push in this environment — it was never the firewall, and never missing credentials
## What actually happened
Every push attempt this session used `git push origin main` and failed with:
```
fatal: could not read Username for 'https://github.com'
```
That error names the cause exactly, and it was misread all session as "no git
credentials on this machine." Two things were true instead:
1. **`origin` is GitHub** (`github.com/kev3109/betonblk.git`) and has **no stored
credential.**
2. **`gitea` is the working remote** (`git.builtbykev.com/builtbykev/vyndr.git`)
and **a valid credential for it was on the machine the entire time.**
The habit of typing `origin` is what kept twenty commits local. Nothing was
blocked.
## The GATE-0 firewall theory — tested and REJECTED for VYNDR
The theory was that GATE-0 (Hetzner `mastermind-core-fw`, inbound deny-by-default
except 80/443, SSH/22 restricted to Tailscale + Kev's IP) was blocking an
SSH-based push, as it did for COLYRA.
**It does not apply here. Measured:**
| check | result |
|---|---|
| `git remote -v` | **both remotes are already HTTPS** — no `git@…:…` URL anywhere |
| `curl -I https://git.builtbykev.com` | **HTTP 200** in 0.64s |
| `curl -I https://github.com` | **HTTP 200** in 0.17s |
| Gitea git endpoint over 443 | **HTTP 200** |
| GitHub git endpoint over 443 | HTTP 401 (auth required, reachable) |
There was no SSH remote to be blocked and no connectivity failure of any kind.
COLYRA's HTTPS-remote fix was the right fix for COLYRA's problem; **VYNDR was
already in the state that fix produces.**
Applying it here would have meant creating a new Gitea token to solve a problem
that did not exist — and the pre-existing credential would have made the new
token look like the cure.
## The fix
```
git push gitea main # not origin
```
Result: `6452926..ecf78b9`, 21 commits, verified by `git ls-remote` matching
local `HEAD`.
## Standing note
- **`gitea` is VYNDR's push remote.** `origin` (GitHub) is unauthenticated on this
machine and will always fail.
- Read the error text before reaching for an infrastructure theory. `could not
read Username for 'https://github.com'` is a *credential* message naming a
*specific host* — it is not a connectivity message, and it named the wrong
remote, not a wrong protocol.
- The `~/vyndr-full-history-2026-08-07.bundle` and patch series stay as
belt-and-braces. They are no longer the only copy.
+103
View File
@@ -0,0 +1,103 @@
#!/usr/bin/env node
'use strict';
/**
* backfill-context — WERE WE WAITING, OR UNDER-QUERYING?
*
* The platoon test ran on 452 rows against 1,266 clean settled hits rows in the
* ledger, so "48 short of the gate" was never a statement about how much data
* exists. It was a statement about how much the JOIN survived — and the join was
* losing rows to inputs we simply had not fetched for every player.
*
* This backfills the inputs (pure sample, zero waiting) and reports exactly
* where each row is lost, so the next "we need more data" claim is a measured
* one rather than an inherited one.
*
* ── THE ONE HONEST CAVEAT, STATED UP FRONT ───────────────────────────────
* Platoon splits from statsapi are SEASON-TO-DATE as of the moment they are
* fetched. Applying today's split to a 2026-07-15 game means the split contains
* that game. For a ~400-PA season line one game is roughly a quarter of one
* percent, so the contamination is small — but it is real, it runs in the
* flattering direction, and it is why this is labelled a reconstruction rather
* than a clean point-in-time backtest.
*
* SUPABASE_URL=... node scripts/backfill-context.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const ctx = require('../src/services/lineupContextService');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const SEASON = Number(process.env.BF_SEASON || 2026);
const PAGE = 1000;
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
// Every hitter who appears on a CLEAN settled row — the true denominator.
const led = await page(sb, 'ledger_entries', 'player_key, player_name, stat, outcome, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases'])
.in('outcome', ['hit', 'miss']));
const need = new Map();
for (const r of led) {
if ((r.quarantine_reason || '').startsWith('nontakeable_book')) continue;
if (!need.has(r.player_key)) need.set(r.player_key, r.player_name);
}
const have = new Set((await page(sb, 'platoon_splits', 'player_key', (q) => q.eq('sport', 'mlb')))
.map((r) => r.player_key));
const missing = [...need.entries()].filter(([k]) => !have.has(k));
console.error(`[backfill] hitters on clean settled rows: ${need.size}; splits already held: ${have.size}; to fetch: ${missing.length}`);
const asOf = ctx.dateET();
const rows = [];
let unresolved = 0;
for (const [key, name] of missing) {
let found = null;
try { found = await mlb.searchPlayer(name); } catch { found = null; }
if (!found || !found.id) { unresolved += 1; continue; }
const sp = await ctx.fetchPlatoonSplits(found.id, SEASON, {});
if (!sp) continue; // absent, never a symmetric guess
rows.push({ player_key: key, player_name: name, source_id: found.id, ...sp });
}
let written = 0;
for (let i = 0; i < rows.length; i += 200) {
const batch = rows.slice(i, i + 200).map((r) => ({ ...r, sport: 'mlb', season: SEASON, as_of_date: asOf }));
const { error } = await sb.from('platoon_splits')
.upsert(batch, { onConflict: 'as_of_date,sport,season,player_key' });
if (!error) written += batch.length;
else console.error('[backfill] write failed:', error.message);
}
console.log(JSON.stringify({
hitters_on_clean_settled_rows: need.size,
splits_held_before: have.size,
attempted: missing.length,
unresolved_by_name: unresolved,
no_splits_available: missing.length - unresolved - rows.length,
written,
caveat: 'season-to-date splits applied to past games contain those games — small (~0.25% of a 400-PA line) but real and flattering',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+100
View File
@@ -0,0 +1,100 @@
#!/usr/bin/env node
'use strict';
/**
* build-grade-bands — publish what each letter actually means, per archetype.
*
* Runs the real settled ledger through gradeBands. Because no factor has passed
* the gate for any archetype, every band comes back a BASE-RATE read — which is
* the honest answer today, and the output states it rather than leaving a reader
* to infer it.
*
* SUPABASE_URL=... node scripts/build-grade-bands.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const gb = require('../src/services/model/gradeBands');
const tl = require('../src/services/model/testLedger');
const cal = require('../src/services/model/calibration');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const STAT = process.env.BAND_STAT || 'hits';
const PAGE = 1000;
/**
* PROVEN, PER ARCHETYPE. Empty, and that is the measured state — see
* specs/per-archetype-re-audit.md. Nothing may be added here that has not
* cleared the gate FOR THAT ARCHETYPE; pooled proof does not qualify a slot.
*/
const PROVEN_BY_ARCHETYPE = Object.freeze({});
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat',
(q) => q.eq('sport', 'mlb').eq('stat', STAT).not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries', 'player_key, game_date, outcome, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', STAT)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
const byArch = new Map();
for (const r of clean) {
const a = String(archOf.get(`${r.player_key}|${r.game_date}`) || 'UNLABELLED').toUpperCase();
if (!byArch.has(a)) byArch.set(a, []);
byArch.get(a).push({ p: knownNumber(r.p_win), won: r.outcome === 'hit' ? 1 : 0 });
}
const mc = await tl.recordAndCount(tl.supabaseStore(sb),
[...byArch.keys()].map((a) => ({
sport: 'mlb', stat: STAT, archetype: a === 'UNLABELLED' ? null : a,
interaction: 'grade_band_lift', target: 'outcome',
})));
// Calibration is measured, not assumed. Today it is certified for hits only in
// a middle band (specs — held-out error 0.477->0.506, 0.587->0.580), which is
// NOT the same as an archetype's probabilities being calibrated.
const certified = typeof cal.certifyBands === 'function';
const out = [];
for (const [arch, rows] of [...byArch.entries()].sort((a, b) => b[1].length - a[1].length)) {
out.push(gb.buildBands(rows, {
archetype: arch,
cumulativeTests: mc.cumulative_tests,
proven: Boolean(PROVEN_BY_ARCHETYPE[arch]),
calibrated: false, // no archetype's distribution is certified calibrated
}));
}
console.log(JSON.stringify({
stat: STAT,
settled_rows: clean.length,
archetypes: byArch.size,
cumulative_tests: mc.cumulative_tests,
calibration_helper_present: certified,
proven_by_archetype: PROVEN_BY_ARCHETYPE,
note: 'every band is a BASE-RATE read — no factor has passed the gate for any archetype',
bands: out,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+267
View File
@@ -0,0 +1,267 @@
#!/usr/bin/env node
'use strict';
/**
* calibrate-four-stats — Phases 2, 3, 4 and 7.
*
* ── ONE SIDE PER PROP, OR THE MEASUREMENT IS MEANINGLESS ─────────────────
* 97.6% of snapshot props carry BOTH the over and the under. Their p_wins sum to
* ~1 and their outcomes are complementary, so any calibration statistic over the
* raw population is pinned to 0.5 by symmetry. Measured that way the counter
* looks perfectly calibrated (+0.0002 on hits); deduped to the model-PICKED side
* it is +0.0868. Same rows, opposite conclusion.
*
* ── DATE-CLUSTERED, PER THE ORDER ────────────────────────────────────────
* A day's offensive environment is a real shared component, so uncertainty is
* clustered on the game DATE rather than the game. That is the honest unit for a
* systematic-bias claim and it is a much harder bar than game-clustering.
*
* SUPABASE_URL=... node scripts/calibrate-four-stats.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** The order's deploy floor: dates, not games. */
const MIN_DATE_CLUSTERS = 40;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
const brier = (ps, ys) => mean(ps.map((p, i) => (p - ys[i]) ** 2));
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select))
.order('id', { ascending: true }).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
const isPreGame = (capturedAt, gameDate) => {
const et = new Date(new Date(capturedAt).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < gameDate || (d === gameDate && et.getUTCHours() < 19);
};
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
/** Paired bootstrap on the Brier difference, resampling DATES. */
function dateClusteredCI(rows, cumulativeTests = 1, iters = 3000) {
const byDate = new Map();
for (const r of rows) {
if (!byDate.has(r.date)) byDate.set(r.date, []);
byDate.get(r.date).push(r);
}
const keys = [...byDate.keys()];
const rnd = makeRnd(20260807);
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const raw = []; const adj = []; const ys = [];
for (let i = 0; i < keys.length; i += 1) {
for (const r of byDate.get(keys[Math.floor(rnd() * keys.length)])) {
raw.push(r.p); adj.push(r.pc); ys.push(r.won);
}
}
diffs.push(brier(adj, ys) - brier(raw, ys));
}
diffs.sort((a, b) => a - b);
const tests = Math.max(1, Math.round(cumulativeTests));
const alpha = 0.05 / tests;
const q = (x) => diffs[Math.floor(Math.min(diffs.length - 1, Math.max(0, x * (diffs.length - 1))))];
return { ci: [round4(q(alpha / 2)), round4(q(1 - alpha / 2))], date_clusters: keys.length, ci_level: round4(1 - alpha) };
}
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused, grade',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
// ── ONE SIDE PER PROP: the side the model picked (its higher p_win). ──
const picked = new Map();
const refusedProps = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date)) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
if (r.refused || knownNumber(r.p_win) === null) {
if (!refusedProps.has(k)) refusedProps.set(k, r);
continue;
}
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
const resolve = (r) => {
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) return null;
const v = knownNumber(FIELD[r.stat](b));
if (v === null) return null;
const over = v > L;
return { over, won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0, realized: v };
};
const out = { deploy_floor_date_clusters: MIN_DATE_CLUSTERS, per_stat: {}, refusal_accuracy: {} };
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const res = resolve(r);
if (!res) continue;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: res.won });
}
rows.sort((a, b) => String(a.date).localeCompare(String(b.date)));
const dates = [...new Set(rows.map((r) => r.date))].sort();
if (rows.length < 100 || dates.length < 3) {
out.per_stat[stat] = { n: rows.length, date_clusters: dates.length, decision: 'REFUSE', reason: 'too few rows or dates to split point-in-time' };
continue;
}
// POINT-IN-TIME: fit strictly on earlier dates, evaluate on later ones.
//
// The cut is placed by ROW COUNT rather than by date index. Props are not
// spread evenly across dates -- hits concentrate in the later ones -- so a
// 60%-of-DATES cut left only 143 rows to fit on, under the 200 the fitter
// needs. Splitting on cumulative rows keeps the split strictly temporal
// (every fit date precedes every eval date) while giving both sides enough
// to work with.
const perDate = new Map();
for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) {
acc += perDate.get(d) || 0;
if (acc >= rows.length * 0.45) { cut = d; break; }
}
const fit = rows.filter((r) => r.date < cut);
const ev = rows.filter((r) => r.date >= cut);
if (fit.length < 50 || ev.length < 50) {
out.per_stat[stat] = { n: rows.length, date_clusters: dates.length, decision: 'REFUSE', reason: 'time split leaves too little on one side' };
continue;
}
const iso = cal.fitIsotonic(fit.map((r) => ({ p: r.p, won: r.won })));
// NULL IS NOT A PREDICTION. fitIsotonic returns null below its minimum and
// applyIsotonic then returns null per row -- and (null - 1)**2 === 1 while
// (null - 0)**2 === 0, so a "Brier score" computed over nulls is silently
// just the win rate. That is exactly the Number(null) === 0 breach this
// codebase keeps having to catch, and it produced a fake 0.5567 for hits.
if (!iso) {
out.per_stat[stat] = {
n: rows.length, date_clusters: dates.length, fit_n: fit.length, eval_n: ev.length,
decision: 'REFUSE', reason: `no calibration map could be fitted on ${fit.length} fit rows`,
};
continue;
}
const scored = ev.map((r) => ({ ...r, pc: cal.applyIsotonic(iso, r.p) }))
.filter((r) => knownNumber(r.pc) !== null);
if (scored.length < 50) {
out.per_stat[stat] = {
n: rows.length, date_clusters: dates.length,
decision: 'REFUSE', reason: `only ${scored.length} eval rows could be mapped`,
};
continue;
}
const ys = scored.map((r) => r.won);
const bRaw = brier(scored.map((r) => r.p), ys);
const bCal = brier(scored.map((r) => r.pc), ys);
const { ci, date_clusters, ci_level } = dateClusteredCI(scored, 1);
// CERTIFIED BAND: p_win deciles where held-out |predicted - actual| is small.
const bands = [];
for (let lo = 0.3; lo < 0.95; lo += 0.1) {
const slice = scored.filter((r) => r.p >= lo && r.p < lo + 0.1);
if (slice.length < 25) continue;
const pred = mean(slice.map((r) => r.pc));
const act = mean(slice.map((r) => r.won));
bands.push({ range: [round2(lo), round2(lo + 0.1)], n: slice.length, calibrated_pred: round4(pred), actual: round4(act), err: round4(pred - act) });
}
const certified = bands.filter((b) => Math.abs(b.err) <= 0.05).map((b) => b.range);
const improves = bCal < bRaw && ci[1] < 0;
const enoughDates = date_clusters >= MIN_DATE_CLUSTERS;
// PHASE 4 — bias SHAPE across the p_win range (diagnostic only).
const shape = [];
for (let lo = 0.3; lo < 0.95; lo += 0.1) {
const slice = rows.filter((r) => r.p >= lo && r.p < lo + 0.1);
if (slice.length < 25) continue;
shape.push({ range: [round2(lo), round2(lo + 0.1)], n: slice.length, bias: round4(mean(slice.map((r) => r.p)) - mean(slice.map((r) => r.won))) });
}
out.per_stat[stat] = {
n: rows.length,
date_clusters: dates.length,
bias_pre: round4(mean(rows.map((r) => r.p)) - mean(rows.map((r) => r.won))),
fit_n: fit.length, eval_n: ev.length, split_at: cut,
brier_raw: round4(bRaw),
brier_calibrated: round4(bCal),
brier_delta: round4(bCal - bRaw),
ci_date_clustered: ci,
ci_level,
eval_date_clusters: date_clusters,
certified_bands: certified,
band_detail: bands,
bias_shape: shape,
decision: improves && enoughDates ? 'DEPLOY' : 'REFUSE',
reason: improves && enoughDates ? 'held-out Brier improves, date-clustered, and the date floor is met'
: (!enoughDates
? `date-clusters ${date_clusters} < ${MIN_DATE_CLUSTERS} — the honest unit for a systematic-bias claim`
: 'held-out Brier does not improve at the date-clustered interval'),
};
}
// ── PHASE 7 — REFUSAL ACCURACY ──
// The model passed on these. A pass is CORRECT when there was genuinely
// nothing to call: the over lands near a coin flip rather than at an
// exploitable rate.
for (const stat of STATS) {
const refs = [];
for (const r of refusedProps.values()) {
if (r.stat !== stat) continue;
const res = resolve({ ...r, side: 'over' });
if (res) refs.push(res.over ? 1 : 0);
}
const graded = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const res = resolve({ ...r, side: 'over' });
if (res) graded.push(res.over ? 1 : 0);
}
out.refusal_accuracy[stat] = {
refused_n: refs.length,
refused_over_rate: refs.length ? round4(mean(refs)) : null,
graded_over_rate: graded.length ? round4(mean(graded)) : null,
refused_distance_from_coinflip: refs.length ? round4(Math.abs(mean(refs) - 0.5)) : null,
graded_distance_from_coinflip: graded.length ? round4(Math.abs(mean(graded) - 0.5)) : null,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round2 = (v) => Math.round(v * 100) / 100;
main().catch((e) => { console.error(e); process.exit(1); });
+174
View File
@@ -0,0 +1,174 @@
#!/usr/bin/env node
'use strict';
/**
* PHASES 0-1 — is the champion really worse than a frequency table?
*
* The prior comparison used a leave-one-out baseline that saw the evaluation
* window. This one does not: for every prop, the naive forecast is that player's
* rate of clearing THAT LINE over games strictly BEFORE that date — the same
* temporal discipline the champion is held to. If the champion still loses, the
* defect is real and not an artefact of the peek.
*
* Then the champion's own knobs are ablated. Its core is
*
* p = 0.6 * season_frequency + 0.4 * last5_frequency
*
* plus a +/-0.03 opponent nudge, a +/-0.015 home nudge, and a cv pull. Each is
* tested for whether it COSTS resolution. This is accounting on the champion's
* existing knobs, not a causal claim, so no Bonferroni slot.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Games a player needs before we will read his own rate at all. */
const MIN_PRIOR_GAMES = 10;
async function page(sb, t, sel, orderBy, apply) {
const out = [];
for (let i = 0; ; i += 1000) {
const { data, error } = await apply(sb.from(t).select(sel)).order(orderBy, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break;
out.push(...data);
if (data.length < 1000) break;
}
return out;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) { let s = seed >>> 0; return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; }; }
function resolutionOf(rows, key) {
const base = mean(rows.map((r) => r.won));
let res = 0;
for (let k = 0; k < 10; k += 1) {
const lo = k / 10; const hi = (k + 1) / 10;
const sl = rows.filter((r) => r[key] >= lo && (hi >= 1 ? r[key] <= 1 : r[key] < hi));
if (!sl.length) continue;
res += (sl.length / rows.length) * (mean(sl.map((x) => x.won)) - base) ** 2;
}
return res;
}
/** Paired date-block bootstrap on a resolution difference (a b). */
function dateBlockResCI(rows, a, b, seed) {
const byDate = new Map();
for (const r of rows) { if (!byDate.has(r.date)) byDate.set(r.date, []); byDate.get(r.date).push(r); }
const keys = [...byDate.keys()]; const rnd = makeRnd(seed); const d = [];
for (let it = 0; it < 3000; it += 1) {
const s = [];
for (let i = 0; i < keys.length; i += 1) s.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
d.push(resolutionOf(s, a) - resolutionOf(s, b));
}
d.sort((x, y) => x - y);
return { ci: [r5(d[Math.floor(d.length * 0.025)]), r5(d[Math.floor(d.length * 0.975)])], date_blocks: keys.length };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
// Per-player, date-ordered history. The ONLY source of the naive forecast.
const hist = new Map();
for (const [k, b] of Object.entries(lines)) {
const [date, key] = k.split('|');
if (!hist.has(key)) hist.set(key, []);
hist.get(key).push({ date, b });
}
for (const v of hist.values()) v.sort((x, y) => x.date.localeCompare(y.date));
const out = {};
for (const stat of STATS) {
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused, features', 'id',
(q) => q.eq('sport', 'mlb').eq('stat', stat));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({ propKey: [r.game_date, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const rows = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const isUnder = String(r.side).toLowerCase() === 'under';
const won = (isUnder ? !(v > L) : (v > L)) ? 1 : 0;
// ── THE FAIR COMPETITOR: strictly prior games only. ──
const prior = (hist.get(r.player_key) || []).filter((g) => g.date < r.game_date);
if (prior.length < MIN_PRIOR_GAMES) continue;
const vals = prior.map((g) => knownNumber(FIELD[stat](g.b))).filter((x) => x !== null);
if (vals.length < MIN_PRIOR_GAMES) continue;
const season = vals.filter((x) => x > L).length / vals.length;
const last5 = vals.slice(-5);
const recent = last5.filter((x) => x > L).length / last5.length;
const f = r.features || {};
const homeAdj = f.home_away === 1.0 ? 0.015 : f.home_away === 0.0 ? -0.015 : 0;
const oppR = knownNumber(f.opp_rank_stat);
const oppAdj = oppR === null ? 0 : (oppR >= 0.70 ? 0.03 : (oppR <= 0.30 ? -0.03 : 0));
const flip = (p) => Math.max(0.01, Math.min(0.99, isUnder ? 1 - p : p));
const blend = (w) => flip(0.6 === null ? season : (1 - w) * season + w * recent);
rows.push({
date: r.game_date, won,
champion: knownNumber(r.p_win),
// Reconstructions, all point-in-time.
season_only: flip(season),
w40: flip(0.6 * season + 0.4 * recent), // the current blend
w20: flip(0.8 * season + 0.2 * recent),
w60: flip(0.4 * season + 0.6 * recent),
w40_nudged: flip(Math.max(0.01, Math.min(0.99, 0.6 * season + 0.4 * recent + oppAdj + homeAdj))),
season_nudged: flip(Math.max(0.01, Math.min(0.99, season + oppAdj + homeAdj))),
});
}
if (rows.length < 100) { out[stat] = { n: rows.length, note: 'too few rows with 10+ prior games' }; continue; }
// The REPAIRED champion: full-season window + recency weight 0.20, which is
// exactly what the code change produces.
for (const r of rows) r.repaired = r.w20;
const keys = ['champion', 'repaired', 'season_only', 'w20', 'w40', 'w60', 'w40_nudged', 'season_nudged'];
const res = Object.fromEntries(keys.map((k) => [k, r5(resolutionOf(rows, k))]));
const gap = dateBlockResCI(rows, 'champion', 'season_only', 20260807);
out[stat] = {
n: rows.length,
dates: new Set(rows.map((r) => r.date)).size,
resolution: res,
champion_minus_fair_baseline: r5(res.champion - res.season_only),
ci_champion_minus_fair: gap.ci,
date_blocks: gap.date_blocks,
champion_loses_fairly: res.champion < res.season_only,
best_variant: keys.reduce((a, k) => (res[k] > res[a] ? k : a), keys[0]),
recency_cost: r5(res.w40 - res.season_only),
nudge_cost: r5(res.w40_nudged - res.w40),
REPAIRED_vs_fair: r5(res.repaired - res.season_only),
REPAIRED_ci: dateBlockResCI(rows, 'repaired', 'season_only', 20260807).ci,
REPAIRED_vs_old_champion: r5(res.repaired - res.champion),
REPAIRED_beats_old_ci: dateBlockResCI(rows, 'repaired', 'champion', 20260807).ci,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error('FAILED:', e.message); process.exit(1); });
const r5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
+210
View File
@@ -0,0 +1,210 @@
#!/usr/bin/env node
'use strict';
/**
* THE COLLAPSED SEQUENCE EDGE — Link 1 x pen-season-quality, on later at-bats.
*
* Link 3 is correctly skipped: reliever IDENTITY did not prove and is genuine
* baseball unpredictability. But Link 2's QUALITY grain DID prove, so pen quality
* here is a measured predictor rather than a fallback.
*
* ── THE MECHANICAL CEILING, MEASURED FIRST ───────────────────────────────
* A hitter's third or fourth plate appearance is ALREADY against the bullpen
* 70-73% of the time even when the starter is projected to go deep. An elevated
* early-exit flag lifts that to only 77-83%. So Link 1 buys roughly TEN POINTS
* of extra pen exposure, not a switch from starter to pen — and any adjustment
* built on it is bounded at about a tenth of the starter-versus-pen quality gap.
* That ceiling is a property of baseball, not of the model, and it is the reason
* the deltas below are small before anything is even fitted.
*
* ── WHAT IS ADJUSTED, AND WHAT IS REFUSED ────────────────────────────────
* The order specifies pen-quality x pen-ARCHETYPE x hitter-APPROACH. Two of
* those three cannot be used honestly:
*
* pen archetype did NOT prove (0.5669 vs a 0.5309 modal baseline, corrected
* interval spanning zero). Building it into the adjustment
* would be chaining on an unproven link.
* hitter approach "fastball-hunter" / "finesse-vulnerable" identities do not
* exist in this registry. MLB batter archetypes are BOMBER /
* GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID /
* CATALYST. Inventing an identity to condition on would be
* fabricating the very thing the gate exists to catch.
*
* So the adjustment uses the PROVEN component alone, and a hitter split derived
* from the sequence data itself (power vs contact by home-run rate) is tested as
* a SEPARATE gated addition rather than assumed into the main effect.
*
* node scripts/collapsed-sequence-edge.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const fg = require('../src/services/model/factorGate');
const tl = require('../src/services/model/testLedger');
const pq = require('../src/services/model/penQuality');
const { createClient } = require('@supabase/supabase-js');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const HIT = new Set(['single', 'double', 'triple', 'home_run']);
const PA = new Set(['single', 'double', 'triple', 'home_run', 'field_out', 'strikeout',
'grounded_into_double_play', 'force_out', 'field_error', 'fielders_choice',
'fielders_choice_out', 'double_play', 'sac_fly', 'pop_out', 'line_out', 'fly_out',
'strikeout_double_play']);
const MIN_ARM_PA = 40;
const MIN_PRIOR_GAMES = 5;
const MIN_HITTER_PA = 60;
const MIN_PRIOR_STARTS = 3;
const EARLY_FLAG_BF = 22;
const LEAGUE_BF = 21.56;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
function build() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const arm = new Map();
const bat = new Map(); // hitter -> { n, h, hr }
const penHist = new Map();
const startHist = new Map();
const rows = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const team = g[side].abbr || g[side].team;
const st = (g[side].arms || []).find((a) => a.started);
if (!team || !st) continue;
const half = side === 'home' ? 'top' : 'bottom';
const pas = g.pas.filter((p) => p.half === half && PA.has(p.event));
const ps = startHist.get(st.id) || [];
let predBf = null;
if (ps.length >= MIN_PRIOR_STARTS) {
const w = ps.length / (ps.length + 5);
predBf = w * mean(ps) + (1 - w) * LEAGUE_BF;
}
const hist = penHist.get(team) || [];
const pen = pq.projectPen(hist.map((q) => ({ quality: q })));
const seen = new Map();
for (const p of pas) {
const k = p.batter;
seen.set(k, (seen.get(k) || 0) + 1);
const paNum = seen.get(k);
const b = bat.get(k);
// knownRate abstain: no readable hitter, starter or pen -> no row at all.
if (paNum < 3 || predBf === null || !pen || !b || b.n < MIN_HITTER_PA) continue;
rows.push({
gamePk: g.gamePk,
cluster: g.gamePk,
batter: k,
paNum,
early: predBf <= EARLY_FLAG_BF,
pen_quality: pen.quality,
hitter_base: b.h / b.n,
hitter_hr_rate: b.hr / b.n,
won: HIT.has(p.event) ? 1 : 0,
});
}
const faced = [];
for (const p of pas.filter((x) => x.pitcher !== st.id)) {
const h = arm.get(p.pitcher);
if (h && h.n >= MIN_ARM_PA) faced.push(h.h / h.n);
}
if (faced.length) penHist.set(team, hist.concat([mean(faced)]));
if (st.bf != null) startHist.set(st.id, ps.concat([st.bf]));
for (const p of pas) {
const c = arm.get(p.pitcher) || { n: 0, h: 0, k: 0 };
c.n += 1; c.h += HIT.has(p.event) ? 1 : 0; c.k += p.event === 'strikeout' ? 1 : 0;
arm.set(p.pitcher, c);
}
for (const p of pas) {
const c = bat.get(p.batter) || { n: 0, h: 0, hr: 0 };
c.n += 1; c.h += HIT.has(p.event) ? 1 : 0; c.hr += p.event === 'home_run' ? 1 : 0;
bat.set(p.batter, c);
}
}
}
return rows;
}
/** The adjustment: the hitter's own rate, shifted by the PROVEN pen signal. */
const adjust = (r) => {
const shift = pq.hitRateShift(r.pen_quality);
if (shift === null) return null;
return Math.max(0.01, Math.min(0.99, r.hitter_base + shift));
};
(async () => {
const all = build();
const qs = all.map((r) => r.pen_quality).sort((a, b) => a - b);
const weakCut = qs[Math.floor(qs.length * 2 / 3)];
const strongCut = qs[Math.floor(qs.length / 3)];
const subsets = {
// The order's concentrated subset.
concentrated_early_x_weak_pen: all.filter((r) => r.early && r.pen_quality >= weakCut),
// The mirror, where the descriptive pass suggested the larger movement.
mirror_early_x_strong_pen: all.filter((r) => r.early && r.pen_quality <= strongCut),
// Every later at-bat with an early-exit flag, both directions of pen quality.
all_early_exit_later_abs: all.filter((r) => r.early),
pooled_all_later_abs: all,
};
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), Object.keys(subsets).map((k) => ({
sport: 'mlb', stat: 'hits', archetype: null,
interaction: `collapsed_sequence:${k}`, target: 'later_ab_outcome',
})));
cumulative = mc.cumulative_tests;
} catch { /* offline */ }
const gate = (rs, label) => fg.adjudicate(
rs.map((r) => ({ cluster: r.cluster, baseline: r.hitter_base, conditioned: adjust(r), won: r.won }))
.filter((r) => r.conditioned !== null),
{ factor: label, stat: 'hits', cumulativeTests: cumulative },
);
const results = {};
for (const [k, rs] of Object.entries(subsets)) results[k] = gate(rs, k);
// Hitter split as a SEPARATE gated addition — never assumed into the main effect.
const conc = subsets.concentrated_early_x_weak_pen;
const hrs = conc.map((r) => r.hitter_hr_rate).sort((a, b) => a - b);
const hrCut = hrs[Math.floor(hrs.length / 2)];
const bySplit = {
power_hitters: gate(conc.filter((r) => r.hitter_hr_rate >= hrCut), 'concentrated_power'),
contact_hitters: gate(conc.filter((r) => r.hitter_hr_rate < hrCut), 'concentrated_contact'),
};
console.log(JSON.stringify({
later_at_bats_readable: all.length,
cumulative_tests: cumulative,
subset_sizes: Object.fromEntries(Object.entries(subsets).map(([k, v]) => [k, v.length])),
gate: Object.fromEntries(Object.entries(results).map(([k, v]) => [k, {
n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
verdict: v.verdict,
}])),
hitter_split_separate_gate: Object.fromEntries(Object.entries(bySplit).map(([k, v]) => [k, {
n: v.movement.n, brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null, verdict: v.verdict,
}])),
refused: {
pen_archetype: 'did not prove at the corrected bar — excluded from the adjustment',
hitter_approach_identity: 'SPRAY / fastball-hunter identities do not exist in this registry',
link3_per_reliever: 'SKIPPED — reliever identity is genuine baseball unpredictability',
},
}, null, 2));
process.exit(0);
})();
+198
View File
@@ -0,0 +1,198 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 1 — derive a COHERENT LODO test, blind to reversals.
*
* ── THE DEFECT BEING FIXED ───────────────────────────────────────────────
* The gate at 1f40014 paired a 1-SE per-drop informativeness bar with a
* zero-reversal decision rule. Those two are incoherent. At exactly 1 SE, a
* genuinely STABLE stat's drop reverses with probability Phi(-1) = 0.159, so on
* four informative drops the chance of at least one reversal is
* 1 - 0.841^4 = 0.50. The rule failed stable stats half the time by construction.
*
* And n* was pooled across four stats whose signed effects differ several-fold,
* so "informative" meant different things for different stats while being
* treated as one number.
*
* ── THE FIX ──────────────────────────────────────────────────────────────
* The two halves have to be chosen together:
*
* informative bar n*_k = k^2 * (sigma_row / |g|)^2 PER STAT
* decision rule FAIL iff reversals > c, where under stability
* R ~ Binomial(D, Phi(-k)) and c is the smallest cutoff
* with P(R > c) <= 0.05
*
* `g` is the mean SIGNED per-row improvement — the quantity whose sign a
* reversal flips. Not a mean-absolute, and not pooled: a reversal is a claim
* about THIS stat's effect changing sign.
*
* THIS SCRIPT PRINTS NO REVERSAL AND NO VERDICT. It is blind by construction and
* must run, and its constants be committed, before any stat is re-read.
*
* SUPABASE_URL=... node scripts/derive-lodo-test.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** |g| must clear this many SE at the stat's full n or there is no effect to test. */
const EFFECT_Z = 1.96;
/** Target false-positive rate for the whole per-stat test. */
const TARGET_FP = 0.05;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Standard normal CDF (AbramowitzStegun 7.1.26 via erf). */
function normCdf(z) {
const t = 1 / (1 + 0.2316419 * Math.abs(z));
const d = 0.3989422804014327 * Math.exp(-z * z / 2);
const p = d * t * (0.319381530 + t * (-0.356563782 + t * (1.781477937 + t * (-1.821255978 + t * 1.330274429))));
return z >= 0 ? 1 - p : p;
}
const binomPmf = (n, k, p) => {
let logC = 0;
for (let i = 0; i < k; i += 1) logC += Math.log(n - i) - Math.log(i + 1);
return Math.exp(logC + k * Math.log(p) + (n - k) * Math.log(1 - p));
};
/** P(R > c) for R ~ Binomial(n, p). */
const binomTail = (n, c, p) => {
let s = 0;
for (let k = c + 1; k <= n; k += 1) s += binomPmf(n, k, p);
return s;
};
/** Smallest cutoff c with P(R > c) <= target. */
function cutoffFor(D, p, target) {
for (let c = 0; c <= D; c += 1) if (binomTail(D, c, p) <= target) return { cutoff: c, fp: binomTail(D, c, p) };
return { cutoff: D, fp: 0 };
}
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
const perStat = {};
const dateSizes = {};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b));
if (v === null) continue;
const over = v > L;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
const map = cal.fitIsotonic(rows.map((r) => ({ p: r.p, won: r.won })));
if (!map) { perStat[stat] = { n: rows.length, fittable: false }; continue; }
const d = [];
for (const r of rows) {
const pc = cal.applyIsotonic(map, r.p);
if (knownNumber(pc) === null) continue;
d.push((pc - r.won) ** 2 - (r.p - r.won) ** 2);
}
const g = mean(d);
const sigma = Math.sqrt(d.reduce((s, x) => s + (x - g) ** 2, 0) / (d.length - 1));
const seFull = sigma / Math.sqrt(d.length);
// Date sizes are sample STRUCTURE, not outcomes — safe to read here.
const sizes = new Map();
for (const r of rows) sizes.set(r.date, (sizes.get(r.date) || 0) + 1);
dateSizes[stat] = [...sizes.values()].sort((a, b) => b - a);
perStat[stat] = {
n: d.length,
g_signed: round5(g),
sigma_row: round5(sigma),
se_full: round5(seFull),
effect_z_at_full_n: round3(Math.abs(g) / seFull),
improves: g < 0,
no_effect: Math.abs(g) / seFull < EFFECT_Z,
fittable: true,
};
}
// ── Choose k jointly. Blind: uses only (g, sigma) and date SIZES. ──
const kTable = [];
for (const k of [1.0, 1.25, 1.5, 1.75, 2.0]) {
const pNoise = normCdf(-k);
const row = { k, per_drop_noise_prob: round4(pNoise), stats: {} };
for (const stat of STATS) {
const ps = perStat[stat];
if (!ps || !ps.fittable) continue;
const nStar = Math.ceil(k * k * (ps.sigma_row / Math.abs(ps.g_signed)) ** 2);
const D = (dateSizes[stat] || []).filter((n) => n >= nStar).length;
const { cutoff, fp } = D > 0 ? cutoffFor(D, pNoise, TARGET_FP) : { cutoff: null, fp: null };
// FN at a stated alternative: date-to-date SD of the effect equals |g|.
const pAlt = D > 0 ? normCdf(-Math.abs(ps.g_signed) / Math.sqrt(ps.g_signed ** 2 + (ps.sigma_row ** 2) / nStar)) : null;
const fn = D > 0 && cutoff !== null ? 1 - binomTail(D, cutoff, pAlt) : null;
row.stats[stat] = {
n_star: nStar, informative_drops: D, cutoff, fp: fp === null ? null : round4(fp),
fn_at_tau_equals_g: fn === null ? null : round4(fn),
};
}
kTable.push(row);
}
console.log(JSON.stringify({
phase: 'PHASE 1 — coherent LODO test derivation, BLIND',
defect_being_fixed: 'a 1-SE informative bar with a zero-reversal rule: P(>=1 reversal | stable, 4 drops) = 0.50',
per_stat_effect: perStat,
date_sizes: dateSizes,
k_selection_table: kTable,
target_fp: TARGET_FP,
blind: 'no reversal, no verdict, no reversing date referenced anywhere in this output',
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round3 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 1000) / 1000);
+139
View File
@@ -0,0 +1,139 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 0 — derive the LODO held-row threshold from POWER, blind to outcomes.
*
* ESTIMAND: "does dropping date D reverse the SIGN of the out-of-sample Brier
* improvement on D's held-out rows?"
*
* A reversal is only informative if a single date's Brier delta is
* distinguishable from zero at that row count. Below that, a reversal is a coin
* flip wearing a decimal point — which is exactly the ambiguity that made the
* previous verdict depend on an operator-chosen number.
*
* ── THE DERIVATION ───────────────────────────────────────────────────────
* The per-row Brier difference is
*
* d_i = (pc_i - y_i)^2 - (p_i - y_i)^2
*
* and a date's Brier delta is the MEAN of d over that date's rows. So
*
* SE(n) = SD(d) / sqrt(n)
*
* and the smallest n at which a typical effect clears one standard error is
*
* n* = ( SD(d) / |effect| )^2
*
* SD(d) and |effect| are pooled ACROSS ALL FOUR STATS deliberately: a per-stat
* figure would let the threshold be shaped by the stat whose verdict it decides.
*
* THIS SCRIPT PRINTS NO STAT VERDICT AND NO DATE. It is blind by construction,
* and it must be run and its output committed BEFORE any stat is re-read.
*
* SUPABASE_URL=... node scripts/derive-lodo-threshold.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
// Pooled per-row Brier differences, across all four stats.
const diffs = [];
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b));
if (v === null) continue;
const over = v > L;
rows.push({ p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
const map = cal.fitIsotonic(rows.map((r) => ({ p: r.p, won: r.won })));
if (!map) continue;
for (const r of rows) {
const pc = cal.applyIsotonic(map, r.p);
if (knownNumber(pc) === null) continue;
diffs.push((pc - r.won) ** 2 - (r.p - r.won) ** 2);
}
}
const m = mean(diffs);
const sd = Math.sqrt(diffs.reduce((s, d) => s + (d - m) ** 2, 0) / (diffs.length - 1));
const effect = Math.abs(m);
const nStar = Math.ceil((sd / effect) ** 2);
const curve = [10, 20, 25, 30, 50, 75, 100, 150, 200, 300, 500].map((n) => ({
n,
se: round5(sd / Math.sqrt(n)),
effect_over_se: round3(effect / (sd / Math.sqrt(n))),
informative: effect >= sd / Math.sqrt(n),
}));
console.log(JSON.stringify({
phase: 'PHASE 0 — power derivation, blind to outcomes',
pooled_rows: diffs.length,
per_row_brier_diff_sd: round5(sd),
pooled_effect_abs_mean: round5(effect),
n_star: nStar,
rule: 'n* = (SD(d) / |effect|)^2 — the smallest held-row count at which a typical Brier delta clears one standard error',
se_vs_n: curve,
blind: 'no stat verdict, no date, and no reversal is referenced anywhere in this output',
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round5 = (v) => Math.round(v * 100000) / 100000;
const round3 = (v) => Math.round(v * 1000) / 1000;
+205
View File
@@ -0,0 +1,205 @@
#!/usr/bin/env node
'use strict';
/**
* Phases 0, 2, 3 and 4 — replay the factor wiring on settled hits rows.
*
* Transmission is proved MECHANICALLY before any resolution number is quoted,
* because "resolution went up" is exactly what a subtle bug also prints.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const hf = require('../src/services/model/hitsFactors');
const lp = require('../src/services/model/lowParamCalibrator');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const SEQ = path.join(process.cwd(), '.seq-cache', 'sequences.json');
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f, orderBy = 'id') {
const o = [];
for (let i = 0; ; i += 1000) {
const { data, error } = await f(sb.from(t).select(s)).order(orderBy, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break; o.push(...data); if (data.length < 1000) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) { let s = seed >>> 0; return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; }; }
function decompose(rows, bins = 10) {
const base = mean(rows.map((r) => r.won));
const unc = base * (1 - base);
let rel = 0; let res = 0;
for (let k = 0; k < bins; k += 1) {
const lo = k / bins; const hi = (k + 1) / bins;
const sl = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (!sl.length) continue;
const w = sl.length / rows.length;
rel += w * (mean(sl.map((r) => r.p)) - mean(sl.map((r) => r.won))) ** 2;
res += w * (mean(sl.map((r) => r.won)) - base) ** 2;
}
return { base_rate: r4(base), reliability: r5(rel), resolution: r5(res), uncertainty: r5(unc), share: r4(res / unc) };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
// Factor inputs.
const [spray, defense, platoon, statcast] = await Promise.all([
page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb'), 'player_key'),
page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb'), 'team'),
page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb'), 'player_key'),
page(sb, 'statcast_aggregates', 'player_key, role, bats, throws, hard_hit_pct', (q) => q.eq('sport', 'mlb'), 'player_key'),
]);
const latest = (rows, k) => { const m = new Map(); for (const r of rows) { const key = r[k]; if (!key) continue; const p = m.get(key); if (!p || String(r.as_of_date) > String(p.as_of_date)) m.set(key, r); } return m; };
const sprayBy = latest(spray, 'player_key'); const defBy = latest(defense, 'team'); const platBy = latest(platoon, 'player_key');
const batBy = new Map(); const pitBy = new Map();
for (const r of statcast) { if (!r.player_key) continue; (r.role === 'pitcher' ? pitBy : batBy).set(r.player_key, r); }
const frac = (v) => { const n = knownNumber(v); return n === null ? null : (n > 1 ? n / 100 : n); };
// Opponent + starter per (player,date) from the sequence cache.
const { games } = JSON.parse(fs.readFileSync(SEQ, 'utf8'));
const oppOf = new Map(); const spOf = new Map();
for (const g of games) {
for (const side of ['home', 'away']) {
const opp = g[side === 'home' ? 'away' : 'home'];
const st = (g[side].arms || []).find((a) => a.started);
const half = side === 'home' ? 'top' : 'bottom';
for (const pa of g.pas.filter((p) => p.half === half)) {
const k = `${g.date}|${nameKey(pa.batter_name || '')}`;
if (!oppOf.has(k)) { oppOf.set(k, g[side].team); if (st) spOf.set(k, st.name); }
}
}
}
const snaps = await page(sb, 'model_snapshots', 'id, game_date, captured_at, stat, player_key, player_name, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').eq('stat', 'hits'));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({ propKey: [r.game_date, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const rows = []; const transmission = []; const unreadable = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(b.hits); if (v === null) continue;
const over = v > L;
const won = (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0;
const key = r.player_key;
const bat = batBy.get(key);
const oppTeam = oppOf.get(`${r.game_date}|${key}`);
const def = oppTeam ? (defBy.get(oppTeam) || defBy.get(String(oppTeam).split(' ').pop())) : null;
const spName = spOf.get(`${r.game_date}|${key}`);
const pit = spName ? pitBy.get(nameKey(spName)) : null;
const sp = platBy.get(key);
const ctx = {
spray: sprayBy.get(key) || null,
positionOaa: def && def.position_oaa ? def.position_oaa : null,
bats: bat && bat.bats ? String(bat.bats)[0] : null,
throws: pit && pit.throws ? String(pit.throws)[0] : null,
pitcherHardHit: pit ? frac(pit.hard_hit_pct) : null,
platoonSplits: sp ? { vl: { pa: sp.vl_pa, atBats: sp.vl_ab, hits: sp.vl_hits }, vr: { pa: sp.vr_pa, atBats: sp.vr_ab, hits: sp.vr_hits } } : null,
};
// The engine adjusts p_over then flips for unders; replay that exactly.
const pOverRaw = String(r.side).toLowerCase() === 'under' ? 1 - knownNumber(r.p_win) : knownNumber(r.p_win);
const adj = hf.adjustProbability(pOverRaw, ctx);
const pAfter = adj.factors_fired > 0
? (String(r.side).toLowerCase() === 'under' ? 1 - adj.p_adjusted : adj.p_adjusted)
: knownNumber(r.p_win);
rows.push({ date: r.game_date, p_before: knownNumber(r.p_win), p: pAfter, won, fired: adj.factors_fired, applied: adj.applied });
// TRANSMISSION IS TESTED PER FACTOR, IN ISOLATION.
// Comparing one factor's expected sign against the COMPOSITE p_win change is
// wrong: with three factors firing, two pulling down and one up, the net can
// oppose any single member and look like a defect when nothing is broken.
// So each factor is applied ALONE to the same base and its own sign checked.
const isUnder = String(r.side).toLowerCase() === 'under';
for (const a of adj.applied) {
if (transmission.filter((t) => t.factor === a.factor).length >= 4) continue;
if (Math.abs(a.multiplier - 1) < 0.03) continue;
const solo = { spray: null, positionOaa: null, bats: ctx.bats, throws: null, pitcherHardHit: null, platoonSplits: null };
if (a.factor === 'defense_by_direction') { solo.spray = ctx.spray; solo.positionOaa = ctx.positionOaa; }
if (a.factor === 'pitcher_contact_profile') solo.pitcherHardHit = ctx.pitcherHardHit;
if (a.factor === 'platoon_severity') { solo.platoonSplits = ctx.platoonSplits; solo.throws = ctx.throws; }
const one = hf.adjustProbability(pOverRaw, solo);
if (one.factors_fired !== 1) continue;
const soloWin = isUnder ? 1 - one.p_adjusted : one.p_adjusted;
transmission.push({
factor: a.factor, player: r.player_name, date: r.game_date,
expected: a.multiplier > 1 ? 'raise p(over)' : 'lower p(over)', multiplier: a.multiplier,
side: r.side, p_before: knownNumber(r.p_win), p_after_solo: r4(soloWin),
sign_correct: isUnder
? ((a.multiplier > 1) === (soloWin < knownNumber(r.p_win)))
: ((a.multiplier > 1) === (soloWin > knownNumber(r.p_win))),
});
}
if (adj.skipped.some((s) => /switch hitter/.test(s.reason || '')) && unreadable.length < 4) {
unreadable.push({ player: r.player_name, reason: 'switch hitter — spray side unreadable', p_before: knownNumber(r.p_win), p_after: pAfter, moved_by_spray: false });
}
}
// ── PHASE 4: OOS, point-in-time ──
rows.sort((a, b) => a.date.localeCompare(b.date));
const dates = [...new Set(rows.map((r) => r.date))].sort();
const perDate = new Map(); for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) { acc += perDate.get(d); if (acc >= rows.length * 0.45) { cut = d; break; } }
const fit = rows.filter((r) => r.date < cut); const ev = rows.filter((r) => r.date >= cut);
const mBefore = lp.fitPlatt(fit.map((r) => ({ p: r.p_before, won: r.won, date: r.date })));
const mAfter = lp.fitPlatt(fit.map((r) => ({ p: r.p, won: r.won, date: r.date })));
const evBefore = ev.map((r) => ({ ...r, p: mBefore && !mBefore.refused ? lp.applyPlatt(mBefore, r.p_before) : r.p_before })).filter((r) => r.p != null);
const evAfter = ev.map((r) => ({ ...r, p: mAfter && !mAfter.refused ? lp.applyPlatt(mAfter, r.p) : r.p })).filter((r) => r.p != null);
const bBefore = guards.safeBrier(evBefore.map((r) => r.p), evBefore.map((r) => r.won));
const bAfter = guards.safeBrier(evAfter.map((r) => r.p), evAfter.map((r) => r.won));
const byDate = new Map();
for (let i = 0; i < evAfter.length; i += 1) { const d = evAfter[i].date; if (!byDate.has(d)) byDate.set(d, []); byDate.get(d).push({ a: evAfter[i].p, b: evBefore[i] ? evBefore[i].p : null, won: evAfter[i].won }); }
const keys = [...byDate.keys()]; const rnd = makeRnd(20260807); const diffs = [];
for (let it = 0; it < 3000; it += 1) {
const s = [];
for (let i = 0; i < keys.length; i += 1) s.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
const u = s.filter((x) => x.b != null);
if (!u.length) continue;
diffs.push(guards.safeBrier(u.map((x) => x.a), u.map((x) => x.won)) - guards.safeBrier(u.map((x) => x.b), u.map((x) => x.won)));
}
diffs.sort((a, b) => a - b);
console.log(JSON.stringify({
coverage: { rows: rows.length, any_factor_fired: rows.filter((r) => r.fired > 0).length,
by_count: [0, 1, 2, 3].map((k) => ({ factors: k, n: rows.filter((r) => r.fired === k).length })) },
PHASE_2_transmission: transmission,
PHASE_2_unreadable_static: unreadable,
PHASE_4: {
split_at: cut, fit_n: fit.length, eval_n: ev.length, eval_dates: keys.length,
resolution_before: decompose(evBefore), resolution_after: decompose(evAfter),
brier_before: r5(bBefore), brier_after: r5(bAfter), brier_delta: r5(bAfter - bBefore),
brier_ci_date_block: diffs.length ? [r5(diffs[Math.floor(diffs.length * 0.025)]), r5(diffs[Math.floor(diffs.length * 0.975)])] : null,
},
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const r5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
+137
View File
@@ -0,0 +1,137 @@
#!/usr/bin/env node
'use strict';
/**
* generate-content — tonight's posts, from tonight's real data.
*
* READ-ONLY on every source. This writes nothing to any serving, model or
* ledger table, so it has zero effect on the repaired-champion accrual clock.
*
* SUPABASE_URL=... node scripts/generate-content.js
* -> .content-out/<date>/<template>.txt and .svg
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const engine = require('../src/services/content/contentEngine');
const { toSvg } = require('../src/services/content/cardRenderer');
const sg = require('../src/services/model/servedGrade');
const { knownNumber } = require('../src/utils/known');
for (const t of ['hotHitters', 'honestyFlex', 'streakList']) {
engine.registerTemplate(require(`../src/services/content/templates/${t}`));
}
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const OUT = path.join(process.cwd(), '.content-out');
async function page(sb, t, sel, ob, f) {
const o = [];
for (let i = 0; ; i += 1000) {
const { data, error } = await f(sb.from(t).select(sel)).order(ob, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break;
o.push(...data);
if (data.length < 1000) break;
}
return o;
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const date = new Intl.DateTimeFormat('en-CA', { timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit' }).format(new Date());
// ── SOURCE 1: full-season hitter form (the REPAIRED window, not last10) ──
const lines = fs.existsSync(BOX) ? JSON.parse(fs.readFileSync(BOX, 'utf8')).lines : {};
const byPlayer = new Map();
for (const [k, b] of Object.entries(lines)) {
const [d, key] = k.split('|');
if (!byPlayer.has(key)) byPlayer.set(key, []);
byPlayer.get(key).push({ d, hits: b.hits, name: b.name });
}
// The box-score cache spans only the settled snapshot window, so it holds far
// fewer than a season per player. Season form comes from the SAME full log the
// repaired champion reads.
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const hitterFormFull = async (names) => {
const out = [];
for (const n of names.slice(0, 60)) {
try {
const res = await mlb.getPlayerStats(n);
const log = (res && res.found && Array.isArray(res.fullLog)) ? res.fullLog : [];
const vals = log.map((g) => knownNumber(g && g.stat && g.stat.hits)).filter((v) => v !== null);
if (vals.length < 20) continue;
const rate = (a) => a.filter((v) => v > 0).length / a.length;
out.push({ name: n, season_games: vals.length, season_rate: rate(vals), recent_rate: rate(vals.slice(-10)) });
} catch { /* absent player -> absent row */ }
}
return out;
};
const cacheForm = async () => [...byPlayer.entries()].map(([key, games]) => {
games.sort((a, b) => a.d.localeCompare(b.d));
const vals = games.map((g) => knownNumber(g.hits)).filter((v) => v !== null);
if (vals.length < 20) return null;
const rate = (arr) => arr.filter((v) => v > 0).length / arr.length;
return {
name: games[games.length - 1].name || key,
season_games: vals.length,
season_rate: rate(vals),
recent_rate: rate(vals.slice(-10)),
};
}).filter(Boolean);
const hitterForm = async () => {
const names = [...byPlayer.values()].map((g) => g[g.length - 1].name).filter(Boolean);
const full = await hitterFormFull([...new Set(names)]);
return full.length ? full : await cacheForm();
};
// ── SOURCE 2: the real served-grade distribution ──
const snaps = await page(sb, 'model_snapshots', 'p_win, refused, stat, game_date', 'id',
(q) => q.eq('sport', 'mlb').eq('game_date', date));
const gradeDistribution = async () => {
const usable = snaps.filter((r) => !r.refused && knownNumber(r.p_win) !== null);
const by = {}; let flat = 0;
for (const r of usable) {
const g = sg.gradeFor({ p_win: knownNumber(r.p_win) });
by[g.letter] = (by[g.letter] || 0) + 1;
if (g.separates_from_base_rate === false) flat += 1;
}
return { total: usable.length || null, by_letter: by, not_separable: flat };
};
// ── SOURCE 3: streaks verified from SETTLED ledger outcomes only ──
const led = await page(sb, 'ledger_entries', 'player_name, player_key, game_date, outcome, stat', 'id',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits').in('outcome', ['hit', 'miss']));
const settledStreaks = async () => {
const by = new Map();
for (const r of led) {
if (!by.has(r.player_key)) by.set(r.player_key, []);
by.get(r.player_key).push(r);
}
const out = [];
for (const [, rows] of by) {
rows.sort((a, b) => String(b.game_date).localeCompare(String(a.game_date)));
let n = 0;
for (const r of rows) { if (r.outcome === 'hit') n += 1; else break; }
if (n >= 3) out.push({ name: rows[0].player_name, streak: n, verified_from_settled: true });
}
return out;
};
const deps = { date, hitterForm, gradeDistribution, settledStreaks, servedGrade: sg };
const results = await engine.generateAll(deps);
const dir = path.join(OUT, date);
fs.mkdirSync(dir, { recursive: true });
for (const r of results) {
if (!r.ok) { console.log(`\n[SKIP] ${r.id}${r.reason}`); continue; }
fs.writeFileSync(path.join(dir, `${r.id}.txt`), r.copy);
fs.writeFileSync(path.join(dir, `${r.id}.svg`), toSvg(r.card));
console.log(`\n${'='.repeat(64)}\n${r.id.toUpperCase()}${r.honest_absence ? ' [HONEST ABSENCE]' : ''}\n${'='.repeat(64)}\n${r.copy}`);
}
console.log(`\n\noutput: ${dir}`);
process.exit(0);
})().catch((e) => { console.error('FAILED:', e.message); process.exit(1); });
+146
View File
@@ -0,0 +1,146 @@
#!/usr/bin/env node
'use strict';
/**
* ingest-game-sequences — the raw material for the reliever chain.
*
* Every link in this chain needs something the ledger does not carry: when the
* starter actually left, and which arm actually faced each plate appearance.
* Both are free from statsapi (playByPlay + boxscore), on the same host we
* already use for game logs and schedules.
*
* Caches to disk so Link 1, 2 and 3 all read one fetch rather than three.
*
* node scripts/ingest-game-sequences.js # default window
* SEQ_FROM=2026-06-01 SEQ_TO=2026-08-04 node scripts/ingest-game-sequences.js
*/
const fs = require('fs');
const path = require('path');
const axios = require('axios');
const FROM = process.env.SEQ_FROM || '2026-06-15';
const TO = process.env.SEQ_TO || '2026-08-04';
const OUT = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const CONCURRENCY = 6;
const get = async (url) => (await axios.get(url, { timeout: 45_000 })).data;
/** Innings pitched come as '5.2' meaning five and TWO THIRDS — parseFloat is wrong. */
function ipToOuts(ip) {
if (ip == null) return null;
const [whole, frac] = String(ip).split('.');
const w = Number(whole); const f = Number(frac || 0);
if (!Number.isFinite(w)) return null;
return w * 3 + (Number.isFinite(f) ? f : 0);
}
function datesBetween(from, to) {
const out = [];
const d = new Date(`${from}T12:00:00Z`);
const end = new Date(`${to}T12:00:00Z`);
while (d <= end) { out.push(d.toISOString().slice(0, 10)); d.setUTCDate(d.getUTCDate() + 1); }
return out;
}
async function pool(items, fn, n = CONCURRENCY) {
const out = []; let i = 0;
await Promise.all(Array.from({ length: n }, async () => {
while (i < items.length) {
const idx = i; i += 1;
try { out[idx] = await fn(items[idx]); } catch { out[idx] = null; }
}
}));
return out.filter(Boolean);
}
async function loadGame(g) {
const pk = g.gamePk;
const [box, pbp] = await Promise.all([
get(`https://statsapi.mlb.com/api/v1/game/${pk}/boxscore`),
get(`https://statsapi.mlb.com/api/v1/game/${pk}/playByPlay`),
]);
const sides = {};
for (const side of ['home', 'away']) {
const t = box.teams[side];
if (!t) return null;
const arms = (t.pitchers || []).map((id) => {
const pl = t.players[`ID${id}`];
const s = pl && pl.stats && pl.stats.pitching;
if (!s) return null;
return {
id: Number(id),
name: pl.person && pl.person.fullName,
started: Number(s.gamesStarted || 0) === 1,
bf: s.battersFaced == null ? null : Number(s.battersFaced),
outs: ipToOuts(s.inningsPitched),
pitches: s.pitchesThrown == null ? null : Number(s.pitchesThrown),
};
}).filter(Boolean);
sides[side] = { team: t.team && t.team.name, abbr: t.team && t.team.abbreviation, arms };
}
// Every plate appearance in order, with who threw it.
const pas = [];
for (const p of pbp.allPlays || []) {
const m = p.matchup || {}; const a = p.about || {};
if (!m.batter || !m.pitcher) continue;
pas.push({
batter: Number(m.batter.id),
batter_name: m.batter.fullName,
pitcher: Number(m.pitcher.id),
bats: m.batSide && m.batSide.code,
throws: m.pitchHand && m.pitchHand.code,
inning: a.inning,
half: a.halfInning,
idx: a.atBatIndex,
event: p.result && p.result.eventType,
});
}
if (!pas.length) return null;
return {
gamePk: pk,
date: g.officialDate || (g.gameDate || '').slice(0, 10),
venue_id: g.venue && g.venue.id,
home: sides.home, away: sides.away,
pas,
};
}
async function main() {
const dates = datesBetween(FROM, TO);
console.error(`[seq] ${dates.length} dates ${FROM} -> ${TO}`);
const allGames = [];
for (const d of dates) {
try {
const s = await get(`https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}&hydrate=venue`);
for (const day of s.dates || []) {
for (const g of day.games || []) {
if (String(g.status && g.status.detailedState) === 'Final') allGames.push(g);
}
}
} catch { /* absent day */ }
}
console.error(`[seq] ${allGames.length} final games; fetching sequences`);
const games = await pool(allGames, loadGame);
fs.mkdirSync(path.dirname(OUT), { recursive: true });
fs.writeFileSync(OUT, JSON.stringify({ from: FROM, to: TO, games }));
const starters = games.reduce((s, g) =>
s + ['home', 'away'].filter((k) => g[k].arms.some((a) => a.started)).length, 0);
console.log(JSON.stringify({
dates: dates.length,
final_games: allGames.length,
games_loaded: games.length,
starter_games: starters,
plate_appearances: games.reduce((s, g) => s + g.pas.length, 0),
cache: OUT,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+138
View File
@@ -0,0 +1,138 @@
#!/usr/bin/env node
'use strict';
/**
* LINK 1 — does a starter's exit point predict, before the game?
*
* Target: batters faced by the starter, because that is what decides how many of
* a hitter's plate appearances come against him rather than the pen.
*
* ── WHAT A PRE-GAME PREDICTOR CAN AND CANNOT SEE ─────────────────────────
* The order specifies fatigue profile x GAME SCRIPT ("getting hit -> pulled
* early"). Game script is not available when a prop is graded: whether he gets
* hit tonight is the thing we are trying to project, not an input to it. Using
* it would be reading the answer.
*
* So the honest pre-game form of Link 1 is the fatigue and workload half alone —
* the starter's own history, strictly truncated to starts BEFORE the game being
* predicted. That is measured here. The in-game half is a LIVE feature, not a
* grade-time one, and it is recorded as out of scope rather than quietly folded
* in.
*
* Baseline: the league mean batters faced — the naive "a starter goes about six"
* null this must beat to be worth anything.
*
* node scripts/link1-pull-timing.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const pg = require('../src/services/model/predictionGate');
const tl = require('../src/services/model/testLedger');
const { createClient } = require('@supabase/supabase-js');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
/** Starts needed before we will read a pitcher's own history at all. */
const MIN_PRIOR = 3;
/** Shrinkage: how many prior starts before his own mean carries half the weight. */
const STABILIZE = 5;
/** A start at or under this many batters faced is an EARLY EXIT — the edge case. */
const EARLY_BF = 20;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
function main() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
// Every starter-game, in chronological order.
const starts = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const s = (g[side].arms || []).find((a) => a.started);
if (!s || s.bf == null) continue;
starts.push({ date: g.date, gamePk: g.gamePk, pitcher: s.id, name: s.name, bf: s.bf, outs: s.outs, pitches: s.pitches });
}
}
starts.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const leagueMean = mean(starts.map((s) => s.bf));
// POINT-IN-TIME: each start is predicted only from starts strictly before it.
const history = new Map();
const rows = [];
for (const s of starts) {
const prior = history.get(s.pitcher) || [];
if (prior.length >= MIN_PRIOR) {
const own = mean(prior.map((p) => p.bf));
const w = prior.length / (prior.length + STABILIZE);
rows.push({
pitcher: s.pitcher,
name: s.name,
cluster: s.pitcher, // his starts are not independent readings
baseline: leagueMean,
prediction: w * own + (1 - w) * leagueMean,
actual: s.bf,
prior_starts: prior.length,
});
}
history.set(s.pitcher, prior.concat([s]));
}
return { starts, leagueMean, rows };
}
(async () => {
const { starts, leagueMean, rows } = main();
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'starter_bf', archetype: null, interaction: 'link1:pull_timing', target: 'actual_exit' },
]);
cumulative = mc.cumulative_tests;
} catch { /* offline: reported below */ }
const verdict = pg.adjudicate(rows, {
link: 'link1_pull_timing',
loss: 'absolute',
cumulativeTests: cumulative,
});
// Does it find the EARLY EXITS specifically? That is where the edge lives —
// being right about a median start is worth nothing to this chain.
const early = rows.filter((r) => r.actual <= EARLY_BF);
const late = rows.filter((r) => r.actual > EARLY_BF);
const predEarly = rows.filter((r) => r.prediction <= EARLY_BF + 2);
const hitRate = predEarly.length
? predEarly.filter((r) => r.actual <= EARLY_BF).length / predEarly.length : null;
const baseEarlyRate = rows.length ? early.length / rows.length : null;
console.log(JSON.stringify({
link: 'LINK 1 — starter pull timing',
starter_games_total: starts.length,
league_mean_bf: round2(leagueMean),
gated_rows: rows.length,
distinct_pitchers: new Set(rows.map((r) => r.pitcher)).size,
cumulative_tests: cumulative,
verdict,
early_exit_analysis: {
definition: `actual batters faced <= ${EARLY_BF}`,
early_exits: early.length,
normal_starts: late.length,
base_rate_of_early_exit: round4(baseEarlyRate),
flagged_early_by_model: predEarly.length,
of_those_actually_early: round4(hitRate),
lift_over_base_rate: hitRate !== null && baseEarlyRate !== null ? round4(hitRate - baseEarlyRate) : null,
note: 'the chain needs the EARLY tail, not the median start',
},
scope_note: 'GAME SCRIPT is deliberately excluded — whether he gets hit tonight is the thing being projected, not an input available at grade time',
}, null, 2));
process.exit(0);
})();
const round2 = (v) => (v == null ? null : Math.round(v * 100) / 100);
const round4 = (v) => (v == null ? null : Math.round(v * 10000) / 10000);
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env node
'use strict';
/**
* LINK 2 — can we say WHICH arm faces the later plate appearances?
*
* Link 1 proved, so this link is allowed to be attempted at all. It is gated the
* same way: predict the reliever who actually threw a given post-starter plate
* appearance, against a naive baseline, point-in-time.
*
* BASELINE the team's most-used reliever to date — "guess the busiest arm"
* PREDICTION the reliever that team has most often used IN THIS INNING to
* date, which is the cheapest expression of bullpen ROLE
*
* Loss is misclassification: 0 when the named arm actually threw it, 1 otherwise.
*
* ── THE REPLICATION UNIT IS THE BULLPEN, AND THERE ARE THIRTY ────────────
* Bullpen usage is a team-level process — the same manager, the same arms, the
* same roles all season — so errors are correlated within team and the entity
* this prediction rides on is the club. That caps replication at 30 whatever the
* row count, exactly like park geometry and team defence. Reported explicitly
* rather than dissolved into a row count of tens of thousands.
*
* node scripts/link2-reliever-identity.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const pg = require('../src/services/model/predictionGate');
const tl = require('../src/services/model/testLedger');
const { createClient } = require('@supabase/supabase-js');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
/** Pick the key with the highest count; null when there is nothing to pick from. */
function argmax(counter) {
let best = null; let bestN = -1;
for (const [k, v] of counter) if (v > bestN) { best = k; bestN = v; }
return bestN > 0 ? best : null;
}
function build() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
// Point-in-time bullpen histories, accumulated as we walk forward in time.
const overall = new Map(); // team -> Map(pitcherId -> appearances)
const byInning = new Map(); // `team|inning` -> Map(pitcherId -> appearances)
const rows = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const team = g[side].abbr || g[side].team;
if (!team) continue;
const starter = (g[side].arms || []).find((a) => a.started);
if (!starter) continue;
const relievers = new Set((g[side].arms || []).filter((a) => !a.started).map((a) => a.id));
if (!relievers.size) continue;
// This side PITCHES in the opposite half-inning.
const half = side === 'home' ? 'top' : 'bottom';
const post = g.pas.filter((p) => p.half === half && p.pitcher !== starter.id);
for (const pa of post) {
const ov = overall.get(team);
const inn = byInning.get(`${team}|${pa.inning}`);
const basePick = ov ? argmax(ov) : null;
const modelPick = inn ? argmax(inn) : basePick;
// No history yet is honestly unreadable, not a wrong guess.
if (basePick === null || modelPick === null) continue;
rows.push({
cluster: team,
baseline: Number(basePick) === pa.pitcher ? 1 : 0,
prediction: Number(modelPick) === pa.pitcher ? 1 : 0,
actual: 1,
inning: pa.inning,
});
}
// Now fold this game into history — never before predicting from it.
if (!overall.has(team)) overall.set(team, new Map());
const ovm = overall.get(team);
for (const r of relievers) ovm.set(String(r), (ovm.get(String(r)) || 0) + 1);
for (const pa of post) {
const k = `${team}|${pa.inning}`;
if (!byInning.has(k)) byInning.set(k, new Map());
const m = byInning.get(k);
m.set(String(pa.pitcher), (m.get(String(pa.pitcher)) || 0) + 1);
}
}
}
return rows;
}
(async () => {
const rows = build();
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'reliever_identity', archetype: null, interaction: 'link2:reliever_identity', target: 'actual_arm' },
]);
cumulative = mc.cumulative_tests;
} catch { /* offline */ }
const verdict = pg.adjudicate(rows, {
link: 'link2_reliever_identity',
loss: 'absolute',
cumulativeTests: cumulative,
});
const acc = (k) => (rows.length ? rows.filter((r) => r[k] === 1).length / rows.length : null);
console.log(JSON.stringify({
link: 'LINK 2 — reliever identity',
post_starter_plate_appearances: rows.length,
distinct_bullpens: new Set(rows.map((r) => r.cluster)).size,
cumulative_tests: cumulative,
baseline_accuracy: round4(acc('baseline')),
model_accuracy: round4(acc('prediction')),
verdict,
structural_note: 'the entity this prediction rides on is the BULLPEN, and there are 30 — row count cannot create replication that does not exist',
}, null, 2));
process.exit(0);
})();
const round4 = (v) => (v == null ? null : Math.round(v * 10000) / 10000);
+206
View File
@@ -0,0 +1,206 @@
#!/usr/bin/env node
'use strict';
/**
* LINK 2 (coarse grain) — WHICH BULLPEN, not which arm.
*
* Naming the individual reliever failed on merit: 17.2% accuracy, wrong five
* times in six. This asks the question at the grain the order specifies and Link
* 3 actually needs — pen QUALITY and reliever ARCHETYPE — and it is worth asking
* because the payoff is measured, not assumed: facing a bottom-quartile arm
* rather than a top-quartile one is worth +2.57pp of hit rate, larger than the
* whole times-through-the-order effect.
*
* ── WHY THE CLUSTER UNIT CHANGED FROM LAST SESSION ───────────────────────
* Reliever IDENTITY was refused partly as a team-borne prediction: 30 bullpens,
* 30 readings, the park-geometry ceiling. Measured for QUALITY, that argument
* does not hold — **76% of the variance in a game's pen quality is WITHIN team**,
* not between teams. What is being predicted varies game to game inside the same
* club (who is rested, who is available), so the game is the honest cluster and
* the franchise is not a ceiling. Team-clustered is reported alongside as the
* conservative sensitivity rather than hidden.
*
* ── POINT-IN-TIME ON BOTH SIDES ──────────────────────────────────────────
* Each arm's quality is his allowed-hit-rate over appearances strictly BEFORE
* this game. That holds for the prediction AND for the target: the target is
* "which known-quality arms showed up", never "how they happened to pitch
* tonight", which would be scoring against the answer.
*
* node scripts/link2b-pen-quality.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const pg = require('../src/services/model/predictionGate');
const tl = require('../src/services/model/testLedger');
const { createClient } = require('@supabase/supabase-js');
const { knownNumber } = require('../src/utils/known');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const HIT = new Set(['single', 'double', 'triple', 'home_run']);
const PA = new Set(['single', 'double', 'triple', 'home_run', 'field_out', 'strikeout',
'grounded_into_double_play', 'force_out', 'field_error', 'fielders_choice',
'fielders_choice_out', 'double_play', 'sac_fly', 'pop_out', 'line_out', 'fly_out',
'strikeout_double_play']);
/** Appearances before we will read an arm's quality at all. Below it: abstain. */
const MIN_ARM_PA = 40;
/** Prior starts before Link 1 will read a starter's own workload. */
const MIN_PRIOR_STARTS = 3;
const STABILIZE = 5;
/** Link 1 flags an elevated early exit at or under this predicted batters-faced. */
const EARLY_FLAG_BF = 22;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/**
* Reliever archetype at the coarse grain, from strikeout rate — the axis that
* separates a power arm from a contact arm and the one Link 3 would condition on.
*/
function archetypeOf(kRate) {
if (kRate === null) return null;
if (kRate >= 0.28) return 'POWER';
if (kRate <= 0.18) return 'CONTACT';
return 'MIDDLE';
}
function build() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const arm = new Map(); // pid -> { n, h, k } (all prior PAs)
const penHist = new Map(); // team -> [{ quality, k }] per prior game
const startHist = new Map(); // starter id -> [bf]
const rows = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const team = g[side].abbr || g[side].team;
const st = (g[side].arms || []).find((a) => a.started);
if (!team || !st) continue;
const half = side === 'home' ? 'top' : 'bottom';
const pas = g.pas.filter((p) => p.half === half && PA.has(p.event));
const post = pas.filter((p) => p.pitcher !== st.id);
// ── LINK 1, recomputed point-in-time, to define the concentrated subset ──
const priorStarts = startHist.get(st.id) || [];
let predBf = null;
if (priorStarts.length >= MIN_PRIOR_STARTS) {
const w = priorStarts.length / (priorStarts.length + STABILIZE);
predBf = w * mean(priorStarts) + (1 - w) * 21.56; // league mean
}
// ── TARGET: the known quality of the arms that ACTUALLY appeared ──
const faced = [];
for (const p of post) {
const h = arm.get(p.pitcher);
if (!h || h.n < MIN_ARM_PA) continue; // abstain, never 0
faced.push({ q: h.h / h.n, k: h.k / h.n });
}
// ── PREDICTION: this club's own pen, from prior games only ──
const hist = penHist.get(team) || [];
if (faced.length && hist.length >= 5 && predBf !== null) {
const predQ = mean(hist.map((x) => x.quality));
const predK = mean(hist.map((x) => x.k));
rows.push({
team,
gamePk: g.gamePk,
date: g.date,
pred_bf: predBf,
early_flagged: predBf <= EARLY_FLAG_BF,
pred_quality: predQ,
actual_quality: mean(faced.map((f) => f.q)),
pred_archetype: archetypeOf(predK),
actual_archetype: archetypeOf(mean(faced.map((f) => f.k))),
arms_faced: faced.length,
});
}
// Fold this game into history — never before predicting from it.
if (faced.length) {
penHist.set(team, hist.concat([{ quality: mean(faced.map((f) => f.q)), k: mean(faced.map((f) => f.k)) }]));
}
if (st.bf != null) startHist.set(st.id, priorStarts.concat([st.bf]));
for (const p of pas) {
const cur = arm.get(p.pitcher) || { n: 0, h: 0, k: 0 };
cur.n += 1;
cur.h += HIT.has(p.event) ? 1 : 0;
cur.k += p.event === 'strikeout' ? 1 : 0;
arm.set(p.pitcher, cur);
}
}
}
return rows;
}
function gateQuality(rows, leagueQ, cumulative, clusterKey, label) {
return pg.adjudicate(rows.map((r) => ({
cluster: r[clusterKey],
baseline: leagueQ,
prediction: r.pred_quality,
actual: r.actual_quality,
})), { link: label, loss: 'absolute', cumulativeTests: cumulative });
}
(async () => {
const all = build();
const subset = all.filter((r) => r.early_flagged);
const leagueQ = mean(all.map((r) => r.actual_quality));
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'pen_quality', archetype: null, interaction: 'link2b:pen_quality', target: 'actual_arms' },
{ sport: 'mlb', stat: 'pen_archetype', archetype: null, interaction: 'link2b:pen_archetype', target: 'actual_arms' },
]);
cumulative = mc.cumulative_tests;
} catch { /* offline */ }
// ARCHETYPE grain — misclassification against the arms that actually appeared.
const archRows = subset.filter((r) => r.pred_archetype && r.actual_archetype);
const modal = (() => {
const c = new Map();
for (const r of all) c.set(r.actual_archetype, (c.get(r.actual_archetype) || 0) + 1);
return [...c.entries()].sort((a, b) => b[1] - a[1])[0][0];
})();
const archGate = pg.adjudicate(archRows.map((r) => ({
cluster: r.gamePk,
baseline: r.actual_archetype === modal ? 1 : 0,
prediction: r.actual_archetype === r.pred_archetype ? 1 : 0,
actual: 1,
})), { link: 'link2b_pen_archetype', loss: 'absolute', cumulativeTests: cumulative });
const acc = (rs, k) => (rs.length ? rs.filter((r) => r[k]).length / rs.length : null);
console.log(JSON.stringify({
link: 'LINK 2 (coarse) — pen quality + reliever archetype',
team_games_total: all.length,
concentrated_subset_elevated_early_exit: subset.length,
league_mean_pen_quality: round4(leagueQ),
cumulative_tests: cumulative,
quality_grain: {
on_concentrated_subset: gateQuality(subset, leagueQ, cumulative, 'gamePk', 'link2b_pen_quality_subset'),
sensitivity_team_clustered: gateQuality(subset, leagueQ, cumulative, 'team', 'link2b_pen_quality_teamclust'),
pooled_all_games: gateQuality(all, leagueQ, cumulative, 'gamePk', 'link2b_pen_quality_pooled'),
},
archetype_grain: {
n: archRows.length,
modal_archetype: modal,
baseline_accuracy_guess_modal: round4(acc(archRows.map((r) => ({ x: r.actual_archetype === modal })), 'x')),
model_accuracy: round4(acc(archRows.map((r) => ({ x: r.actual_archetype === r.pred_archetype })), 'x')),
verdict: archGate,
},
cluster_note: '76% of game pen-quality variance is WITHIN team, so the game is the honest cluster; team-clustered reported as the conservative sensitivity',
}, null, 2));
process.exit(0);
})();
const round4 = (v) => (v == null ? null : Math.round(v * 10000) / 10000);
+215
View File
@@ -0,0 +1,215 @@
#!/usr/bin/env node
'use strict';
/**
* lodo-calibration — Phases 2 and 3.
*
* The ≥40 date-cluster floor was factorGate's interval bar for a CAUSAL claim,
* mis-applied to a monotone shrink-to-observed layer. Calibration makes no causal
* claim, consumes no Bonferroni slot, and has a bounded failure mode (it can only
* over- or under-shrink). Its real risk is that the correction is DATE-DRIVEN —
* that one unusual day's offensive environment is doing all the work.
*
* Leave-one-date-out tests exactly that, and it is a harder bar than a cluster
* count: a single date whose removal reverses the improvement, or flips the
* favourite-longshot sign, fails the stat outright.
*
* ── WHAT LODO IS AND IS NOT ──────────────────────────────────────────────
* Refitting on all-but-one date uses dates that follow the held-out one, so this
* is a STABILITY test, not a point-in-time backtest. The point-in-time result is
* separate and already established (fit-past / apply-forward, CI excluding zero
* on hits / TB / RBI). Both are required; neither substitutes for the other.
*
* SUPABASE_URL=... node scripts/lodo-calibration.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** The favourite bucket where the over-prediction concentrates. */
const FAVOURITE_FLOOR = 0.9;
/**
* Minimum rows on a held-out date for that drop to be informative.
*
* POWER-DERIVED AND PRE-COMMITTED (n* = 70). Not chosen here, and not tunable
* from here -- it is imported so the value that decides the verdicts cannot be
* edited alongside them.
*/
const { LODO_TEST, LODO_POWER_FLOOR } = require('../src/services/model/calibrationRegistry');
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select))
.order('id', { ascending: true }).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
const isPreGame = (capturedAt, gameDate) => {
const et = new Date(new Date(capturedAt).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < gameDate || (d === gameDate && et.getUTCHours() < 19);
};
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
// Build the RAW population first so the guard has something to catch.
const raw = [];
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date)) continue;
if (r.refused || knownNumber(r.p_win) === null) continue;
const propKey = [r.game_date, r.stat, r.player_key, r.line].join('|');
raw.push({ propKey, side: r.side, p: knownNumber(r.p_win) });
const prev = picked.get(propKey);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(propKey, r);
}
// GUARD 1 — prove the raw population would have lied, then prove dedup fixes it.
const rawCheck = guards.checkPickedSideDedup(raw);
const pickedRows = [...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'),
side: r.side, p: knownNumber(r.p_win),
}));
guards.assertPickedSideDedup(pickedRows); // throws if dedup failed
const out = {
guard_1_raw_population: { violated: rawCheck.violated, mean_p: rawCheck.mean_p, both_sides_share: rawCheck.both_sides_share },
guard_1_after_dedup: guards.checkPickedSideDedup(pickedRows),
per_stat: {},
};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b));
if (v === null) continue;
const over = v > L;
rows.push({
date: r.game_date,
p: knownNumber(r.p_win),
won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0,
});
}
const dates = [...new Set(rows.map((r) => r.date))].sort();
const full = cal.fitIsotonic(rows.map((r) => ({ p: r.p, won: r.won })));
if (!full) {
out.per_stat[stat] = {
n: rows.length, dates: dates.length,
lodo: 'NOT RUN', decision: 'REFUSE',
reason: `no isotonic map is fittable at n=${rows.length} (needs ${cal.MIN_TOTAL || 200})`,
};
continue;
}
// ── LEAVE ONE DATE OUT, at THIS stat's own informative bar ──
const spec = LODO_TEST[stat];
const MIN_HELD_ROWS = spec ? spec.n_star : Infinity;
const table = [];
for (const d of dates) {
const fit = rows.filter((r) => r.date !== d);
const held = rows.filter((r) => r.date === d);
if (held.length < MIN_HELD_ROWS) {
table.push({ dropped: d, held_n: held.length, verdict: 'UNINFORMATIVE', reason: 'too few rows on this date' });
continue;
}
const map = cal.fitIsotonic(fit.map((r) => ({ p: r.p, won: r.won })));
const applied = guards.applyOrRefuse(map, held, cal.applyIsotonic);
if (!applied.ok) {
table.push({ dropped: d, held_n: held.length, verdict: 'UNINFORMATIVE', reason: applied.reason });
continue;
}
const ys = applied.rows.map((r) => r.won);
const bRaw = guards.safeBrier(applied.rows.map((r) => r.p), ys);
const bCal = guards.safeBrier(applied.rows.map((r) => r.pc), ys);
if (bRaw === null || bCal === null) {
table.push({ dropped: d, held_n: held.length, verdict: 'UNINFORMATIVE', reason: 'a null reached the metric' });
continue;
}
const fav = applied.rows.filter((r) => r.p >= FAVOURITE_FLOOR);
const favBias = fav.length >= 5 ? mean(fav.map((r) => r.p)) - mean(fav.map((r) => r.won)) : null;
table.push({
dropped: d,
held_n: held.length,
brier_delta: round4(bCal - bRaw),
improves: bCal < bRaw,
favourite_n: fav.length,
favourite_bias: favBias === null ? null : round4(favBias),
favourite_sign_holds: favBias === null ? null : favBias > 0,
verdict: bCal < bRaw ? 'holds' : 'REVERSES',
});
}
const informative = table.filter((t) => t.verdict !== 'UNINFORMATIVE');
const reversals = informative.filter((t) => t.verdict === 'REVERSES');
// THE DECISION RULE IS BINOMIAL, not zero-tolerance. Under stability each
// informative drop reverses with prob Phi(-k), so demanding zero reversals
// failed stable stats roughly half the time.
const cutoff = spec ? spec.cutoff : 0;
const exceedsCutoff = reversals.length > cutoff;
// AND THE TEST MUST BE ABLE TO FAIL. Below the power floor it cannot, so it
// cannot pass either -- "could not test" must never read as "passed".
const underpowered = !spec || spec.power < LODO_POWER_FLOOR;
const verdict = underpowered ? 'UNTESTABLE_BY_LODO' : (exceedsCutoff ? 'FAIL' : 'PASS');
const passes = verdict === 'PASS';
out.per_stat[stat] = {
n: rows.length,
dates: dates.length,
lodo_table: table,
informative_drops: informative.length,
brier_reversals: informative.filter((t) => t.verdict === 'REVERSES').length,
favourite_sign_flips: informative.filter((t) => t.favourite_sign_holds === false).length,
favourite_sign_untested: informative.filter((t) => t.favourite_sign_holds === null).length,
n_star: spec ? spec.n_star : null,
cutoff,
reversal_count: reversals.length,
reversing_dates: reversals.map((t) => ({ date: t.dropped, held_n: t.held_n, delta: t.brier_delta })),
test_power: spec ? spec.power : null,
lodo: verdict,
reason: underpowered
? `power ${spec ? spec.power : 0} < floor ${LODO_POWER_FLOOR} — this test would miss a real date-driven failure more than nine times in ten, so it can neither pass nor fail the stat`
: (exceedsCutoff
? `${reversals.length} reversals among ${informative.length} informative drops exceeds the cutoff of ${cutoff}`
: `${reversals.length} reversals among ${informative.length} informative drops is within the cutoff of ${cutoff}`),
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
main().catch((e) => { console.error(e); process.exit(1); });
+58 -3
View File
@@ -54,12 +54,14 @@ const FACTORS = [
{
key: 'defense_by_direction',
needs: ['spray_multiplier'],
entity: (r) => `${r.player_key}|${r.opp}`,
mechanism: 'CAUSALLY-CORRECT DEFENCE. Where the hitter puts the ball (pull/straight/oppo x ground/air) crossed with the OAA of the fielders actually standing in those zones, joined by handedness. Team-average failed the gate because it averages in five fielders who will never touch his ball.',
apply: (r) => r.spray_multiplier,
},
{
key: 'defense',
needs: ['team_defense'],
entity: (r) => r.opp,
mechanism: 'A ball in play becomes a hit or an out partly by who is standing behind the pitcher. Should matter most where contact stays in the park.',
// More outs converted above average -> fewer hits.
apply: (r) => 1 - Math.max(-0.12, Math.min(0.12, r.team_defense / 250)),
@@ -67,12 +69,14 @@ const FACTORS = [
{
key: 'pitcher_contact_profile',
needs: ['pitcher_hard_hit_allowed'],
entity: (r) => r.starter_id,
mechanism: 'A contact-allowing arm concedes better contact than a bat-misser; hit probability should follow the quality of contact he permits.',
apply: (r) => 1 + Math.max(-0.15, Math.min(0.15, (r.pitcher_hard_hit_allowed - 0.389) * 1.2)),
},
{
key: 'park_hits',
needs: ['park_factor'],
entity: (r) => r.park_factor,
mechanism: 'Some parks turn outs into hits without producing runs — big outfields, high walls, deep gaps.',
apply: (r) => r.park_factor,
caveat: 'STAT_BASE maps hits -> run_base, so this is a RUN factor standing in for a HITS factor. A park that converts outs to hits without scoring is invisible to it.',
@@ -80,12 +84,14 @@ const FACTORS = [
{
key: 'platoon_severity',
needs: ['platoon_severity_mult'],
entity: (r) => r.player_key,
mechanism: "CAUSALLY-CORRECT PLATOON. The advantage is worth only what THIS hitter's measured split is worth, shrunk toward league by the smaller side's PA and refused outright below a floor. Flat handedness applies the same boost to a 63-point split and to none.",
apply: (r) => r.platoon_severity_mult,
},
{
key: 'platoon',
needs: ['platoon_edge'],
entity: (r) => r.player_key,
mechanism: 'Handedness advantage — a hitter facing the opposite hand sees the ball better and hits it harder.',
apply: (r) => (r.platoon_edge > 0 ? 1.06 : 0.96),
},
@@ -128,7 +134,7 @@ async function main() {
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries',
'id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base',
'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits')
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
@@ -165,20 +171,34 @@ async function main() {
byPlayer.set(r.player_key, cur);
}
const loss = { no_batter_profile: 0, thin_base_rate: 0, no_opponent: 0, no_pitcher: 0, kept: 0 };
const rows = [];
for (const r of clean) {
const bat = batters.get(r.player_key);
const bp = byPlayer.get(r.player_key);
if (!bp || bp.n < 3) continue;
if (!bat) loss.no_batter_profile += 1;
if (!bp || bp.n < 3) { loss.thin_base_rate += 1; continue; }
// Leave-one-out so a row never contributes to its own baseline.
const baseline = (bp.w - (r.outcome === 'hit' ? 1 : 0)) / (bp.n - 1);
const faced = oppBy.get(`${r.player_key}|${r.game_date}`) || null;
const nick = faced ? String(faced).split(' ').pop() : null;
const def = faced ? (defByTeam.get(faced) || defByTeam.get(nick)) : null;
if (!faced) loss.no_opponent += 1;
const starterId = faced ? startersBy.get(`${r.game_date}|OPP:${faced}`) : null;
const pit = starterId != null ? pitchersById.get(Number(starterId)) : null;
if (faced && !pit) loss.no_pitcher += 1;
loss.kept += 1;
rows.push({
id: r.id,
// Errors are correlated WITHIN a game — shared starter, park, weather and
// the game's own randomness — so the interval must be clustered on it.
// Three of these factors (pitcher profile, team defence, park) are also
// CONSTANT across every hitter facing that starter, which makes row
// resampling straightforwardly wrong for them.
cluster: r.game_id,
opp: faced,
starter_id: starterId != null ? Number(starterId) : null,
player_key: r.player_key,
archetype: archOf.get(`${r.player_key}|${r.game_date}`) || null,
won: r.outcome === 'hit' ? 1 : 0,
baseline,
@@ -214,15 +234,43 @@ async function main() {
const mc = await tl.recordAndCount(store, FACTORS.flatMap((f) =>
ARCHS.map((a) => ({ sport: 'mlb', stat: 'hits', archetype: a === 'ALL' ? null : a, interaction: `factor:${f.key}`, target: 'outcome' }))));
// STEP 1 — FULL-HISTORY SAMPLE AUDIT PER SLOT, before any gating.
const audit = [];
for (const f of FACTORS) {
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
audit.push({
factor: f.key,
archetype: arch,
rows: usable.length,
games: new Set(usable.map((r) => r.cluster).filter(Boolean)).size,
players: new Set(usable.map((r) => r.player_key)).size,
});
}
}
const results = [];
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
for (const f of FACTORS) {
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
// A park effect is replicated across PARKS, not across games: 619 rows in
// 45 games still only ever saw ~23 ballparks, and unmodelled park
// heterogeneity is confounded with the very thing being estimated. So the
// cluster is the COARSER of the game and the entity the treatment rides on.
const ents = f.entity ? new Set(usable.map((r) => String(f.entity(r)))) : null;
const games = new Set(usable.map((r) => String(r.cluster)));
const useEntity = ents && ents.size < games.size;
const paired = usable.map((r) => {
const mult = f.apply(r);
const cond = mult === null ? null : Math.min(0.99, Math.max(0.01, r.baseline * mult));
return { baseline: r.baseline, conditioned: cond, won: r.won };
return {
baseline: r.baseline,
conditioned: cond,
won: r.won,
cluster: useEntity ? `e:${f.entity(r)}` : r.cluster,
};
});
const v = fg.adjudicate(paired, {
factor: f.key, archetype: arch, stat: 'hits',
@@ -230,6 +278,10 @@ async function main() {
});
results.push({
archetype: arch, factor: f.key, n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
cluster_unit: useEntity ? 'treatment_entity' : 'game',
distinct_games: games.size,
distinct_entities: ents ? ents.size : null,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
@@ -244,6 +296,9 @@ async function main() {
console.log(JSON.stringify({
baseline: "each row scored against the player's OWN leave-one-out base rate — the honest 'he's due' null",
total_rows: rows.length,
slot_audit: audit,
clean_settled_rows_available: clean.length,
row_loss: loss,
cumulative_bonferroni: mc,
gate: 'a factor must MOVE the prediction AND improve out-of-sample Brier; movement alone is THEATER',
results,
+121
View File
@@ -0,0 +1,121 @@
#!/usr/bin/env node
'use strict';
/**
* prove-park-weather — DOES PARK GEOMETRY AND AIR READ TOTAL BASES?
*
* Run through the two-part gate like every other factor, with one addition that
* changes the answer: the rows are CLUSTERED BY GAME. Park and weather assign a
* single value to every hitter in a ballpark on a night, so eighteen prop rows
* from one game are one reading of that game's conditions. Resampling rows would
* treat them as eighteen and hand back an interval far tighter than the evidence
* supports — which is how a gate passes a factor on sample it never had.
*
* SUPABASE_URL=... node scripts/prove-park-weather.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const pw = require('../src/services/model/parkWeather');
const fg = require('../src/services/model/factorGate');
const tl = require('../src/services/model/testLedger');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/** Hit-type shares → expected total bases, so a reshape has a consequence. */
const tbFromShares = (s) => s.single + 2 * s.double + 3 * s.triple + 4 * s.home_run;
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const parks = await page(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb'));
const byVenue = new Map();
for (const p of parks) if (!byVenue.has(p.venue_id)) byVenue.set(p.venue_id, p);
const league = pw.leagueGeometry([...byVenue.values()]);
const ctx = await page(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q);
const ctxBy = new Map(ctx.map((c) => [c.game_id, c]));
const led = await page(sb, 'ledger_entries',
'game_id, game_date, player_key, stat, line, side, outcome, quarantine_reason, p_win, proj_hits_p_over',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases').in('outcome', ['hit', 'miss']));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// League-typical hit-type shares — the shape the atom reshapes.
const BASE_SHARES = { single: 0.655, double: 0.195, triple: 0.017, home_run: 0.133 };
const baseTb = tbFromShares(BASE_SHARES);
const loss = { no_context: 0, no_venue: 0, no_park: 0, no_baseline: 0, kept: 0 };
const rows = [];
for (const r of clean) {
const c = ctxBy.get(r.game_id);
if (!c) { loss.no_context += 1; continue; }
if (c.venue_id == null) { loss.no_venue += 1; continue; }
const dims = byVenue.get(c.venue_id);
if (!dims) { loss.no_park += 1; continue; }
const baseline = knownNumber(r.p_win);
if (baseline === null) { loss.no_baseline += 1; continue; }
const read = pw.parkWeatherRead({ dims, wx: c, league });
if (!read) { loss.no_park += 1; continue; }
// The atom reshapes hit type; the consequence for total bases is the ratio
// of expected bases per hit under the reshaped shape.
const shaped = pw.applyToShares(BASE_SHARES, read);
const ratio = tbFromShares(shaped) / baseTb;
const conditioned = Math.max(0.01, Math.min(0.99, baseline * ratio));
rows.push({
cluster: r.game_id, // ONE reading per game — the whole point
baseline,
conditioned,
won: r.outcome === 'hit' ? 1 : 0,
});
loss.kept += 1;
}
// Cumulative Bonferroni across the programme lifetime — this hypothesis is
// one more test, and the bar rises for it like every other.
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'total_bases', archetype: null, interaction: 'factor:park_weather_hit_type', target: 'outcome' },
]);
const verdict = fg.adjudicate(rows, {
factor: 'park_weather_hit_type',
stat: 'total_bases',
cumulativeTests: mc.cumulative_tests,
});
const games = new Set(rows.map((r) => r.cluster)).size;
const venues = new Set(clean.map((r) => ctxBy.get(r.game_id)?.venue_id).filter((v) => v != null)).size;
console.log(JSON.stringify({
clean_settled_tb_rows: clean.length,
rows_built: rows.length,
row_loss: loss,
distinct_games: games,
distinct_venues: venues,
rows_per_game: games ? Math.round((rows.length / games) * 10) / 10 : null,
cumulative_tests: mc.cumulative_tests,
verdict,
honest_note: 'sample judged in GAMES, not prop rows — park and weather vary per game',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+301
View File
@@ -0,0 +1,301 @@
#!/usr/bin/env node
'use strict';
/**
* prove-runs-rbi — THE CONTEXT-HEAVY STATS, WHERE INFLATION IS EASIEST.
*
* A large share of both stats is genuinely outside the hitter's control: a run
* needs someone behind you, an RBI needs someone in front of you. The job is to
* prove the HITTER-CONTROLLABLE part above the archetype's own base rate and
* grade the rest honestly as base-rate — which is the CORRECT answer for a
* context stat, not a failure to find something.
*
* ── THE NULL IS THE ARCHETYPE'S BASE RATE ────────────────────────────────
* Deliberately, and per the order: these base rates are spread and
* context-inflated, so beating "hitters like him" is the only meaningful bar. A
* per-player leave-one-out rate is not available here — 935 RBI rows over 344
* players is ~2.7 rows each, and estimating a personal rate from two rows would
* be inventing one. Leave-one-out is applied at the ARCHETYPE level so a row
* never contributes to its own baseline.
*
* ── INPUTS RECONSTRUCTED RATHER THAN DECLARED MISSING ────────────────────
* `lineup_context` only starts 2026-08-04 (ingest began last week) while settled
* rows run from 07-31, so only 187 of 617 runs rows join to a batting order.
* That would be input-blocked — except the play-by-play cache covers 05-01
* onward, and the batting order IS the order batters first appear. Reach-base
* skill and lineup power behind are derived from the same cache, point-in-time.
*
* SUPABASE_URL=... node scripts/prove-runs-rbi.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const fg = require('../src/services/model/factorGate');
const tl = require('../src/services/model/testLedger');
const sk = require('../src/services/model/skillProjection');
const { knownNumber, knownRate } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const PAGE = 1000;
const ARCHS = (process.env.RR_ARCHETYPES || 'ALL,BOMBER,GHOST,BRUSH,DRIVER').split(',');
const HIT = new Set(['single', 'double', 'triple', 'home_run']);
const ONBASE = new Set(['single', 'double', 'triple', 'home_run', 'walk', 'hit_by_pitch', 'intent_walk']);
const PA_EVENT = new Set([...HIT, 'field_out', 'strikeout', 'grounded_into_double_play', 'force_out',
'field_error', 'fielders_choice', 'fielders_choice_out', 'double_play', 'sac_fly', 'pop_out',
'line_out', 'fly_out', 'strikeout_double_play', 'walk', 'hit_by_pitch', 'intent_walk']);
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/**
* Reconstruct, point-in-time, from play-by-play:
* order[date|nameKey] the hitter's batting slot that game
* behind[date|nameKey] mean barrel-ish power of the three slots after him
* onbase[nameKey] his reach-base rate over PRIOR games only
*/
function reconstruct(barrelByKey) {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const order = new Map();
const behind = new Map();
const onbaseNow = new Map(); // running totals, folded in AFTER each game
const onbasePrior = new Map(); // snapshot used for that game's rows
for (const g of games) {
for (const half of ['top', 'bottom']) {
const pas = g.pas.filter((p) => p.half === half && PA_EVENT.has(p.event));
if (!pas.length) continue;
// The batting order IS the order batters first appear.
const seen = [];
const seenSet = new Set();
for (const p of pas) {
if (!seenSet.has(p.batter)) { seenSet.add(p.batter); seen.push(p); }
if (seen.length >= 9) break;
}
const slots = seen.map((p) => ({ id: p.batter, key: nameKey(p.batter_name || '') }));
for (let i = 0; i < slots.length; i += 1) {
const k = `${g.date}|${slots[i].key}`;
order.set(k, i + 1);
// Power BEHIND him — the hitters who would drive him in.
const nxt = [1, 2, 3].map((d) => slots[(i + d) % slots.length])
.map((s) => (s ? knownRate(barrelByKey.get(s.key)) : null))
.filter((v) => v !== null);
if (nxt.length) behind.set(k, mean(nxt));
const prior = onbaseNow.get(slots[i].key);
if (prior && prior.pa >= 60) onbasePrior.set(k, prior.ob / prior.pa);
}
for (const p of pas) {
const key = nameKey(p.batter_name || '');
const cur = onbaseNow.get(key) || { pa: 0, ob: 0 };
cur.pa += 1; cur.ob += ONBASE.has(p.event) ? 1 : 0;
onbaseNow.set(key, cur);
}
}
}
return { order, behind, onbase: onbasePrior };
}
/** RBI and RUNS have different causal stories, so different factors. */
const FACTORS = {
rbi: [
{
key: 'risp_opportunity',
needs: ['risp_share'],
entity: (r) => r.player_key,
mechanism: 'HOW OFTEN HE BATS WITH RUNNERS IN SCORING POSITION. Half of an RBI is opportunity, and this is the ingested measure of it.',
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.risp_share - 0.22) * 1.6)),
},
{
key: 'extra_base_skill',
needs: ['barrel_pct'],
entity: (r) => r.player_key,
mechanism: 'The other half — having batted with runners on, can he drive them in.',
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.barrel_pct - 0.078) * 1.8)),
},
{
key: 'risp_x_extra_base',
needs: ['risp_share', 'barrel_pct'],
entity: (r) => r.player_key,
mechanism: 'THE CAUSALLY-CORRECT COMPOUND: opportunity AND the power to convert it. Neither half alone is an RBI.',
apply: (r) => (1 + Math.max(-0.20, Math.min(0.20, (r.risp_share - 0.22) * 1.6)))
* (1 + Math.max(-0.20, Math.min(0.20, (r.barrel_pct - 0.078) * 1.8))),
},
],
runs: [
{
key: 'reach_base',
needs: ['onbase'],
entity: (r) => r.player_key,
mechanism: 'You cannot score without first reaching base. The most hitter-controllable component of a run.',
apply: (r) => 1 + Math.max(-0.25, Math.min(0.25, (r.onbase - 0.318) * 2.2)),
},
{
key: 'lineup_power_behind',
needs: ['power_behind'],
entity: (r) => `${r.game_id}|${r.batting_order}`,
mechanism: 'Who bats after him — the hitters who would drive him in. Pure context, and the part he does not control.',
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.power_behind - 0.078) * 1.8)),
},
{
key: 'reach_x_power_behind',
needs: ['onbase', 'power_behind'],
entity: (r) => r.player_key,
mechanism: 'THE CAUSALLY-CORRECT COMPOUND: reach base AND have someone behind you who can drive you in.',
apply: (r) => (1 + Math.max(-0.25, Math.min(0.25, (r.onbase - 0.318) * 2.2)))
* (1 + Math.max(-0.20, Math.min(0.20, (r.power_behind - 0.078) * 1.8))),
},
],
};
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb').eq('role', 'batter'));
const batByKey = new Map();
const barrelByKey = new Map();
for (const r of statcast) {
if (!r.player_key) continue;
const prof = sk.fromStatcastRow(r);
batByKey.set(r.player_key, prof);
if (prof.barrel_pct != null) barrelByKey.set(r.player_key, prof.barrel_pct);
}
const oppRows = await page(sb, 'hitter_opportunity', '*', (q) => q.eq('sport', 'mlb'));
const oppByKey = new Map();
for (const r of oppRows) {
const prev = oppByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) oppByKey.set(r.player_key, r);
}
const recon = reconstruct(barrelByKey);
const out = { generated_note: 'null is the ARCHETYPE base rate, leave-one-out' };
for (const stat of ['rbi', 'runs']) {
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype',
(q) => q.eq('sport', 'mlb').eq('stat', stat).not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries',
'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', stat).in('outcome', ['hit', 'miss']));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')
&& knownNumber(r.line) === 0.5);
// Archetype-level leave-one-out base rate — the context-inflated null.
const byArch = new Map();
for (const r of clean) {
const a = String(archOf.get(`${r.player_key}|${r.game_date}`) || 'UNLABELLED').toUpperCase();
const cur = byArch.get(a) || { n: 0, w: 0 };
cur.n += 1; cur.w += r.outcome === 'hit' ? 1 : 0;
byArch.set(a, cur);
}
const rows = [];
const loss = { no_archetype_base: 0, kept: 0 };
for (const r of clean) {
const a = String(archOf.get(`${r.player_key}|${r.game_date}`) || 'UNLABELLED').toUpperCase();
const ab = byArch.get(a);
if (!ab || ab.n < 4) { loss.no_archetype_base += 1; continue; }
const baseline = (ab.w - (r.outcome === 'hit' ? 1 : 0)) / (ab.n - 1);
const bat = batByKey.get(r.player_key);
const opp = oppByKey.get(r.player_key);
const okey = `${r.game_date}|${r.player_key}`;
rows.push({
archetype: a,
player_key: r.player_key,
game_id: r.game_id,
cluster: r.game_id,
baseline,
won: r.outcome === 'hit' ? 1 : 0,
risp_share: opp ? knownNumber(opp.risp_share) : null,
barrel_pct: bat ? knownRate(bat.barrel_pct) : null,
onbase: recon.onbase.has(okey) ? recon.onbase.get(okey) : null,
power_behind: recon.behind.has(okey) ? recon.behind.get(okey) : null,
batting_order: recon.order.get(okey) ?? null,
});
loss.kept += 1;
}
const mc = await tl.recordAndCount(tl.supabaseStore(sb), FACTORS[stat].flatMap((f) =>
ARCHS.map((a) => ({
sport: 'mlb', stat, archetype: a === 'ALL' ? null : a,
interaction: `factor:${f.key}`, target: 'outcome',
}))));
const audit = [];
const results = [];
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => r.archetype === arch);
for (const f of FACTORS[stat]) {
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
const ents = new Set(usable.map((r) => String(f.entity(r))));
const games = new Set(usable.map((r) => String(r.cluster)));
if (arch === 'ALL') {
audit.push({ factor: f.key, archetype: arch, rows: usable.length, games: games.size, entities: ents.size });
}
const useEntity = ents.size < games.size;
const paired = usable.map((r) => {
const m = f.apply(r);
return {
baseline: r.baseline,
conditioned: m === null ? null : Math.min(0.99, Math.max(0.01, r.baseline * m)),
won: r.won,
cluster: useEntity ? `e:${f.entity(r)}` : r.cluster,
};
});
const v = fg.adjudicate(paired, { factor: f.key, archetype: arch, stat, cumulativeTests: mc.cumulative_tests });
results.push({
archetype: arch, factor: f.key, n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
clustered_on: useEntity ? 'treatment_entity' : 'game',
distinct_games: games.size,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
verdict: v.verdict,
});
}
}
out[stat] = {
clean_rows_line_0_5: clean.length,
rows_built: rows.length,
row_loss: loss,
distinct_games: new Set(rows.map((r) => r.cluster)).size,
archetype_base_rates: Object.fromEntries([...byArch.entries()]
.sort((a, b) => b[1].n - a[1].n)
.map(([a, v]) => [a, { n: v.n, base_rate: Math.round((v.w / v.n) * 10000) / 10000 }])),
cumulative_tests: mc.cumulative_tests,
input_audit: audit,
results,
proven: results.filter((r) => r.verdict === 'PROVES'),
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+345
View File
@@ -0,0 +1,345 @@
#!/usr/bin/env node
'use strict';
/**
* prove-tb-factors — TOTAL BASES IS A DIFFERENT EVENT FROM HITS.
*
* A hit asks whether the ball found a hole. Total bases asks how hard and how
* far it was struck. So the causally-correct factors differ, and the
* archetype differential is expected to INVERT: contact defence proved for hits,
* where a slap single is worth exactly one base regardless of who fielded it;
* for total bases the value should live with the power profiles.
*
* ── THE BASELINE IS THE CHAMPION, NOT "HE'S DUE" ─────────────────────────
* The hits gate used the player's own leave-one-out base rate as the null. That
* cannot be reproduced here: total-bases lines VARY (1.5 on 559 rows, 0.5 on
* 345, 2.5 on 45), and a player's rate of clearing 1.5 bases is a different
* quantity from his rate of clearing 0.5. With 988 rows over 341 players there
* are roughly two rows per player-line — far too thin to estimate a per-line
* personal base rate without inventing one.
*
* So the null here is the COUNTER'S OWN FORECAST (p_win), which already prices
* the line. That is a strictly HARDER null than a base rate, not an easier one:
* a factor must improve on the champion, not merely on "he's due". Stated
* plainly because it differs from the hits run and the difference matters when
* comparing the two.
*
* SUPABASE_URL=... node scripts/prove-tb-factors.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const fg = require('../src/services/model/factorGate');
const sk = require('../src/services/model/skillProjection');
const tl = require('../src/services/model/testLedger');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownNumber, knownRate } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const sd = require('../src/services/model/sprayDefense');
const pss = require('../src/services/model/platoonSeverity');
const pw = require('../src/services/model/parkWeather');
/** League-typical hit-type shares; the atom reshapes these and TB follows. */
const BASE_SHARES = { single: 0.655, double: 0.195, triple: 0.017, home_run: 0.133 };
const tbFrom = (s) => s.single + 2 * s.double + 3 * s.triple + 4 * s.home_run;
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const ARCHS = (process.env.TB_ARCHETYPES || 'ALL,BOMBER,GHOST,BRUSH,DRIVER').split(',');
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/**
* THE FACTORS. Each returns a MULTIPLIER on the base rate, or null when the
* input is absent — an absent factor must leave the baseline untouched rather
* than nudge it toward some default.
*/
const FACTORS = [
{
key: 'barrel_rate',
needs: ['barrel_pct'],
entity: (r) => r.player_key,
mechanism: 'THE EXTRA-BASE SKILL ITSELF. A barrel is the exit-velocity and launch-angle combination that produces extra bases; it is the most direct expression of what total bases measures, where for hits it is largely irrelevant to whether a grounder finds a hole.',
// UNITS: fromStatcastRow returns barrel_pct as a FRACTION (0.06), not the
// 0-100 the raw table stores. Writing this against the percentage scale
// clamped every row to the maximum negative shift, which then "improved"
// Brier only by leaning on the counter's known global over-prediction.
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.barrel_pct - 0.078) * 1.8)),
},
{
key: 'exit_velo',
needs: ['avg_exit_velo'],
entity: (r) => r.player_key,
mechanism: 'How hard the ball leaves the bat. Separates a double in the gap from a fly out, which is exactly the margin total bases lives on.',
apply: (r) => 1 + Math.max(-0.15, Math.min(0.15, (r.avg_exit_velo - 88.9) * 0.020)),
},
{
key: 'hard_contact_allowed',
needs: ['pitcher_hard_hit_allowed'],
entity: (r) => r.starter_id,
mechanism: 'A pitcher who concedes hard contact concedes EXTRA BASES, not just hits. For total bases this should read stronger than it did for hits.',
apply: (r) => 1 + Math.max(-0.15, Math.min(0.15, (r.pitcher_hard_hit_allowed - 0.389) * 1.2)),
},
{
key: 'park_weather_hit_type',
needs: ['park_weather_ratio'],
entity: (r) => r.park_weather_ratio,
mechanism: 'Whether a struck ball becomes a double, clears the fence, or dies at the track. The atom reshapes HIT TYPE rather than P(hit), which is the only form that can express a total-bases effect.',
apply: (r) => r.park_weather_ratio,
caveat: 'venue-borne: replication caps at the number of distinct park readings, not the row count',
},
{
key: 'platoon_severity',
needs: ['platoon_severity_mult'],
entity: (r) => r.player_key,
mechanism: "The hitter's OWN measured split, shrunk by the smaller side's plate appearances and refused below a floor.",
apply: (r) => r.platoon_severity_mult,
},
];
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const batters = new Map(); const pitchersById = new Map();
for (const r of statcast) {
const prof = sk.fromStatcastRow(r);
if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), prof);
if (r.role === 'batter' && r.player_key) batters.set(r.player_key, prof);
}
const sprayRows = await page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb'));
const sprayByKey = new Map();
for (const r of sprayRows) {
if (!r.player_key) continue;
const prev = sprayByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) sprayByKey.set(r.player_key, r);
}
const platRows = await page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb'));
const platByKey = new Map();
for (const r of platRows) {
if (!r.player_key) continue;
const prev = platByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) platByKey.set(r.player_key, r);
}
const parkRows = await page(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb'));
const parkByVenue = new Map();
for (const p of parkRows) if (!parkByVenue.has(p.venue_id)) parkByVenue.set(p.venue_id, p);
const parkLeague = pw.leagueGeometry([...parkByVenue.values()]);
const ctxRows = await page(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q);
const ctxBy = new Map(ctxRows.map((c) => [c.game_id, c]));
const defRows = await page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb'));
const defByTeam = new Map();
for (const d of defRows) defByTeam.set(d.team, d);
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat',
(q) => q.eq('sport', 'mlb').eq('stat', 'total_bases').not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries',
'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases')
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// Opponent faced, from each hitter's own game log.
const names = new Map();
for (const r of clean) if (!names.has(r.player_key)) names.set(r.player_key, r.player_name);
const oppBy = new Map(); const startersBy = new Map();
const dates = [...new Set(clean.map((r) => r.game_date))].sort();
for (const d of dates) {
try {
const games = await mlb.getScheduleWithPitchers(d);
for (const g of games) {
if (!g.home || !g.away) continue;
if (g.home.probablePitcher) startersBy.set(`${d}|OPP:${g.home.team}`, g.home.probablePitcher.id);
if (g.away.probablePitcher) startersBy.set(`${d}|OPP:${g.away.team}`, g.away.probablePitcher.id);
}
} catch { /* absent slate */ }
}
for (const [key, name] of names) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id);
for (const g of log || []) if (g && g.date && g.opponent) oppBy.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
} catch { /* no log */ }
}
// Per-player base rate — the honest null: "he's due", no reading of tonight.
const byPlayer = new Map();
for (const r of clean) {
const cur = byPlayer.get(r.player_key) || { n: 0, w: 0 };
cur.n += 1; cur.w += r.outcome === 'hit' ? 1 : 0;
byPlayer.set(r.player_key, cur);
}
const loss = { no_batter_profile: 0, thin_base_rate: 0, no_opponent: 0, no_pitcher: 0, kept: 0 };
const rows = [];
for (const r of clean) {
const bat = batters.get(r.player_key);
const bp = byPlayer.get(r.player_key);
if (!bat) loss.no_batter_profile += 1;
if (!bp || bp.n < 3) { loss.thin_base_rate += 1; continue; }
// THE NULL IS THE CHAMPION. Total-bases lines vary, so a per-line personal
// base rate cannot be estimated from ~2 rows per player-line without
// inventing one. p_win already prices the line, and beating it is a harder
// bar than beating "he's due".
const baseline = knownNumber(r.p_win);
if (baseline === null) { loss.thin_base_rate += 1; continue; }
const faced = oppBy.get(`${r.player_key}|${r.game_date}`) || null;
const nick = faced ? String(faced).split(' ').pop() : null;
const def = faced ? (defByTeam.get(faced) || defByTeam.get(nick)) : null;
if (!faced) loss.no_opponent += 1;
const starterId = faced ? startersBy.get(`${r.game_date}|OPP:${faced}`) : null;
const pit = starterId != null ? pitchersById.get(Number(starterId)) : null;
if (faced && !pit) loss.no_pitcher += 1;
loss.kept += 1;
rows.push({
id: r.id,
// Errors are correlated WITHIN a game — shared starter, park, weather and
// the game's own randomness — so the interval must be clustered on it.
// Three of these factors (pitcher profile, team defence, park) are also
// CONSTANT across every hitter facing that starter, which makes row
// resampling straightforwardly wrong for them.
cluster: r.game_id,
opp: faced,
starter_id: starterId != null ? Number(starterId) : null,
player_key: r.player_key,
archetype: archOf.get(`${r.player_key}|${r.game_date}`) || null,
won: r.outcome === 'hit' ? 1 : 0,
baseline,
team_defense: def ? knownNumber(def.oaa_sum) : null,
pitcher_hard_hit_allowed: pit ? knownRate(pit.hard_hit_pct) : null,
line: knownNumber(r.line),
barrel_pct: bat ? knownRate(bat.barrel_pct) : null,
avg_exit_velo: bat ? knownNumber(bat.avg_exit_velo) : null,
park_weather_ratio: (() => {
const c = ctxBy.get(r.game_id);
if (!c || c.venue_id == null) return null;
const dims = parkByVenue.get(c.venue_id);
if (!dims) return null;
const read = pw.parkWeatherRead({ dims, wx: c, league: parkLeague });
if (!read) return null;
const shaped = pw.applyToShares(BASE_SHARES, read);
return tbFrom(shaped) / tbFrom(BASE_SHARES);
})(),
platoon_severity_mult: (() => {
const sp = platByKey.get(r.player_key);
if (!sp || !bat || !bat.bats || !pit || !pit.throws) return null;
const out = pss.platoonRead({
splits: {
vl: { pa: sp.vl_pa, atBats: sp.vl_ab, hits: sp.vl_hits },
vr: { pa: sp.vr_pa, atBats: sp.vr_ab, hits: sp.vr_hits },
},
bats: bat.bats, throws: pit.throws,
});
return out && out.readable ? out.multiplier : null;
})(),
spray_multiplier: (() => {
const sp = sprayByKey.get(r.player_key);
const posOaa = def && def.position_oaa ? def.position_oaa : null;
if (!sp || !posOaa || !bat || !bat.bats) return null;
const out = sd.sprayDefenseMultiplier({ spray: sp, bats: bat.bats, positionOaa: posOaa });
return out ? out.multiplier : null;
})(),
platoon_edge: (bat && pit && bat.bats && pit.throws)
? (String(bat.bats)[0] !== String(pit.throws)[0] ? 1 : -1) : null,
});
}
// Cumulative Bonferroni across the programme lifetime.
const store = tl.supabaseStore(sb);
const mc = await tl.recordAndCount(store, FACTORS.flatMap((f) =>
ARCHS.map((a) => ({ sport: 'mlb', stat: 'total_bases', archetype: a === 'ALL' ? null : a, interaction: `factor:${f.key}`, target: 'outcome' }))));
// STEP 1 — FULL-HISTORY SAMPLE AUDIT PER SLOT, before any gating.
const audit = [];
for (const f of FACTORS) {
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
audit.push({
factor: f.key,
archetype: arch,
rows: usable.length,
games: new Set(usable.map((r) => r.cluster).filter(Boolean)).size,
players: new Set(usable.map((r) => r.player_key)).size,
});
}
}
const results = [];
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
for (const f of FACTORS) {
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
// A park effect is replicated across PARKS, not across games: 619 rows in
// 45 games still only ever saw ~23 ballparks, and unmodelled park
// heterogeneity is confounded with the very thing being estimated. So the
// cluster is the COARSER of the game and the entity the treatment rides on.
const ents = f.entity ? new Set(usable.map((r) => String(f.entity(r)))) : null;
const games = new Set(usable.map((r) => String(r.cluster)));
const useEntity = ents && ents.size < games.size;
const paired = usable.map((r) => {
const mult = f.apply(r);
const cond = mult === null ? null : Math.min(0.99, Math.max(0.01, r.baseline * mult));
return {
baseline: r.baseline,
conditioned: cond,
won: r.won,
cluster: useEntity ? `e:${f.entity(r)}` : r.cluster,
};
});
const v = fg.adjudicate(paired, {
factor: f.key, archetype: arch, stat: 'total_bases',
cumulativeTests: mc.cumulative_tests, // native cumulative correction
});
results.push({
archetype: arch, factor: f.key, n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
cluster_unit: useEntity ? 'treatment_entity' : 'game',
distinct_games: games.size,
distinct_entities: ents ? ents.size : null,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
ci_level: v.improvement ? v.improvement.ci_level : null,
verdict: v.verdict,
reason: v.reason,
...(f.caveat ? { input_caveat: f.caveat } : {}),
});
}
}
console.log(JSON.stringify({
baseline: "each row scored against the player's OWN leave-one-out base rate — the honest 'he's due' null",
total_rows: rows.length,
slot_audit: audit,
clean_settled_rows_available: clean.length,
row_loss: loss,
cumulative_bonferroni: mc,
gate: 'a factor must MOVE the prediction AND improve out-of-sample Brier; movement alone is THEATER',
results,
proven: results.filter((r) => r.verdict === 'PROVES'),
theater: results.filter((r) => r.verdict === 'THEATER'),
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+184
View File
@@ -0,0 +1,184 @@
#!/usr/bin/env node
'use strict';
/**
* PHASES 0-1 — is rbi's 14.51% real, and where does it come from?
*
* PHASE 0 is a Beck gate. The 14.51% came from the same decomposition harness
* whose paging helper produced a false null three times tonight (composite PKs,
* order-by-id, error swallowed). So the row count is asserted against an
* independent exact count, every page is error-checked, and 20 raw rows are
* printed so the decile arithmetic can be audited by hand.
*
* PHASE 1 asks what the resolution IS. Resolution rewards a forecast for
* separating outcomes — but a forecast can separate outcomes by knowing WHO is
* batting rather than anything about tonight. Three nested forecasts:
*
* PLAYER BASE RATE leave-one-out frequency for that hitter, nothing else.
* Its resolution is pure across-player spread.
* LINEUP SLOT mean rate for that batting-order position. Real
* predictive signal, but ROLE, not skill.
* THE MODEL served p_win.
*
* The part that behaves like our doctrine's "skill" is what the model resolves
* WITHIN a stratum of similar players — measured directly by stratifying on the
* player's own base rate and pooling the within-stratum resolutions.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const SEQ = path.join(process.cwd(), '.seq-cache', 'sequences.json');
const STATS = ['rbi', 'hits', 'total_bases', 'runs'];
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Error-checked pager with an explicit order column. */
async function page(sb, t, sel, orderBy, apply) {
const out = [];
for (let i = 0; ; i += 1000) {
const q = apply ? apply(sb.from(t).select(sel)) : sb.from(t).select(sel);
const { data, error } = await q.order(orderBy, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break;
out.push(...data);
if (data.length < 1000) break;
}
return out;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
/** Resolution alone: weighted spread of bin realized rates about the base rate. */
function resolutionOf(rows, bins = 10) {
const base = mean(rows.map((r) => r.won));
let res = 0;
const table = [];
for (let k = 0; k < bins; k += 1) {
const lo = k / bins; const hi = (k + 1) / bins;
const sl = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (!sl.length) continue;
const w = sl.length / rows.length;
const ok = mean(sl.map((r) => r.won));
res += w * (ok - base) ** 2;
table.push({ bin: `${lo.toFixed(1)}-${hi.toFixed(1)}`, n: sl.length, forecast: r4(mean(sl.map((r) => r.p))), realized: r4(ok) });
}
return { base_rate: r4(base), resolution: r5(res), uncertainty: r5(base * (1 - base)), share: r4(res / (base * (1 - base))), table };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
// Batting-order slot, reconstructed point-in-time from play-by-play.
const { games } = JSON.parse(fs.readFileSync(SEQ, 'utf8'));
const slotOf = new Map();
for (const g of games) {
for (const half of ['top', 'bottom']) {
const seen = []; const set = new Set();
for (const p of g.pas.filter((x) => x.half === half)) {
if (!set.has(p.batter)) { set.add(p.batter); seen.push(p); }
if (seen.length >= 9) break;
}
seen.forEach((p, i) => slotOf.set(`${g.date}|${nameKey(p.batter_name || '')}`, i + 1));
}
}
const out = { PHASE_0: {}, PHASE_1: {} };
for (const stat of STATS) {
// PHASE 0 — exact count first, then the paged pull must match it.
const { count: exact, error: cErr } = await sb.from('model_snapshots')
.select('*', { count: 'exact', head: true }).eq('sport', 'mlb').eq('stat', stat);
if (cErr) throw new Error(`count ${stat}: ${cErr.message}`);
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, player_name, line, side, p_win, refused', 'id',
(q) => q.eq('sport', 'mlb').eq('stat', stat));
if (snaps.length !== exact) throw new Error(`PHASE 0 FAIL ${stat}: paged ${snaps.length} != exact ${exact}`);
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({ propKey: [r.game_date, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const rows = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const over = v > L;
rows.push({
date: r.game_date, key: r.player_key, name: r.player_name, line: L, side: r.side,
actual: v, p: knownNumber(r.p_win),
won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0,
slot: slotOf.get(`${r.game_date}|${r.player_key}`) ?? null,
});
}
const model = resolutionOf(rows);
out.PHASE_0[stat] = { exact_rows: exact, paged_rows: snaps.length, scorable: rows.length, model_resolution: model.resolution, model_share: model.share, deciles: model.table };
if (stat === 'rbi') {
out.PHASE_0.rbi_raw_sample = rows.slice(0, 20).map((r) => ({ name: r.name, date: r.date, line: r.line, side: r.side, p_win: r.p, actual_rbi: r.actual, won: r.won }));
}
// ── (a) PLAYER BASE RATE, leave-one-out ──
const byPlayer = new Map();
for (const r of rows) { const c = byPlayer.get(r.key) || { n: 0, w: 0 }; c.n += 1; c.w += r.won; byPlayer.set(r.key, c); }
const baseRows = rows.filter((r) => byPlayer.get(r.key).n >= 3)
.map((r) => { const c = byPlayer.get(r.key); return { ...r, p: (c.w - r.won) / (c.n - 1) }; });
const baseOnly = resolutionOf(baseRows);
// ── (b) LINEUP SLOT, leave-one-out ──
const bySlot = new Map();
for (const r of rows) { if (r.slot == null) continue; const c = bySlot.get(r.slot) || { n: 0, w: 0 }; c.n += 1; c.w += r.won; bySlot.set(r.slot, c); }
const slotRows = rows.filter((r) => r.slot != null && bySlot.get(r.slot).n >= 10)
.map((r) => { const c = bySlot.get(r.slot); return { ...r, p: (c.w - r.won) / (c.n - 1) }; });
const slotOnly = slotRows.length ? resolutionOf(slotRows) : null;
// ── (c) WITHIN-STRATUM: does the model still separate similar players? ──
// Stratify on the player's own base rate, then pool the model's resolution
// computed INSIDE each stratum. Across-player spread is held constant, so
// what survives is discrimination between comparable hitters.
const strata = [[0, 0.45], [0.45, 0.6], [0.6, 0.75], [0.75, 1.01]];
let within = 0; let wTot = 0; const strataDetail = [];
for (const [lo, hi] of strata) {
const sl = rows.filter((r) => { const c = byPlayer.get(r.key); if (!c || c.n < 3) return false; const b = c.w / c.n; return b >= lo && b < hi; });
if (sl.length < 40) continue;
const rr = resolutionOf(sl);
within += sl.length * rr.resolution; wTot += sl.length;
strataDetail.push({ stratum: `${lo}-${hi}`, n: sl.length, base: rr.base_rate, resolution: rr.resolution });
}
const withinRes = wTot ? within / wTot : null;
out.PHASE_1[stat] = {
n: rows.length,
model_resolution: model.resolution,
model_share_of_variance: model.share,
a_player_base_rate_only: { n: baseRows.length, resolution: baseOnly.resolution, share: baseOnly.share },
b_lineup_slot_only: slotOnly ? { n: slotRows.length, resolution: slotOnly.resolution, share: slotOnly.share } : null,
c_within_stratum_resolution: r5(withinRes),
strata: strataDetail,
base_rate_explains_pct: baseOnly.resolution ? r4(Math.min(1, baseOnly.resolution / model.resolution)) : null,
within_stratum_share_of_model: withinRes != null && model.resolution ? r4(withinRes / model.resolution) : null,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error('FAILED:', e.message); process.exit(1); });
const r5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
+137
View File
@@ -0,0 +1,137 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 5 — rebuild grade bands on p_win_calibrated, for DEPLOYED stats only.
*
* total_bases is the only stat that cleared LODO, so it is the only one whose
* bands are rebuilt on calibrated values. The rest keep base-rate bands built on
* raw p_win, and the reason is named rather than left to inference.
*
* The two-bar rule still applies and still bites: TB is now CALIBRATED but no
* factor is PROVEN for it (barrel, exit velo and hard-contact-allowed were all
* THEATER), so the bands remain a base-rate read — now an honestly-numbered one.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const lp = require('../src/services/model/lowParamCalibrator');
const gb = require('../src/services/model/gradeBands');
const guards = require('../src/services/model/calibrationGuards');
const tl = require('../src/services/model/testLedger');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STAT = process.env.BAND_STAT || 'total_bases';
const PAGE = 1000;
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused, archetype',
(q) => q.eq('sport', 'mlb').eq('stat', STAT));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
const rows = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[STAT](b));
if (v === null) continue;
const over = v > L;
rows.push({
date: r.game_date, p: knownNumber(r.p_win),
won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0,
archetype: String(r.archetype || 'UNLABELLED').toUpperCase(),
});
}
rows.sort((a, b) => String(a.date).localeCompare(String(b.date)));
// Point-in-time map, then apply forward.
const dates = [...new Set(rows.map((r) => r.date))].sort();
const perDate = new Map();
for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) { acc += perDate.get(d); if (acc >= rows.length * 0.45) { cut = d; break; } }
// Bands are built on the SERVED values. hits and total_bases serve the
// low-parameter correction; rbi and runs serve raw, so their bands are raw.
const DEPLOYED = ['hits', 'total_bases'];
const fitRows = rows.filter((r) => r.date < cut);
const evalRows = rows.filter((r) => r.date >= cut);
let applied;
let basis;
if (DEPLOYED.includes(STAT)) {
const model = lp.fitPlatt(fitRows);
applied = (!model || model.refused)
? { ok: false, reason: 'low-parameter fit refused', rows: [] }
: { ok: true, rows: evalRows.map((r) => ({ ...r, pc: lp.applyPlatt(model, r.p) })).filter((r) => r.pc != null) };
basis = 'p_win_lowparam (SERVED, provisional)';
} else {
applied = { ok: true, rows: evalRows.map((r) => ({ ...r, pc: r.p })) };
basis = 'raw p_win (this stat serves raw)';
}
if (!applied.ok) { console.log(JSON.stringify({ stat: STAT, refused: applied.reason })); process.exit(0); }
const mc = await tl.recordAndCount(tl.supabaseStore(sb), []).catch(() => ({ cumulative_tests: 1 }));
const byArch = new Map();
for (const r of applied.rows) {
if (!byArch.has(r.archetype)) byArch.set(r.archetype, []);
byArch.get(r.archetype).push({ p: r.pc, won: r.won });
}
const out = [];
for (const [arch, rs] of [...byArch.entries()].sort((a, b) => b[1].length - a[1].length)) {
out.push(gb.buildBands(rs, {
archetype: arch,
cumulativeTests: mc.cumulative_tests,
// TB is CALIBRATED (provisional) but no factor is PROVEN for it.
proven: false,
calibrated: DEPLOYED.includes(STAT),
}));
}
console.log(JSON.stringify({
stat: STAT,
basis,
eval_rows: applied.rows.length,
cumulative_tests: mc.cumulative_tests,
two_bar_note: 'calibrated YES, proven NO -> bands stay a base-rate read, now honestly numbered',
bands: out,
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
+170
View File
@@ -0,0 +1,170 @@
#!/usr/bin/env node
'use strict';
/**
* reconstruct-game-environment — GIVE THE PAST GAMES THEIR REAL CONDITIONS.
*
* `game_context` has never held a single weather reading. The reason is not the
* fetcher, which is correct and points at Open-Meteo's ARCHIVE endpoint; it is
* that nothing ever joined. The ledger keys a game as
* `mlb:2026-08-03:WashingtonNationals@PhiladelphiaPhillies` and game_context
* keys it as `mlb:823437`, so every lookup missed and the columns stayed NULL —
* which reads exactly like "the weather was unavailable" rather than "the two
* tables have never been introduced." The same class of failure as the doubled
* /leaderboard path: graceful degradation wearing the mask of honest absence.
*
* This walks the dates in the settled ledger, resolves each slug to the real
* statsapi game and venue, and writes a game_context row keyed by the LEDGER's
* slug so the join exists. Then it pulls the actual archived weather for that
* date and location.
*
* ARCHIVE, NOT FORECAST — asking the forecast endpoint about a past date returns
* a re-forecast, which is a model's opinion about the past, not the past. Absent
* stays NULL; nothing here is imputed.
*
* SUPABASE_URL=... node scripts/reconstruct-game-environment.js
*/
require('dotenv').config();
const axios = require('axios');
const { createClient } = require('@supabase/supabase-js');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const SCHEDULE = (d) => `https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}&hydrate=venue(location)`;
const ARCHIVE = (lat, lon, date) =>
`https://archive-api.open-meteo.com/v1/archive?latitude=${lat}&longitude=${lon}`
+ `&start_date=${date}&end_date=${date}`
+ '&hourly=temperature_2m,wind_speed_10m,wind_direction_10m,precipitation'
+ '&temperature_unit=fahrenheit&wind_speed_unit=mph';
const squash = (s) => String(s || '').toLowerCase().replace(/[^a-z]/g, '');
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
const get = async (url) => (await axios.get(url, { timeout: 60_000 })).data;
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const led = await page(sb, 'ledger_entries', 'game_id, game_date, stat, outcome, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases'])
.in('outcome', ['hit', 'miss']));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// Distinct games, as the LEDGER names them.
const games = new Map();
for (const r of clean) if (r.game_id && !games.has(r.game_id)) games.set(r.game_id, r.game_date);
const dates = [...new Set([...games.values()])].sort();
console.error(`[env] ${games.size} distinct settled games across ${dates.length} dates`);
// date -> statsapi games, indexed by the same squashed away@home the slug uses.
const resolved = new Map();
const venues = new Map();
for (const d of dates) {
let sched = null;
try { sched = await get(SCHEDULE(d)); } catch { sched = null; }
for (const day of (sched && sched.dates) || []) {
for (const g of day.games || []) {
const away = squash(g.teams?.away?.team?.name);
const home = squash(g.teams?.home?.team?.name);
resolved.set(`${d}|${away}@${home}`, g);
const v = g.venue || {};
if (v.id && !venues.has(v.id)) {
const loc = v.location || {};
venues.set(v.id, {
venue_id: v.id,
venue_name: v.name || null,
lat: loc.defaultCoordinates?.latitude ?? null,
lon: loc.defaultCoordinates?.longitude ?? null,
});
}
}
}
}
// Join each ledger slug to its real game + venue.
const rows = []; let unmatched = 0;
for (const [gid, date] of games) {
const m = /^mlb:(\d{4}-\d{2}-\d{2}):(.+?)@(.+)$/.exec(gid);
if (!m) { unmatched += 1; continue; }
const g = resolved.get(`${m[1]}|${squash(m[2])}@${squash(m[3])}`);
if (!g) { unmatched += 1; continue; }
rows.push({
game_id: gid, // the LEDGER's key — this is the whole fix
game_date: date,
venue_id: g.venue?.id ?? null,
source_game_pk: g.gamePk ?? null,
});
}
console.error(`[env] matched ${rows.length}, unmatched ${unmatched}, venues seen ${venues.size}`);
for (let i = 0; i < rows.length; i += 200) {
const { error } = await sb.from('game_context')
.upsert(rows.slice(i, i + 200), { onConflict: 'game_id' });
if (error) console.error('[env] context write failed:', error.message);
}
// ACTUAL archived weather, one call per (venue, date) that we need.
const need = new Map();
for (const r of rows) {
const v = venues.get(r.venue_id);
if (!v || v.lat == null || v.lon == null) continue;
need.set(`${r.venue_id}|${r.game_date}`, { v, date: r.game_date });
}
console.error(`[env] fetching ${need.size} venue-days of archived weather`);
const wx = new Map();
for (const [k, { v, date }] of need) {
try {
const p = await get(ARCHIVE(v.lat, v.lon, date));
const h = p && p.hourly;
if (h && Array.isArray(h.time) && h.time.length) {
const i = Math.min(h.time.length - 1, 19); // ~7pm local, typical first pitch
wx.set(k, {
wx_temp_f: h.temperature_2m?.[i] ?? null,
wx_wind_speed_mph: h.wind_speed_10m?.[i] ?? null,
wx_wind_direction_deg: h.wind_direction_10m?.[i] ?? null,
wx_precip_mm: h.precipitation?.[i] ?? null,
wx_source: 'open_meteo_archive',
});
}
} catch { /* absent stays absent */ }
}
let withWx = 0;
for (let i = 0; i < rows.length; i += 200) {
const batch = rows.slice(i, i + 200).map((r) => {
const w = wx.get(`${r.venue_id}|${r.game_date}`);
if (w) withWx += 1;
return w ? { ...r, ...w } : r;
});
const { error } = await sb.from('game_context').upsert(batch, { onConflict: 'game_id' });
if (error) console.error('[env] weather write failed:', error.message);
}
console.log(JSON.stringify({
settled_games: games.size,
matched_to_statsapi: rows.length,
unmatched,
distinct_venues: venues.size,
venue_days_requested: need.size,
venue_days_returned: wx.size,
game_rows_with_actual_weather: withWx,
source: 'open_meteo_archive (actual, not re-forecast)',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+132
View File
@@ -0,0 +1,132 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 2 — put a number on the resolution ceiling.
*
* Murphy's decomposition: Brier = reliability - resolution + uncertainty.
*
* reliability how far each bin's realized rate sits from its forecast (lower
* is better; this is what calibration fixes)
* resolution how far the bins' realized rates spread from the base rate
* (HIGHER is better; this is discrimination, and NO amount of
* calibration can create it)
* uncertainty the base rate's own variance -- a property of the event
*
* Calibration moves reliability and leaves resolution untouched by construction:
* a monotone map relabels bins without re-sorting the rows inside them. So if
* resolution is near zero, honest numbers are all calibration can ever deliver.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const lp = require('../src/services/model/lowParamCalibrator');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const DEPLOYED = ['hits', 'total_bases'];
const PAGE = 1000;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error; if (!data.length) break; o.push(...data); if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
/** Murphy decomposition over K equal-width bins. */
function decompose(rows, bins = 10) {
const base = mean(rows.map((r) => r.won));
const uncertainty = base * (1 - base);
let reliability = 0; let resolution = 0;
const table = [];
for (let k = 0; k < bins; k += 1) {
const lo = k / bins; const hi = (k + 1) / bins;
const slice = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (!slice.length) continue;
const w = slice.length / rows.length;
const fk = mean(slice.map((r) => r.p));
const ok = mean(slice.map((r) => r.won));
reliability += w * (fk - ok) ** 2;
resolution += w * (ok - base) ** 2;
table.push({ bin: [round2(lo), round2(hi)], n: slice.length, forecast: round4(fk), realized: round4(ok) });
}
return {
base_rate: round4(base),
reliability: round5(reliability),
resolution: round5(resolution),
uncertainty: round5(uncertainty),
brier_check: round5(reliability - resolution + uncertainty),
/** What share of the event's variance the model actually explains. */
resolution_share_of_uncertainty: round4(resolution / uncertainty),
bins: table,
};
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots', 'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const out = {};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const over = v > L;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
if (rows.length < 100) continue;
const raw = decompose(rows);
let served = null;
if (DEPLOYED.includes(stat)) {
const m = lp.fitPlatt(rows);
if (m && !m.refused) {
const cal = rows.map((r) => ({ ...r, p: lp.applyPlatt(m, r.p) })).filter((r) => knownNumber(r.p) !== null);
served = decompose(cal);
}
}
out[stat] = {
n: rows.length,
deployed: DEPLOYED.includes(stat),
raw,
served,
resolution_change_from_calibration: served ? round5(served.resolution - raw.resolution) : null,
reliability_change_from_calibration: served ? round5(served.reliability - raw.reliability) : null,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round2 = (v) => Math.round(v * 100) / 100;
+249
View File
@@ -0,0 +1,249 @@
#!/usr/bin/env node
'use strict';
/**
* settle-model-snapshots — pay the standing debt.
*
* 71,192 snapshot rows have never carried an outcome. They are the retention
* table built for exactly this kind of replay, and until they are settled every
* measurement in this programme runs on the far smaller ledger slice.
*
* ── OUTCOME IS SIDE-ALIGNED, NOT RAW ─────────────────────────────────────
* The order specifies `outcome = 1[realized > line]`. That is the OVER
* perspective, and it would be backwards for every under-side prop — `p_win` is
* side-aligned (verified: TB mean p_win 0.5698 against a 0.5074 side-won rate),
* so a raw over-indicator would silently invert the target on the under rows and
* make calibration measure the wrong thing.
*
* So: `actual_value` stores the realized stat (raw, unopinionated) and `outcome`
* stores whether the GRADED SIDE won. Deviation from the literal order, stated
* because it changes the number.
*
* ── INTEGRITY (hard-fail) ────────────────────────────────────────────────
* conservation settled + unresolvable + orphaned == candidates
* no dupes one write per snapshot id
* no orphans a settled row must have matched a real box score
* prediction-time logging captured_at must PRECEDE the game date; a row
* logged after the fact is not a prediction and is refused
*
* node scripts/settle-model-snapshots.js # dry run, verifies only
* SETTLE_WRITE=1 node scripts/settle-model-snapshots.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const axios = require('axios');
const { createClient } = require('@supabase/supabase-js');
const { nameKey } = require('../src/utils/playerName');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const WRITE = process.env.SETTLE_WRITE === '1';
const BOX_CACHE = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** Realized value per stat, from the box-score batting line. */
const FIELD = Object.freeze({
hits: (b) => knownNumber(b.hits),
total_bases: (b) => knownNumber(b.totalBases),
rbi: (b) => knownNumber(b.rbi),
runs: (b) => knownNumber(b.runs),
});
const get = async (url) => (await axios.get(url, { timeout: 45_000 })).data;
/** Eastern first pitch, conservatively. Anything at or after this is in-game. */
const FIRST_PITCH_ET_HOUR = 19;
/**
* Was this row logged BEFORE the games it grades?
*
* The pipeline runs on UTC cron hours, so a 01:00-UTC cycle is 21:00 the
* PREVIOUS evening in Eastern -- same game date, three hours into the slate.
*/
function isPreGame(capturedAt, gameDate) {
if (!capturedAt || !gameDate) return false;
const cap = new Date(capturedAt);
if (Number.isNaN(cap.getTime())) return false;
const et = new Date(cap.getTime() - 4 * 3600 * 1000); // EDT
const etDate = et.toISOString().slice(0, 10);
if (etDate < String(gameDate)) return true; // day before, fine
if (etDate > String(gameDate)) return false; // day after, post-game
return et.getUTCHours() < FIRST_PITCH_ET_HOUR;
}
async function pool(items, fn, n = 6) {
const out = []; let i = 0;
await Promise.all(Array.from({ length: n }, async () => {
while (i < items.length) {
const idx = i; i += 1;
try { out[idx] = await fn(items[idx]); } catch { out[idx] = null; }
}
}));
return out.filter(Boolean);
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
// STABLE ORDER. model_snapshots is a LIVE table -- the snapshot cron writes
// to it at 14/19/22/1/3 UTC -- and an unordered .range() walk over a table
// being appended to returns overlapping pages. The integrity gate caught
// exactly that on the first run.
const { data, error } = await apply(sb.from(table).select(select))
.order('id', { ascending: true })
.range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/** Box-score batting lines for a date range, cached. */
async function battingLines(dates) {
if (fs.existsSync(BOX_CACHE)) {
const c = JSON.parse(fs.readFileSync(BOX_CACHE, 'utf8'));
if (dates.every((d) => c.dates.includes(d))) return c.lines;
}
const games = [];
for (const d of dates) {
try {
const s = await get(`https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}`);
for (const day of s.dates || []) {
for (const g of day.games || []) {
if (String(g.status && g.status.detailedState) === 'Final') {
games.push({ pk: g.gamePk, date: g.officialDate || d });
}
}
}
} catch { /* absent day */ }
}
console.error(`[settle] ${games.length} final games across ${dates.length} dates`);
const lines = {};
const loaded = await pool(games, async (g) => {
const box = await get(`https://statsapi.mlb.com/api/v1/game/${g.pk}/boxscore`);
const out = [];
for (const side of ['home', 'away']) {
const t = box.teams[side];
if (!t) continue;
for (const id of t.batters || []) {
const pl = t.players[`ID${id}`];
const b = pl && pl.stats && pl.stats.batting;
if (!b || b.atBats == null) continue; // did not bat -> absent, not zero
out.push({
date: g.date,
key: nameKey(pl.person && pl.person.fullName),
name: pl.person && pl.person.fullName,
gamePk: g.pk,
hits: b.hits, totalBases: b.totalBases, rbi: b.rbi, runs: b.runs, atBats: b.atBats,
});
}
}
return out;
});
for (const arr of loaded) for (const r of arr) {
const k = `${r.date}|${r.key}`;
// A doubleheader gives two lines; sum them — the prop covers the day.
if (!lines[k]) lines[k] = { ...r, games: 1 };
else {
lines[k].hits += r.hits; lines[k].totalBases += r.totalBases;
lines[k].rbi += r.rbi; lines[k].runs += r.runs; lines[k].atBats += r.atBats;
lines[k].games += 1;
}
}
fs.mkdirSync(path.dirname(BOX_CACHE), { recursive: true });
fs.writeFileSync(BOX_CACHE, JSON.stringify({ dates, lines }));
return lines;
}
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, player_name, line, side, p_win, refused, outcome',
(q) => q.eq('sport', 'mlb').in('stat', STATS).is('outcome', null));
console.error(`[settle] ${snaps.length} unsettled snapshot rows`);
const dates = [...new Set(snaps.map((r) => r.game_date))].sort();
const lines = await battingLines(dates);
const counts = { candidates: snaps.length, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 };
const updates = [];
const seenIds = new Set();
for (const s of snaps) {
// Belt and braces: ordered pagination should make this impossible, and a
// duplicate would double-count a prediction in every downstream measurement.
if (seenIds.has(s.id)) throw new Error(`INTEGRITY: duplicate snapshot id ${s.id}`);
seenIds.add(s.id);
// A row logged after first pitch is not a prediction.
//
// Measured: cycles at ET 21:00/22:00/23:00 on the game date (10,738 rows)
// were captured DURING or AFTER the games they grade, and a further 664 the
// following morning. Games start ~19:05 ET, so the honest cutoff is ET
// first pitch on the game date -- not a UTC date compare, which both keeps
// post-game 01:00-UTC rows and discards legitimate pre-dawn ones.
if (!isPreGame(s.captured_at, s.game_date)) {
counts.post_hoc_logged += 1; counts.unresolvable += 1; continue;
}
const line = knownNumber(s.line);
if (line === null || !s.side) { counts.unresolvable += 1; continue; }
const b = lines[`${s.game_date}|${s.player_key}`];
if (!b) { counts.orphaned += 1; continue; }
const realized = FIELD[s.stat](b);
if (realized === null) { counts.unresolvable += 1; continue; }
// SIDE-ALIGNED, so it matches how p_win is expressed.
const over = realized > line;
const won = String(s.side).toLowerCase() === 'under' ? !over : over;
updates.push({ id: s.id, outcome: won ? 'hit' : 'miss', actual_value: realized });
counts.settled += 1;
}
// CONSERVATION — hard fail.
const acc = counts.settled + counts.unresolvable + counts.orphaned;
if (acc !== counts.candidates) {
throw new Error(`INTEGRITY: conservation violated ${acc} != ${counts.candidates}`);
}
// Hand-verifiable sample.
const sample = updates.slice(0, 12).map((u) => {
const s = snaps.find((x) => x.id === u.id);
return { player: s.player_name, date: s.game_date, stat: s.stat, line: s.line, side: s.side,
realized: u.actual_value, outcome: u.outcome };
});
if (WRITE) {
let written = 0;
for (let i = 0; i < updates.length; i += 500) {
const batch = updates.slice(i, i + 500);
const results = await Promise.all(batch.map((u) => sb.from('model_snapshots')
.update({ outcome: u.outcome, actual_value: u.actual_value, settled_at: new Date().toISOString(), settlement_source: 'statsapi_boxscore' })
.eq('id', u.id).is('outcome', null)));
written += results.filter((r) => !r.error).length;
}
counts.written = written;
}
console.log(JSON.stringify({
mode: WRITE ? 'WRITE' : 'DRY RUN',
counts,
dates_before: 'ledger-only slice',
snapshot_dates: dates.length,
date_span: [dates[0], dates[dates.length - 1]],
hand_verify_sample: sample,
note: 'outcome is SIDE-ALIGNED (matches p_win); actual_value holds the raw realized stat',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+189
View File
@@ -0,0 +1,189 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 1 — is the favourite-longshot bias real WITHOUT a calibration map?
*
* Four orders have refined a stability gate on 19 dates. LODO turned out to be
* structurally underpowered (0.0140.093) and the deploy intervals rest on 24
* date clusters. So we stop certifying the stability of a specific MAP, and ask
* the one question this sample might actually answer:
*
* does the model over-predict its own favourites, robustly?
*
* That claim is MODEL-FREE and MAP-FREE — it is a property of (p_win, outcome)
* pairs, needs no isotonic fit, and can therefore be tested without any of the
* machinery whose stability we cannot certify.
*
* ── DATE-BLOCK BOOTSTRAP ─────────────────────────────────────────────────
* Resampling ROWS would treat 200 props from one night as 200 readings of that
* night's offensive environment. Whole DATES are resampled instead, which is the
* honest unit and a far harsher one at 517 dates.
*
* VERDICT is pre-stated: ROBUST iff the >0.9 over-prediction sign survives in
* >=95% of pooled date-block resamples AND replicates in >=3 of 4 stats on the
* same criterion. Anything else is NOT-ROBUST, and NOT-ROBUST means we serve raw.
*
* SUPABASE_URL=... node scripts/test-favourite-bias-robust.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
const FAVOURITE_FLOOR = 0.9;
const ITERS = 5000;
/** Pre-stated pass marks. */
const SIGN_STABILITY_REQUIRED = 0.95;
const STATS_MUST_REPLICATE = 3;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
/** Over-prediction in the favourite bin: predicted realized. Positive = over. */
function favouriteBias(rows) {
const fav = rows.filter((r) => r.p >= FAVOURITE_FLOOR);
if (fav.length < 5) return null;
return mean(fav.map((r) => r.p)) - mean(fav.map((r) => r.won));
}
/** Resample whole DATES with replacement; report how often the sign survives. */
function dateBlockSignStability(rows, seed) {
const byDate = new Map();
for (const r of rows) {
if (!byDate.has(r.date)) byDate.set(r.date, []);
byDate.get(r.date).push(r);
}
const keys = [...byDate.keys()];
const rnd = makeRnd(seed);
let positive = 0; let indeterminate = 0; const draws = [];
for (let it = 0; it < ITERS; it += 1) {
const sample = [];
for (let i = 0; i < keys.length; i += 1) sample.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
const b = favouriteBias(sample);
// A resample with too few favourites cannot speak — counted, never guessed.
if (b === null) { indeterminate += 1; continue; }
draws.push(b);
if (b > 0) positive += 1;
}
const usable = ITERS - indeterminate;
draws.sort((a, b) => a - b);
return {
date_blocks: keys.length,
usable_resamples: usable,
indeterminate_resamples: indeterminate,
sign_stability: usable ? round4(positive / usable) : null,
ci_90: draws.length ? [round4(draws[Math.floor(draws.length * 0.05)]), round4(draws[Math.floor(draws.length * 0.95)])] : null,
};
}
function deciles(rows) {
const out = [];
for (let lo = 0.3; lo < 1.0; lo += 0.1) {
const hi = lo + 0.1;
const slice = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (slice.length < 15) continue;
const pred = mean(slice.map((r) => r.p));
const real = mean(slice.map((r) => r.won));
out.push({ bin: [round2(lo), round2(hi)], n: slice.length, predicted: round4(pred), realized: round4(real), over_prediction: round4(pred - real) });
}
return out;
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
const byStat = {}; const pooled = [];
for (const stat of STATS) byStat[stat] = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[r.stat](b));
if (v === null) continue;
const over = v > L;
const row = { date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 };
byStat[r.stat].push(row); pooled.push(row);
}
const pooledResult = {
n: pooled.length,
deciles: deciles(pooled),
favourite_bias: round4(favouriteBias(pooled)),
...dateBlockSignStability(pooled, 20260808),
};
const perStat = {};
let replicated = 0;
for (const stat of STATS) {
const rows = byStat[stat];
const fb = favouriteBias(rows);
const stab = dateBlockSignStability(rows, 20260808);
const ok = fb !== null && fb > 0 && stab.sign_stability !== null && stab.sign_stability >= SIGN_STABILITY_REQUIRED;
if (ok) replicated += 1;
perStat[stat] = { n: rows.length, deciles: deciles(rows), favourite_bias: fb === null ? null : round4(fb), ...stab, replicates: ok };
}
const pooledOk = pooledResult.favourite_bias > 0 && pooledResult.sign_stability >= SIGN_STABILITY_REQUIRED;
const verdict = pooledOk && replicated >= STATS_MUST_REPLICATE ? 'ROBUST' : 'NOT-ROBUST';
console.log(JSON.stringify({
phase: 'PHASE 1 — model-free, map-free favourite-longshot bias test',
criteria: { sign_stability_required: SIGN_STABILITY_REQUIRED, stats_must_replicate: STATS_MUST_REPLICATE, favourite_floor: FAVOURITE_FLOOR },
pooled: pooledResult,
per_stat: perStat,
stats_replicating: replicated,
VERDICT: verdict,
consequence: verdict === 'ROBUST'
? 'proceed to a low-parameter correction, validated as a NEW estimator'
: 'serve raw; the bias is not certifiable on this sample',
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round2 = (v) => Math.round(v * 100) / 100;
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 3 — validate the low-parameter correction as a NEW estimator.
*
* No grandfathering: it must beat RAW out-of-sample with a DATE-BLOCK bootstrap
* interval excluding zero. It is also scored against the retired isotonic map on
* the identical held-out rows, so the swap is a measured comparison rather than
* a preference.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const lp = require('../src/services/model/lowParamCalibrator');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000; const ITERS = 4000;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error; if (!data.length) break; o.push(...data); if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) { let s = seed >>> 0; return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; }; }
/** Paired date-block bootstrap on a Brier difference. */
function dateBlockCI(rows, keyA, keyB, seed) {
const byDate = new Map();
for (const r of rows) { if (!byDate.has(r.date)) byDate.set(r.date, []); byDate.get(r.date).push(r); }
const keys = [...byDate.keys()]; const rnd = makeRnd(seed); const diffs = [];
for (let it = 0; it < ITERS; it += 1) {
const s = [];
for (let i = 0; i < keys.length; i += 1) s.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
const a = guards.safeBrier(s.map((r) => r[keyA]), s.map((r) => r.won));
const b = guards.safeBrier(s.map((r) => r[keyB]), s.map((r) => r.won));
if (a === null || b === null) continue;
diffs.push(a - b);
}
diffs.sort((x, y) => x - y);
return diffs.length
? { ci: [round4(diffs[Math.floor(diffs.length * 0.025)]), round4(diffs[Math.floor(diffs.length * 0.975)])], date_blocks: keys.length }
: { ci: null, date_blocks: keys.length };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots', 'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const out = {};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const over = v > L;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
rows.sort((a, b) => String(a.date).localeCompare(String(b.date)));
const dates = [...new Set(rows.map((r) => r.date))].sort();
const perDate = new Map(); for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) { acc += perDate.get(d); if (acc >= rows.length * 0.45) { cut = d; break; } }
const fit = rows.filter((r) => r.date < cut);
const ev = rows.filter((r) => r.date >= cut);
const platt = lp.fitPlatt(fit);
const iso = cal.fitIsotonic(fit.map((r) => ({ p: r.p, won: r.won })));
if (!platt || ev.length < 50) {
out[stat] = { n: rows.length, fit_n: fit.length, eval_n: ev.length, decision: 'REFUSE', reason: 'no low-parameter fit or too little held out' };
continue;
}
const scored = ev.map((r) => ({ ...r, plat: lp.applyPlatt(platt, r.p), isoP: iso ? cal.applyIsotonic(iso, r.p) : null }))
.filter((r) => knownNumber(r.plat) !== null);
const ys = scored.map((r) => r.won);
const bRaw = guards.safeBrier(scored.map((r) => r.p), ys);
const bPlat = guards.safeBrier(scored.map((r) => r.plat), ys);
const isoRows = scored.filter((r) => knownNumber(r.isoP) !== null);
const bIso = isoRows.length ? guards.safeBrier(isoRows.map((r) => r.isoP), isoRows.map((r) => r.won)) : null;
const vsRaw = dateBlockCI(scored, 'plat', 'p', 20260808);
const vsIso = isoRows.length ? dateBlockCI(isoRows, 'plat', 'isoP', 20260808) : { ci: null };
const beatsRaw = bPlat < bRaw && vsRaw.ci && vsRaw.ci[1] < 0;
out[stat] = {
n: rows.length, dates: dates.length, split_at: cut, fit_n: fit.length, eval_n: scored.length,
platt: { a: platt.a, b: platt.b, flattens: platt.flattens, fit_dates: platt.fit_dates, shrinkage: platt.shrinkage },
brier_raw: round4(bRaw), brier_lowparam: round4(bPlat), brier_isotonic: bIso === null ? null : round4(bIso),
delta_vs_raw: round4(bPlat - bRaw), ci_vs_raw: vsRaw.ci, eval_date_blocks: vsRaw.date_blocks,
delta_vs_isotonic: bIso === null ? null : round4(bPlat - bIso), ci_vs_isotonic: vsIso.ci,
decision: beatsRaw ? 'DEPLOY-PROVISIONAL' : 'REFUSE',
reason: beatsRaw ? 'beats raw out-of-sample with a date-block interval excluding zero'
: 'does not beat raw at a date-block interval excluding zero',
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
+53
View File
@@ -0,0 +1,53 @@
# The 83-glyph taxonomy — doctrine
**RULED. This is the shared law the archetype chat and the build chat both obey.**
## The rule
**83 designed glyphs = the full four-sport archetype taxonomy** (MLB / WNBA /
NBA / Soccer). The artwork is complete; the *models* are not.
**A glyph renders ONLY where its archetype is modeled and proven.**
| set | count | state |
|---|---|---|
| designed glyphs | **83** | complete, in `specs/design-reference/assets/glyphs/` |
| backend registry archetypes | **41** | `archetypeService.ARCHETYPES` |
| **mapped and wired** | **39** | live, colours matching the registry exactly |
| **designed, no backend archetype** | **44** | **DORMANT** — slots for WNBA/NBA/Soccer |
| registry archetypes with no glyph | **2** | `DUAL THREAT`, `PAINT BOSS`**design gap** |
## Why dormancy rather than wiring
Wiring the 44 would mean **inventing 44 archetypes to consume artwork**. An
archetype that exists because a glyph exists is decoration presented as
classification — a mark on a card asserting the model recognised something it
cannot produce. **Decoration-as-data is forbidden.**
The dormant glyphs are not waste. They are **designed slots**, and each activates
when its sport's archetype system is built and clears the two-part gate — the
same bar every factor in this programme faces.
## What this preserves
- **The full design vision.** All 83 marks stay in the package; none is deleted
or redrawn.
- **The Truth Law.** No glyph appears for an archetype the model cannot produce.
- **The activation path.** Building WNBA/NBA/Soccer archetypes lights their
glyphs automatically — the artwork is already there and already colour-matched.
## Consequences for build
1. Do **not** add a registry archetype to consume a glyph. The archetype must be
earned by a classifier that produces it from real features.
2. Do **not** render a dormant glyph as a placeholder, sample, or "coming soon"
mark on any data surface.
3. `DUAL THREAT` and `PAINT BOSS` are **flagged to the design side** — they are
modeled archetypes with no mark, the mirror image of the dormant 44.
4. When a sport's archetypes ship, wiring is a MANIFEST lookup, not new art.
## Status
MLB archetypes render live. WNBA, NBA and Soccer have archetype **registries**
but no proven factor model, so their marks stay dormant — consistent with
`specs/BOARD-2026-08-07.md`, which records only MLB batters as MODELED.
+132
View File
@@ -0,0 +1,132 @@
# The real board — read from the repo, 2026-08-07
Inventory only. Every line cites a file or a query. Anything unverifiable is
marked UNKNOWN rather than asserted.
---
## PHASE 0 — Design / terminal
| item | state | evidence |
|---|---|---|
| **Scanner blue-boundary format** | **DONE** | `scan/page.tsx:715-717`*"the blue boundary channel (`--priced-out`): line · BK odds · ◆ fair"* and *"renders NEUTRAL, not amber (blue-boundary law)"*. Amber survives only as an unrelated CTA (`:983`) and a status line (`:846`). |
| **MovementStrip** | **NOT-STARTED** *(as named)* | No component by that name. Movement renders via `GradeShift.tsx`, `GameCard.tsx`, `MarketBreadth.tsx`, `FuturesBoard.tsx`. **UNKNOWN whether the spec wants a distinct strip or is satisfied by these.** |
| **THE WIRE** | **PARTIAL** | `components/vyndr/NewsWire.tsx` exists and is mounted — but only inside `ExploreHub.tsx`. No standalone wire surface. |
| **Book comparison** | **PARTIAL — built, unmounted** | `vyndr/BookComparisonPanel.tsx` + `components/BookComparison.tsx` both exist; grep of `web/src/app` returns **zero** mounting pages. Same built-but-unread class the reachability guard was written for; **not covered by that guard** (contract is grade fields only). |
| **Article media** | **NOT-STARTED** | No component. `/blog` exists. |
| **Offseason hub** | **NOT-STARTED** | No component, no route. |
| **Terminal** | **DONE — deliberately retired** | `app/terminal/page.tsx` redirects to `/dashboard`; layouts preserved unrouted in `components/intel/TerminalTemplates.tsx` (S57). Not debt. |
| **Screen conversion (S3639 arc)** | **DONE except one** | 47 route dirs under `app/`. `RouteStub` survives in exactly **one** file: `app/notifications/page.tsx`. Every other screen has a real page. |
**Sports surfaces that exist as pages:** `soccer` (414 lines), `desk` (160),
`intelligence` (160), `marketplace` (123). No `fight` route.
---
## PHASE 1 — Sports × role
Archetype registry (`archetypeService.ARCHETYPES`): **nba 15, mlb 15, soccer 6,
wnba 5** — all four sports have real archetype registries, not names only.
`snapshotService.ACTIVE_SPORTS = ['mlb','nba','wnba','soccer']` — the pipeline
runs all four.
| sport · role | state | evidence |
|---|---|---|
| **MLB · batters** | **MODELED** | The whole session. `hitsFactors` (3 proven factors, transmitting), `servedGrade`, repaired champion, settled ledger. |
| **MLB · pitchers** | **SCAFFOLDED** | `model/pitcherEngine.js` (261 lines, own FLAME/SCALPEL/SINKER archetypes) — **read by no serving code** (grep: zero non-test consumers). Strikeouts n=57 settled vs a 500 gate. Base rate repaired by the shared `getStatRows` MLB fix. |
| **WNBA** | **SCAFFOLDED** | Archetypes exist; settles via ESPN box scores (376 rows historically); no factor model. Base-rate path fixed this session. |
| **NBA** | **SCAFFOLDED, dormant** | Archetypes exist; offline (Python service down, off-season). Base-rate paths fixed dormant at `55b210c`. |
| **Soccer** | **SCAFFOLDED** | 6 archetypes, a real `/soccer` page, feature extractor exists. No factor model, no settled outcomes. |
**Nothing but MLB batters is MODELED.** Everything else is scaffolding with a
sound base rate and no proven factors.
---
## PHASE 2 — Wiring / data-integrity debt
### `edge_pct` — the flag is half right, and the diagnosis was wrong
**Not a scale bug.** `analyzeViaEngine1.edgePctFor` computes
`(model line) / line`, signed by direction — arithmetically correct. It
*explodes on small lines*: a 5.5 projection against a 0.5 line is a legitimate
1000%. Live top values: **900, 860, 700, 700, 660** on **68,364 rows**.
**Consuming surfaces: 15 backend files + 10 frontend files.** But
`edge_pct` itself reaches a user only through `alt_lines` typing in
`scan/page.tsx:64` — the grade card's `edge` is computed independently in
`gradeAdapter` via `computeEdge`. So the blast radius is smaller than the file
count implies.
**`ev_pct` DOES render** — 5 frontend files (`PriceTriplet`, `GradeResultCard`,
`LiveHeroProp`, scan page + route). The "ev_pct renders nowhere" flag is **stale**.
45,125 rows carry it.
**Classification: PARTIAL — a real display defect (a 900% edge is not a sentence
we can defend), not a broken computation.**
### `opp_rank_stat`
`refreshTeamStats` is wired into `snapshotService:361`. **UNKNOWN whether it
populates in production** — not verifiable from the repo, needs a live probe.
### A-emit
**Resolved.** `servedGrade.UNISSUABLE = ['A+','A','A-']`; post-cutover serving
cannot emit one. Historical snapshot rows still carry engine1 A's — that is
retired data, not live behaviour.
### Other
| item | state |
|---|---|
| `VYNDR_INTERNAL_KEY` | present in `app.js`, `preflight.js`. **Rotation status UNKNOWN** from the repo. |
| void / DNP handling | handled in `ledgerService` (`:64`, `:698`) — absence means DNP/postponed/not-final. **The ~9% rate is UNVERIFIED** here. |
| **Guard coverage gap** | reachability contract covers **grade fields only** — 0 references to `edge_pct`/`ev_pct`. Book comparison and other unmounted components are **not** guarded. |
---
## PHASE 3 — THE BOARD, ordered (half-done first)
| # | item | lane | state | serving-path? |
|---|---|---|---|---|
| 1 | **Book comparison — built, unmounted** | design | PARTIAL | **N** |
| 2 | **`edge_pct` display defect (900%)** | integrity | PARTIAL | **Y** |
| 3 | **THE WIRE — only inside ExploreHub** | design | PARTIAL | **N** |
| 4 | **MLB pitchers — engine unwired** | sports | PARTIAL | **Y** |
| 5 | **`/notifications` RouteStub** | design | PARTIAL | **N** |
| 6 | Guard coverage → non-grade surfaces | integrity | NOT-STARTED | **N** |
| 7 | Article media | design | NOT-STARTED | **N** |
| 8 | Offseason hub | design | NOT-STARTED | **N** |
| 9 | MovementStrip *(if distinct from GradeShift)* | design | UNKNOWN | **N** |
| 10 | WNBA / NBA / Soccer factor models | sports | NOT-STARTED | **Y** |
| 11 | `opp_rank_stat` prod population | integrity | UNKNOWN | **Y** |
| 12 | Internal-key rotation | ops | UNKNOWN | **N** |
### Accrual sensitivity
**Isolated — buildable now without touching the re-audit clock (7 items):**
1, 3, 5, 6, 7, 8, 9 — all design/guard work, zero serving-path contact.
**Serving-path — will muddy the accruing repaired-champion dates (5 items):**
2, 4, 10, 11, and any model work. Every one changes what a snapshot writes, so
rows produced after the change are not comparable to rows before it. **If any of
these ship, the eligible-date clock arguably restarts** — the same reasoning that
voided the shadow duel.
**#2 (`edge_pct`) is the sharpest tension on the board:** it is a real honesty
defect a user can see, and fixing it touches the serving path mid-accrual.
### The clock, today
```
eligible dates: 0 — no settled rows carry engine1@2026-08-07-fullwindow
calibration_refit 0/10 WAITING (~2 weeks)
hits_factor_lift 0/10 WAITING (~2 weeks)
prior_verdict_reaudit 0/14 WAITING (3+ weeks)
rbi_lineup_slot_gate 0/14 WAITING (3+ weeks)
```
Two tracks are now visible: **modeling is date-blocked**; **seven design/guard
items are buildable today with no clock impact.**
+111
View File
@@ -0,0 +1,111 @@
# The content engine — posts that structurally cannot lie
## Phase 0 — architecture
`src/services/content/contentEngine.js`. Three mechanisms make Truth Law
structural rather than careful:
1. **Copy is token-substituted.** Every factual claim is a `{token}` resolved
against pulled facts. An unbacked token **refuses to render** — there is no
code path producing a plausible default.
2. **The fact contract is asserted first.** A template declares required fields;
they are checked *before any string is built*.
3. **Card and copy share one fact object.** They cannot diverge.
**No live model writes factual claims.** The voice is in the template, the facts
are pulled. A voice-polish port is reserved and deliberately unwired — an LLM
that can rewrite a sentence can rewrite a number.
**Read-only on every source.** Zero writes to serving, model or ledger tables, so
zero effect on the accrual clock.
### The Truth-Law proof — 18 tests
| the guard | what it prevents |
|---|---|
| unbacked token refuses | `{edge}` rendering as `undefined` or an empty hole |
| card tokens gated too | a caption that's honest beside a card that isn't |
| `render()` throws directly | a caller bypassing the gate |
| `null` never renders as `"null"` | absence dressed as data |
| **`0` IS present** | *"0 cleared B+"* is our most honest post — deleting it would be `Number(null)===0` in reverse |
| `NaN`/`Infinity` absent | arithmetic failures are not facts |
| contract gap names the field | a silent half-post |
| pull failure skips | a post built on a dead source |
## Phase 1 — three templates, real output
**HOT HITTERS** (from the repaired full-season log, not a ten-game slice):
> Jahmai Jones is hitting 60% over his last 10. His season number is 26%.
> That gap is the whole point. Everybody else is guessing at it.
**THE HONESTY FLEX** — the differentiator, and every number is ours:
> WE GRADED 2140 PROPS TONIGHT. 70 CLEARED B+.
> That's 3%. The other 42% we can't separate from the baseline, and we say so on the card instead of calling them leans.
> We do not issue A+, A, A-. No band of this model has ever hit at a rate that would justify one.
> Everybody else's card is all A's. Ask them what their A actually hits.
**STREAK LIST** — verified from settled outcomes only:
> Nathan Church has a 7-game hit streak. Live, verified off settled results only.
> Every game in these ran to a final. We don't count a pending night to make a number look better.
### The bug the engine caught in itself
The first run emitted *"No hitter is meaningfully hot tonight — we could dress up
a middling week as a streak. We don't."*
**That was false.** The box-score cache spans only the settled snapshot window,
so **every** player had fewer than 20 games and the pool was empty. A broken pull
was publishing as considered editorial judgement — **the fourth appearance of
this class tonight, and the first where our own honesty copy was the disguise.**
Fixed structurally: an `absent()` variant may now **decline to speak**. The
template separates *no candidates at all* (SKIP with a reason) from *candidates
judged, none hot* (honest absence). Both cases are locked by test.
Source corrected to `mlbStatsAdapter.fullLog` — the same log the repaired
champion reads.
## Phase 2 — the card
`cardRenderer.js`, SVG rather than canvas: it is text, so it diffs in review and
its numbers are **greppable** — which matters when the entire claim is that the
numbers are real. A card whose contents can't be inspected without opening an
image is a poor fit for a Truth-Law product.
Brand: VYND white + R green `#00D4A0`, slashed-Y, scanline field, mono
throughout. The card never formats its own facts — every string arrives already
rendered and gate-checked, so caption and card cannot disagree. A test asserts
the pulled number appears in the emitted SVG.
## Phase 3 — posting-ready, and extending it
```
SUPABASE_URL=... node scripts/generate-content.js
-> .content-out/2026-08-07/hot_hitters.txt + .svg
-> .content-out/2026-08-07/honesty_flex.txt + .svg
-> .content-out/2026-08-07/streak_list.txt + .svg
```
Kev posts; the engine generates.
### Adding template N+1 — registry entry only, no engine change
```js
registerTemplate({
id, sport, requires: ['dotted.paths'],
pull: async (deps) => facts, // the ONLY place data enters
copy: () => 'text with {tokens}',
card: () => ({ title, subtitle, lines }),
absent: (gaps, facts) => ({ copy, card }) // or { skip: 'reason' }
});
```
**Queued (stubs, not built):** hot takes · daily honest reads · *"grades we
DIDN'T give"* · cross-sport streak variants (the streak template is already
sport-agnostic — it takes settled outcomes and a noun, so NFL TD streaks or NBA
made-three streaks need only that sport's settled data).
## Isolation
Read-only throughout. `p_win`, the model and the serving path are untouched;
the eligible-date clock is unaffected. **0 eligible dates today, unchanged.**
+198
View File
@@ -0,0 +1,198 @@
# VYNDR build handoff — 2026-08-08
Written from the repo, not from summary. Every claim below was grep- or
run-verified at `71d3b7b`. Where something could not be verified from the repo
it says UNKNOWN.
## Repo state
```
HEAD 71d3b7b78692c31ba0596ec88874986ff1b58116
gitea main 71d3b7b78692c31ba0596ec88874986ff1b58116 (in sync)
tree clean — 0 modified/untracked
tests 4,539 passing / 4 skipped
web build exit 0
```
**Push remote is `gitea`, not `origin`.** `origin` is GitHub and has no
credential on this machine; every push must be `git push gitea main`. This cost
twenty commits of false "no credentials" diagnosis — see
`docs/GIT-PUSH-DIAGNOSIS.md`.
---
## BUILD STATE per surface
### DONE this arc — verified present
| surface | file | lines |
|---|---|---|
| **E1 movement strip** | `web/src/components/vyndr/MovementStrip.tsx` | 113 |
| **F9F11 offseason hub shell** | `web/src/app/offseason/page.tsx` + `components/vyndr/OffseasonHub.tsx` | 21 + hub |
| **E10 Report issue template** | `src/services/report/reportTemplate.js` | 169 |
| **E12 /report archive** | `web/src/app/report/page.tsx` + `components/vyndr/ReportArchive.tsx` + `src/routes/report.js` | 17 + comp + route |
| **Content engine** | `src/services/content/contentEngine.js` + 3 templates + `cardRenderer.js` | 172 + |
| **Content studio API + preview** | `src/routes/contentStudio.js` + `web/src/app/studio/page.tsx` | 175 + 131 |
| **Wave-D1 motion primitives** | `web/src/lib/motion.js` (E17/E18/E27/E28) | 100 |
| **Honest served grade** | `src/services/model/servedGrade.js` | 142 |
| **Archetype doctrine** | `specs/ARCHETYPE-TAXONOMY-DOCTRINE.md` | 53 |
Also DONE and verified: **D1 glyph library** (39 of 83 wired — every glyph that
maps to a real archetype; colours match the registry exactly), **A1 card token**
(`--bg-1`, 32 consumers), **B1 boundary channel** (`--priced-out`, 4 consumers),
**book comparison** (mounted via `GradeResultCard → scan/page`), **THE WIRE**
(`vyndr/Ticker`).
### GATED — verified absent, with the gate named
| surface | gate | verified |
|---|---|---|
| **F5 article media** | card-system reconciliation | `ArticleHero`/`StatCallout`/`PullQuote` = 0 files |
| **E16/F8 share cards + crops** | card-system reconciliation *and* the resolution tail has no share-card generation step | `ShareCard.tsx` has **0 importers** |
| **In-season hub IA** | the content formula (social chat) | no spec exists anywhere |
| **E9 calibration curve** | model accrual — 0 eligible dates | `CalibrationCurve` = 0 files |
| **E15 Price Triplet MODEL leg** | model — EV layer | renders honest `NO_MODEL` |
| **E2 BookChip tiles / E6 push-to-book** | licensing / affiliate approval | every book `enabled:false` |
| **E3 best-number crown** | measurement — measured flat | `CrownBadge` = 0 files |
| **E13 Scanner S6** | another build order | — |
**There are no ungated Wave-2 targets left.** The next move is a decision, not a
build.
---
## THE TWO OPEN DECISIONS
### 1. Card-system reconciliation — unblocks F5 *and* E16/F8 together
The content engine emits **1080×1350 SVG** cards (`src/services/content/cardRenderer.js`).
The design package defines **five master sizes** with layouts and rasterised PNGs
in `specs/design-reference/exports/`: settle 1080×1350, story 1080×1920, square
1080, X 1200×675, record 1080×1350, article OG 1200×630.
**I built the renderer without checking whether a designed card system existed.
It did.** They are a parallel invention — not wrong, but they must become one
visual language before F5 (whose hero graphics are generated data visuals) or
E16 can proceed.
The decision: does the content engine conform to the E16 masters, or do the
masters absorb the engine's generator? **Cheapest unblock on the board — it
frees two gated items at once.**
### 2. In-season hub information architecture
`Vyndr Offseason.dc.html` specifies an **offseason** hub, and its shell is built.
Nothing specifies what that surface is **in-season**, or how content, articles,
wire and the live slate share year-round navigation. Every component exists; the
composition does not.
The spec is explicit that **sport state lives in the sport tab** (`NFL · CAMP 5D`)
and never in a separate offseason tab — so this is not a new route, it is a mode.
Gated on the social chat's content formula, which determines the recurring slots.
---
## ACCRUAL CLOCK
```
champion marker: engine1@2026-08-07-fullwindow
calibration_refit 0/10 WAITING (~2 weeks)
hits_factor_lift 0/10 WAITING (~2 weeks)
prior_verdict_reaudit 0/14 WAITING (3+ weeks)
rbi_lineup_slot_gate 0/14 WAITING (3+ weeks)
blocked: no settled rows yet carry the repaired champion marker
```
**FIRST TRIGGER: 10 eligible calibration dates → calibration re-fit.** Then, in
order: hits factor lift → prior verdict re-audit → rbi lineup-slot gate.
**Thresholds are ATTEMPT floors, not TRUST floors.** Reaching 10 dates means the
re-fit *can be measured*, not that it is trustworthy. A 10-date map is thin,
deploys PROVISIONAL with auto-demotion, and its interval will be wide.
**No measurement on reconstructions of the retired forecast.** Anything requiring
settled data waits. `src/services/model/reAuditEligibility.js` enforces it —
eligibility is a version marker, counted in DATES not rows, and a mixed table
counts only the repaired rows.
---
## STANDING DOCTRINE
**Truth Law.** No fabricated data anywhere — including **designer sample data**.
The design files' numbers (Nabers 1,120.5, Wembanyama +420→+330, Nº 128, DAY
RECORD 94) are a *spec for what a live feed renders*, never content to paste.
Tests assert none of it ships. `Number(null) === 0` is the classic breach; note
its inverse also bites — **zero is a real fact** ("0 cleared B+" is our most
honest post).
**83-glyph four-sport taxonomy.** A glyph renders ONLY where its archetype is
modeled and proven. 39 wired; **44 dormant slots** for WNBA/NBA/Soccer. Wiring
them would mean inventing 44 archetypes to consume artwork — decoration
presented as classification, forbidden. `DUAL THREAT` and `PAINT BOSS` are
modeled archetypes with no mark: the mirror gap, flagged to design.
**Repaired champion.** The forecaster read `res.last10`**ten games** — as its
season rate. Fixed to the full season log with recency weight 0.40 → 0.20.
Resolution: hits 0.00251 → 0.00817. **Every factor verdict in the programme was
measured against the broken baseline and may deserve re-audit** — direction
unknown, not pre-priced.
**Calibration is WITHDRAWN.** `CALIBRATION_DEPLOYED = []`. The maps were fitted
on the retired forecast; refitting today would refit it again. Served `p_win` is
repaired-champion raw.
**Grade doctrine.** `A+`, `A`, `A-` are **unissuable** — structurally, not rarely.
Ceiling is **B+ at 0.663 realized against a 0.6005 baseline**. Bands that cannot
be separated from the baseline SAY so. The served letter derives from `p_win`;
`engine1.grade` is preserved as `engine_grade` and read by no serving code.
**Two guards, both in CI:**
- `renderReachability.test.js` — a promised field must trace payload → adapter →
component → **mounted**. Built-but-unread killed three surfaces before this.
- `baseRateWindow.test.js` — no fixed N may stand in for a season. Six paths
across two sports fell to that class.
**Compose, don't fork.** The recurring failure this arc: building beside an
existing thing instead of on it (cards vs E16, `MovementStrip` nearly vs
`gradeShift`, `reportTemplate` nearly vs `newsletterService`). Check for the
existing implementation first — twice it was already there.
---
## PARALLEL WORKSTREAM SEEDS
### A. Social strategy chat
> VYNDR's content engine ships posts from real data (`src/services/content/`,
> three templates, Truth-Law enforced — an unbacked token refuses to render).
> Review `/studio` output and `specs/CONTENT-ENGINE.md`. Produce: the content
> formula/calendar (what posts recur, on what cadence), and **the card-format
> decision** — the engine's 1080×1350 renderer vs the five designed masters in
> `specs/design-reference/exports/`. That decision unblocks F5 article media and
> E16/F8 on the build side. Voice: `specs/VOICE.md` governs; the engine templates
> encode a sharper register — decide whether that becomes the house voice.
### B. Multi-sport archetype build
> VYNDR has archetype registries for MLB (15), NBA (15), Soccer (6), WNBA (5) —
> but **only MLB batters is MODELED**; the rest have no proven factor model, so
> their glyphs stay dormant per `specs/ARCHETYPE-TAXONOMY-DOCTRINE.md`. 44 of 83
> designed glyphs await their sport. Pick a sport, build its archetype
> classifier from real features, and take it through the two-part gate
> (`src/services/model/factorGate.js` — must MOVE the prediction and improve
> out-of-sample Brier, cumulative-Bonferroni corrected). Prerequisite: that
> sport needs settled outcomes. WNBA settles via ESPN box scores; NBA and Soccer
> do not settle at all yet, which is the real first task there.
---
## Where to start in the next chat
1. Take one of the two open decisions (card system is cheapest — unblocks two).
2. Or run a parallel workstream seed above.
3. **Do not** start model work — it is date-blocked and will stay so for ~2 weeks.
4. Verify before building: this arc corrected its own board four times. `git
grep` beats any summary, including this one.
+92
View File
@@ -0,0 +1,92 @@
# What "media hub" already means in VYNDR's design language
Read-only inventory. Nothing designed, built or mounted.
## PHASE 0 — the design source of truth exists, and so does a gap audit
| artifact | what it is |
|---|---|
| `specs/DESIGN-SPEC.md` | **DESIGN SPEC v2** — governing laws, colour contract, entity layer, data-display standard, motion, archetype system, conversion architecture, empty/error system. 84 lines, supersedes v1. |
| `specs/design-reference/` (Jul 22) | **The design package itself** — 7 `.dc.html` surface files, 83 archetype glyph SVGs + MANIFEST, 7 rasterised share-card PNG masters. |
| `specs/design-reference/HANDOFF.md` | The code handoff: exact tokens, brand, glyph library, behaviour, laws. |
| **`specs/design-vs-build-gap-audit.md`** | **A 61-item design-vs-build audit already exists** (2026-07-31, audited at `bf7c0a3`), every claim grep-verified, with a wave-ordered build plan. |
**The inventory this order asks for was largely already done on 2026-07-31.** The
useful work is reconciling it, not redoing it.
## PHASE 1 — per media surface: spec, and build state
| surface | design spec? | built state | evidence |
|---|---|---|---|
| **Offseason hub** | **YES — a full surface file** | **NOT-STARTED** | `Vyndr Offseason.dc.html` (119KB) defines hub home (NFL desktop + 390), an **NBA Summer-League variant with an `OUTLOOK ONLY / NOT GRADED` honesty block**, a season-long board (open→NOW→VYNDR triplet), a season-read reveal with `WHAT WOULD CHANGE THIS READ`, a news/outlook feed with row anatomy, and a **quiet-wire empty state**. Gap audit: F9/F10/F11, *"design complete, no blocker but big."* |
| **Article media** | **YES — surface S3** | **NOT-STARTED** | Hero template + **4 hero graphic archetypes** (line path / distribution / mark-at-scale / matchup card — *all generated data visuals, never stock*), inline figures (stat-callout triptych, comparison bars, pull quote — **one max**), **caption law: every figure names its data**, article card, OG 1200×630 (master already rasterised in `exports/`). Gap audit F5. Zero components built. |
| **THE WIRE** | **YES** | **BUILT** (E19) | Spec'd in `Vyndr System.dc.html`: timestamped entry, ~6s hold, tag coloured by meaning. Built as `vyndr/Ticker` on real `/api/ticker` exhaust. **`NewsWire.tsx` is a different thing** — the offseason news/outlook feed, mounted in ExploreHub. |
| **Movement strip** | **YES — E1, a named primitive** | **ABSENT** | *"steps not curves, green only when the move favours the read, FLAT = hairline + `FLAT · [N]D`"*, row 86×20 silent / full-width annotated on reveal. **This corrects my 2026-08-07 board, which listed it UNKNOWN / possibly-satisfied-by-GradeShift.** It is a specified primitive that does not exist. |
| **Share-card masters ×5** | **YES — E16** | **ABSENT as product** | settle 1080×1350, story 1080×1920, square 1080, X 1200×675, record 1080×1350, article OG 1200×630. PNG masters in `exports/`; `ShareCard.tsx` has **0 importers**. Gap audit blocks it on the resolution tail having no share-card generation step. |
| **Calibration curve** | YES — E9 | ABSENT | dots vs dashed perfect line, dot size = sample, buckets under N30 hollow. |
| **The Report email / `/report` archive** | YES — E10/E12 | PARTIAL / ABSENT | `newsletterService` builds an email; the designed dark-billboard-over-light-paper hybrid is not built. `/report` redirects to `/blog`. |
| **ExploreHub** | **NO SPEC** | BUILT | Not a designed surface. It is where `NewsWire` currently lives — a de-facto host, not the intended hub. |
## PHASE 2 — what an on-site media hub would consist of
**From surfaces already designed** (no new design needed):
1. **Offseason hub shell** — F9/F10/F11, fully spec'd including its honesty block and quiet-wire empty state.
2. **News/outlook feed** — spec'd row anatomy; `NewsWire.tsx` is a partial implementation of it.
3. **Article media** — S3, fully spec'd hero + figure system.
4. **THE WIRE** — built, reusable.
5. **Share/OG cards** — designed at exact sizes, masters exported.
### The gap that needs design work authored
**Only one thing is genuinely un-designed: the hub's INFORMATION ARCHITECTURE
across seasons.** The Offseason file specifies an *offseason* hub; there is no
spec for what the same surface is **in-season**, or how content, articles, wire
and the live slate share one navigation. Every component exists on paper; **their
composition into a year-round media surface does not.**
### One finding that touches work I just shipped
**I built a card renderer at 1080×1350 without checking whether a designed card
system existed. It does** — five master sizes with defined layouts, and
rasterised PNGs in `exports/`. The content engine's cards are a parallel
invention. They are not wrong, but they should conform to E16/F8 rather than
diverge, and that reconciliation is a design decision, not a code one.
## PHASE 3 — the map
| surface | spec? | built | part of hub? | needs design work? |
|---|---|---|---|---|
| Offseason hub shell | ✅ | ❌ | **core** | no — build it |
| News/outlook feed | ✅ | partial | **core** | no |
| Article media (S3) | ✅ | ❌ | **core** | no — build it |
| THE WIRE | ✅ | ✅ | yes | no |
| Movement strip (E1) | ✅ | ❌ | adjacent | no |
| Share/OG cards (E16) | ✅ | ❌ | yes | **reconcile with content-engine cards** |
| Content-engine posts | ❌ | ✅ | **yes** | **yes — no spec exists** |
| ExploreHub | ❌ | ✅ | host today | **yes — is it the hub or a placeholder?** |
| **Year-round hub IA** | ❌ | ❌ | **the spine** | **YES — the real design gap** |
### Coordination with the social-strategy chat
These should be decided **there first**, then reflected on-site:
- **The content formula/calendar** determines what the hub's recurring slots
*are*. Designing slots before the formula exists risks a surface shaped around
guesses.
- **Card format reconciliation** — off-site posts and on-site share cards should
be one system. E16's five masters vs the content engine's renderer is the same
decision on both sides.
- **Voice**: `specs/VOICE.md` governs; the content engine templates encode a
sharper register. Whether the hub speaks in that register is a brand call.
### Correction to the still-buildable list
The order lists content API, preview page, book-comparison mount and guard
widening as pending. **All four shipped at `c575a70`.** Book comparison was never
unmounted — my earlier board grepped only `web/src/app` and missed
component-level mounting.
**Genuinely buildable now, design-complete, no clock impact:** Offseason hub
(F9F11), article media (F5), movement strip (E1), calibration curve (E9), the
Report email + archive (E10/E12). The gap audit's own wave ordering already
sequences these.
+120
View File
@@ -0,0 +1,120 @@
# The accrual watch — the program is idle on modeling, and that is correct
## PHASE 0 — live-surface integrity
| check | result |
|---|---|
| Grade bands not derived from the retired champion | **PASS**`gradeBands` is required by *no* serving code. Built across several orders, never wired. No stale band can reach a user because none reaches a user at all. |
| No withdrawn-map leak | **PASS**`CALIBRATION_DEPLOYED` is `[]` and the calibrate loop iterates it, so `calibrate()` is never called. The only `p_win_calibrated` assignment sits inside that empty loop. Served `p_win` is repaired-champion raw. |
| Refusals on the repaired reference | **PASS**`projectionFor` reads `l20_avg`, which `mlbGameLogFeatures` now builds from `fullLog`. |
| Factors still fire post-repair | **PASS** — the repair moved the base they adjust, so sign was re-verified across it: at base 0.35/0.50/0.65/0.80, defence lowers, pitcher-contact raises, platoon raises at every point. Firing check only; **not** a lift re-measurement. |
### One defect found, and it was mine
The grade card rendered **"Last 20 games average: X"** from `l20_avg` — a field
that, after the repair, holds a **full-season** average. The number moved and the
label did not, so the surface asserted a window that no longer existed. Fixed to
"Season average"; `trapDetection`'s L20 explanations likewise.
Same class as everything else tonight, one layer out: **a correct-looking string
describing data that moved underneath it.**
---
## PHASE 1 — the existence risk, resolved
Nineteen commits (`7b85934``2391574`) were local-only with no push
credentials. Three retrievable artifacts now exist:
| artifact | path | size |
|---|---|---|
| session bundle | `~/vyndr-session-2026-08-07.bundle` | 244K |
| **full-history bundle (self-contained)** | `~/vyndr-full-history-2026-08-07.bundle` | 6.8M |
| patch series (20 files) | `~/vyndr-session-patches/` | 1.5M |
**Use the full-history bundle** — the session bundle verifies as requiring ref
`6452926…`, so it only applies onto a repo that already has this history. The
full-history one clones standalone:
```
git clone ~/vyndr-full-history-2026-08-07.bundle vyndr-recovered
```
**This is the top non-accrual action item.** Copy one of these off the machine.
---
## PHASE 2 — the accrual watch
```
repaired champion marker: engine1@2026-08-07-fullwindow
eligible rows: 0 | eligible dates: 0
blocked: no settled rows yet carry the repaired champion marker
```
| item | need | have | status |
|---|---|---|---|
| calibration re-fit | 10 dates | 0 | WAITING |
| hits factor lift | 10 dates | 0 | WAITING |
| prior verdict re-audit | 14 dates | 0 | WAITING |
| rbi lineup-slot gate | 14 dates | 0 | WAITING |
Zero is the correct starting line: the repair ships in *this* session's commits,
so no settled row can carry the marker yet.
### These are ATTEMPT floors, not TRUST floors
**Reaching 10 dates means "the calibration re-fit can now be measured." It does
NOT mean the re-fit is trustworthy.** We lived this distinction the hard way
tonight: date-block CIs on 24 clusters, a LODO gate with 1.49.3% power, and a
≥40 date-cluster bar that was correct as a promotion bar and wrong as a deploy
bar. A 10-date map is thin. It deploys **PROVISIONAL with auto-demotion**, like
everything else, and its interval will still be wide.
**No future session may read "threshold met" as "answer certified."**
### Real-time estimates
Roughly one MLB slate per day, but per-stat settled volume differs — hits props
are far more numerous than rbi, and a "date" only counts once its props settle.
| item | dates | realistic wall-clock |
|---|---|---|
| calibration re-fit, hits lift | 10 | **~2 weeks** |
| verdict re-audit, rbi gate | 14 | **3+ weeks** (rbi accrues slowest) |
The wait is **designed, not a stall.** The alternative — measuring on
reconstructions of a retired forecast — is the trap this program has now refused
by name three times.
---
## PHASE 3 — the pre-registered resumption order
Triggered purely by eligible-date thresholds. No re-litigation, no re-deciding:
1. **@10 dates — calibration re-fit on the repaired champion.** Low-parameter,
LODO where powered, PROVISIONAL, auto-demotion armed, shadow duel restarts on
the repaired forecast. The favourite-longshot bias must be **re-measured**,
not assumed to have survived the repair.
2. **@10 dates — hits factor composed-lift re-measure.** The 1.39% figure is
**void** — measured on the broken baseline. Direction unknown.
3. **@14 dates — prior factor verdict re-audit.** Every null and every THEATER
was scored against a champion worse than a frequency table. Not pre-priced;
some may pass, some may still fail.
4. **@14 dates — rbi lineup-slot / RISP through the two-part gate.** World A
~90%, within-role residual 0.01908 real, `lineup_context` prod-verified (S89).
**FIRST TRIGGER:** when `reAuditEligibility.assess()` reports **10 eligible
calibration dates**, the next order is the calibration re-fit.
**Until then the program is honestly IDLE on modeling.** That is the correct
state, not a gap to fill.
---
## Invariants
No measurement on reconstructions — held. Phase 0 verified live integrity and
factor *firing*; it did not re-measure lift or re-fit calibration. Serving change
for the label fix, fingerprinted. `p_win` never mutated. No Bonferroni slot.
+124
View File
@@ -0,0 +1,124 @@
# The champion was reading ten games — repaired
## PHASE 0 — the defect is real past the peek
The prior baseline peeked at the evaluation window. This one does not: for every
prop the naive forecast is **that player's rate of clearing that line over games
strictly before that date**, from box scores back to 2026-05-01, requiring ≥10
prior games. Same temporal discipline the champion is held to.
| stat | n | champion | **fair PIT baseline** | gap | CI | loses |
|---|---|---|---|---|---|---|
| hits | 799 | 0.00251 | **0.00774** | 0.00523 | [0.0074, 0.0011] | **yes** |
| total_bases | 832 | 0.00393 | **0.00619** | 0.00226 | [0.0055, 0.0003] | **yes** |
| rbi | 501 | 0.02481 | **0.03133** | 0.00652 | [0.0153, 0.0005] | **yes** |
| runs | 473 | 0.00181 | **0.00683** | 0.00502 | [0.0114, +0.0008] | yes (CI touches) |
**Confirmed, not an artefact of the peek.** Three of four CIs exclude zero. The
served forecast was reliably worse than a frequency table.
---
## PHASE 1 — the cause: the window, not the weights
`estimateProbability` computes its base rate as the frequency over **every row it
is handed**. It was handed ten:
```js
// featureCache.getStatRows, MLB branch
const logs = res.last10; // <- the "season rate" was a TEN-GAME rate
```
So the forecast was `0.6 × (ten-game frequency) + 0.4 × (last five OF THOSE TEN)`
— a five-game read carrying 40% of the weight, on top of a ten-game base.
Resolution by variant, all point-in-time:
| stat | champion | season only | w=0.20 | w=0.40 | w=0.60 | best |
|---|---|---|---|---|---|---|
| hits | 0.00251 | 0.00774 | **0.00817** | 0.00688 | 0.00647 | w=0.20 |
| total_bases | 0.00393 | 0.00619 | **0.00734** | 0.00512 | 0.00485 | w=0.20 |
| rbi | 0.02481 | **0.03133** | 0.02727 | 0.02571 | 0.02559 | season only |
| runs | 0.00181 | **0.00683** | 0.00436 | 0.00318 | 0.00180 | season+nudge |
**The 0.40 recency weight costs resolution on all four stats** (0.00086,
0.00107, 0.00562, 0.00365). The nudges are mixed and small: harmful on hits
(0.00157) and rbi (0.00284), marginally helpful on TB (+0.00091) and runs
(+0.00056) — left alone, since the evidence does not support removing them.
---
## PHASE 2 — the repair
Two lines, no new data, no extra API call — **`fullLog` was already being fetched
by the same adapter call that produced `last10`**:
1. `featureCache.getStatRows` MLB branch reads `fullLog`, falling back to
`last10`.
2. `RECENCY_WEIGHT` 0.40 → **0.20**, set at the value the measurement supports.
| stat | OLD | **REPAIRED** | fair baseline | vs baseline | CI | vs old champion |
|---|---|---|---|---|---|---|
| hits | 0.00251 | **0.00817** | 0.00774 | **+0.00043** | [0.0030, +0.0025] | +0.00566 |
| total_bases | 0.00393 | **0.00734** | 0.00619 | **+0.00115** | [0.0014, +0.0046] | +0.00341, **CI [0.0020, 0.0067]** |
| rbi | 0.02481 | 0.02727 | 0.03133 | 0.00406 | [0.0091, +0.0020] | +0.00246 |
| runs | 0.00181 | 0.00436 | 0.00683 | 0.00247 | [0.0088, +0.0014] | +0.00255 |
**Hits resolution tripled; total_bases and runs roughly doubled.**
**Gate assessment, stated exactly:** hits and total_bases now exceed the fair
baseline on the point estimate; rbi and runs remain below it but **every CI now
includes zero.** So no stat *reliably loses* to a frequency table any more, which
satisfies "beat or tie, never lose" in the only sense this sample can support. It
is a tie on rbi/runs, not a win, and it is reported as one. Only total_bases'
improvement over the old champion is CI-confirmed; the rest are directional.
### The stale-fit gate — calibration is OFF
The low-parameter maps were fitted on the retired forecast, and `fromLedger`
cannot rescue them: settled ledger rows still carry OLD `p_win` values, so
refitting today would fit the retired forecast again.
**`CALIBRATION_DEPLOYED` is now empty.** Nothing is served calibrated until
enough dates settle under the repaired champion, and the favourite-longshot bias
must be **re-measured** on the new forecast rather than assumed to have survived.
The shadow duel is likewise void. Serving the raw repaired number is the honest
state, not a regression.
---
## PHASE 3 — the hits factor lift, NOT re-measured
Honest answer: it **cannot** be measured yet. The three proven hits factors were
measured against the old baseline, and re-measuring their lift on the repaired
champion requires settled rows produced *by* the repaired champion. Those do not
exist — the repair ships in this commit. Replaying it would score the factors
against a reconstruction rather than the served forecast.
**Deferred to the first order after the repaired champion has settled dates.**
The factors remain wired and transmitting (43f65d3, sign-verified, 75% coverage);
only their *lift* is unquantified on the new baseline.
---
## PHASE 4 — log and re-queue
**Standing flag, and it is a large one:** every factor verdict in this
programme — every null, every THEATER — was measured against a champion that was
worse than a frequency table. Signal added to noise reads as noise. **Prior
verdicts may deserve re-audit on the repaired champion.** Not re-run here; logged
as standing.
**Re-queued, not built — rbi lineup-slot / RISP opportunity** through the
two-part gate, now landing on a repaired champion. World A ~90%, within-role
residual 0.01908 real, `lineup_context` ingested and prod-verified (S89). That is
the next factor order.
---
## Invariants
Serving-path change by design — the byte-identical invariant inverted again, and
all four stats' numbers move. Nine frozen model modules verified unchanged.
`p_win` is the forecast itself, not mutated post-hoc. No Bonferroni slot: this is
resolution accounting on the champion's own knobs, not a causal claim.
+111
View File
@@ -0,0 +1,111 @@
# The collapsed sequence edge — two proven links whose product is too small to use
**Both components are real. Their product is 0.37pp, and detecting it would take
52 seasons.** This is the most instructive negative in the programme so far,
because nothing in it failed: link-by-link proof did not produce a usable edge.
---
## One correction to the framing
The order states Link 2 proved you *can't* predict the reliever. Half true, and
the other half matters: the **individual** grain failed (17.2% accuracy), but the
**quality** grain PROVED — predicted pen quality separates 2.70pp of realized hit
rate. Pen-season-quality here is a measured predictor, not a fallback after a
failure. That strengthened the plan going in.
## Two of the three specified inputs could not be used honestly
| specified | status |
|---|---|
| pen **quality** | PROVEN (Link 2 coarse grain) — used |
| pen **archetype** | did NOT prove (0.5669 vs 0.5309 modal baseline, corrected interval spanning zero) — **excluded**, building it in would chain on an unproven link |
| hitter **approach identity** ("fastball-hunter", "finesse-vulnerable") | **does not exist** in this registry. MLB batter archetypes are BOMBER / GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID / CATALYST. Inventing an identity to condition on is the fabrication the gate exists to catch |
A hitter power/contact split derived from the sequence data itself was tested as a
**separate gated addition** rather than assumed into the main effect. Neither half
proved (power 0.0002, contact 0.0001, both intervals spanning zero).
---
## The gate — two-part, on the concentrated subset, 114 cumulative tests
| subset | n | games | mean shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|
| concentrated (early-exit × WEAK pen) | 1,931 | 141 | 0.0193 | 0.0001 | [0.0014, +0.0010] | NOT_PROVEN |
| mirror (early-exit × STRONG pen) | 2,574 | 189 | 0.0186 | 0.0000 | [0.0011, +0.0010] | **THEATER** |
| all early-exit later ABs | 6,869 | 451 | 0.0140 | 0.0001 | [0.0007, +0.0005] | NOT_PROVEN |
| pooled all later ABs | 17,891 | 803 | 0.0141 | 0.0000 | [0.0004, +0.0003] | **THEATER** |
Not pooled-diluted: the concentrated subset was gated on its own and is no
better. Two subsets are THEATER by the gate's own definition — the adjustment
moves the number ~1.9pp and improves accuracy by essentially nothing.
---
## Why: the mechanical ceiling
The descriptive pass found the direction the order predicted (early-exit + weak
pen +0.74pp, early-exit + strong pen 0.79pp vs a deep-starter baseline). The
signs are right. The magnitude is the problem, and it is structural:
```
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925 <- NOT a switch to the pen
hit-rate swing across pen quality 0.0394 (weak 0.2491 vs strong 0.2038)
MAX JUSTIFIABLE ADJUSTMENT = 0.0925 x 0.0394 = 0.00365 (0.37pp)
adjustment actually applied (mean |shift|) = 0.01930 (1.93pp)
OVER-MOVEMENT FACTOR = 5.3x
```
**A hitter's third or fourth plate appearance is ALREADY against the bullpen 71%
of the time even when the starter is projected to go deep.** Link 1 lifts that to
81%. It buys nine points of extra pen exposure, not a change of opponent — so any
adjustment riding on it is capped at about a tenth of the pen-quality swing.
The 5.3× over-movement is exactly why the mirror subset reads as THEATER rather
than as a small true effect: the adjustment asserts five times more than the
mechanism can support.
### And a correctly-scaled version is not detectable either
```
concentrated subset n = 1,931 SE of hit rate = 0.00985
max justifiable effect 0.00365 = 0.37 SE
to detect at the corrected bar (z~3.46 for 114 tests): n = 168,488
shortfall 87x -> ~52 seasons of concentrated-subset accrual
```
**This line is structurally closed, not sample-blocked.** Waiting does not fix it.
---
## Not wired, and the self-check deliberately not wired either
The adjustment does not prove, so it feeds nothing. The order also asks for a
self-check flagging where our sequence read diverges from the line's
starter-script, as an opportunity signal. **That is not wired**, because flagging
divergence on an adjustment measured as absent would advertise an edge we have
just shown does not exist — the same failure as fabricated reasoning, one layer up.
## The lesson worth keeping
Link 1 proved (MAE 3.22 → 2.80 batters faced). Link 2's quality grain proved
(2.70pp of realized separation). Both are real, both are point-in-time, both
survived cumulative correction. **Their product is still too small to use.**
Link-by-link validation guarantees each link is real. It does not guarantee the
chain transmits anything. The multiplicative structure has to be sized BEFORE
building — one exposure term of 0.09 is enough to reduce a genuine 3.94pp signal
to noise, and no amount of downstream care recovers it.
## Parallel track — total_bases per-archetype (logged, not run)
Unchanged: `total_bases` settled n=948 pooled, BOMBER × TB **340**, short by 160.
Sample-readiness only, not a verdict. The `specs/per-archetype-grade-bands.md`
blocker still stands — the grade does not yet separate within any archetype.
Link 3 confirmed SKIPPED. Counter and frozen clusters byte-identical.
+142
View File
@@ -0,0 +1,142 @@
# Wiring the three proven hits factors — transmission proven, gain inconclusive
## The bug this order nearly shipped as a finding
The first audit run reported **0 factors fired on all 1,140 rows**. Not a
modelling result — my paging helper ordered by `id`, and `batter_spray`,
`team_defense`, `platoon_splits` and `statcast_aggregates` have **composite
primary keys with no `id` column**. The query errored, the loop broke on error,
and four fully-populated tables read as empty.
`hitsFactorContext.js` — the *production* loader — had the identical defect, so
the live wiring would have loaded nothing and served unadjusted forecasts while
logging success.
**Third occurrence of this class in one session** (doubled `/leaderboard`, the
silent settlement outage, this). Both loaders now order by a real column and
**throw** rather than degrade, because a wiring fault must not be able to wear
the costume of an honest absence.
The Phase 2 transmission gate is what caught it: no resolution number was quoted
until transmission was proved mechanically.
---
## PHASE 1 — the pipeline, in order
```
base rate -> FACTORS (pre-grade) -> CALIBRATE -> GRADE
```
- `model/hitsFactors.js` — the three proven factors composed, each applying only
where it proved, bounded at ±0.25 combined.
- `model/hitsFactorContext.js` — loads the inputs **once per slate**, indexed.
- `snapshotService` builds the context **before** `gradeAndCacheSlate`; it was
previously computed at line 640+, downstream of the grade at 454.
- `gradeSlateService` threads it per prop; `analyzeViaEngine1` applies it to
`p_over` **before** `p_win` is set, recording `p_win_prefactor` and a full
`factor_adjustment` trace.
Hits only. TB/rbi/runs have no proven factors and are untouched.
**Coverage: 859 of 1,140 rows (75%) have at least one factor fire** — 474 with
all three, 256 with two, 129 with one, 281 with none.
---
## PHASE 2 — transmission, proven mechanically
Each factor applied **in isolation** to the same base. (My first table compared
each factor's expected sign against the *composite* change and showed 3 false
failures — with three factors firing, the net can oppose any single member. That
was a flaw in the test, not the wiring.)
| factor | player | side | expected | mult | p before | p after (solo) | sign |
|---|---|---|---|---|---|---|---|
| pitcher_contact | Travis Bazzana | over | raise | 1.0396 | 0.604 | 0.6279 | ✓ |
| platoon_severity | Travis Bazzana | over | raise | 1.0860 | 0.604 | 0.6559 | ✓ |
| pitcher_contact | Patrick Bailey | over | raise | 1.0396 | 0.604 | 0.6279 | ✓ |
| pitcher_contact | Kyle Manzardo | **under** | raise p(over) | 1.0396 | 0.684 | 0.6715 | ✓ |
| platoon_severity | Kyle Manzardo | **under** | raise p(over) | 1.0350 | 0.684 | 0.6729 | ✓ |
| defense_by_direction | Royce Lewis | over | lower | 0.9690 | 0.577 | 0.5591 | ✓ |
| platoon_severity | Royce Lewis | over | lower | 0.9650 | 0.577 | 0.5568 | ✓ |
| pitcher_contact | Petey Halpin | over | raise | 1.0396 | 0.662 | 0.6882 | ✓ |
| platoon_severity | Chase DeLauter | over | lower | 0.9620 | 0.838 | 0.8062 | ✓ |
| defense_by_direction | Gabriel Arias | over | raise | 1.0400 | 0.536 | 0.5574 | ✓ |
| defense_by_direction | Austin Hedges | over | raise | 1.0390 | 0.685 | 0.7117 | ✓ |
| defense_by_direction | Ryan Kreidler | over | lower | 0.9560 | 0.523 | 0.5000 | ✓ |
**12/12 sign-correct — 4/4 for each of the three factors.** The two `under` rows
confirm the flip is handled: a factor raising p(over) correctly *lowers* p_win.
**Unreadables static:** Patrick Bailey, Josh Bell and Brayan Rocchio are switch
hitters — spray applied to none of them, while their other factors fired
normally. The refusal is selective, not a blanket skip.
**TRANSMISSION PASSES.** Resolution may now be quoted.
---
## PHASE 3/4 — refit and OOS measurement
Both calibration maps refit on the factor-adjusted forecast (the un-factored
distribution no longer exists). **The shadow-duel baseline is VOID and restarts**
— it accumulated against a different forecast.
Point-in-time, fit on dates < 2026-08-02, evaluated on 765 held-out rows:
| | reliability | **resolution** | uncertainty | **variance explained** |
|---|---|---|---|---|
| before (four-input) | 0.00795 | 0.00229 | 0.2476 | **0.93%** |
| **after (factor-adjusted)** | 0.00828 | **0.00345** | 0.2476 | **1.39%** |
**Resolution rose 51% relative (+0.00116).** Held-out Brier 0.25398 → 0.25305,
**delta 0.00093, date-block CI [0.00225, +0.00002]**.
**The CI touches zero on 4 eval dates. The composition does NOT earn a proven
keep.** The point estimate favours the factors and the resolution gain is real in
sample, but the honest verdict is **INCONCLUSIVE** — three isolated passes did
not grant a composed pass, exactly as the order anticipated.
**Double-counting note:** the gain is far below the sum of the isolated factor
effects. Expected — defence, pitcher contact and platoon all run through the same
pitcher-batter confrontation and share signal.
---
## PHASE 5 — bands and the honest headline
Resolution moved from 0.93% to 1.39% of variance. **Both are far below what band
separation requires** — a forecast explaining 1.4% of an outcome's variance
cannot produce archetype bands that clear their own base rate.
**The headline, landed as it fell:** the pivot was correct and incomplete. The
plumbing defect was real and is fixed — three proven factors now reach the served
number for the first time, verified sign-by-sign. But **transmission alone did
not buy grade separation.** The factors are real and too weak *in combination* at
current strength.
So the next arc is **factor STRENGTH and BREADTH, not more plumbing.** The wiring
is now a working conduit with three things flowing through it; it needs more, and
stronger.
---
## PHASE 6 — the rbi anomaly, logged only
`rbi` shows **14.51% variance explained vs hits 1.03%** — 13×, on the stat we do
*not* serve corrected and which has **no proven factors**. Open question for the
next order: real counter structure (its four inputs happen to discriminate on a
stat where opportunity is lumpier), or an artefact of line placement and
base-rate spread? **Not investigated here.** It is either the biggest lever on
the board or a mirage, and it deserves its own order.
---
## Invariants
The byte-identical invariant **inverted** for hits by design — the served hits
number should move, and does. TB/rbi/runs paths and all frozen non-hits modules
verified unchanged. `p_win_prefactor` preserves the un-factored forecast in the
trace. No new Bonferroni slot (the factors were already proven); the composed
OOS claim is reported with its CI and is **not** claimed as a pass.
+130
View File
@@ -0,0 +1,130 @@
# Link 2 at the coarse grain — pen QUALITY proves, archetype does not
**The refinement was right.** Naming the individual reliever failed; asking the
same question at the grain the chain actually needs passes, and the payoff it
transmits is larger than anything else measured in this chain.
---
## Why this was worth re-asking
Last session closed Link 2 on the individual grain and separately measured that
the bullpen is, on average, no softer than the starter (+0.0010 on 35,760 plate
appearances). That null does **not** rule this out, and conflating the two would
have been an error: an average washing out is entirely consistent with QUALITY
VARIATION mattering a great deal.
It does. Measured on 18,809 post-starter plate appearances against arms with ≥40
prior appearances:
| arm faced (prior allowed-hit-rate quartile) | n | realized hit rate |
|---|---|---|
| Q1 — best arms | 4,702 | **0.2244** ±0.0119 |
| Q2 | 4,702 | 0.2293 |
| Q3 | 4,702 | 0.2410 |
| Q4 — worst arms | 4,702 | **0.2501** ±0.0124 |
Monotone, spread **+2.57pp** — larger than the whole times-through-the-order
effect (+1.6pp) and far larger than the pen-vs-starter difference (0.7pp).
---
## The cluster unit, corrected — and then checked rather than argued
Last session refused Link 2 partly as a team-borne prediction: 30 bullpens, the
park-geometry ceiling. For QUALITY that argument needed re-testing, and the first
number I reached for was the wrong one.
- **Treatment variance:** 76% of a game's pen-quality variance is WITHIN team.
That says pen quality is not chiefly a club property — but it is a statement
about the treatment, not about where ERRORS correlate, and those are different
claims. Stopping there would have been picking the convenient answer.
- **Measured directly:** ICC of the prediction ERROR by team = **0.0261**. With
16.7 rows per club that is a design effect of 1.41, inflating standard errors
~19% — small, but not nothing on a marginal interval.
So the verdict was checked under all three treatments rather than resting on the
most favourable:
| inference treatment | CI on loss delta | |
|---|---|---|
| unclustered | [0.0067, 0.0010] | excludes zero |
| team-clustered (30 clusters, floor overridden — indicative only) | [0.0086, 0.0003] | excludes zero |
| design-effect adjusted (deff 1.41) | [0.0072, 0.0005] | excludes zero |
It survives all three. Note the team-clustered run sits below this codebase's own
40-cluster floor and is reported as indicative, not as a pass.
---
## The gate
Concentrated subset as instructed — team-games where Link 1's point-in-time
early-exit signal is elevated (predicted ≤22 batters faced), i.e. where the pen
actually enters for the later plate appearances.
### QUALITY grain — **PROVES**
```
n=501 team-games · 426 game clusters · 110 cumulative tests
MAE 0.0294 (league-average baseline) -> 0.0260 delta -0.0034
CI [-0.0063,-0.0005] at 0.9995 VERDICT: PROVES
```
Pooled across all games it also proves (n=1,305, delta 0.0030, CI [0.0049,
0.0015]), so the result is not an artefact of the subset.
### ARCHETYPE grain — **NOT PROVEN**
```
n=501 · modal-guess baseline 0.5309 -> model 0.5669
corrected interval [-0.1073, +0.0268] spans zero
VERDICT: NOT_PROVEN_AT_CORRECTED_BAR
```
Two grains were tested; one earned a place. `penQuality.js` deliberately exposes
no archetype, and a test asserts it.
---
## What Link 3 actually receives
The number that matters is not the MAE gain but how much real outcome separation
the prediction buys — measured on realized outcomes, prediction strictly
point-in-time:
| our prediction | games | PAs | realized hit rate |
|---|---|---|---|
| predicted BEST pen (bottom tercile) | 167 | 2,044 | **0.2231** ±0.0180 |
| predicted WORST pen (top tercile) | 167 | 1,799 | **0.2501** ±0.0200 |
**2.70pp of realized separation**, intervals non-overlapping — capturing nearly
all of the 2.57pp available at the quartile grain. corr(predicted, actual pen
quality) = 0.393.
Caveat stated rather than buried: the tercile split point is chosen in-sample.
The prediction driving the separation is point-in-time, so this is a forward
measurement, but the cut is not.
---
## Built
`src/services/model/penQuality.js` (+ 9 tests) — the proven half, ready for Link
3. `projectPen` abstains below 5 prior club games; `armQuality` abstains below 40
appearances. A league-average stand-in would assert "this is an ordinary
bullpen", which is a claim, and usually the wrong one for exactly the clubs whose
pens have just turned over.
`hitRateShift` carries the measured consequence, bounded — it was measured over a
range and is not extrapolated past one.
**Link 3 is now unblocked** on a proven Link 2 at the quality grain only. It is
not run here; the order scopes this session to building and gating Link 2.
## Parallel track — total_bases per-archetype (logged, not run)
Unchanged from last session: `total_bases` settled n=948 pooled, BOMBER × TB
**340**, short by 160. Sample-readiness only, not a verdict. The second blocker
from `specs/per-archetype-grade-bands.md` still applies — the grade does not yet
separate within any archetype.
Counter and frozen clusters byte-identical.
+170
View File
@@ -0,0 +1,170 @@
# Auditing the LODO instrument — it cannot evaluate any stat
**All four stats are UNTESTABLE-BY-LODO. Not one exceeds its reversal cutoff, so
rbi's and runs' prior FAILs were both false. And no stat may claim LODO
stability, because at this date count the test cannot fail.**
---
## PHASE 0 — the defect, on record
The gate at `1f40014` was mine, and it was an incoherent pair:
**1. A 1-SE informativeness bar coupled to a ZERO-reversal rule.** At exactly 1
SE, a genuinely STABLE stat's drop reverses with probability Φ(1) = 0.1587. On
four informative drops:
```
P(>= 1 reversal | perfectly stable) = 1 - 0.8413^4 = 0.50
```
**The rule failed stable stats half the time by construction.** A test cannot
have a 1-SE noise floor and a zero-tolerance decision rule; the two have to be
chosen together.
**2. A pooled n\*.** The four stats' signed effects differ several-fold, so one
threshold meant four different things. Measured, the pooled 70 was:
| stat | own n\* | pooled 70 was |
|---|---|---|
| hits | 77 | **too low** |
| total_bases | 60 | too high |
| rbi | 54 | too high |
| runs | 81 | **too low** |
It mis-credited **every** stat, in both directions.
Neither defect touched the counter or `p_win`. Both touched only which
calibrations were judged stable.
---
## PHASE 1 — the coherent test, derived blind
### (a)(b)(c) per-stat effect
| stat | n | g (signed) | σ_row | SE_full | effect z | NO-EFFECT? |
|---|---|---|---|---|---|---|
| hits | 1,140 | 0.01288 | 0.11251 | 0.00333 | 3.87 | no |
| total_bases | 1,050 | 0.01380 | 0.10680 | 0.00330 | 4.19 | no |
| rbi | 630 | 0.00884 | 0.06459 | 0.00257 | 3.44 | no |
| runs | 597 | 0.00902 | 0.08080 | 0.00331 | 2.73 | no |
All four have a real effect at full n. None is NO-EFFECT — there is something for
stability to be tested *of* in every case.
### The committed test pair
`k = 1`, chosen because informative drops (D) are the binding scarcity here and
k=1 maximises them while the binomial cutoff holds the false-positive rate.
| stat | n\*=k²(σ/\|g\|)² | informative drops D | cutoff | FP | **power at τ=\|g\|** |
|---|---|---|---|---|---|
| hits | 77 | 5 | 2 | 0.031 | **0.093** |
| total_bases | 60 | 5 | 2 | 0.031 | **0.093** |
| rbi | 54 | 4 | 2 | 0.014 | **0.045** |
| runs | 81 | 3 | 2 | 0.004 | **0.014** |
Per-drop noise probability under stability Φ(1) = 0.1587. FAIL iff reversals > cutoff.
### The finding that dominates everything else: the test has no power
Against a **strong** instability — date-to-date SD of the effect equal to the
effect itself — this test detects a failure between **1.4% and 9.3%** of the
time. Across every k examined (1.0 → 2.0), the best any stat reaches is 0.337,
and reaching even that costs all but two informative drops.
**A gate that cannot fail cannot pass.** `LODO_POWER_FLOOR = 0.50` makes that
structural: below it a stat is UNTESTABLE-BY-LODO regardless of its reversal
count, so "could not test" can never be read as "passed".
Committed as `LODO_K`, `LODO_TEST`, `LODO_POWER_FLOOR`; a test recomputes each
n\* from (σ, g) and each cutoff from the binomial tail, and asserts the old
zero-reversal rule's ~0.50 false-fail rate. The stale pooled constant is nulled
so nothing can read it.
---
## PHASE 2/3 — cold re-read at each stat's own n\*
| stat | own n\* | informative | reversals | cutoff | power | verdict |
|---|---|---|---|---|---|---|
| hits | 77 | 5 | 0 | 2 | 0.093 | **UNTESTABLE** |
| total_bases | 60 | 5 | 0 | 2 | 0.093 | **UNTESTABLE** |
| rbi | 54 | 4 | 1 | 2 | 0.045 | **UNTESTABLE** |
| runs | 81 | 3 | 2 | 2 | 0.014 | **UNTESTABLE** |
**Setting the power floor aside entirely, not one stat exceeds its cutoff.**
---
## PHASE 4 — reconcile against 1f40014
| stat | 1f40014 | now | why it changed |
|---|---|---|---|
| hits | PASS | UNTESTABLE | the PASS was from a test that cannot fail; 0 reversals is uninformative at 9% power |
| total_bases | PASS | UNTESTABLE | same — and it "passed" at a pooled n\*=70 above its own 60, so its drop count was under-credited too |
| **rbi** | **FAIL** | UNTESTABLE (1 reversal, cutoff 2) | **FALSE FAIL.** One reversal on four drops is a ~16%-per-drop coin flip, not evidence |
| **runs** | **FAIL** | UNTESTABLE (2 reversals, cutoff 2) | **ALSO A FALSE FAIL** under the coherent rule — this was not anticipated |
Answering the order's three questions directly:
- **Is rbi's FAIL a false fail?** Yes. And so is runs' — which the order did not
anticipate, having classified runs as DATE-DRIVEN on the strength of a 244-row
reversal. Under a rule with a stated error rate, two reversals in three drops
does not clear the cutoff.
- **Was TB's PASS real?** No. It was vacuous: the test could not have failed it.
- **Does hits still pass at its own smaller n\*?** Its own n\* is *larger* (77 vs
the pooled 70), it still shows zero reversals, and it is still untestable.
---
## PHASE 5 — deploy, withdraw, route
| stat | basis | decision |
|---|---|---|
| hits | date-clustered CI [0.0139, 0.0097], **4 date clusters** | DEPLOY-PROVISIONAL, relabelled `ci_only_lodo_untestable` |
| total_bases | CI [0.0061, 0.0045], **2 date clusters** | **RELABELLED** — kept, no longer claims LODO stability |
| **rbi** | CI [0.0092, 0.0010], **2 date clusters** | **NEWLY DEPLOYED** — its FAIL was false |
| runs | **no fittable map** at its point-in-time split | REFUSE — no CI to stand on either |
Every deployed stat now carries `calibration_basis: 'ci_only_lodo_untestable'`.
**Auto-demotion is the sole stability guard**, not a backstop to a passed test.
The honest weakness, stated rather than buried: those intervals rest on **2 to 4
date clusters**. That is thin support, and it is now the *only* support.
**rbi chainAcross stackability is newly granted**; hits' remains from `1f40014`.
rbi bands rebuilt on `p_win_calibrated` (425 eval rows): every archetype
indistinguishable from its base rate, two-bar rule keeping them `base_rate`.
runs is **not** routed to the low-parameter calibrator on a date-driven finding —
that finding was an artefact. It is queued for the ordinary reason: no isotonic
map is fittable at its sample.
---
## PHASE 6 — logged
**The deploy set at `1f40014` was set by a coin-flip-power ruler.** It is now set
by a per-stat, power-coherent, pre-committed test with a stated false-positive
rate — and that test's first act was to report that it cannot evaluate anything,
which is a more useful answer than either verdict it replaced.
**The audit was permitted to wound the live deploy, and did**: total_bases lost
its LODO claim and now stands on a two-date-cluster interval. That it could is
the integrity property.
**Standing question unchanged.** Calibrated `p_win` still separates within
archetype no better than raw — 18 archetype slots across three deployed stats,
every one a single band indistinguishable from its base rate. Per-archetype
separation comes from proven factors or it does not exist.
---
## Invariants
Blind ordering held: (g, σ_row) and the committed (n\*, cutoff, power floor) were
derived and locked with no reversal or verdict in view, before any stat was
re-read. `p_win` never mutated. No Bonferroni slot consumed. Counter and frozen
clusters verified file-by-file.
+184
View File
@@ -0,0 +1,184 @@
# LODO-gated provisional calibration — total_bases deploys, three stats refuse
## PHASE 0 — Record correction
**The ≥40 date-cluster deploy floor applied to CALIBRATION was the wrong
instrument, and I applied it without challenging the binding.**
It is `factorGate`'s cluster-robust interval floor, built for a factor making a
CAUSAL claim, where the risk is a false positive dressed as mechanism. A
calibration layer is different in kind:
- it makes **no causal claim** — it is a monotone shrink toward observed
- its failure mode is **bounded**: it can only over- or under-shrink
- it consumes **no Bonferroni slot**
Its real risk is that the correction is **date-driven**, and leave-one-date-out
tests that directly. The replacement bar is **stricter on stability**, not looser
on standard: LODO fails a stat if removing any single day reverses the
improvement, which a cluster count cannot detect at all.
The prior order's date premise was also wrong (05-01→08-04, "~90 dates"); the
snapshots span 07-19→08-06 = 19 dates. That was corrected in the settlement
session. The mis-bound instrument is mine.
**The ≥40 floor is retained, correctly scoped as the PROMOTION bar** — the point
at which a stat leaves provisional status.
---
## PHASE 1 — Both guards codified (11 tests)
`src/services/model/calibrationGuards.js`
**GUARD 1 — the both-sides tell.** Picked-side dedup is now mandatory
preprocessing, asserted. The guard fires on the CONJUNCTION of both sides being
present AND mean p_win pinned near 0.5 — either alone is unremarkable, and
flagging a genuinely balanced one-sided book would be a false alarm.
Demonstrated on live data in this run:
```
raw population violated=true mean_p 0.4962 both_sides_share 0.9763
after dedup violated=false mean_p 0.6694
```
**GUARD 2 — a null that scores itself.** `safeBrier` refuses when any prediction
is null; `applyOrRefuse` drops unmappable rows rather than passing nulls
downstream. A test demonstrates the trap explicitly — `(null1)² === 1` and
`(null0)² === 0`, so a Brier over nulls silently equals the win rate.
---
## PHASE 2 — LODO table
Refit dropping each date; measure held-out Brier delta and the sign of the >0.9
favourite bias. Drops with fewer than 20 held rows are marked UNINFORMATIVE
rather than counted either way.
**A note on what LODO is:** refitting on all-but-one date uses dates that follow
the held-out one, so this is a STABILITY test, not a point-in-time backtest. The
point-in-time result is separate and already established. Both are required.
| stat | n | dates | informative drops | reversals | sign flips | LODO |
|---|---|---|---|---|---|---|
| hits | 1,140 | 17 | 7 | **2** (07-22 n=20, 07-26 n=25) | 0 | **FAIL** |
| **total_bases** | 1,050 | 7 | 5 | 0 | 0 | **PASS** |
| rbi | 630 | 5 | 5 | **1** (08-01 n=99) | 0 | **FAIL** |
| runs | 597 | 5 | 5 | **2** (08-01 n=86, 08-05 n=244) | 0 | **FAIL** |
### Threshold sensitivity — reported because the verdict moves
| min held rows | hits | total_bases | rbi | runs |
|---|---|---|---|---|
| **20** (applied) | FAIL | **PASS** | FAIL | FAIL |
| 30 | PASS | **PASS** | FAIL | FAIL |
| 50 / 75 | PASS | **PASS** | FAIL | FAIL |
| 100 | PASS | **PASS** | PASS | FAIL |
- **total_bases passes at every threshold** — the only unambiguous result.
- **runs fails at every threshold**, reversing on a 244-row date.
- **hits' failure is threshold-fragile**: it fails only when 20- and 25-row dates
are admitted, and those are the two smallest informative drops in the set.
I chose `MIN_HELD_ROWS = 20` before seeing which stats passed, and did not move
it afterwards to preserve a deploy. The honest caveat: a per-date Brier delta on
20 rows has a standard error several times the effect being tested, so the LODO
instrument is underpowered per-drop at this sample size. That argues for
pre-registering a higher threshold — a Roundtable decision, not one to make while
holding the results.
---
## PHASE 3 — Deploy decisions
| stat | LODO | point-in-time CI | decision |
|---|---|---|---|
| **total_bases** | PASS | [0.0061, 0.0045] | **DEPLOY-PROVISIONAL** |
| hits | FAIL | [0.0139, 0.0097] | REFUSE — improvement reverses on 07-22 / 07-26 |
| rbi | FAIL | [0.0092, 0.0010] | REFUSE — improvement reverses on 08-01 |
| runs | FAIL | no fittable map at the point-in-time split | REFUSE — honest null |
Certified band for total_bases: **[0.60.8]**. Outside it → refuse, fall to base
rate.
### hits was being served calibrated, and is not any more
`snapshotService` hardcoded hits calibration since S91. hits fails LODO, so it
has been removed from the deployed set. **A stat that cannot survive dropping one
day was never calibrated — it was fitted to that day.** The consequence is real:
hits props become unstackable again for `chain.chainAcross`. That is the honest
result of measuring it, not a regression to route around, and it errs toward
withdrawing a claim rather than preserving one on a fragile verdict.
Deployment is now driven by `CALIBRATION_DEPLOYED` (frozen, tested), not a
hardcoded stat name.
---
## PHASE 4 — Auto-demotion (14 tests)
`src/services/model/calibrationRegistry.js`
- **Deploy needs BOTH gates** — LODO pass AND a point-in-time CI excluding zero.
Neither is waivable.
- **`reverify` demotes on the first breach**: the CI ceasing to exclude zero, or
the favourite over-prediction flipping sign (which would mean the correction is
now pushing the wrong way). The breaking date is logged.
- **Promotion to non-provisional** requires the original ≥40 date-cluster bar,
with the interval still holding.
A provisional deploy that cannot be taken away is just a deploy; `reverify` is
what makes the label mean something.
---
## PHASE 5 — Bands rebuilt on p_win_calibrated (total_bases only)
625 eval rows on calibrated values. **The two-bar rule still bites**: TB is now
CALIBRATED but no factor is PROVEN for it (barrel, exit velo and
hard-contact-allowed were all THEATER), so bands remain a base-rate read — now an
honestly-numbered one.
| archetype | n | base rate | bands | separation |
|---|---|---|---|---|
| UNLABELLED | 275 | 0.6255 | 1 | indistinguishable from base rate |
| BOMBER | 200 | 0.6100 | 1 | indistinguishable |
| GHOST | 87 | 0.5747 | 1 | indistinguishable |
| DRIVER | 24 | 0.7917 | 1 (PROVISIONAL) | indistinguishable |
| BRUSH | 19 | 0.4737 | 1 (PROVISIONAL) | indistinguishable |
| MIRROR | 6 | — | REFUSED | insufficient outcomes |
Calibration compressed the served range to 0.42861.0. Every archetype still
collapses to a single band — calibrated p_win does not separate within archetype
any better than raw p_win did. Refused stats keep base-rate bands on raw p_win.
---
## PHASE 6 — Logged, not acted on
**The dead gradient is buried.** Over-prediction ordering on the fuller settled
set is **hits ≈ TB > runs > RBI**, not TB > RBI > runs. The skill-driven-gradient
mechanism did not survive — **RBI has the SMALLEST bias** (+0.0164). Descriptive
only; no mechanism claimed.
**Refusal coverage.** Refused props are predictable-but-input-less rather than
genuinely uncertain (3.20 vs 3.39 AB rules out playing time). The refused set is
a MAP OF MISSING INPUTS and feeds the input-coverage roadmap. Not this order.
**Queued candidate:** a low-parameter calibrator (Platt / beta) fits a
favourite-longshot shape on far fewer points than isotonic needs, which is
exactly the constraint that refused runs. It is a NEW estimator requiring its own
out-of-sample validation. Not built here.
**Programme-level finding:** calibration beats every factor tried on TB / RBI /
runs, and the defect is systematic over-prediction **concentrated in favourites**
(+0.21 to +0.28 above p_win 0.9 on all four stats) rather than a uniform shift.
---
## Invariants
`p_win` never mutated — calibration rides as `p_win_calibrated` with
`calibration_status: 'provisional'`. No Bonferroni slot consumed; testLedger
factor count untouched. Counter and frozen clusters byte-identical.
+146
View File
@@ -0,0 +1,146 @@
# The LODO threshold, derived from power — hits restored, rbi/runs routed
## PHASE 0 — the threshold, derived blind
**Estimand:** does dropping date D reverse the SIGN of the out-of-sample Brier
improvement on D's held-out rows? A reversal is informative only if that date's
Brier delta is distinguishable from zero at its row count.
The per-row Brier difference is `d_i = (pc_i y_i)² (p_i y_i)²`, so a date's
delta is `mean(d)` and `SE(n) = SD(d)/√n`. The smallest n at which a typical
effect clears one standard error is `n* = (SD(d)/|effect|)²`.
Pooled across all four stats — deliberately, so no single stat's verdict could
shape the threshold that decides it:
```
pooled rows 3,417
SD(per-row Brier diff) 0.09816
|pooled effect| 0.01175
n* = (0.09816 / 0.01175)^2 = 69.8 -> 70
```
### SE-vs-n
| n | SE | effect / SE | informative |
|---|---|---|---|
| 10 | 0.0310 | 0.38 | no |
| **20** (previously chosen by hand) | 0.0220 | **0.54** | **no** |
| 25 | 0.0196 | 0.60 | no |
| 30 | 0.0179 | 0.66 | no |
| 50 | 0.0139 | 0.85 | no |
| **70 (n\*)** | 0.0117 | **1.00** | **yes** |
| 100 | 0.0098 | 1.20 | yes |
| 244 | 0.0063 | 1.87 | yes |
The hand-chosen 20 sat at 0.54 SE — a coin flip. That is the defect this
derivation removes, and it is why the previous verdict moved with the number.
**Committed as `calibrationRegistry.LODO_MIN_HELD_ROWS = 70`** with
`LODO_THRESHOLD_BASIS` recording the inputs. A test recomputes `(SD/effect)²` and
asserts it equals the constant, so the value cannot drift from the basis that
justifies it, and cannot be silently tuned. The derivation script prints no stat
verdict, no date and no reversal; it ran and the constant was committed before
any stat was re-read.
---
## PHASE 1 — LODO at n\*, applied cold
| stat | n | dates | informative drops | reversals | LODO |
|---|---|---|---|---|---|
| **hits** | 1,140 | 17 | 5 | **0** | **PASS** |
| **total_bases** | 1,050 | 7 | 4 | **0** | **PASS** |
| rbi | 630 | 5 | 4 | 1 — 2026-08-01 (**n=99**) | **FAIL** |
| runs | 597 | 5 | 3 | 2 — 2026-08-01 (**n=86**), 2026-08-05 (**n=244**) | **FAIL** |
hits' held-out deltas at n\*: 0.0041 / 0.0080 / 0.0192 / 0.0140 / 0.0139
across 123272 row dates. Every drop holds, and the favourite over-prediction
holds sign on every drop where it is testable (+0.52 / +0.318 / +0.272).
**This is the instrument finally being powered, not vindication of a prediction.**
The withdrawal at `6ae11f1` was correct on the instrument available then, which
admitted 20- and 25-row dates as evidence. Nothing about hits changed; what
changed is that the threshold is now derived rather than chosen.
---
## PHASE 2 — Failure classification
Both failures are **DATE-DRIVEN**, not underpowered-per-drop:
| stat | deciding date | held n | vs n\*=70 | classification |
|---|---|---|---|---|
| rbi | 2026-08-01 | 99 | **above** | DATE-DRIVEN |
| runs | 2026-08-01 | 86 | **above** | DATE-DRIVEN |
| runs | 2026-08-05 | 244 | **far above** | DATE-DRIVEN |
Every reversal sits comfortably above the powered threshold, so **no threshold
choice and no further date accrual rescues either stat.** Isotonic is fitting
day-structure on both.
**Routed to the low-parameter calibrator queue** (Platt / beta), which fits a
favourite-longshot shape on far fewer free parameters and is therefore much
harder to bend to one day. Not built here — it is a new estimator and needs its
own out-of-sample validation.
---
## PHASE 3 — Deploy and bands
| stat | status | certified band |
|---|---|---|
| **hits** | **DEPLOY-PROVISIONAL (restored)** | [0.50.7] |
| **total_bases** | DEPLOY-PROVISIONAL (unchanged from 6ae11f1) | [0.60.8] |
| rbi | REFUSE — date-driven | — |
| runs | REFUSE — date-driven | — |
`CALIBRATION_DEPLOYED` is now `['hits', 'total_bases']`, frozen and tested. Both
carry `calibration_status: 'provisional'` with auto-demotion armed; promotion bar
remains the original ≥40 date-clusters.
**hits stackability is RESTORED.** It was withdrawn at `6ae11f1`, which removed
hits props from `chain.chainAcross`. They are stackable again — and the record
shows it came back **through the powered gate, not by fiat**. A test asserts the
restoration alongside the threshold's provenance.
### hits bands on p_win_calibrated (765 eval rows)
| archetype | n | base rate | bands | lift bands |
|---|---|---|---|---|
| UNLABELLED | 284 | 0.5211 | 1 | 0 |
| BOMBER | 271 | 0.5351 | 1 | 0 |
| GHOST | 110 | 0.6182 | 1 | 0 |
| BRUSH | 35 | 0.5714 | 1 | 0 |
| DRIVER | 34 | 0.6765 | 1 | 0 |
| CATALYST | 21 | 0.5238 | 1 | 0 |
| MIRROR | 10 | — | REFUSED | — |
Two-bar rule still bites: hits is CALIBRATED but its only proven factor
(`defense_by_direction`) is pooled, not per-archetype, so the bands remain a
base-rate read — now honestly numbered.
---
## PHASE 4 — Logged
**The deploy set is now determined by a power-derived, pre-committed, tested
constant rather than an operator-chosen number.** That property matters more than
either verdict: at `6ae11f1` the rule moved the live path *against* the operator,
withdrawing a stat that was already serving. It has now moved it back, on the
same evidence, because the instrument changed. Both directions are the rule
working. Keep it.
**Standing question for the chain model:** calibrated `p_win` separates within
archetype no better than raw — every archetype collapses to a single band on both
deployed stats (TB and now hits), across 13 archetype slots. Per-archetype grade
separation is **not** going to come from calibration. It comes from proven
per-archetype factors or it does not exist. Descriptive only; no action here.
---
## Invariants
`p_win` never mutated — calibration rides as `p_win_calibrated`. No Bonferroni
slot consumed. Counter and frozen clusters byte-identical. Threshold fixed from
power before any stat was re-read; no post-hoc movement, enforced by test.
+126
View File
@@ -0,0 +1,126 @@
# Per-archetype grade bands — built, gated, and not yet shippable
**The rescale does not ship, for two independent reasons — and the second one is
new and more interesting than the first.**
---
## 1. The premise, checked against the ledger
`scripts/proven-status.js` (run fresh, not recalled):
```
PROVEN_SET: EMPTY — no stat has beaten the counter out-of-sample
archetype_x_stat_at_or_above_gate: []
closest: BOMBER x hits n=471 (short 29) · GHOST x hits n=196 (short 304)
```
| the order states | measured |
|---|---|
| three proven causally-correct factors | **one.** `pitcher_contact_profile` CI upper bound is exactly `0.0000` (not proven); `platoon_severity` passes but is held on 4.5%-median-contaminated season-to-date splits |
| several now archetype-conditioned | **zero.** No archetype slot reaches n≥500 |
| "defense strong for GHOST/BRUSH, honestly-null for BOMBER" | measured the **opposite direction** — BOMBER 0.0036, GHOST 0.0024, both noise-dominated (specs/per-archetype-re-audit.md) |
Pooled proof does not qualify a slot. `PROVEN_BY_ARCHETYPE` is `{}` and the band
builder reads it directly.
---
## 2. The new finding: the grade does not separate within archetype
This is the reason that matters, because it would block the rescale even if the
factors had proved.
Published bands, full clean settled hits history (n=1,312, 106 cumulative tests):
**every archetype collapses to ONE band.** Bands merge when their corrected
intervals overlap, because publishing two letters we cannot tell apart is a
distinction we have not measured.
Uncorrected, so the underlying ranking is visible rather than hidden by the bar:
| archetype | n | base rate | quintiles (hi→lo) | corr(p_win, outcome) | bands @95% |
|---|---|---|---|---|---|
| **BOMBER** | 466 | 0.588 | 0.75 · 0.62 · 0.60 · 0.48 · 0.48 | **+0.207** | 2 — A 0.660 (n=279) vs B 0.481 (n=187), A shows lift |
| UNLABELLED | 490 | 0.539 | 0.66 · 0.53 · 0.58 · 0.51 · 0.41 | +0.155 | 2, no lift |
| **GHOST** | 192 | 0.578 | 0.47 · 0.63 · 0.74 · 0.58 · 0.45 | **0.007** | 1 |
| BRUSH | 80 | 0.537 | 0.81 · 0.63 · 0.44 · 0.50 · 0.31 | +0.350 | 1 |
### This inverts the order's design
The order assigns **contact types the factor-rich treatment** and **power types
the honest base-rate treatment**, on the reasoning that single-game hits are
variance for a power profile.
Measured, it is the other way round. **BOMBER is the one archetype where the
model ranks** (+0.207, monotone across quintiles, and it separates into a genuine
A/B at 95%). **GHOST is where it is flat** — corr 0.007, and its quintiles are
non-monotone: the model's most confident GHOST reads hit 47% while its middle
reads hit 74%.
That ordering is mechanically sensible in hindsight. A power hitter's chance of a
hit tracks whether he can damage the arm he is facing, which the counter's
frequency question partly captures. A contact hitter's hits depend on batted
balls finding holes, which is much closer to luck — the same reason the pooled
hits negative closed (S83) and the reason `defense_by_direction` is the factor
that survives.
**Had the rescale shipped as specified, it would have given the factor-rich
treatment to the archetype the model reads worst, and left base-rate on the one
it reads best.**
### BOMBER still cannot publish two letters
BOMBER's A/B split is real at 95% but does not survive the cumulative correction
at 106 tests — the intervals merge. So even the best archetype gets one honest
band today. Lowering the correction to expose it would be exactly the "curve to
make more A's" the order forbids.
---
## 3. What was built
`src/services/model/gradeBands.js` — per-archetype bands from realized outcomes:
- **Lift, not raw rate.** A band is credited only when its interval clears *that
archetype's own* base rate. Locked by test: the same 62% realized rate is lift
for a 45%-base profile and a deficit for a 68%-base one.
- **The two-bar rule is structural.** `factor_informed` requires `proven` AND
`calibrated` for that archetype. Tests assert that proven-alone, calibrated-
alone, and neither all return `basis: 'base_rate'` with the reason stated —
so with nothing proven, which is today, no factor-informed band can be produced
at all. Same shape as `featureRegistry.liveFeatures()`: the honest state is the
default and the richer claim has to be earned past a gate.
- **Indistinguishable neighbours merge** rather than becoming different letters.
- **Thin bands are PROVISIONAL, not dropped** — "still counting" and "nothing
here" are different claims.
- **Wilson intervals**, widened by the cumulative correction. Wilson because
these bands are small and rates sit near the edges, where a normal
approximation runs past 0 and 1 and implies impossible rates.
`reasoning()` is built and tested: a base-rate band says *"base-rate read … no
matchup factor is proven for this profile yet"*, and a factor-informed band with
no named proven factors returns **nothing** rather than inventing a why.
It is **not wired to the card**, deliberately. There is no per-archetype band
being served, so attaching per-archetype copy now would ship product language for
a rescale that does not exist.
### One piece of the order I did not build
The specified power-type reasoning — *"the matchup edge is in total_bases"* — is
**not supported by any measurement**. `proven-status.js` records total_bases as
INCONCLUSIVE (delta +0.0038, CI [0.068, +0.075]). Wiring that sentence would
assert an edge we have measured as indistinguishable from zero, which is the
fabricated-reason failure the rest of this module exists to prevent.
---
## 4. What would unblock the rescale
1. **BOMBER × hits is 29 rows short** of the gate — days away, and it is the
archetype the model actually reads. That is the first slot to test, not GHOST.
2. **Point-in-time platoon splits** would convert two held passes into real ones.
3. **Separation must survive the correction**, not just 95%. More sample tightens
the intervals; nothing else legitimately does.
Counter and frozen clusters untouched. No letter was moved.
+134
View File
@@ -0,0 +1,134 @@
# Per-archetype re-audit — and the replication unit that decided it
**The premise this order opened with does not hold, and the query it blames was
never windowed.** `prove-hit-factors.js` selects on sport / stat / outcome only —
there is no date filter anywhere in it, and it pages the full table. Nothing was
being clipped.
| claimed | measured |
|---|---|
| full clean history 2,715 rows | **1,266** |
| platoon "tripled to 1,208 and PROVED" | 1,056 rows; passes the gate, **not promoted** |
| three proven pooled factors | **one** — see below |
`platoon` and `platoon_severity` were explicitly held last session, not promoted:
they ride season-to-date splits containing the games they predict (4.5% median
contamination, 12.4% p90) and passed with an upper bound of 0.0001. That still
stands. `pitcher_contact_profile` was **demoted** last session. So the proven set
going in was one factor, not three.
---
## STEP 1 — Full-history sample audit per slot
Run against full clean settled history, deduped, non-quarantined:
| factor | ALL | BOMBER | GHOST | BRUSH | DRIVER | CATALYST |
|---|---|---|---|---|---|---|
| pitcher_contact_profile | 1,059 | 408 | 173 | 64 | 43 | 16 |
| platoon | 1,056 | 408 | 173 | 64 | 43 | 16 |
| defense | 912 | 357 | 150 | 57 | 37 | 12 |
| defense_by_direction | 782 | 319 | 117 | 49 | 32 | 12 |
| platoon_severity | 700 | 343 | 121 | 24 | 38 | 16 |
| park_hits | 619 | 239 | 96 | 36 | 22 | 9 |
**No archetype slot reaches n≥500 on full history.** The best is BOMBER at 408,
and BOMBER is by far the most common archetype on the board. These are
**CONFIRMED genuinely short — not windowed-query artifacts.**
---
## STEP 2 — The replication unit, which changed every verdict
Errors are correlated within a game (shared starter, park, weather, game state),
so the interval must be clustered. But clustering on the *game* is still wrong
for some factors, and the audit exposed it: `park_hits` initially "PROVED" at 619
rows across 45 games — yet those 45 games only ever visited **14 distinct park
values**. A park effect is replicated across parks. Unmodelled park heterogeneity
is confounded with the very thing being estimated.
So each factor is now clustered on the **coarser of the game and the entity its
treatment rides on**:
| factor | rows | games | treatment entities | clustered on | k | Brier Δ | CI (corrected, 99 tests) | verdict |
|---|---|---|---|---|---|---|---|---|
| `defense_by_direction` | 782 | 84 | **442** hitter×team | game | 84 | 0.0031 | [0.0054, 0.0012] | **PROVES** |
| `platoon_severity` | 700 | 85 | 145 hitters | game | 85 | 0.0038 | [0.0070, 0.0007] | PROVES\* |
| `platoon` | 1,056 | 85 | 228 hitters | game | 85 | 0.0033 | [0.0061, 0.0006] | PROVES\* |
| `pitcher_contact_profile` | 1,059 | 85 | 130 pitchers | game | 85 | 0.0034 | [0.0067, **0.0000**] | NOT_PROVEN |
| `defense` | 912 | 84 | **26** teams | entity | 26 | 0.0038 | [0.0075, 0.0006] | PENDING — k<40 |
| `park_hits` | 619 | 45 | **14** park values | entity | 14 | 0.0037 | [0.0093, **+0.0029**] | PENDING — k<40 |
\* held, not promoted — contaminated inputs, unchanged from last session.
### This is Kev's causal-correctness thesis confirmed from a new direction
The causally-correct atom is not merely more accurate — **it is the only one that
is measurable at all.** `defense_by_direction` has **442** units of replication
because spray direction varies per hitter; crude team `defense` has **26**,
because there are 26 teams. The crude factor cannot be validated no matter how
long the ledger runs, and its apparently-tight interval was pseudo-replication
across teams that does not exist.
`park_hits` losing its pass is the same finding as last session's park-dimensions
result, arrived at independently: **venue- and team-borne factors cap at ~30
units, permanently.**
### Gate change: two floors, not one transplanted bar
Last session I applied the 500 bar to clusters. That was wrong in a way worth
naming: it refused a factor with 1,059 rows over 85 games — ample observations
*and* ample clusters — while answering neither question. The floors are now
separate, because they answer different things:
- **rows ≥ 500** — is the point estimate stable?
- **clusters ≥ 40** — can the interval around it be believed?
40 is the conventional floor below which cluster-robust inference under-covers.
This is not a lowered bar: `park_hits` (14) and `defense` (26) are still refused,
and park geometry is still permanently unvalidatable.
---
## STEP 3 — What gets wired
**Nothing new.** No archetype slot earns a wire, and none is grandfathered in
from pooled proof.
- `defense_by_direction`**PROVES pooled, stays POOLED-ONLY.** Its BOMBER
(n=319) and GHOST (n=117) slots are short, so no per-archetype reasoning is
wired. The card must not say "GHOST: defence matchup strong" — we have not
earned that sentence.
- `platoon`, `platoon_severity` — pass the gate, **held** pending point-in-time
splits.
- `pitcher_contact_profile`, `defense`, `park_hits` — honest null / confirmed
short.
### The predicted fingerprint did NOT appear
The order expected `defense_by_direction` strong for GHOST and ~zero for BOMBER.
Measured point estimates run the other way — BOMBER 0.0036, GHOST 0.0024 — and
at n=319/117 both are noise-dominated. **Recorded so it is not claimed later.**
This is not evidence against the theory; it is evidence we cannot see it yet.
---
## STEP 4 — Rescale readiness: NOT READY
One proven factor, worth 0.0031 Brier, clustered-honest. Two more held behind a
contaminated input. The counter still supplies essentially all of the model's
resolution (S78 ablation).
Rescaling the grade distribution on that would be **relabelling** — the same
error as minting A's by moving thresholds, which is a permanent founder ruling.
The distribution is not factor-rich enough.
**What would change the answer**, in order of cost:
1. **Point-in-time platoon splits** — would convert two held passes into real
ones. Cheapest, no waiting; needs per-game split reconstruction.
2. **More games** — every archetype slot is short, and slots grow with games, not
rows. BOMBER needs ~92 more rows to reach 500.
3. **A factor with high replication** — the lesson of this audit is that new
factors should be chosen for *causal correctness first*, which also buys
measurability. Anything venue- or team-borne is dead on arrival.
+141
View File
@@ -0,0 +1,141 @@
# The rbi anomaly, decomposed — the model is out-resolved by a batting-order integer
## PHASE 0 — the 14.51% is REAL
The figure came from the harness that produced a false null three times tonight,
so it was re-derived with a paged pull **asserted against an exact count**:
| stat | exact rows | paged rows | scorable | resolution | share |
|---|---|---|---|---|---|
| rbi | 7,930 | 7,930 ✓ | 630 | **0.03268** | **14.51%** |
| hits | 11,690 | 11,690 ✓ | 1,140 | 0.00252 | 1.03% |
| total_bases | 12,086 | 12,086 ✓ | 1,050 | 0.00442 | 1.82% |
| runs | 6,440 | 6,440 ✓ | 597 | 0.00130 | 0.56% |
**Reproduces exactly.** rbi deciles are monotone through the middle (0.552→0.577,
0.646→0.691, 0.746→0.784) with a genuinely low bin at 0.13→0.122, and 20 raw
rows are printed in the artifact for hand audit.
---
## The caveat that governs every number below
The naive forecasts are **leave-one-out on the evaluation window itself** — they
see that player's performance in the very rows being scored, while the model is
strictly point-in-time. They are therefore **upper bounds on available
resolution, not fair competitors.** Every comparison is read that way.
---
## PHASE 1 — the three-way split
| stat | MODEL | (a) player base rate | (b) lineup slot | (c) within-stratum |
|---|---|---|---|---|
| **rbi** | **0.03268** | 0.01167 | **0.03608** | 0.01908 |
| hits | 0.00252 | **0.00446** | 0.00100 | 0.00473 |
| total_bases | 0.00442 | **0.01331** | **0.03448** | 0.00607 |
| runs | 0.00130 | **0.00262** | **0.01156** | 0.01170 |
### Finding 1 — rbi's resolution is LINEUP ROLE, almost exactly
**Batting-order slot alone resolves 0.03608 against the model's 0.03268.** A
single integer — where he hits in the order — accounts for the entire anomaly and
slightly more. That is real predictive signal and it is **opportunity, not
skill**: the cleanup hitter bats with runners on, the 8-hole hitter does not.
Player base rate alone gives 0.01167, so ~36% of the model's rbi resolution is
matched by knowing only *who* is batting.
Within strata of similar-base-rate players the model still resolves **0.01908**
58% of its total, and **higher than any other stat's entire model resolution.**
So rbi does carry genuine within-role discrimination on top of the role effect.
### Finding 2 — on three of four stats the model is beaten by "he's a .270 hitter"
| stat | model | player base rate alone | |
|---|---|---|---|
| hits | 0.00252 | **0.00446** | base-rate-only resolves **1.8×** the model |
| total_bases | 0.00442 | **0.01331** | **3.0×** |
| runs | 0.00130 | **0.00262** | **2.0×** |
| rbi | **0.03268** | 0.01167 | model wins, 2.8× |
Even allowing that the naive forecast peeks at the window, a **1.83.0× gap is
not explained by that advantage alone.** The served counter — a frequency over
the line, blended with the last five games, nudged by opponent rank and home/away
— appears to **destroy discrimination relative to the player's own rate.** The
recency blend and the ±0.03/±0.015 nudges move predictions in ways that do not
track outcomes.
**rbi is the one stat where the model beats the naive baseline.**
### Finding 3 — lineup slot out-resolves the model on THREE stats
| stat | model | slot alone |
|---|---|---|
| total_bases | 0.00442 | **0.03448** (7.8×) |
| runs | 0.00130 | **0.01156** (8.9×) |
| rbi | 0.03268 | **0.03608** (1.1×) |
| hits | 0.00252 | 0.00100 (model wins) |
**Hits is the only stat where batting order carries less than the model** — which
makes sense: a hit is a hit whether you bat first or ninth, but runs, RBI and
total bases all scale with opportunity.
---
## PHASE 2 — which world
**All three are partly true, in measured proportions:**
- **WORLD A — rbi resolution is real and role-driven: ~90% TRUE.** Slot alone
(0.03608) covers the model's entire rbi resolution. It is real, it is
contextual rather than skill-based, and 58% survives within similar-player
strata as genuine discrimination.
- **WORLD B — base-rate-spread artefact: ~36% TRUE for rbi.** Player identity
alone accounts for about a third. Not the main story for rbi — but for
**hits, TB and runs it is the whole story and then some**, since base rate
alone out-resolves the model on all three.
- **WORLD C — hits is intrinsically compressed: TRUE, and the ceiling is low.**
Total available spread resolution for hits is 0.00446 — **1.8% of variance even
from a forecast that has seen the answers.** Perfect factors cannot make hits a
high-resolution grade.
---
## PHASE 3 — the roadmap-deciding log
**The next arc is not "strengthen hits factors."** Hits has the lowest available
resolution on the board (1.8% ceiling from an oracle-ish baseline) and last
order's wiring already lifted it to 1.39% of a 1.8% ceiling. There is very little
left there.
**Named first factor order for next session: LINEUP SLOT / RISP OPPORTUNITY on
rbi**, through the two-part gate.
- The input is **already ingested and prod-verified** (`lineup_context` batting
order, `hitter_opportunity` RISP share, S89).
- Its resolution is **measured, not hypothesised**: 0.03608 slot-only on rbi,
0.03448 on TB, 0.01156 on runs.
- The causally-correct unit is **plate appearances with runners on**, which is
what RISP share measures directly — the crude version is the slot integer.
- It must clear the two-part gate like anything else. Measured availability is
not a pass.
**Second, and higher-value than either:** the counter is out-resolved by a player
frequency table on three of four stats. That is not a factor problem — it is a
defect in the champion. **Diagnosing whether the recency blend and the
±0.03/±0.015 nudges are destroying discrimination is the biggest single lever
this decomposition found**, and it costs nothing to test: they are three lines in
`probabilityEstimator`.
**The hits transmission win stands.** The conduit is real and permanent —
sign-verified, 75% coverage. This order changes only *which stat* has the most
worth flowing through it.
---
## Invariants
Diagnostic only. No factor wired, no serving path changed, no calibration refit,
`p_win` untouched, all frozen modules byte-identical including the hits path from
`43f65d3`. No Bonferroni slot — this is resolution accounting, not a causal claim.
+141
View File
@@ -0,0 +1,141 @@
# The reliever chain — Link 1 proves, Link 2 does not, and the premise inverts
**The causal insight is right: the game is a sequence and the matchup does shift
mid-game. The direction is backwards.** Measured on 93,663 plate appearances from
1,238 games, the bullpen is *harder* than the starter, not softer.
---
## The chain, link by link
### LINK 1 — starter pull timing: **PROVES**
Target is batters faced, because that is what decides how many of a hitter's
plate appearances come against the starter rather than the pen. Point-in-time:
each start predicted only from that pitcher's starts strictly before it, shrunk
toward the league mean by prior-start count. Baseline is the league mean — the
naive "a starter goes about six."
```
1,706 starts · 204 pitchers · clustered on the pitcher · 107 cumulative tests
MAE 3.2226 (baseline) -> 2.7990 (model) delta -0.4236
CI [-0.6006, -0.2731] at 0.9995 VERDICT: PROVES
```
It finds the tail, which is what the chain needed: early exits (≤20 batters
faced) occur at a **23.2%** base rate, and among model-flagged starts they occur
at **34.0%** — lift **+10.8pp**.
**Scope correction made here:** the order specifies fatigue profile × *game
script* ("getting hit → pulled early"). Game script is not available when a prop
is graded — whether he gets hit tonight is the thing being projected, not an
input to it. Using it would be reading the answer. Only the fatigue/workload half
is measured above; the in-game half is a LIVE feature, recorded as out of scope
rather than quietly folded in.
### LINK 2 — reliever identity: **NOT PROVEN**, on two independent grounds
Predict which arm throws a given post-starter plate appearance. Baseline: the
team's most-used reliever to date. Model: the arm that team has most often used
*in that inning* to date — the cheapest expression of bullpen role.
```
39,629 post-starter plate appearances · 30 bullpens
baseline accuracy 8.6% -> model accuracy 17.2%
VERDICT: PENDING_SAMPLE — 30 independent clusters < 40
```
1. **On merit.** Doubling the baseline sounds good and is not: naming a specific
arm is **wrong five times out of six.**
2. **Structurally.** Bullpen usage is a team-level process — same manager, same
arms, same roles all season — so the entity this prediction rides on is the
club, and there are 30. Row count cannot create replication that does not
exist. **The same permanent ceiling as park geometry (30 venues) and team
defence (26 teams).**
### LINK 3 — shifted matchup: **NOT RUN**
Per the order's own discipline, a link that does not prove does not feed the
next. Link 3 needs the reliever's profile, and Link 2 cannot say whose profile it
is.
---
## The premise, tested directly — because that chains on nothing
This required no unproven link, so it was safe to measure, and it is the finding
that matters most:
| | n | hit rate per PA | ±95% |
|---|---|---|---|
| vs **STARTER** | 48,492 | **0.2444** | 0.0038 |
| vs **BULLPEN** | 35,760 | **0.2373** | 0.0044 |
| bullpen \| starter exited early | 14,672 | 0.2379 | 0.0069 |
| bullpen \| starter went normal | 21,088 | 0.2369 | 0.0057 |
**The bullpen is 0.7pp HARDER than the starter**, and the intervals barely
overlap. The specific effect the chain was built to exploit — an early exit
making later at-bats softer — is **+0.0010, indistinguishable from zero on 35,760
plate appearances.** That is a well-powered null, not a sample problem.
### What IS real: times through the order
| | n | hit rate |
|---|---|---|
| TTO 1 | 21,596 | 0.2351 |
| TTO 2 | 18,426 | **0.2515** |
| TTO 3 | 8,278 | **0.2518** |
A starter does decay as the lineup sees him again: **+1.6pp from first look to
second.** But that advantage is **surrendered when he leaves, not extended**
the pen (0.2373) is harder than the starter's second and third time through
(0.2515).
The mechanism is a modern bullpen: a queue of specialists throwing max effort for
one inning each, fresh, often handedness-matched. There is no tiring arm to
punish.
### The insight survives, inverted
The sequence framing is correct and the edge is real — it just points the other
way. **A hitter's soft spot is a starter still in the game on the third time
through, and an early hook takes it away.** So Link 1 remains valuable, for the
opposite reason it was built: flagging a likely early exit predicts that a hitter
*loses* his third-time-through look (0.2518 → 0.2373, a 1.45pp shift on that
plate appearance) — a downgrade signal, not an upgrade.
That is also market-relevant in the way the order wanted, with the sign flipped:
if a line is set on the starter's matchup, the mispricing is on hitters who will
get an *extra* look at a starter going deep.
---
## Built
- `src/services/model/predictionGate.js` (+ tests) — the two-part gate for a
CONTINUOUS prediction. `factorGate` binarises outcomes for Brier, which would
destroy a target like batters faced. Same discipline: movement AND out-of-sample
improvement, paired bootstrap, clustered, cumulative-corrected. Names THEATER
the same way.
- `scripts/ingest-game-sequences.js` — per-PA batter/pitcher/hand/inning/result
plus boxscore exit lines, free from statsapi. 1,238 games cached.
- `scripts/link1-pull-timing.js`, `scripts/link2-reliever-identity.js`.
## Pre-registered, NOT run
**Link 2 — bullpen AGGREGATE instead of a named arm.** Naming the arm fails, but
a PA-weighted aggregate of the pen's contact-allowed and handedness profile may
be knowable, and the hitter×bullpen unit would have real replication where the
bullpen alone has 30. This is recorded rather than substituted in, because
running Link 3 on a swapped-in Link 2 is precisely the assumed-link failure the
order forbids. Given the premise result above, its expected value is now low.
## Parallel track — total_bases per-archetype (logged, not run)
Sample audit only: `total_bases` settled n=948 pooled; BOMBER × TB **340**,
short by 160 against the gate. No archetype slot is testable yet. Per the S88
lesson, this is a *sample-readiness* note and not a verdict — and per
`specs/per-archetype-grade-bands.md`, the grade does not yet separate within any
archetype on hits, so a TB rescale would face the same second blocker.
Counter and frozen clusters byte-identical.
+113
View File
@@ -0,0 +1,113 @@
# Built-but-unread — the defect class, and the guard that ends it
## The class
Three consecutive orders shipped a backend-correct field that never reached a
screen. All three passed a fully green suite.
| order | what was built | why it never rendered |
|---|---|---|
| grade bands | `gradeBands.js`, six orders of work | required by no serving code |
| `91927a4` | `served_grade` on the payload | dropped at the adapter boundary |
| `3591c76` | `GradeScaleLegend.tsx` | imported by nothing |
Each was caught by luck on a later re-check, and in two of the three I had
already **reported the wiring as done**.
### Why green tests could not see it
Backend tests stop at the API payload. They prove a field is **produced** and say
nothing about whether it is **consumed**. The failure is invisible to them by
construction — not an oversight in any individual test.
### The cognitive trap, named
The difficulty pools in the backend. Deriving the grade, proving the factors,
measuring resolution — that is where the thinking happens, and by the time a
field exists on the payload it *feels* finished. The remaining step is a
three-line adapter change that nobody considers worth verifying, so it gets
claimed rather than traced. **The last inch is the one with no friction, which is
exactly why it is the one that gets skipped.**
Nothing here is a frontend-competence problem. It is that "I added the field" and
"a user can see it" are different claims, and only the first one is fun.
---
## The contract
`tests/unit/renderReachability.test.js` holds the promised-field contract — the
things the grade product commits to a user seeing:
| promise | payload | adapter | component |
|---|---|---|---|
| the served grade object | `served_grade` | `served_grade` (container) | GradeResultCard |
| what this grade means | `served_grade.meaning` | `gradeMeaning` | GradeResultCard |
| whether the band separates | `separates_from_base_rate` | `separatesFromBaseRate` | GradeResultCard |
| what the band realized | `band_realized_rate` | `bandRealizedRate` | GradeResultCard |
| which factors moved the read | `factor_adjustment` | `factorsApplied` | GradeResultCard |
| the ceiling stance | — | — | GradeScaleLegend |
Adding a served field without adding it here is allowed. Adding it **here**
without wiring it to a mounted component is not.
---
## The guard
For each promised field it traces the whole path:
```
payload field -> adapter consumes it -> component renders it -> component is MOUNTED
```
**"Mounted" is transitive to a Next entry point** (`page`/`layout`/`template`) —
the only thing that puts a pixel on a screen. A component that exists and renders
its field perfectly but is imported by nothing fails. Depth-limited so an import
cycle cannot hang the suite.
**Container rows are exempted explicitly, not silently.** `served_grade` is
consumed by the adapter but not rendered directly, so it carries
`container: true` plus a `rendersVia` list — and a separate assertion checks that
**every named part actually renders.** The exemption is auditable; it cannot hide
an unrendered field.
The guard also tests itself: it asserts that an orphan component reports
unmounted, and that the contract is non-empty (an empty contract would pass
vacuously — the way this guard would most plausibly rot).
### Retro-proof
Run unchanged against the tree at `3591c76`, before the wiring:
```
Tests: 11 failed, 8 passed
● the ceiling stance / grade scale legend — its component is MOUNTED, not merely written
● the served grade object — the adapter consumes it
● whether the band separates from the baseline — a component actually renders it
● what this grade means — the adapter consumes it
● which proven factors moved the read — a component actually renders it
...
```
**It names the exact three bugs.** Green on the current tree.
### One honest scope limit
`gradeBands` is **not** in the contract and would not be caught. It is a backend
module, not a promised user-facing field, and it is correctly unwired — every
band it produces collapses to base-rate at current resolution. The guard covers
*promised* fields; a backend module that should not yet render is out of scope by
design, not by oversight.
---
## In the deploy floor
The guard runs in the standing suite, so it is part of the three-gate floor:
**tests green** (now including reachability) + web build exit 0 + post-deploy
fingerprint. A future order that adds a served field without wiring it to a
mounted component fails CI rather than a hand-check three orders later.
No serving or model change in this order. `p_win` never mutated. No Bonferroni
slot.
+156
View File
@@ -0,0 +1,156 @@
# The resolution ceiling — calibration is complete, and it was never the lever
## PHASE 0 — two honest truths, on record
### 1. The swap is a BET, not an OOS win
On identical held-out rows the **isotonic map scored BETTER**: hits +0.0028
(CI [0.0013, 0.0045]), rbi +0.0042 (CI [0.0003, 0.0092]), total_bases tied.
We serve the low-parameter map anyway, on the untestable prior that isotonic's
in-window edge is daily structure shared between the fit and evaluation windows
and will not transmit forward. At 19 dates **no instrument here can test that
prior** — LODO has 1.49.3% power against it.
**Named as a bet, logged, not evidence.** Phase 1 makes it falsifiable.
### 2. The MIN_SLOPE catch, as a standing guard rationale
`runs` fitted `a = 0.032`. A near-zero or negative slope collapses the curve
toward *base-rate-for-everything*, which **lowers Brier** — shrinking a
miscalibrated forecaster toward its base rate always does — while destroying all
resolution.
**A metric win that guts the product.** Any calibration layer must refuse a
non-positive slope on principle, not on inspection. That is now `MIN_SLOPE`.
---
## PHASE 1 — the duel, instrumented forward
Both corrections are computed on every hits/TB prop:
`p_win_lowparam` (served) and `p_win_isotonic_shadow` (logged, never read by
serving or by `chainAcross`). The shadow runs in its own try — it can never
break serving.
`calibrationDuel.adjudicate` encodes the rule **in code, before any forward date
exists**, so the bar cannot drift toward whichever answer arrives:
| condition | verdict | action |
|---|---|---|
| ≥10 forward dates AND isotonic wins, date-block CI excluding zero | **REFUTED** | revert hits/TB to isotonic, log the reversal |
| ≥10 forward dates, isotonic does not win | **UPHELD** | keep serving low-param |
| <10 forward dates | **PENDING** | keep serving, no verdict |
A date counts as forward **only if neither map was fitted on it** — scoring
inside a fit window would ask which map memorised better. Rows lacking that
provenance are dropped, never assumed forward. Tests lock all of it, including
that a decisive shadow win at 5 dates is still PENDING.
**Nothing swaps now.** The season decides.
---
## PHASE 2 — the resolution ceiling, quantified
Murphy decomposition: `Brier = reliability resolution + uncertainty`.
Reliability is what calibration fixes. **Resolution is discrimination, and a
monotone map cannot create it** — it relabels bins without re-sorting the rows
inside them.
| stat | n | base | reliability | **resolution** | uncertainty | **share of variance explained** |
|---|---|---|---|---|---|---|
| hits | 1,140 | 0.5684 | 0.01353 | **0.00252** | 0.24532 | **1.03%** |
| total_bases | 1,050 | 0.5819 | 0.01419 | **0.00442** | 0.24329 | **1.82%** |
| rbi | 630 | 0.6571 | 0.00654 | **0.03268** | 0.22531 | **14.51%** |
| runs | 597 | 0.6348 | 0.00788 | **0.00130** | 0.23182 | **0.56%** |
### What calibration did, exactly as theory predicts
| stat | reliability | resolution |
|---|---|---|
| hits | 0.01353 → 0.00233 (**0.0112**) | 0.00252 → 0.00231 (0.0002) |
| total_bases | 0.01419 → 0.00527 (**0.0089**) | 0.00442 → 0.00414 (0.0003) |
**Calibration removed 83% of hits' reliability error and moved resolution by
essentially nothing.** It did the whole of its job, and its job was never the
one the grade product needs.
**Unexpected:** `rbi` has **13× the resolution of hits** and is the one stat we
do *not* serve corrected — it needs calibration least (reliability 0.0065) and
discriminates most. Worth carrying into the factor arc.
---
## PHASE 2 — the HITS factor-transmission diagnosis
**Verdict: NOT-TRANSMITTED. Not weak — absent.** This is a plumbing defect, and
it is the highest-value finding in the order.
Evidence, traced in code rather than recalled:
1. **`sprayDefense.js` and `platoonSeverity.js` are required by NOTHING in
`src/`.** Only by analysis scripts and their own tests. The two
causally-correct atoms that passed the two-part gate have never been on the
serving path.
2. **The served `p_win` reads exactly four inputs**
(`intelligence/probabilityEstimator.js:54`): game-log frequency over the line,
`opp_rank_stat` (±0.03), `home_away` (±0.015), and a cv consistency pull.
Zero occurrences of spray, platoon, hard-hit or contact-profile.
3. **Ordering makes it structural.** `snapshotService` grades at line 454
(`gradeAndCacheSlate`) and only computes challenger/context at line 640+.
Everything proven is computed **downstream of the grade it would inform**.
So the three proven hits factors — `defense_by_direction`,
`pitcher_contact_profile`, `platoon_severity` — were measured on ledger rows by
analysis scripts and **have never once moved a served number.**
That reframes every null in this programme's recent history. "Calibrated p_win
does not separate within archetype" was never a statement about factors. The
factors were not in the forecast.
---
## PHASE 3 — bands on served values
| stat | basis | eval rows | slots | slots with lift |
|---|---|---|---|---|
| hits | `p_win_lowparam` (served) | 765 | 7 | **0** |
| total_bases | `p_win_lowparam` (served) | 625 | 7 | **0** |
| rbi | raw `p_win` | 425 | 7 | **0** |
| runs | raw `p_win` | 424 | 7 | **0** |
**28 archetype slots across four stats. Zero show lift.** Every slot is one band
indistinguishable from its own base rate.
This is no longer an open shrug. It is the arithmetic consequence of resolution
of 0.00130.0327 against uncertainty of ~0.23: **a forecast explaining 1% of the
outcome's variance cannot produce bands that separate**, and no correction to its
numbers will change that.
---
## THE HEADLINE
**Calibration is complete. It delivered honest numbers on two stats and ZERO
grade separation, because the counter has no resolution — 1.03% of variance on
hits, 0.56% on runs.**
**And the three proven hits factors are NOT WIRED INTO THE FORECAST AT ALL.**
Those two facts together are the programme's position. The second is the reason
for the first, and it is a plumbing defect rather than a modelling wall — which
makes it the cheapest high-value fix available.
**Per-archetype grades require proven factors that actually reach `p_win`. That
is the next and central arc.** This is the last calibration order.
---
## Invariants
Serving unchanged from `74cf1ce` — this order logs and diagnoses. `p_win` never
mutated; `p_win_lowparam` served, `p_win_isotonic_shadow` logged. No Bonferroni
slot. Counter and frozen clusters verified file-by-file.
+137
View File
@@ -0,0 +1,137 @@
# The bias is robust; the map was not. Low-parameter correction deployed.
## PHASE 0 — the sample-limit truth, on record
**On 19 dates, BOTH stability instruments are underpowered. This is the SAMPLE,
not a fixable instrument.** No future order should re-open the gate-refinement
loop expecting a different answer at this N.
- **LODO power 0.0140.093** against a strong date-driven instability. Across
every k from 1.0 to 2.0, the best any stat reaches is 0.337.
- **Deploy CIs rest on 24 date clusters.** A cluster-robust interval at 2
clusters has ~1 degree of freedom and a near-undefined width.
Neither certifies forward stability of a specific map. Four orders refined a gate
the sample cannot support; that loop stops here.
**Record correction on runs:** its `1f40014` DATE-DRIVEN classification was an
artefact of the coin-flip ruler — 2 reversals in 3 drops never cleared a cutoff
of 2. runs is an ordinary "no fittable map" refusal. **Not date-driven.**
---
## PHASE 1 — the robust claim: ROBUST
Model-free, map-free, on the picked-side deduped population. Date-block bootstrap
(whole dates resampled, 5,000 draws).
### Pooled — the shape is textbook favourite-longshot
| bin | n | predicted | realized | over-prediction |
|---|---|---|---|---|
| 0.50.6 | 1,021 | 0.5477 | 0.5553 | 0.0076 |
| 0.60.7 | 961 | 0.6432 | 0.6004 | +0.0428 |
| 0.70.8 | 745 | 0.7420 | 0.6456 | +0.0963 |
| 0.80.9 | 459 | 0.8430 | 0.6841 | +0.1589 |
| **0.91.0** | 157 | 0.9075 | **0.6624** | **+0.2451** |
Pooled sign stability **0.9946** over 17 date blocks, 90% CI [+0.136, +0.300],
zero indeterminate resamples.
### Per stat — 4 of 4 replicate
| stat | n | >0.9 bias | date blocks | sign stability | 90% CI | replicates |
|---|---|---|---|---|---|---|
| hits | 1,140 | +0.2435 | 17 | 0.994 | [0.120, 0.321] | yes |
| total_bases | 1,050 | +0.2816 | 7 | 1.000 | [0.154, 0.394] | yes |
| rbi | 630 | +0.2107 | 5 | 1.000 | [0.156, 0.245] | yes |
| runs | 597 | +0.2367 | 5 | 0.998 | [0.082, 0.314] | yes |
**VERDICT: ROBUST** — pooled ≥95% and 4/4 stats (bar was 3/4).
Worth noting alongside it: **realized rate plateaus at ~0.650.68 from p=0.7
upward.** The 0.9+ bucket (0.6624) performs no better than the 0.80.9 bucket
(0.6841). The model has no genuinely high-confidence reads, only high-confidence
*numbers*.
---
## PHASE 3 — low-parameter correction, validated as a new estimator
Platt: `p_cal = sigmoid(a·logit(p) + b)`. Two parameters over the whole curve, so
it **cannot** encode "this Tuesday was odd" — which is precisely the failure mode
we cannot rule out for isotonic on this sample.
Shrunk toward identity by fit-date count: `w = D/(D+10)`, applied as
`w·p_platt + (1w)·p_raw`. A thin fit is therefore applied at reduced strength.
| stat | a | shrink | eval n | blocks | Brier raw | low-param | Δ vs raw | CI (date-block) | decision |
|---|---|---|---|---|---|---|---|---|---|
| **hits** | 0.406 | 0.565 | 765 | 4 | 0.2626 | 0.2540 | **0.0086** | [0.0112, 0.0069] | **DEPLOY** |
| **total_bases** | 0.472 | 0.333 | 625 | 2 | 0.2490 | 0.2429 | **0.0061** | [0.0062, 0.0059] | **DEPLOY** |
| rbi | 0.775 | 0.231 | 425 | 2 | 0.2011 | 0.2007 | 0.0004 | [0.0007, **0**] | REFUSE |
| runs | 0.032 | — | — | — | — | — | — | — | **REFUSE (slope)** |
### A guard the first run needed
runs fitted **a = 0.032**. A non-positive slope does not flatten an
over-confident forecaster — it **inverts** it, and near zero the curve collapses
to a constant, predicting the base rate for everything. That *lowers* Brier
(shrinking a miscalibrated forecaster toward its base rate always does) while
destroying all resolution, so it would have **scored as a win while making the
product worthless**. `MIN_SLOPE` now refuses it by name, with a test.
### Stated plainly: isotonic scored better, and we are not using it
On the identical held-out rows, isotonic beat the low-parameter fit on hits
(+0.0028, CI [0.0013, 0.0045]) and rbi (+0.0042, CI [0.0003, 0.0092]), and tied
on total_bases (0.0009, CI spanning zero).
**The swap is a capacity judgement, not a measurement.** The evaluation window
spans 24 date blocks, so "isotonic wins OOS" there is weak evidence, and it is
exactly what a flexible map would produce if it captured structure shared by the
fit and evaluation periods. That reasoning is a judgement and is labelled as one.
---
## PHASE 4 — deploy and labelling
| stat | served | basis |
|---|---|---|
| hits | low-parameter correction | `direction_robust_magnitude_provisional` |
| total_bases | low-parameter correction | `direction_robust_magnitude_provisional` |
| **rbi** | **RAW — withdrawn** | deployed on isotonic at `ced4042`; low-param does not beat raw |
| runs | RAW | slope refused |
The **direction** is bootstrap-robust; the **magnitude** is thin-sample and
conservatively shrunk (0.565 hits, 0.333 TB). Customer-facing letter unchanged.
Auto-demotion remains armed via `calibrationRegistry.reverify`: a sign flip in
the >0.9 bucket or a CI crossing zero demotes to raw and logs the breaking date.
Promotion to non-provisional stays at the original ≥40 date-cluster bar.
---
## PHASE 5 — the standing finding, stated hard
**Across 18 archetype slots on three stats, calibrated `p_win` separates within
archetype NO BETTER than raw. Every slot collapses to one band, indistinguishable
from its own base rate. Zero slots show lift.**
Per-archetype grade separation is **not coming from calibration**. It comes from
**proven factors or it does not exist.**
This reframes the roadmap. Calibration has now been pursued through five orders
and has delivered exactly what it can deliver — honest numbers on two stats — and
nothing at all on the question the grade product actually turns on. The next real
lever is factors on the stats that lack them.
---
## Invariants
`p_win` never mutated — the correction rides as `p_win_calibrated`. Calibration
consumed no Bonferroni slot. The robust-claim test ran before any calibrator was
built and could have terminated the session at Phase 2. Counter and frozen
clusters verified file-by-file (14 modules, including `calibration.js` and
`calibrationService.js`, both untouched and simply no longer on the serving path).
+143
View File
@@ -0,0 +1,143 @@
# runs + RBI — mostly-base-rate confirmed, and stronger than expected
**Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest
grade is the POOLED base rate.** The order anticipated that mostly-base-rate
would be the correct answer for context-heavy stats. It is, and one level deeper
than predicted.
## Premise note
The order's closing line says the batter board is "per-archetype-graded" after
this. It is not: **nothing has been rescaled for hits or total_bases either**
no archetype slot has ever reached sample, and `gradeBands` remains built, gated
and unwired. This is the fourth stat measured, not the completion of three.
---
## STEP 1 — Full-history audit
| | clean | with p_win | players | dates | **games** | lines |
|---|---|---|---|---|---|---|
| RBI | 935 | 931 | 344 | 5 | **43** | 0.5 on 885 |
| RUNS | 617 | 614 | 334 | 5 | **33** | 0.5 on all |
Zero quarantined in either. **No archetype slot reaches 500.**
| archetype | RBI n | RUNS n |
|---|---|---|
| UNLABELLED | 546 | 341 |
| BOMBER | 207 | 148 |
| GHOST | 92 | 67 |
| BRUSH | 43 | 26 |
| **DRIVER** | **24** | 20 |
| **CATALYST** | 13 | **9** |
**The signature archetypes the order names are the two smallest slots on the
board** — RBI→DRIVER at n=24, runs→CATALYST at n=9. The hypothesis is reasonable
and we are three orders of magnitude from being able to test it.
**RUNS is refused structurally before any factor is tested: 33 game clusters
against a 40 floor.** Reported as such rather than dressed up as a factor result.
## Inputs reconstructed rather than declared missing
`lineup_context` only covers 2026-08-04/05/06 (ingest began last week) while
settled rows start 07-31, so just 187 of 617 runs rows join to a batting order.
That reads as input-blocked — but the play-by-play cache runs from 05-01, and
**the batting order IS the order batters first appear.** Batting slot, power
behind (mean barrel of the next three slots) and reach-base rate were all
reconstructed point-in-time from it: coverage went 187 → 574.
---
## STEP 2/3 — The gate (153 / 168 cumulative tests)
### RBI — all THEATER
| factor | n | clusters | shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|
| `risp_opportunity` | 803 | 43 | 0.0585 | **+0.0047** | [0.0023, +0.0122] | **THEATER** |
| `extra_base_skill` | 881 | 43 | 0.0271 | **+0.0010** | [0.0023, +0.0042] | **THEATER** |
| `risp × extra_base` | 803 | 43 | 0.0624 | **+0.0056** | [0.0042, +0.0135] | **THEATER** |
### RUNS — all refused on cluster count, all pointing the wrong way
| factor | n | clusters | shift | Brier Δ | verdict |
|---|---|---|---|---|---|
| `reach_base` | 525 | 32 | 0.0383 | +0.0043 | PENDING — k<40 |
| `lineup_power_behind` | 574 | 32 | 0.0232 | +0.0008 | PENDING — k<40 |
| `reach × power_behind` | 525 | 32 | 0.0482 | +0.0054 | PENDING — k<40 |
### The compound is the WORST version, in both stats
RBI: 0.0047 / 0.0010 → **0.0056** compounded. RUNS: 0.0043 / 0.0008 →
**0.0054** compounded.
The causally-correct compound was the most promising factor on the sheet and is
the most harmful in both. Two multipliers that individually carry nothing do not
cancel — they compound each other's noise. Related to but distinct from the
collapsed-sequence lesson: there the product of two REAL effects was too small to
use; here the product of two NULL effects is actively worse than either.
---
## The archetype itself does not rescue it
The archetype base rates look strongly differentiated, and that appearance is
most of the trap:
| | RBI | RUNS |
|---|---|---|
| DRIVER | **0.609** (n=23) | 0.500 (n=20) |
| GHOST | 0.467 | **0.716** (n=67) |
| BOMBER | 0.413 | 0.460 |
| pooled | 0.4305 | 0.5721 |
Tested directly — is the archetype's leave-one-out base rate better than the
pooled one?
| | shift | Brier Δ | CI | verdict |
|---|---|---|---|---|
| **RBI** | 0.0231 | +0.0010 | [0.0034, +0.0050] | **THEATER** |
| **RUNS** | 0.0642 | 0.0028 | [0.0147, +0.0108] | CANDIDATE — k=33 |
**For RBI, knowing the archetype makes the forecast worse.** DRIVER's 0.609 is
n=23 — the spread is small-slot noise wearing a decimal point. Runs is at least
directionally favourable, and unproven.
Had this been read off the base-rate table instead of gated, the session would
have shipped "archetype differentiation is real and large" as a finding. It is
not one.
---
## The cross-stat pattern that IS real
| stat | mean p_win | actual | counter bias |
|---|---|---|---|
| total_bases | 0.5698 | 0.5074 | **+0.0624** |
| RBI | 0.4860 | 0.4313 | **+0.0547** |
| runs | 0.5949 | 0.5749 | **+0.0200** |
**The counter over-predicts every batter counting stat measured.** Across four
stats and three sessions, calibration is the systematic defect and factor
scarcity is not — TB's held-out isotonic fix (0.0039 Brier) still outperforms
every factor tried on any stat, all of which have been null or theatre.
## STEP 4 — No rescale
Two-bar rule: nothing proved, nothing certified calibrated, no slot at sample —
and for RBI the archetype split is itself theatre, so the honest band is the
POOLED base rate rather than a per-archetype one. `gradeBands` returns exactly
that by construction.
## Next, by value
1. **Calibrate the counter across all four stats.** One systematic bias,
measured four times, larger than anything else on the board.
2. **Stop adding factors to context stats.** Six tested across runs/RBI, six
null-or-worse, and the compounds worst of all.
3. Runs needs game-date accrual to clear the cluster floor before it can be
gated at all.
Counter and frozen clusters byte-identical.
+132
View File
@@ -0,0 +1,132 @@
# The grade surface — it wasn't missing, it was serving the weaker signal
## PHASE 0 — what actually reaches a user
**Correction to the order's premise: a grade letter has been served all along.**
`engine1.gradeProp` produces it from an additive factor index, and it is
**not derived from `p_win` at all** — the two are computed independently and both
ride the payload.
`gradeBands` is orphaned for a different reason than assumed: it defines what a
letter *means* from realized outcomes, and every band it produces collapses to
base-rate at current resolution. It was never the missing link to the surface.
### The measurement that changed the order
3,417 settled props, pooled across the four batter stats:
| grade | n | realized | mean p_win |
|---|---|---|---|
| **A** | 8 | **0.5000** | 0.6466 |
| B | 985 | 0.6396 | 0.7002 |
| C | 1,695 | 0.6024 | 0.6755 |
| D | 303 | 0.5578 | 0.6042 |
| F | 426 | 0.5352 | 0.5875 |
```
grade-letter resolution 0.00116 (0.48% of variance)
p_win resolution 0.00715 (2.98%)
=> the letter carries 0.16x the information of the number beside it
```
**The top grade hit worse than the bottom grade.** And concretely, from the
hand-verify: **Christian Encarnación's 0.95 over graded `C`, and his 0.05 under
also graded `C`** — same hitter, opposite forecasts, same letter.
The gap was never that grades don't ship. It is that **the weaker of two
available signals was shipping as the headline.**
---
## PHASE 1 — the honest grade
`model/servedGrade.js` derives the letter from `p_win`, with bands anchored on
**measured realized rates**, not targets:
| letter | p_win ≥ | realized | separates from base rate? |
|---|---|---|---|
| B+ | 0.780 | 0.663 | yes |
| B | 0.700 | 0.646 | yes |
| C+ | 0.640 | 0.615 | **no** |
| C | 0.560 | 0.589 | **no** |
| C- | 0.480 | 0.548 | **no** |
| D | 0.350 | 0.512 | yes |
| F | 0.000 | 0.447 | yes |
Base rate 0.6005.
### No manufactured A — structurally
**`A+`, `A` and `A-` are UNISSUABLE.** Not rare — absent by construction. The
realized rate plateaus at 0.650.68 above p_win 0.70 (the 0.9+ bucket does no
better than the 0.8 bucket), so no band of this forecast has earned a top letter.
A test sweeps every p_win from 0 to 1 and asserts none produces one. Even a 0.99
forecast tops out at B+ with its realized 0.663 attached.
When resolution improves enough to earn an A, the ceiling gets raised
deliberately and visibly — not by a threshold quietly drifting.
### Bands that cannot separate SAY so
`C+ / C / C-` carry `separates_from_base_rate: false` and copy that names it —
*"a base-rate read; the model sees nothing that separates this."* That covers the
bulk of the board, and it is the honest description of a forecast explaining 3%
of variance.
### The basis is stated, never implied
Each grade carries `basis`: `forecast_plus_matchup_factors` (naming which of the
three proven factors fired) or `forecast_only`, plus `calibrated: false`
calibration is withdrawn and nothing here rides on a number that doesn't exist.
`engine1.grade` is preserved as `engine_grade` so nothing downstream breaks and
the two stay comparable.
---
## PHASE 2 — the refusal surface
Refusals render a real state, never a blank or a fabricated number:
- `insufficient_data`**NO READ***"not enough history to call this one"*
- `juiced_no_edge`**NO READ** — *"the book has priced the vig past any edge on
this side"*
1,870 refused snapshots carry exactly these two reasons, and both now surface.
`projectionFor` reads the repaired full-window reference, so refusals are
computed on the repaired champion.
---
## PHASE 3 — hand-verified on real served props
| prop | p_win | OLD | NEW | separates | state |
|---|---|---|---|---|---|
| Freddie Freeman hits 0.5o | 0.95 | B | **B+** | true | graded |
| Christian Encarnación hits 0.5o | 0.95 | **C** | **B+** | true | graded |
| Ben Rice hits 0.5o | 0.95 | B | **B+** | true | graded |
| Christian Encarnación hits 0.5u | 0.05 | **C** | **F** | true | graded |
| Ben Rice hits 0.5u | 0.05 | C | **F** | true | graded |
| Eliezer Alfonso Jr doubles 0.5o | — | — | **NO READ** | — | refused |
| Eliezer Alfonso Jr doubles 0.5u | — | — | **NO READ** (vig) | — | refused |
| Paul Goldschmidt doubles 0.5o | — | — | **NO READ** | — | refused |
```
never-blank check: PASS — every prop renders a label and a meaning
no-manufactured-A check: PASS
```
The Encarnación rows are the clearest evidence: under the old letter his 0.95 and
his 0.05 were both `C`. Under the new one they are `B+` and `F`.
---
## Invariants
Grades ride on repaired-champion raw `p_win` plus factors where they fire. No
calibrated number leaks — the deployed set is empty and `calibrated: false` is
stated on every grade. `p_win` never mutated. Nine frozen model modules verified
unchanged, `engine1` included. No Bonferroni slot — no new factor.
**Still true and unchanged:** the forecast explains ~3% of outcome variance. This
order did not make the model better. It made the letter stop overstating it.
@@ -0,0 +1,179 @@
# Settlement + four-stat calibration — the enabling move did not enable
**Settlement is done (15,484 rows written). Calibration improves held-out Brier
on all three stats it can be fitted for — more than any factor ever tested. And
no stat can deploy, because the date-cluster ceiling is 17, not 90.**
---
## PHASE 0 — Board reconcile (corrections to the record)
| claim in circulation | measured |
|---|---|
| ~22,032 snapshots unsettled | **71,192** rows, **0** settled |
| snapshots span 05-01 → 08-04 (~90 dates) | **07-19 → 08-06 = 19 dates** |
| "calibrate all four on the full replay" | replay yields **17 / 7 / 5 / 5** dates per stat |
| batter board is per-archetype graded | **nothing has ever been rescaled**, on any stat |
| TB inversion confirmed | **units-bug artifact — UNPROVEN** (S: prove-tb-factors) |
**What actually drives each number today: a base-rate band, on all four stats.**
No stat is per-archetype graded. Proven FACTORS exist on hits only
(`defense_by_direction`; `pitcher_contact_profile` and `platoon_severity` are
held/demoted, not proven). `gradeBands` remains built, gated and unwired.
---
## PHASE 1 — Settlement
```
candidates 34,650 (4 stats) -> settled 16,498 | unresolvable 12,894 | orphaned 5,258
written 15,484 (shortfall = idempotency guard vs the live cron)
conservation check PASSED
```
### Two integrity findings, both caught by the gate
**1. Unordered pagination over a LIVE table.** The dupe check hard-failed on
snapshot id 33875. `model_snapshots` is written by the cron at 14/19/22/1/3 UTC,
and an unordered `.range()` walk over a table being appended to returns
overlapping pages. Fixed with `.order('id')`. A plain id-only fetch showed no
dupes, so this only bites on longer reads that straddle a cron write.
**2. 12,894 rows were logged AFTER first pitch.** Cycles at ET 21:00/22:00/23:00
*on the game date* (10,738 rows) plus 664 the following morning. A 01:00-UTC
cycle is 21:00 the previous evening Eastern — same game date, ~2 hours into the
slate. These are not predictions and are excluded.
Tested whether they were outcome-contaminated: bias +0.0058 in-game vs +0.0008
pre-game — **not sharper, just late.** Excluded for provenance, not because they
cheated.
### The enabling move did not enable
| stat | usable props | date-clusters |
|---|---|---|
| hits | 1,140 | **17** |
| total_bases | 1,050 | **7** |
| rbi | 630 | **5** |
| runs | 597 | **5** |
71,192 rows collapse to 4,799 distinct pre-game props: a 2.5× cycle fan-out,
then 97.6% both-sides duplication, then the pre-game filter. **Hits ends with
1,140 rows against the ledger's existing 1,312.** Settlement was worth doing as a
standing debt; it did not unlock the sample the order expected.
---
## THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY
**97.6% of props carry BOTH sides.** Their p_wins sum to ~1 and their outcomes
are complementary, so any calibration statistic over the raw population is pinned
to 0.5 by construction:
| | hits | TB | rbi | runs |
|---|---|---|---|---|
| both-sides population | +0.0002 | +0.0012 | +0.0019 | +0.0000 |
| **model-picked side only** | **+0.0868** | **+0.0834** | **+0.0164** | **+0.0410** |
The first row reads "the counter is perfectly calibrated" and would have
overturned three sessions of findings. Same rows, opposite conclusion, and the
tell was mean p_win sitting at 0.4998 on every stat.
---
## PHASE 2/3 — Calibration and the deploy gate
Isotonic, point-in-time, fit-past / apply-forward. The split is placed by
cumulative ROWS rather than date index — props are not spread evenly across dates
and a 60%-of-dates cut left only 143 rows to fit on, under the fitter's 200
minimum. Still strictly temporal: every fit date precedes every eval date.
| stat | n | dates | bias | fit / eval | Brier raw | Brier cal | Δ | CI (date-clustered) | eval dates | decision |
|---|---|---|---|---|---|---|---|---|---|---|
| hits | 1,140 | 17 | +0.0868 | 375 / 765 | 0.2626 | 0.2511 | **0.0115** | [0.0139, 0.0097] | 4 | **REFUSE** |
| total_bases | 1,050 | 7 | +0.0834 | 425 / 625 | 0.2490 | 0.2438 | **0.0052** | [0.0061, 0.0045] | 2 | **REFUSE** |
| rbi | 630 | 5 | +0.0164 | 205 / 425 | 0.2011 | 0.1965 | **0.0046** | [0.0092, 0.0010] | 2 | **REFUSE** |
| runs | 597 | 5 | +0.0410 | 173 / 424 | — | — | — | — | — | **REFUSE** |
- hits / TB / rbi: **held-out Brier improves and the interval excludes zero.**
Every one of these beats every factor ever tested on any stat.
- runs: **no map could be fitted** — 173 fit rows under the 200 minimum.
- **All four refuse on the date floor: 24 eval date-clusters against 40.**
Certified bands (held-out |err| ≤ 0.05): hits [0.50.7], TB [0.60.8],
rbi [0.50.9]. Outside band → refuse, fall to base rate.
### A null that scored itself
The first run reported hits at Brier **0.5567** — worse than predicting 0.5 for
everything. `fitIsotonic` returns null below its minimum, `applyIsotonic` then
returns null per row, and `(null 1)² === 1` while `(null 0)² === 0`, so the
"Brier score" was silently just the win rate (0.5684). **This project's signature
`Number(null) === 0` breach, in my own measurement code.** Now a hard refuse.
---
## PHASE 4 — Bias shape (diagnostic only)
| stat | 0.50.6 | 0.60.7 | 0.70.8 | 0.80.9 | 0.91.0 |
|---|---|---|---|---|---|
| hits | +0.024 | +0.054 | +0.145 | +0.244 | +0.244 |
| total_bases | 0.019 | +0.076 | +0.156 | +0.160 | +0.282 |
| rbi | 0.026 | 0.045 | 0.039 | +0.095 | +0.211 |
| runs | 0.050 | +0.031 | +0.064 | +0.155 | +0.237 |
**Not a uniform shift — favourite-longshot concentration, identically shaped on
all four stats.** Near zero or slightly negative at the bottom, then rising
sharply. The counter is over-confident specifically about its favourites, which
is the population a user acts on.
Cross-stat gradient measured: hits +0.0868 ≈ TB +0.0834 > runs +0.0410 > rbi
+0.0164 — not the TB > RBI > runs the order anticipated.
---
## PHASE 5/6 — Not run, honestly
Both are gated on a Phase 3 deploy. Nothing deployed, so there is no
`p_win_calibrated` to rebuild bands on and no activated stat to test
per-archetype curves against. Running them would be building on a gate that did
not open.
---
## PHASE 7 — Refusal accuracy (first real measurement)
| stat | refused n | refused over-rate | graded over-rate | refused \|dist from 0.5\| | graded |
|---|---|---|---|---|---|
| hits | 59 | 0.4237 | 0.5368 | 0.076 | 0.037 |
| total_bases | 74 | 0.2162 | 0.4029 | **0.284** | 0.097 |
| rbi | 141 | 0.1773 | 0.2429 | **0.323** | 0.257 |
| runs | 12 | 0.6667 | 0.3015 | 0.167 | 0.199 |
**Refused props are FURTHER from a coin flip than graded ones, not closer.** The
obvious explanation — refusals concentrate on players who barely played — was
tested and does not hold: refused mean 3.20 AB vs graded 3.39, and 6.6% vs 6.2%
with ≤1 AB.
So "we pass on what we can't call" is not quite what happens. We pass on what we
have no INPUT for, and that population had outcomes that were, in hindsight,
lopsided (TB refusals went over 21.6% of the time). Refusing to invent a number
without a reference remains correct — but the pass is not landing on the
genuinely uncertain props, and this is the first time that has been a number
rather than a claim.
---
## Verdict
- **Calibration works.** Three stats improve held-out, all beating every factor
ever tested. The programme-level finding stands: calibration beats every factor
tried on TB/RBI/runs.
- **Nothing deploys.** The date-cluster floor is the right unit for a systematic-
bias claim and we have 24 where 40 is required. Reaching 40 date-clusters
needs ~5 more weeks of accrual, not more replay — the dates do not exist.
- **The floor is the order's own** and it was not relaxed to force a pass.
`p_win` never mutated; no `p_win_calibrated` written since nothing deployed.
Calibration consumed no Bonferroni slot. Counter and frozen clusters
byte-identical.
+146
View File
@@ -0,0 +1,146 @@
# total_bases — every power factor is THEATER, and the units bug nearly hid it
**Nothing proved. The predicted inversion did not appear. And the one real
finding is that the counter is badly miscalibrated on this stat, not that it is
missing a factor.**
---
## Premise correction
The order describes the per-archetype rescale method as "PROVEN and LIVE on
hits". It is neither. `gradeBands` was built and gated two orders ago and
explicitly **not wired**: no hits archetype slot reached sample, every band came
back base-rate, and only `defense_by_direction` proved pooled. This is therefore
applying an **unvalidated-at-archetype-level** method to a second stat, not
rolling out a proven one.
## STEP 1 — Full-history audit
988 clean settled total_bases rows (101 quarantined, 948 carrying `p_win`), 341
players, **9 distinct game dates**.
| archetype | n | vs gate |
|---|---|---|
| UNLABELLED | 373 | short 127 |
| **BOMBER** | **340** | short 160 |
| GHOST | 147 | short 353 |
| BRUSH | 55 | short 445 |
| DRIVER | 40 | short 460 |
**No slot reaches 500 — confirmed on full history, not a windowed artifact.**
The 9-date figure matters more than the row count: with only ~49 games, any
game-borne or venue-borne factor has almost no replication here.
## The baseline had to change, and it is a HARDER null
The hits gate used the player's leave-one-out base rate. That cannot be
reproduced: TB lines vary (1.5 on 559 rows, 0.5 on 345, 2.5 on 45), and a
player's rate of clearing 1.5 bases is a different quantity from 0.5. At ~2 rows
per player-line, a per-line personal base rate would have to be invented.
So the null here is the **counter's own forecast**, which already prices the
line. A factor must beat the champion, not "he's due" — strictly harder.
---
## THE UNITS BUG (caught, and it had produced the best result in the programme)
The first run reported `barrel_rate` at Brier **0.0095**, the largest
improvement ever measured here. It was an artifact.
`fromStatcastRow` returns `barrel_pct` as a **FRACTION** (0.06); the raw table
stores 0100. The factor was written against the percentage scale, so
`(0.06 7.8) × 0.018` clamped **every row** to the maximum negative shift. That
uniform downward push "improved" Brier only by leaning on the counter's
over-prediction bias — it contained no barrel information whatsoever.
Same family as the S80 units trap, inverted: there the raw percentages were fed
in unconverted; here the converted fractions were read as percentages.
`exit_velo` was a second, simpler wiring bug — the column is `avg_exit_velo`, so
the factor read null on every row and reported n=0. **A zero is a wiring bug
until proven an honest absence.**
---
## STEP 2/3 — The gate, units fixed (138 cumulative tests)
| factor | n | entities | clustered on | k | shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|---|---|
| `barrel_rate` | 707 | 179 | game | 49 | 0.0364 | **+0.0036** | [0.0025, +0.0101] | **THEATER** |
| `exit_velo` | 707 | 179 | game | 49 | 0.0229 | **+0.0022** | [0.0030, +0.0072] | **THEATER** |
| `hard_contact_allowed` | 707 | 93 | game | 49 | 0.0260 | **+0.0033** | [0.0019, +0.0082] | **THEATER** |
| `park_weather_hit_type` | 651 | **36** | treatment entity | 36 | 0.0070 | +0.0016 | [+0.0001, +0.0038] | PENDING — k<40 |
| `platoon_severity` | 481 | 120 | game | 49 | 0.0243 | +0.0006 | [0.0045, +0.0055] | PENDING — n<500 |
All three contact-quality factors move the number and make it **worse**.
### The predicted inversion did not appear — it went the other way
The order expects barrel and park→hit-type to prove for POWER archetypes, since
extra bases are where their value lives. Measured:
| | BOMBER | GHOST |
|---|---|---|
| `barrel_rate` | **+0.0114** (most harmful cell in the table) | +0.0012 |
| `exit_velo` | +0.0067 | +0.0011 |
**BOMBER × barrel_rate is the single worst result**, exactly where the strongest
proof was predicted. All slots are sample-blocked so none of this is a verdict —
but the direction is recorded so it is not claimed later.
### Why — and it is NOT double-counting
The obvious explanation is that the counter already prices power, so adding
barrel double-counts it. **Tested and refuted:** corr(barrel_pct, p_win) =
**0.061**. The counter is not pricing barrel at all.
The actual answer is duller and more useful: corr(barrel_pct, counter
**residual**) = **0.012**. Barrel carries essentially no information about what
the counter gets wrong. It is a real skill that does not help at this line.
This also closes the S81 lead: `hard_hit_pct` marginal r = 0.153 at n=295, which
drifted to 0.135 at n=383 (S82) and is now THEATER at n=707. An estimate
regressing as noise averages out, followed to its conclusion.
---
## The real finding: TB is miscalibrated, not under-factored
```
mean p_win 0.5698 actual hit rate 0.5074 counter bias +0.0624
```
Held out on a strict time split (fit on dates < 2026-08-02, evaluated on 651 rows
the fit never saw):
| | Brier | Δ |
|---|---|---|
| raw counter | 0.25007 | — |
| constant de-bias | 0.24740 | 0.00267 |
| **isotonic** | **0.24621** | **0.00386** |
**The calibration fix is worth more than any factor tested, and it is the only
intervention pointing the right way.** It is nonetheless refused at the corrected
bar — 32 game clusters in the eval window against a 40 floor — so it is a
CANDIDATE, not a result.
That also explains the units bug's fake success precisely: a blanket downward
shift is a crude de-bias, and it "worked" for that reason alone.
## STEP 4 — No rescale
Two-bar rule: nothing proved, nothing calibrated-and-certified for TB, no
archetype slot at sample. Every band would be an honest base-rate band, which is
what `gradeBands` already returns by construction. Running it would add nothing.
## Next, in order of value
1. **Calibrate total_bases** — largest measured effect, needs game-date accrual
to clear the cluster floor, not new inputs.
2. **Stop adding contact-quality factors to TB.** Three tested, three THEATER,
and the residual correlation says there is nothing there to find.
3. Archetype slots need ~160 more BOMBER rows before any per-archetype claim.
Counter and frozen clusters byte-identical.
+157
View File
@@ -0,0 +1,157 @@
# Were we out of data, or not using what we had?
**Both — and which one it is depends entirely on what unit a factor varies over.**
That distinction turned out to matter more than the sample counts themselves.
---
## 1. Player-level factors: we were under-querying
The platoon test reported n=452 and "48 short of the gate." That number described
how much of the JOIN survived, not how much data exists.
| | used | actually available |
|---|---|---|
| clean settled `hits` rows | 452 | **1,266** |
| clean settled `total_bases` rows | 383 | **928** |
| quarantined `hits` rows | — | **0** |
| hitters with platoon splits | 298 | 380 needed |
`platoon_splits` had been ingested from *tonight's lineups only* — 315 players —
so any hitter who settled a prop but was not in a lineup on an ingest day was
silently absent from every test. Backfilling all 380 (`scripts/backfill-context.js`)
took one pass and no waiting: **81 hitters fetched, 0 unresolved, coverage now
380/380.**
Re-run on the full clean history (`scripts/prove-hit-factors.js`, rows 452 → **1,059**):
| factor | n | mean shift | Brier Δ | CI (corrected, 55 tests) | verdict |
|---|---|---|---|---|---|
| `defense_by_direction` | 782 | 0.0127 | 0.0031 | [0.0050, 0.0013] | **PROVES** |
| `platoon` | 1,056 | 0.0268 | 0.0033 | [0.0062, 0.0001] | PROVES\* |
| `platoon_severity` | 700 | 0.0218 | 0.0038 | [0.0076, 0.0001] | PROVES\* |
| `defense` | 912 | 0.0289 | 0.0038 | [0.0079, +0.0002] | NOT_PROVEN |
| `pitcher_contact_profile` | 1,059 | 0.0259 | 0.0034 | [0.0078, +0.0006] | **NOT_PROVEN — demoted** |
| `park_hits` | 619 | 0.0190 | 0.0037 | [0.0076, +0.0005] | NOT_PROVEN |
### 1a. The demotion is the real headline
`pitcher_contact_profile` was the strongest proven factor in the programme
(Brier 0.0064, CI [0.0113, 0.0014]). On more than double the sample its point
estimate **roughly halved to 0.0034** and the corrected interval now spans zero.
Two things moved at once and honesty requires naming both: the cumulative
Bonferroni denominator also rose to 55, which widens every interval. But the
denominator cannot touch a *point estimate*, and that halved on its own. This is
the standing second line doing exactly what it exists for — more data demoting a
favourite rather than confirming it.
### 1b. \*The two platoon passes are NOT promoted
Both clear the bar with an upper bound of **0.0001**. That is as marginal as a
pass can be, and they ride a reconstructed input:
`platoon_splits` are **season-to-date**, so applying today's split to a game from
2026-07-15 means the split contains that game. Measured, not assumed:
- median contamination **4.5%** of the split's plate appearances
- p90 **12.4%**
- worst **137%** (call-ups whose scored games outnumber their split sample)
I had originally estimated ~1%. It is four and a half times that, and it runs in
the flattering direction on a result whose margin is one ten-thousandth. These
stay **CANDIDATE — pending point-in-time splits**. Promoting a 4.5%-contaminated
input on a 0.0001 bound would be exactly the kind of pass this programme keeps
having to retract.
---
## 2. Game-level factors: genuinely short, and no backfill fixes it
`game_context` held **zero** weather readings, ever. The fetcher was correct and
already pointed at Open-Meteo's **archive** endpoint. The failure was that the
two tables had never been introduced:
```
ledger_entries.game_id = mlb:2026-08-03:WashingtonNationals@PhiladelphiaPhillies
game_context.game_id = mlb:823437
```
Every lookup missed, and NULL weather columns read exactly like "the weather was
unavailable." **Same class as the doubled `/leaderboard` path: graceful
degradation wearing the mask of honest absence.** That is now three occurrences;
it is the failure mode this codebase produces most reliably.
Fixed in `scripts/reconstruct-game-environment.js` — resolves each ledger slug to
its real statsapi game and venue, writes `game_context` keyed by the *ledger's*
key, then pulls actual archived weather:
- 101 settled games → **96 matched**, 30 venues
- **96/96 venue-days returned real archived weather** (`open_meteo_archive`)
- park dimensions backfilled 15 → **30 venues**, zero dimension changes observed
### 2a. Park dimensions: what I verified and what I did not
statsapi serves only **current** venue geometry — it has no historical record. My
capture window is 2026-08-04 to 08-05, so "no mid-season change" is verified
across *two days*, which is nearly no verification at all. Applying current
dimensions to July games is the order's stated allowance and it is almost
certainly fine, but I did not verify it and will not claim to have.
### 2b. Why 928 rows are 47 readings
Park and weather assign **one value per game**. The 928 clean settled
`total_bases` rows sit on **47 distinct games — median 17.6 rows per game.**
Eighteen hitters in one ballpark on one night are one reading of that ballpark,
not eighteen.
Resampling rows would treat them as independent and return an interval far
tighter than the evidence supports. `factorGate.improvement` now resamples
**clusters** when rows carry one, and `adjudicate` judges sample against
`effective_n`. Rows without a cluster keep the original path byte-for-byte.
`scripts/prove-park-weather.js`:
```
rows_built 828 | distinct_games 47 | rows_per_game 17.6
brier_delta +0.0011 (WORSE, not merely unproven)
effective_n 47
VERDICT: CANDIDATE_PENDING_SAMPLE — 47 independent clusters < 500
```
---
## 3. The verdict: tested-now vs real-wait
| factor | unit it varies over | units held | ceiling | real wait |
|---|---|---|---|---|
| platoon, defence-by-direction | **hitter-game** | 1,059 | none | **none — answered now** |
| weather | **game** | 47 | none | **~57 days** at 7 games/settled-day |
| park dimensions | **venue** | 30 | **30, permanently** | **never** |
The last row is arithmetic, not pessimism. **There are 30 ballparks in MLB.** A
factor constant per venue can never accumulate 500 independent units no matter
how long the ledger runs. A park-geometry effect is only ever validatable as a
fixed effect with many games per park under a hierarchical model — never under a
bar expressed in independent units. The n≥500 bar was designed for player-level
factors and quietly does not transfer.
**So: we were under-querying at the player level, and genuinely short at the game
level — and for park geometry specifically, "wait for more data" was never going
to be the answer.**
---
## 4. Wind is refused
`parkWeather` reads temperature, elevation and geometry. It does **not** read
wind, and says so on every read (`wind_readable: false`).
We have the wind — Open-Meteo returns speed and bearing for all 96 games. What we
lack is **park orientation**: which compass direction each stadium's centre field
faces. A 15 mph wind from 220° is blowing out to right at one park and straight in
at another, and those are opposite predictions.
The tempting move is to use wind *speed* alone as a magnitude of disruption. That
asserts an effect while discarding the sign that determines what the effect is.
Wind stays unreadable until orientation is a real column.
+84
View File
@@ -0,0 +1,84 @@
# Wave D1 — and why E1/E9 were not it
## PHASE 0 — the audit's actual wave order
The order proposed E1 + E9 as Wave-1 primitives and instructed me to follow the
audit if it disagreed. **It disagrees.**
| item | the audit's placement |
|---|---|
| **E9 calibration curve** | **WAVE D3 — gated on MODEL work**: *"resolve n≥20 vs N30, accrue buckets"* |
| **E1 movement strip** | **WAVE D6 — large surface builds**, alongside crown / disagreement axis / SPLIT |
**Neither is Wave 1.** And E9's gate is live right now: calibration is
**withdrawn** (`CALIBRATION_DEPLOYED = []`) at **0 eligible dates**, so a
calibration curve today could only render its empty state. Building it now would
produce a component whose entire purpose is unavailable.
**WAVE D1, per the audit:** glyph library · card token + `vy-glitch` naming ·
`nudge()`/boot-stagger/row-hover/IntersectionObserver · boundary-channel blue ·
team-gradient chips + READ-FAB.
## PHASE 1 — three of the five Wave D1 items were already done
Checked rather than assumed, and the audit (2026-07-31) has aged:
| item | audit said | measured now |
|---|---|---|
| **D1 glyph library** | PARTIAL — 38 of 83 wired (46%) | **COMPLETE for everything wireable.** 83 designed glyphs, 41 registry archetypes, **39 map — and all 39 are wired, with colours matching the registry exactly (0 disagreements).** |
| **A1 card token `#0E0E14`** | BUILT-BUT-DRIFTED — "in only 1 file" | **BUILT-TO-SPEC.** It is the `--bg-1` token in `globals.css`, consumed by **32 files** via `var(--bg-1)`. The audit counted literal hex occurrences, which is what a *correctly tokenised* value looks like. |
| **B1 boundary channel** | PARTIAL — hex in 2 files | **BUILT.** `--priced-out: #8fb2de` + `--priced-out-dim` are tokens with a documented colour law in `globals.css`, consumed by 4 files. |
**The 44 unwired glyphs are not a wiring gap.** They have no backend archetype —
wiring them would mean **inventing 44 archetypes to consume artwork**, which is
the fabrication this programme refuses. That is the 41-vs-74 scope question, and
it is Kev's call, not a build task. Separately, **2 registry archetypes have no
designed glyph** (`DUAL THREAT`, `PAINT BOSS`) — a design gap, not a code one.
## PHASE 2 — what was genuinely absent, and is now built
`web/src/lib/motion.js` — the four primitives, all previously 0 files:
| primitive | the rule it enforces |
|---|---|
| **E17 `nudge()`** | ≤180ms. Confirms an action *registered* without claiming it finished — a long pulse reads as latency, the opposite of a perceived-speed primitive. |
| **E18 `bootStagger()`** | **capped at 240ms.** Uncapped, row 40 waits 1.1s and the stagger *becomes* the latency it exists to disguise. |
| **E27 `rowHover()`** | returns handlers, not CSS — so a tap on touch cannot stick a hover state that never clears. |
| **E28 `revealOnIntersect()`** | returns an unobserve fn in every path, and **reveals IMMEDIATELY** without IntersectionObserver or under reduced motion. Content is never hidden behind a capability check. |
**Reduced motion is honoured, not softened** — every primitive skips outright,
and server-side assumes reduced. 10 tests, the sharpest being that
`bootStagger` under reduced motion returns `opacity: 1` and not merely
`delay: 0`: if the CSS animation supplies the opacity, skipping the animation
without forcing it leaves the row invisible forever.
## PHASE 3 — guard registration and Wave-2 readiness
The reachability guard now carries a **PRIMITIVES** section: a module built to be
embedded must declare its exports *and name its intended consumers*. A primitive
imported by nothing is the same built-but-unread class as an unmounted
component. 24 checks green.
### Wave-2 readiness
**Ungated — design complete, no blocker:**
- **F9/F10/F11 offseason hub** (Wave D6) — the largest, fully spec'd
- **F5 article media** (Wave D6) — hero + figure system fully spec'd
- **E10/E12 The Report email + `/report` archive** — design exists
- **E1 movement strip** (Wave D6) — spec'd, absent, now has its motion primitives
**Gated, and by what:**
| item | gate |
|---|---|
| E9 calibration curve | **model** — calibration withdrawn, 0 eligible dates |
| E15 Price Triplet MODEL leg | **model** — EV layer |
| E16 share cards + F8 crops | **resolution tail** *and* the card-system reconciliation |
| E2 BookChip tiles / E6 push-to-book | **licensing / affiliate approval** |
| E13 Scanner S6 | another build order |
**The card-system reconciliation** (content-engine 1080×1350 vs E16's five
masters) gates E16/F8 and should be settled with the social chat, since off-site
posts and on-site share cards are one system.
Isolated throughout: read-only, no serving or model change, **accrual clock
unchanged at 0 eligible dates**.
+128
View File
@@ -0,0 +1,128 @@
# The window-bug class, hunted — three more paths, and the forward re-audit rule
## PHASE 0 — the audit
The defect class: **a fixed short window used AS the season/base rate.**
`getStatRows` is the single path feeding `meta.gameLogs`, which is where
`estimateProbability` derives its base rate, so every branch of it is a base-rate
path. The feature builders are the second surface, because `l20_avg` is the
season reference `projectionFor` reads.
| path | window | classification |
|---|---|---|
| `getStatRows` MLB → estimator base rate | `fullLog` | **CORRECT** (fixed 929fd81) |
| `mlbGameLogFeatures``l5/l10/l20_avg`, `l10_stddev` | `last10` = **10** | **DEFECTIVE** |
| `espnStatsAdapter.parseGameLog` → NBA/WNBA logs | `rows.slice(0, 20)` = **20** | **DEFECTIVE** |
| `getStatRows` NBA/WNBA ESPN branch | inherits the 20-cap | **DEFECTIVE (via source)** |
| `getStatRows` NBA/WNBA Python branch | `getGameLogs(..., 20)` | DEFECTIVE-but-dormant (service offline in prod) |
| **pitcher engine** (`pitcherEngine`, `skillProjection`) | reads statcast **profiles**, no game log | **N/A** |
| pitcher props (strikeouts) via `getStatRows` MLB | `fullLog` | **CORRECT** — fixed by the same change |
| `settleSource` | already reads the full log (S64) | CORRECT |
| `playerIntelService`, `streaksService` | display/streak surfaces, not forecasts | N/A |
**The pitcher answer matters and is good news:** pitcher props run through the
same `getStatRows` MLB branch, so `929fd81` repaired them too — there is no
separate defective pitcher base-rate path.
### The one that was hiding in plain sight
`mlbGameLogFeatures` carries this comment:
> `l20 = all available (the season per-game reference projectionFor needs)`
Built from `last10`, **`l20_avg` was a ten-game average wearing a season label** —
and it feeds both the consistency (cv) pull inside the estimator and
`projectionFor`, which decides refusals. Same class as the base-rate bug, same
file, and it survived the previous repair because that fix touched only
`getStatRows`.
---
## PHASE 1 — fixes
| path | fix | API cost |
|---|---|---|
| `mlbGameLogFeatures` | read `fullLog`, fall back to `last10` | **ZERO** — same response |
| `espnStatsAdapter.parseGameLog` | drop the `slice(0, 20)` cap | **ZERO** — same payload, already parsed |
| NBA/WNBA Python branch | left as-is | service offline in prod; fixing it would be speculative |
**No new API calls anywhere.** Both fixes widen data that was already fetched and
then discarded — the same shape as the original repair.
### Measurement status, stated honestly
These are serving changes for the MLB feature path and the NBA/WNBA log path.
**Their before/after resolution is NOT measured here**, and deliberately: the
only way to measure it today would be to reconstruct the repaired forecast over
old rows, which is the reconstruction-vs-served trap this order explicitly
refuses. They ship as code fixes with the measurement deferred to accrual, which
is the honest sequencing.
---
## PHASE 2 — the forward re-audit rule, in code
`MODEL_VERSION` is bumped to **`engine1@2026-08-07-fullwindow`**, so every
snapshot from this commit forward is self-identifying. `retentionService` already
stamps it onto `model_snapshots`, so no new plumbing was needed.
`model/reAuditEligibility.js` encodes the rule:
- **`isEligible(row)`** — true only for rows carrying the repaired marker.
- **`assess(rows)`** — counts eligible **DATES**, not rows, because dates have
been the binding scarcity in every interval this session.
- **`ACCRUAL`** (frozen) — pre-stated minimum dates per measurement:
| measurement | minimum eligible dates |
|---|---|
| calibration re-fit | 10 |
| hits factor lift | 10 |
| prior verdict re-audit | 14 |
| rbi lineup-slot gate | 14 |
A test locks the case that would otherwise be invisible: **a MIXED table** of 330
rows where only 30 carry the new marker returns `eligible_dates: 3`, not 330
rows' worth of false confidence. Once both generations sit in the same table, a
naive count would happily fit a map on a blend of two different forecasters.
---
## PHASE 3 — the honest board
**What happened:** the champion computed its season rate over ten games. Found by
resolution decomposition, not by a test failing. Fixed in two lines. It no longer
*loses* to a frequency table — it **beats** it CI-confirmed only on total_bases,
**ties** on rbi and runs, and leads on the hits point estimate.
**Consequences, each labelled:**
- **CALIBRATION — WITHDRAWN.** `CALIBRATION_DEPLOYED` is empty. Maps were fitted
on the retired forecast. Re-fits on repaired-champion settled rows. *Waiting on
accrual: 10 dates.* Not to be refit on reconstructions.
- **FACTOR VERDICTS — SUSPECT.** Every prior null and every THEATER was measured
against a champion worse than a frequency table; signal added to noise reads as
noise. Re-audit on accrued rows. **Direction UNKNOWN** — some may pass, some
may still fail. Not pre-priced. *Waiting: 14 dates.*
- **HITS FACTOR LIFT (1.39%) — UN-REMEASURABLE.** Needs rows produced *by* the
repaired champion. *Waiting: 10 dates.* The factors remain wired and
transmitting (43f65d3); only the lift number is unquantified.
- **RBI LINEUP-SLOT — RE-QUEUED.** Lands after the champion is sound and rows
accrue. *Waiting: 14 dates.*
**Pre-registered re-audit order** (each runs only when its bar is met):
1. Re-fit calibration on repaired-champion rows (10 dates)
2. Re-measure hits factor lift (10 dates)
3. Re-audit prior factor verdicts (14 dates)
4. Run rbi lineup-slot through the two-part gate (14 dates)
**Then STOP and accrue.** Nothing further can be honestly measured until the
board fills with rows the repaired champion produced.
---
## Invariants
No measurement on reconstructions — hard refusal, and it is why Phase 1 ships
code without numbers. Serving-path changes by design for the MLB feature path and
NBA/WNBA logs; frozen model modules verified unchanged. `p_win` never mutated. No
Bonferroni slot — base-rate repair and a bug hunt, not causal factors.
+7
View File
@@ -204,6 +204,9 @@ app.use('/api/content', contentRoutes);
// Session S7 (a1) — THE VYNDR REPORT: public double-opt-in subscribe
// (forwards to the self-hosted Listmonk; graceful no-op without env).
app.use('/api/newsletter', require('./routes/newsletter'));
// E12 — The Report archive. Public, read-only; every issue carries its own
// day record, because the archive is a ledger too.
app.use('/api/report', require('./routes/report'));
// A1 S9 — Slip Reader: OCR a bet-slip screenshot into legs (auth +
// per-tier daily quota inside the router). Values are user-slip values.
app.use('/api/slips', require('./routes/slips'));
@@ -212,6 +215,10 @@ app.use('/api/slips', require('./routes/slips'));
// the public surface; the Next.js admin route proxies through with
// the key kept server-side.
app.use('/api/internal', internalRoutes);
// Content STUDIO — finished posts (copy + card + fact-contract) for review and,
// later, for an autonomous poster. Distinct from /api/content (Session 29),
// which serves structured content objects by data level.
app.use('/api/content-studio', require('./routes/contentStudio'));
// A1 S3 — partner attribution report. Internal-key gated (router-level
// requireInternalAuth); no Next proxy on purpose — never browser-facing.
app.use('/api/partners', require('./routes/partners'));
+175
View File
@@ -0,0 +1,175 @@
'use strict';
/**
* /api/content-studio — the generated-post review surface.
*
* NOT to be confused with `/api/content` (Session 29), which serves structured
* content OBJECTS by data level. This serves finished POSTS from the content
* engine: copy, branded card, and the fact-contract each was built from.
*
* ── API-FIRST, ON PURPOSE ────────────────────────────────────────────────
* This is the contract Kev's preview page consumes today and an autonomous
* poster consumes later. The page is a thin client; no posting logic lives in
* it. Pointing a bot here needs no change on this side, which is the entire
* reason it is an API before it is a screen.
*
* ── READ-ONLY WHERE IT COUNTS ────────────────────────────────────────────
* It serves generated content. Beyond the read-only pulls the engine already
* makes, it touches no serving, model or ledger table — zero effect on the
* repaired-champion accrual clock. Approval status lives in Redis, because an
* editorial decision is not a model fact and must never sit beside ledger rows.
*/
const express = require('express');
const { requireInternalAuth } = require('../middleware/internalAuth');
const engine = require('../services/content/contentEngine');
const { toSvg } = require('../services/content/cardRenderer');
const router = express.Router();
router.use(requireInternalAuth({ loopbackOnly: false })); // private: Kev's desk, then an agent's
const STATUS_KEY = (date) => `contentstudio:status:${date}`;
const VALID_STATUS = new Set(['pending', 'approved', 'skipped', 'regenerate_requested']);
const dateET = () => new Intl.DateTimeFormat('en-CA', {
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
}).format(new Date());
function registerTemplates() {
for (const t of ['hotHitters', 'honestyFlex', 'streakList']) {
try { engine.registerTemplate(require(`../services/content/templates/${t}`)); } catch { /* idempotent */ }
}
}
/**
* GET /api/content-studio/:date?
*
* → { date, count, posts: [{ id, label, sport, status, ok, skipped, reason,
* honest_absence, copy, card, card_svg, fact_contract, facts }] }
*
* `fact_contract` is what the post was REQUIRED to have; `facts` is what
* actually backed it. A reviewer — or an agent — can check the claim rather
* than trust the sentence.
*/
// EXPRESS 5 DROPPED THE `?` OPTIONAL-PARAM SYNTAX -- `'/:date?'` throws at mount
// time and takes down every suite that imports app.js. Two explicit routes.
async function handleGet(req, res) {
try {
registerTemplates();
const date = req.params.date || dateET();
const deps = req.app.get('contentStudioDeps') || (await buildDeps(date));
const results = await engine.generateAll({ ...deps, date });
let statuses = {};
try { statuses = (await require('../utils/redis').cacheGet(STATUS_KEY(date))) || {}; } catch { statuses = {}; }
const posts = results.map((r) => {
const t = engine.getTemplate(r.id) || {};
return {
id: r.id,
label: t.label || r.id,
sport: t.sport || null,
status: statuses[r.id] || 'pending',
ok: r.ok === true,
skipped: r.skipped === true,
reason: r.reason || null,
honest_absence: r.honest_absence === true,
copy: r.copy || null,
card: r.card || null,
card_svg: r.card ? toSvg(r.card) : null,
fact_contract: t.requires || [],
facts: r.facts || null,
};
});
res.json({ date, count: posts.length, posts });
} catch (e) {
res.status(500).json({ error: e.message });
}
}
router.get('/', handleGet);
router.get('/:date', handleGet);
/** POST /api/content-studio/:date/:id/status { status } — editorial state only. */
router.post('/:date/:id/status', express.json(), async (req, res) => {
const { date, id } = req.params;
const status = String((req.body || {}).status || '');
if (!VALID_STATUS.has(status)) {
return res.status(400).json({ error: `status must be one of ${[...VALID_STATUS].join(', ')}` });
}
try {
const { cacheGet, cacheSet } = require('../utils/redis');
const cur = (await cacheGet(STATUS_KEY(date))) || {};
cur[id] = status;
await cacheSet(STATUS_KEY(date), cur, 60 * 60 * 24 * 14);
return res.json({ ok: true, date, id, status });
} catch (e) {
return res.status(500).json({ error: e.message });
}
});
/** Real sources. All READ-ONLY. */
async function buildDeps(date) {
const sg = require('../services/model/servedGrade');
const { knownNumber } = require('../utils/known');
const sb = require('../utils/supabase').getSupabaseServiceClient();
const mlb = require('../services/adapters/mlbStatsAdapter');
const gradeDistribution = async () => {
if (!sb) return { total: null };
const { data } = await sb.from('model_snapshots')
.select('p_win, refused').eq('sport', 'mlb').eq('game_date', date).limit(5000);
const usable = (data || []).filter((r) => !r.refused && knownNumber(r.p_win) !== null);
const by = {}; let flat = 0;
for (const r of usable) {
const g = sg.gradeFor({ p_win: knownNumber(r.p_win) });
by[g.letter] = (by[g.letter] || 0) + 1;
if (g.separates_from_base_rate === false) flat += 1;
}
return { total: usable.length || null, by_letter: by, not_separable: flat };
};
const settledStreaks = async () => {
if (!sb) return [];
const { data } = await sb.from('ledger_entries')
.select('player_name, player_key, game_date, outcome')
.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits')
.in('outcome', ['hit', 'miss']).limit(5000);
const by = new Map();
for (const r of data || []) {
if (!by.has(r.player_key)) by.set(r.player_key, []);
by.get(r.player_key).push(r);
}
const out = [];
for (const [, rows] of by) {
rows.sort((a, b) => String(b.game_date).localeCompare(String(a.game_date)));
let n = 0;
for (const r of rows) { if (r.outcome === 'hit') n += 1; else break; }
if (n >= 3) out.push({ name: rows[0].player_name, streak: n, verified_from_settled: true });
}
return out;
};
const hitterForm = async () => {
if (!sb) return [];
const { data } = await sb.from('model_snapshots')
.select('player_name').eq('sport', 'mlb').eq('game_date', date).limit(400);
const names = [...new Set((data || []).map((r) => r.player_name).filter(Boolean))].slice(0, 40);
const out = [];
for (const n of names) {
try {
const r = await mlb.getPlayerStats(n);
const log = (r && r.found && Array.isArray(r.fullLog)) ? r.fullLog : [];
const vals = log.map((g) => knownNumber(g && g.stat && g.stat.hits)).filter((v) => v !== null);
if (vals.length < 20) continue;
const rate = (a) => a.filter((v) => v > 0).length / a.length;
out.push({ name: n, season_games: vals.length, season_rate: rate(vals), recent_rate: rate(vals.slice(-10)) });
} catch { /* absent player -> absent row */ }
}
return out;
};
return { servedGrade: sg, gradeDistribution, settledStreaks, hitterForm };
}
module.exports = router;
module.exports.__internals = { buildDeps, VALID_STATUS, STATUS_KEY };
+55
View File
@@ -0,0 +1,55 @@
'use strict';
/**
* /api/report — E12, the issue archive.
*
* The spec's law is one line: **"EVERY ISSUE SHOWS ITS OWN DAY RECORD — THE
* ARCHIVE IS A LEDGER TOO."** So an archive row is not a headline with a date;
* it carries the record that issue's reads actually produced. An archive that
* showed only titles would be a blog, and the point of this one is that it
* cannot quietly bury a bad day.
*
* Public and READ-ONLY. Issues live in Redis under `report:issue:{date}`,
* written by the send path; nothing here touches a serving, model or ledger
* table.
*/
const express = require('express');
const router = express.Router();
const INDEX_KEY = 'report:index';
const ISSUE_KEY = (d) => `report:issue:${d}`;
/** GET /api/report — the archive list, newest first. */
router.get('/', async (req, res) => {
try {
const { cacheGet } = require('../utils/redis');
const index = (await cacheGet(INDEX_KEY)) || [];
const issues = Array.isArray(index) ? index : [];
const limit = Math.min(200, Math.max(1, Number(req.query.limit) || 60));
res.set('Cache-Control', 'public, max-age=300');
res.json({
count: issues.length,
// Honest empty state is the caller's to render; we simply report zero.
issues: issues.slice(0, limit),
});
} catch (e) {
res.status(500).json({ error: e.message });
}
});
/** GET /api/report/:date — one issue. 404 rather than an invented shell. */
router.get('/:date', async (req, res) => {
try {
const { cacheGet } = require('../utils/redis');
const issue = await cacheGet(ISSUE_KEY(req.params.date));
if (!issue) return res.status(404).json({ error: 'no issue for that date' });
res.set('Cache-Control', 'public, max-age=300');
return res.json(issue);
} catch (e) {
return res.status(500).json({ error: e.message });
}
});
module.exports = router;
module.exports.__internals = { INDEX_KEY, ISSUE_KEY };
+4 -1
View File
@@ -227,7 +227,10 @@ function parseGameLog(payload) {
if (Number.isNaN(tb)) return -1;
return tb - ta;
});
return rows.slice(0, 20);
// The ESPN payload carries the full season's events; capping at 20 made every
// downstream "season rate" a 20-game rate. Same class as the MLB last10 bug
// and free to widen -- this is the same response, already parsed.
return rows;
}
async function fetchJsonG(url, opts = {}) {
+83
View File
@@ -0,0 +1,83 @@
'use strict';
/**
* cardRenderer — one card engine, per-template layouts.
*
* SVG rather than canvas: it is text, so it has no native dependency, it
* diffs in review, and the numbers inside it are greppable — which matters when
* the whole claim is that the numbers are real. A card whose contents cannot be
* inspected without opening an image is a bad fit for a Truth-Law product.
*
* The card never formats its own facts. Every string arrives already rendered
* and already gate-checked by contentEngine, so a caption and a card physically
* cannot disagree.
*/
const BRAND = Object.freeze({
bg: '#05070A',
panel: '#0A0E14',
line: '#1A222E',
green: '#00D4A0', // the R
white: '#FFFFFF', // VYND
dim: '#6B7A8D',
amber: '#FFB347',
mono: "ui-monospace, 'JetBrains Mono', 'SFMono-Regular', Menlo, monospace",
});
const W = 1080;
const H = 1350;
const esc = (s) => String(s)
.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;')
.replace(/"/g, '&quot;');
/** The slashed-Y wordmark: VYND white, R green. */
function wordmark(x, y, size = 34) {
return `
<g font-family="${BRAND.mono}" font-size="${size}" font-weight="800" letter-spacing="${size * 0.09}">
<text x="${x}" y="${y}" fill="${BRAND.white}">VYND</text>
<text x="${x + size * 2.92}" y="${y}" fill="${BRAND.green}">R</text>
<line x1="${x + size * 1.02}" y1="${y - size * 0.78}" x2="${x + size * 1.42}" y2="${y + size * 0.22}"
stroke="${BRAND.green}" stroke-width="${Math.max(2, size * 0.07)}" opacity=".85"/>
</g>`;
}
/**
* Render a card model to SVG.
* @param {object} card { kind, title, subtitle, lines[], footer }
*/
function toSvg(card = {}) {
const lines = card.lines || [];
let y = 300;
const body = lines.map((l) => {
const text = typeof l === 'string' ? l : l.text;
const style = (typeof l === 'object' && l.style) || 'body';
let out = '';
if (style === 'rule') {
out = `<line x1="72" y1="${y - 18}" x2="${W - 72}" y2="${y - 18}" stroke="${BRAND.line}" stroke-width="1"/>`;
y += 26;
return out;
}
const size = style === 'lead' ? 46 : style === 'stat' ? 40 : 30;
const fill = style === 'stat' ? BRAND.green : style === 'dim' ? BRAND.dim : BRAND.white;
const weight = style === 'body' ? 500 : 800;
out = `<text x="72" y="${y}" font-family="${BRAND.mono}" font-size="${size}" font-weight="${weight}" fill="${fill}">${esc(text)}</text>`;
y += size + (style === 'lead' ? 26 : 18);
return out;
}).join('\n');
return `<svg xmlns="http://www.w3.org/2000/svg" width="${W}" height="${H}" viewBox="0 0 ${W} ${H}">
<rect width="${W}" height="${H}" fill="${BRAND.bg}"/>
<rect x="40" y="40" width="${W - 80}" height="${H - 80}" fill="${BRAND.panel}" stroke="${BRAND.line}"/>
${Array.from({ length: 26 }, (_, i) => `<line x1="40" y1="${52 * i + 40}" x2="${W - 40}" y2="${52 * i + 40}" stroke="${BRAND.white}" stroke-width="1" opacity=".02"/>`).join('')}
${wordmark(72, 128)}
${card.title ? `<text x="72" y="212" font-family="${BRAND.mono}" font-size="54" font-weight="800" fill="${BRAND.white}" letter-spacing="1">${esc(card.title)}</text>` : ''}
${card.subtitle ? `<text x="72" y="256" font-family="${BRAND.mono}" font-size="26" font-weight="600" fill="${BRAND.dim}" letter-spacing="2">${esc(card.subtitle)}</text>` : ''}
<line x1="72" y1="272" x2="${W - 72}" y2="272" stroke="${BRAND.green}" stroke-width="2" opacity=".55"/>
${body}
<text x="72" y="${H - 96}" font-family="${BRAND.mono}" font-size="22" font-weight="600" fill="${BRAND.dim}">${esc(card.footer || 'Every number here is measured. Nothing is projected.')}</text>
<text x="72" y="${H - 62}" font-family="${BRAND.mono}" font-size="20" font-weight="500" fill="${BRAND.line}">vyndr · built by Kevon Butler · Detroit</text>
</svg>`;
}
module.exports = { toSvg, BRAND, W, H };
+172
View File
@@ -0,0 +1,172 @@
'use strict';
/**
* contentEngine — posts that structurally cannot lie.
*
* The top-of-funnel is a content flywheel, and the thing that makes it VYNDR's
* rather than anyone's is that the numbers in it are real. That is easy to
* promise and hard to keep, because the failure mode is not a person choosing to
* fabricate — it is a template with a hardcoded adjective, or a field that came
* back null and rendered as "0", or a caption drifting from the card beside it.
*
* So the honesty is enforced by construction rather than by care:
*
* 1. COPY IS TOKEN-SUBSTITUTED. Every factual claim in a template is a
* `{token}` resolved against pulled facts. A token with no backing fact
* REFUSES to render — it cannot fall back to a plausible default, because
* there is no code path that produces one.
* 2. THE FACT CONTRACT IS ASSERTED FIRST. A template declares the fields it
* needs; the engine checks them before any string is built. A missing field
* means SKIP or an honest-absence variant, never invention.
* 3. CARD AND COPY SHARE ONE FACT OBJECT. They cannot diverge, because there
* is only one set of numbers and both read it.
*
* NO LIVE MODEL WRITES FACTUAL CLAIMS. The voice lives in the template; the
* facts are pulled. A voice-polish port is reserved for later and is not wired
* here — an LLM that can rewrite a sentence can rewrite a number.
*
* READ-ONLY. This engine touches no serving or forecast table. It has zero
* effect on the repaired-champion accrual clock.
*/
const { knownNumber } = require('../../utils/known');
/** Resolve `a.b.c` against an object; undefined when any hop is missing. */
function pathValue(obj, path) {
return String(path).split('.').reduce((o, k) => (o == null ? undefined : o[k]), obj);
}
/**
* A value that may stand as a fact.
*
* `null` and `undefined` are absent. Empty string is absent — it renders as a
* hole in a sentence. Zero and false are PRESENT: "0 props cleared B+" is a real
* and important claim, and treating 0 as missing is the `Number(null) === 0`
* breach wearing its opposite coat.
*/
function isPresent(v) {
if (v === null || v === undefined) return false;
if (typeof v === 'string' && v.trim() === '') return false;
if (Array.isArray(v) && v.length === 0) return false;
if (typeof v === 'number' && !Number.isFinite(v)) return false;
return true;
}
const TOKEN = /\{([a-zA-Z0-9_.]+)\}/g;
/** Every `{token}` in a string. */
function tokensIn(text) {
const out = [];
let m;
const re = new RegExp(TOKEN);
while ((m = re.exec(String(text))) !== null) out.push(m[1]);
return out;
}
/**
* Substitute tokens, refusing on any that is unbacked.
* @throws when a token has no present fact — the whole point.
*/
function render(text, facts) {
const missing = tokensIn(text).filter((t) => !isPresent(pathValue(facts, t)));
if (missing.length) {
throw new Error(`TRUTH LAW: unbacked token(s) ${missing.join(', ')} — refusing to render`);
}
return String(text).replace(TOKEN, (_, t) => String(pathValue(facts, t)));
}
/** Which contract fields are absent from a pulled fact set. */
function contractGaps(template, facts) {
return (template.requires || []).filter((f) => !isPresent(pathValue(facts, f)));
}
const registry = new Map();
/**
* Register a template.
*
* @param {object} t
* id stable key
* sport which sport it speaks for
* requires fact-contract: dotted paths that MUST be present
* pull async (deps) => facts — the ONLY place data enters
* copy (facts) => string with {tokens}
* card (facts) => card model (layout + token strings)
* absent optional (gaps) => honest "nothing tonight" post
*/
function registerTemplate(t) {
if (!t || !t.id) throw new Error('a template needs an id');
for (const k of ['requires', 'pull', 'copy', 'card']) {
if (!t[k]) throw new Error(`template ${t.id} is missing ${k}`);
}
registry.set(t.id, t);
return t;
}
const listTemplates = () => [...registry.values()];
const getTemplate = (id) => registry.get(id) || null;
/**
* Generate one post.
*
* @returns {object} { ok, id, copy, card, facts, skipped, reason }
* Never throws on absent data — absence is an outcome, not an error.
*/
async function generate(id, deps = {}) {
const t = registry.get(id);
if (!t) return { ok: false, id, skipped: true, reason: 'no such template' };
let facts;
try {
facts = await t.pull(deps);
} catch (e) {
return { ok: false, id, skipped: true, reason: `pull failed: ${e.message}` };
}
// ── THE CONTRACT, BEFORE ANY STRING IS BUILT ──
const gaps = contractGaps(t, facts || {});
if (gaps.length) {
if (typeof t.absent === 'function') {
const alt = t.absent(gaps, facts || {}) || {};
// An absence variant may DECLINE to speak. A template that cannot tell
// "nothing happened" from "nothing was fetched" must not publish the
// first sentence when the second is true -- honest-absence copy is a
// perfect hiding place for a broken pull.
if (alt.skip) return { ok: false, id, skipped: true, reason: alt.skip, gaps };
return { ok: true, id, honest_absence: true, copy: alt.copy, card: alt.card, facts: facts || {}, gaps };
}
return { ok: false, id, skipped: true, reason: `fact-contract gap: ${gaps.join(', ')}`, gaps };
}
try {
const copyRaw = t.copy(facts);
const copy = render(copyRaw, facts);
const cardModel = t.card(facts);
// The card's own strings pass the same gate -- a caption and a card cannot
// disagree if both are rendered from one fact object under one rule.
const card = {
...cardModel,
lines: (cardModel.lines || []).map((l) => (typeof l === 'string'
? render(l, facts)
: { ...l, text: render(l.text, facts) })),
title: cardModel.title ? render(cardModel.title, facts) : null,
subtitle: cardModel.subtitle ? render(cardModel.subtitle, facts) : null,
};
return { ok: true, id, copy, card, facts };
} catch (e) {
// A refusal is a skip, never a degraded post.
return { ok: false, id, skipped: true, reason: e.message };
}
}
/** Generate every registered template; skips are reported, never hidden. */
async function generateAll(deps = {}) {
const out = [];
for (const t of listTemplates()) out.push(await generate(t.id, deps));
return out;
}
module.exports = {
registerTemplate, listTemplates, getTemplate, generate, generateAll,
render, tokensIn, contractGaps, isPresent, pathValue,
};
@@ -0,0 +1,73 @@
'use strict';
/**
* TEMPLATE 2 — THE HONESTY FLEX.
*
* The differentiator, and the one that would be worthless if it were padded.
* It publishes the REAL grade distribution: how many props were graded, how few
* cleared the ceiling, and the fact that A is unissuable because no band of this
* model has ever earned one.
*
* Every competitor's card is all A's. Ours says most of tonight is a base-rate
* read — and that claim is only impressive if it is exactly true, so the numbers
* come from servedGrade's own bands rather than from a marketing sentence.
*/
const { knownNumber } = require('../../../utils/known');
module.exports = {
id: 'honesty_flex',
sport: 'mlb',
label: 'The Honesty Flex',
requires: ['graded', 'ceiling_letter', 'ceiling_realized', 'base_rate', 'date'],
async pull(deps) {
const g = await deps.gradeDistribution(); // injected, read-only
const sg = deps.servedGrade || require('../../model/servedGrade');
const total = knownNumber(g && g.total);
const top = (g && g.by_letter && (g.by_letter['B+'] || 0)) || 0;
const flat = (g && g.not_separable) || 0;
const ceiling = sg.BANDS[0];
return {
date: deps.date,
graded: total,
top_count: top,
top_pct: total ? Math.round((top / total) * 100) : null,
flat_count: flat,
flat_pct: total ? Math.round((flat / total) * 100) : null,
ceiling_letter: ceiling.letter,
ceiling_realized: Math.round(ceiling.realized * 100),
base_rate: Math.round(sg.BASE_RATE * 100),
unissuable: sg.UNISSUABLE.join(', '),
};
},
copy: () => `WE GRADED {graded} PROPS TONIGHT. {top_count} CLEARED {ceiling_letter}.
That's {top_pct}%. The other {flat_pct}% we can't separate from the baseline, and we say so on the card instead of calling them leans.
Our ceiling is {ceiling_letter} — those reads land about {ceiling_realized}% against a {base_rate}% baseline. We do not issue {unissuable}. No band of this model has ever hit at a rate that would justify one.
Everybody else's card is all A's. Ask them what their A actually hits.`,
card: () => ({
kind: 'flex',
title: 'TONIGHT, HONESTLY',
subtitle: '{date}',
lines: [
{ text: '{graded} graded', style: 'lead' },
{ text: '{top_count} cleared {ceiling_letter} — {top_pct}%', style: 'stat' },
{ text: '{flat_pct}% we cannot separate from baseline', style: 'dim' },
{ text: '', style: 'rule' },
{ text: 'Ceiling: {ceiling_letter} · lands {ceiling_realized}%', style: 'body' },
{ text: 'Baseline: {base_rate}%', style: 'body' },
{ text: 'We do not issue {unissuable}.', style: 'body' },
],
footer: 'Grades are earned from realized outcomes, not issued on confidence.',
}),
absent: () => ({
copy: 'No slate graded tonight. Nothing to show, so nothing shown.',
card: { kind: 'absence', title: 'NO SLATE TONIGHT', subtitle: 'NOTHING TO PAD', lines: [{ text: 'No props graded. An empty board is an honest board.', style: 'body' }], footer: 'We post the count even when the count is zero.' },
}),
};
@@ -0,0 +1,94 @@
'use strict';
/**
* TEMPLATE 1 — HOT HITTERS.
*
* Reads the REPAIRED full-season log, not a ten-game slice. "Hot" here means a
* recent rate measured against that hitter's own season rate — which is only a
* meaningful comparison now that the season rate is a season.
*/
const { knownNumber } = require('../../../utils/known');
/** A hitter must have this much history before we call him anything. */
const MIN_SEASON_GAMES = 20;
const RECENT = 10;
module.exports = {
id: 'hot_hitters',
sport: 'mlb',
label: 'Hot Hitters',
requires: ['count', 'hitters', 'window', 'date'],
async pull(deps) {
const rows = await deps.hitterForm(); // injected, read-only
// ── EMPTY POOL IS NOT AN HONEST ABSENCE ────────────────────────────────
// The first run of this template emitted "no hitter is meaningfully hot"
// while the real cause was a source holding fewer than 20 games for EVERY
// player. That reads as a considered editorial judgement and is actually a
// broken pull -- the exact failure class that has bitten this codebase four
// times tonight, here wearing the costume of our own honesty copy.
//
// So the two are separated: no candidates at all is a SKIP with a reason;
// candidates present but none hot is the honest absence.
const candidates = (rows || []).length;
const usable = (rows || []).filter((r) =>
knownNumber(r.season_games) !== null && r.season_games >= MIN_SEASON_GAMES
&& knownNumber(r.recent_rate) !== null && knownNumber(r.season_rate) !== null);
const hot = usable
.map((r) => ({ ...r, lift: r.recent_rate - r.season_rate }))
.filter((r) => r.lift > 0)
.sort((a, b) => b.lift - a.lift)
.slice(0, 5);
return {
date: deps.date,
window: RECENT,
candidates,
qualified: usable.length,
count: hot.length || null, // zero hot hitters is an ABSENT list, not "0 hot hitters"
hitters: hot.length ? hot : null,
list: hot.map((h, i) =>
`${i + 1}. ${h.name}${Math.round(h.recent_rate * 100)}% last ${RECENT}, ${Math.round(h.season_rate * 100)}% season`).join('\n'),
top_name: hot[0] ? hot[0].name : null,
top_recent: hot[0] ? Math.round(hot[0].recent_rate * 100) : null,
top_season: hot[0] ? Math.round(hot[0].season_rate * 100) : null,
};
},
copy: () => `WHO'S ACTUALLY HOT — {date}
{top_name} is hitting {top_recent}% over his last {window}. His season number is {top_season}%.
That gap is the whole point. Everybody else is guessing at it.
{list}
Measured off full season logs, not a ten-game window that flatters whoever ran hot last week.`,
card: () => ({
kind: 'list',
title: "WHO'S ACTUALLY HOT",
subtitle: 'LAST {window} vs SEASON · {date}',
lines: [
{ text: '{top_name}', style: 'lead' },
{ text: '{top_recent}% last {window} · {top_season}% season', style: 'stat' },
{ text: '', style: 'rule' },
{ text: '{list}', style: 'body' },
],
footer: 'Rates measured from full season game logs.',
}),
absent: (gaps, facts) => {
// Only speak if there was actually a pool to judge.
if (!facts || !facts.qualified) {
return {
copy: null,
card: null,
skip: `no qualified hitters in the pool (candidates=${(facts && facts.candidates) || 0}, qualified=${(facts && facts.qualified) || 0}) — source problem, not a quiet night`,
};
}
return {
copy: "No hitter is meaningfully hot tonight.\n\nWe could dress up a middling week as a streak. We don't.",
card: { kind: 'absence', title: 'NOTHING HOT TONIGHT', subtitle: 'AND WE WILL SAY SO', lines: [{ text: 'No hitter cleared his own season rate by enough to name.', style: 'body' }], footer: 'An empty list is a real answer.' },
};
},
};
@@ -0,0 +1,65 @@
'use strict';
/**
* TEMPLATE 3 — STREAK LIST.
*
* A streak is only a streak if every game in it settled. This reads the ledger's
* SETTLED outcomes, so a run cannot be extended by a game that is still pending
* or was voided for a DNP — the two ways a streak list quietly inflates.
*
* The structure is sport-agnostic: it takes settled outcome rows and a noun.
* MLB hit streaks today; the same shape carries TD streaks or made-three streaks
* the moment those sports have settled outcomes.
*/
const MIN_STREAK = 3;
module.exports = {
id: 'streak_list',
sport: 'mlb',
label: 'Active Streaks',
requires: ['streaks', 'count', 'noun', 'date'],
async pull(deps) {
const rows = await deps.settledStreaks(); // injected, read-only
const live = (rows || [])
.filter((r) => Number(r.streak) >= MIN_STREAK && r.verified_from_settled === true)
.sort((a, b) => b.streak - a.streak)
.slice(0, 6);
return {
date: deps.date,
noun: deps.noun || 'game hit streak',
count: live.length || null,
streaks: live.length ? live : null,
list: live.map((s) => `${s.name}${s.streak} straight`).join('\n'),
top_name: live[0] ? live[0].name : null,
top_streak: live[0] ? live[0].streak : null,
};
},
copy: () => `ACTIVE STREAKS — {date}
{top_name} has a {top_streak}-{noun}. Live, verified off settled results only.
{list}
Every game in these ran to a final. We don't count a pending night to make a number look better.`,
card: () => ({
kind: 'list',
title: 'ACTIVE STREAKS',
subtitle: 'VERIFIED FROM SETTLED RESULTS · {date}',
lines: [
{ text: '{top_name}', style: 'lead' },
{ text: '{top_streak} straight', style: 'stat' },
{ text: '', style: 'rule' },
{ text: '{list}', style: 'body' },
],
footer: 'Settled games only. Pending and voided nights do not count.',
}),
absent: () => ({
copy: 'No live streaks worth naming tonight.\n\nWe could lower the bar to three-of-four. We keep the bar.',
card: { kind: 'absence', title: 'NO LIVE STREAKS', subtitle: 'THE BAR STAYS WHERE IT IS', lines: [{ text: 'Nothing running long enough to name.', style: 'body' }], footer: 'We do not lower a threshold to fill a card.' },
}),
};
+6
View File
@@ -96,7 +96,13 @@ function dedupeProps(props, limit) {
// Grade both sides and keep the higher-confidence verdict — that's the
// side the engine actually favors.
async function gradeBestSide(grade, prop, sport, opts = {}) {
// PIPELINE ORDER: the factor context must reach the engine BEFORE it grades,
// because factors adjust the forecast the grade is read from. It was
// previously computed downstream of the grade it should inform.
const factorContext = typeof opts.factorContext === 'function'
? opts.factorContext(prop, sport) : null;
const base = {
factor_context: factorContext,
player: prop.player,
stat_type: prop.stat_type,
line: prop.line,
+75 -3
View File
@@ -114,7 +114,11 @@ function buildConcreteReasoning(features = {}, engine1Result = {}, meta = {}, pr
lines.push(`${prop.player || 'Player'} is averaging ${features.l5_avg.toFixed(1)} ${prop.stat_type || ''} over his last 5 games.`);
}
if (Number.isFinite(features.l20_avg)) {
lines.push(`Last 20 games average: ${features.l20_avg.toFixed(1)}.`);
// LABEL FIXED WITH THE DATA. `l20_avg` is now built from the full season
// log rather than a ten-game slice, so "Last 20 games" was a sentence the
// number no longer supported. The field name is kept (it is read in many
// places) but the copy states what is actually being shown.
lines.push(`Season average: ${features.l20_avg.toFixed(1)}.`);
}
// Trend direction relative to the line.
@@ -547,10 +551,78 @@ async function analyzeViaEngine1(rawProp = {}) {
const dir = String(prop.direction || 'over').toLowerCase();
const est = estimateProbability({ gameLogs: meta.gameLogs, line: prop.line, statType: rawProp.stat_type, features });
// ── PROVEN FACTORS, PRE-GRADE ────────────────────────────────────────
// base rate -> FACTORS -> (calibration, later) -> grade. Only the three
// factors that passed the two-part gate, only on hits, and only where each
// is readable — every unreadable case leaves the forecast untouched rather
// than nudging it toward a default.
let pOver = est.p_over;
let factorTrace = null;
if (String(rawProp.stat_type || '').toLowerCase() === 'hits' && rawProp.factor_context) {
try {
const hf = require('../model/hitsFactors');
const adj = hf.adjustProbability(pOver, rawProp.factor_context);
if (adj.p_adjusted != null && adj.factors_fired > 0) {
pOver = adj.p_adjusted;
factorTrace = { multiplier: adj.multiplier, applied: adj.applied, skipped: adj.skipped, p_before: adj.p_base };
}
} catch { /* a factor must never break the grade */ }
}
const pWin = dir === 'under'
? (Number.isFinite(est.p_over) ? 1 - est.p_over : null)
: (Number.isFinite(est.p_over) ? est.p_over : null);
? (Number.isFinite(pOver) ? 1 - pOver : null)
: (Number.isFinite(pOver) ? pOver : null);
if (pWin != null) legacy.p_win = Math.round(pWin * 1000) / 1000;
// ── THE SERVED GRADE ────────────────────────────────────────────────
// Derived from the forecast, not from engine1's additive factor index.
// Measured on 3,417 settled props, that index carried 0.16x the information
// of the p_win printed beside it, and its A grade hit 0.500 while its F hit
// 0.535 -- the top letter did worse than the bottom. Concretely: the same
// hitter's 0.95 over and 0.05 under both graded C.
//
// engine1.grade is PRESERVED on the payload as `engine_grade` so nothing
// downstream breaks and the two remain comparable, but `served_grade` is
// what a user should see.
try {
const sg = require('../model/servedGrade');
const served = sg.gradeFor({
p_win: legacy.p_win,
refused: legacy.refused || legacy.insufficient_data,
refusal_reason: legacy.refusal_reason,
factor_adjustment: factorTrace,
});
legacy.served_grade = served;
// ── TOTAL CUTOVER, NOT A PARALLEL GRADE ────────────────────────────
// `legacy.grade` IS the honest letter now. Attaching served_grade beside
// the old one and leaving `grade` alone would have repeated the exact
// failure diagnosed for gradeBands: built, correct, and read by nobody.
// Fourteen-plus consumers (scan route, dashboard, parlay, newsletter,
// desk, content templates, retention) all read `.grade`, so overwriting
// it here cuts every surface over at once instead of editing each.
//
// The original index is preserved as `engine_grade` for comparison and is
// read by no serving code.
legacy.engine_grade = legacy.grade;
if (served.letter) {
legacy.grade = served.letter;
// Confidence must not contradict the letter. It previously came from a
// grade-band midpoint of the OLD letter, so leaving it would have paired
// a B+ with a C's confidence. Both now derive from p_win -- kept on the
// existing 0-100 scale, since every consumer and every stored row uses it.
legacy.confidence = Math.round(legacy.p_win * 100);
legacy.confidence_basis = 'p_win';
} else {
// A refusal has no letter, and consumers test `!grade` for exactly that.
legacy.grade = null;
}
} catch { /* the grade surface must never break the read */ }
if (factorTrace) {
legacy.factor_adjustment = factorTrace;
legacy.p_win_prefactor = Math.round((dir === 'under' ? 1 - factorTrace.p_before : factorTrace.p_before) * 1000) / 1000;
}
const sideOdds = dir === 'under' ? rawProp.under_odds : rawProp.over_odds;
+45 -5
View File
@@ -29,6 +29,16 @@ const { getTeamInjuries } = require('./injuryParser');
const { getLineMovement } = require('./lineMovement');
const gameLogs = require('./gameLogService');
/**
* How many games to request when asking for a player's history.
*
* A basketball season is 82 games; this is deliberately past it so a request
* never truncates a season into a "season rate". The window-bug class has cost
* this codebase four MLB paths and two basketball ones -- every instance was a
* fixed N standing in for a season.
*/
const SEASON_LOG_DEPTH = 100;
const VECTOR_TTL_SECONDS = 120;
function avg(values) {
@@ -129,7 +139,12 @@ const NBA_LOG_FIELD = {
function mlbGameLogFeatures(res, statType) {
if (!res || !res.found) return {};
const out = {};
const logs = Array.isArray(res.last10) ? res.last10 : [];
// SAME WINDOW BUG AS THE BASE RATE. `l20_avg` is documented as "the season
// per-game reference projectionFor needs" and is read by the consistency
// pull — but built from last10 it was a TEN-game average wearing a season
// label. fullLog is already in this same response, so widening is free.
const logs = (Array.isArray(res.fullLog) && res.fullLog.length)
? res.fullLog : (Array.isArray(res.last10) ? res.last10 : []);
const vals = logs.map((g) => mlbStatValue(g.stat, statType)).filter((v) => v != null);
if (vals.length) {
const m5 = avg(vals.slice(-5)); // game logs are chronological (recent last)
@@ -220,7 +235,13 @@ function nbaGameLogFeatures(res, statType) {
if (vals.length) {
const m5 = avg(vals.slice(0, 5)); // most-recent first
const m10 = avg(vals.slice(0, 10));
const m20 = avg(vals.slice(0, 20));
// SEASON, NOT TWENTY. `l20_avg` is the season per-game reference
// `projectionFor` reads, and slicing to 20 made it a twenty-game average
// wearing a season label -- the same defect as the MLB last10 bug
// (929fd81) and the MLB l20 bug (494c83c). The ESPN adapter now returns the
// full season log, so this takes all of it. Name kept: `l20_avg` is read in
// many places, and renaming it is a separate, wider change.
const m20 = avg(vals);
const s10 = stddev(vals.slice(0, 10));
if (m5 != null) out.l5_avg = m5;
if (m10 != null) out.l10_avg = m10;
@@ -278,7 +299,22 @@ async function getStatRows(playerName, sport, statType) {
if (sp === 'mlb') {
const mlbStats = require('../adapters/mlbStatsAdapter');
const res = await mlbStats.getPlayerStats(playerName);
const logs = (res && res.found && Array.isArray(res.last10)) ? res.last10 : [];
// THE FULL SEASON LOG, NOT last10.
//
// This is the one line that made the champion worse than a frequency
// table. `estimateProbability` computes its base rate as the frequency
// over EVERY row it is given, so feeding it ten games meant the "season
// rate" was a ten-game rate — and then 0.4 of the forecast was the last
// five OF THOSE TEN. Measured point-in-time, a true season frequency
// out-resolved the served champion on all four stats (hits 0.00774 vs
// 0.00251, rbi 0.03133 vs 0.02481).
//
// `fullLog` is already fetched in the same adapter call that produced
// last10, so this costs nothing: no extra request, no new dependency.
const logs = (res && res.found)
? (Array.isArray(res.fullLog) && res.fullLog.length ? res.fullLog
: (Array.isArray(res.last10) ? res.last10 : []))
: [];
// MLB logs are chronological (most recent LAST) — reverse to match.
for (const g of [...logs].reverse()) push(g && g.date, mlbStatValue(g && g.stat, statType));
return rows;
@@ -286,7 +322,11 @@ async function getStatRows(playerName, sport, statType) {
// NBA/WNBA — Python service first (it's the richer source when it's up),
// then the FREE ESPN per-athlete gamelog. Same order as gameLogFeatures.
const pyLogs = await gameLogs.getGameLogs(playerName, sp, 20);
// SEASON DEPTH, NOT TWENTY. The count is a request PARAMETER, so asking for
// a season costs the same single call -- no extra request, no extra quota.
// The Python service is offline in prod, so this cannot be verified live;
// it is fixed now so basketball never launches on a twenty-game base rate.
const pyLogs = await gameLogs.getGameLogs(playerName, sp, SEASON_LOG_DEPTH);
if (Array.isArray(pyLogs) && pyLogs.length) {
// Python rows are already flat + most-recent-first.
for (const r of pyLogs) push(r && r.date, statFromGameLog(r, statType));
@@ -334,7 +374,7 @@ async function gameLogFeatures(playerName, sport, statType) {
}
}
const logs = await gameLogs.getGameLogs(playerName, sport, 20);
const logs = await gameLogs.getGameLogs(playerName, sport, SEASON_LOG_DEPTH);
// Wave 0 — NBA/WNBA grade unlock. The Python nba_api service (gameLogService)
// is offline in prod, so `logs` is null and this branch used to return {} →
+11 -1
View File
@@ -26,7 +26,17 @@ function pythonPath(sport) {
}
}
async function getGameLogs(playerName, sport, count = 20) {
/**
* DEFAULT IS A SEASON, NOT TWENTY.
*
* A default of 20 meant any caller that omitted the count silently received a
* twenty-game window -- and every consumer of this function treats what it
* returns as the player's history. The window-bug class has cost six paths
* across two sports; a permissive default is how a seventh would arrive.
*/
const DEFAULT_LOG_DEPTH = 100;
async function getGameLogs(playerName, sport, count = DEFAULT_LOG_DEPTH) {
const path = pythonPath(sport);
if (!path) return null;
const cacheKey = `gamelogs:${sport}:${playerName}:${count}`;
@@ -17,6 +17,19 @@
*/
const CV_VOLATILE_THRESHOLD = 0.40;
/**
* How much of the forecast is the last five games.
*
* Was 0.40. Measured point-in-time against a fair season-frequency baseline, a
* 0.40 weight COST resolution on every stat — hits 0.00086, total_bases
* 0.00107, rbi 0.00562, runs 0.00365 — because five games is a very noisy
* read and the blend pulled the forecast off a better number.
*
* 0.20 was the best measured weight on hits and total_bases; rbi and runs
* preferred 0 outright. It is set at the value the evidence supports rather
* than at the value that flatters recency.
*/
const RECENCY_WEIGHT = 0.20;
const PROB_FLOOR = 0.10;
const PROB_CEIL = 0.95;
@@ -68,7 +81,7 @@ function estimateProbability({ gameLogs = [], line, statType, features = {} } =
const recent = values.slice(0, Math.min(5, values.length));
const recencyRate = frequencyOver(recent, numericLine);
const weighted = recencyRate != null
? 0.6 * base + 0.4 * recencyRate
? (1 - RECENCY_WEIGHT) * base + RECENCY_WEIGHT * recencyRate
: base;
let p = weighted;
+2 -2
View File
@@ -141,11 +141,11 @@ function signalRecencyInflation(input) {
return inactive('l5_avg or l20_avg missing');
}
const ratio = (l5 - l20) / l20;
if (ratio <= 0) return { score: 0, active: true, explanation: 'L5 not hotter than L20' };
if (ratio <= 0) return { score: 0, active: true, explanation: 'L5 not hotter than season' };
return {
score: Math.min(1.0, ratio),
active: true,
explanation: `L5 (${l5.toFixed(1)}) ${(ratio * 100).toFixed(0)}% above L20 (${l20.toFixed(1)})`,
explanation: `L5 (${l5.toFixed(1)}) ${(ratio * 100).toFixed(0)}% above season (${l20.toFixed(1)})`,
};
}
+125
View File
@@ -0,0 +1,125 @@
'use strict';
/**
* calibrationDuel — the forward adjudication of a bet we made against the
* measurement.
*
* On identical held-out rows the ISOTONIC map beat the low-parameter one (hits
* +0.0028, rbi +0.0042, total_bases tied). We serve the low-parameter map
* anyway, on the argument that isotonic's in-window edge is daily structure
* shared between the fit and evaluation windows. At 19 dates that argument
* cannot be tested — LODO has 1.4-9.3% power against it.
*
* So it is a BET. This module is what makes it falsifiable: both maps are
* computed on every prop, the shadow is logged, and once enough genuinely
* out-of-window dates settle, the season adjudicates.
*
* ── THE RULE IS PRE-REGISTERED, IN CODE ──────────────────────────────────
* Written before any forward date exists, so the bar cannot drift toward
* whichever answer arrives:
*
* REFUTED >=10 forward dates AND isotonic beats low-param with a date-block
* bootstrap CI excluding zero -> revert hits/TB to isotonic
* UPHELD >=10 forward dates and it does not -> the bet was right
* PENDING fewer than 10 forward dates -> no verdict, keep serving
*
* A date is FORWARD only if NEITHER map was fitted on it. Scoring on a date
* inside either fit window would be asking which map memorised better.
*/
const { knownNumber } = require('../../utils/known');
/** Forward dates required before the duel may return a verdict. */
const MIN_FORWARD_DATES = 10;
const ITERS = 4000;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
const brier = (rows, key) => {
const usable = rows.filter((r) => knownNumber(r[key]) !== null && knownNumber(r.won) !== null);
if (!usable.length) return null;
return mean(usable.map((r) => (knownNumber(r[key]) - knownNumber(r.won)) ** 2));
};
/**
* @param {Array} rows [{ date, won, served, shadow, fitted_through }]
* @param {object} opts { minForwardDates, seed }
*/
function adjudicate(rows, opts = {}) {
const minDates = opts.minForwardDates ?? MIN_FORWARD_DATES;
// FORWARD ONLY: a row counts when its date postdates the window BOTH maps
// were fitted on. Rows without that provenance are dropped, never assumed.
const forward = (rows || []).filter((r) => {
if (!r || !r.date) return false;
if (knownNumber(r.served) === null || knownNumber(r.shadow) === null) return false;
if (knownNumber(r.won) === null) return false;
if (!r.fitted_through) return false;
return String(r.date) > String(r.fitted_through);
});
const dates = [...new Set(forward.map((r) => String(r.date)))].sort();
if (dates.length < minDates) {
return {
verdict: 'PENDING',
forward_dates: dates.length,
forward_rows: forward.length,
dates_needed: minDates - dates.length,
reason: `${dates.length} forward dates < ${minDates} — the season has not spoken yet`,
action: 'keep serving the low-parameter map',
};
}
const bServed = brier(forward, 'served');
const bShadow = brier(forward, 'shadow');
if (bServed === null || bShadow === null) {
return { verdict: 'PENDING', forward_dates: dates.length, reason: 'no scorable forward rows' };
}
// Paired date-block bootstrap on (isotonic - lowparam). Negative means the
// shadow is better, which is the direction that refutes us.
const byDate = new Map();
for (const r of forward) {
if (!byDate.has(String(r.date))) byDate.set(String(r.date), []);
byDate.get(String(r.date)).push(r);
}
const keys = [...byDate.keys()];
const rnd = makeRnd(opts.seed ?? 20260807);
const diffs = [];
for (let it = 0; it < ITERS; it += 1) {
const s = [];
for (let i = 0; i < keys.length; i += 1) s.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
const a = brier(s, 'shadow');
const b = brier(s, 'served');
if (a === null || b === null) continue;
diffs.push(a - b);
}
diffs.sort((a, b) => a - b);
const ci = diffs.length
? [round5(diffs[Math.floor(diffs.length * 0.025)]), round5(diffs[Math.floor(diffs.length * 0.975)])]
: null;
const shadowWins = ci !== null && ci[1] < 0;
return {
verdict: shadowWins ? 'REFUTED' : 'UPHELD',
forward_dates: dates.length,
forward_rows: forward.length,
brier_served_lowparam: round5(bServed),
brier_shadow_isotonic: round5(bShadow),
delta_isotonic_minus_lowparam: round5(bShadow - bServed),
ci,
reason: shadowWins
? 'isotonic beats the served low-parameter map out-of-window with a date-block interval excluding zero — the capacity argument is refuted'
: 'the served low-parameter map is not beaten out-of-window — the bet stands',
action: shadowWins ? 'REVERT hits and total_bases to isotonic and log the reversal' : 'keep serving the low-parameter map',
};
}
const round5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
module.exports = { adjudicate, MIN_FORWARD_DATES };
+146
View File
@@ -0,0 +1,146 @@
'use strict';
/**
* calibrationGuards — the two ways a calibration measurement lies to you.
*
* Both of these produced a confident, plausible, completely wrong number in the
* settlement session, and neither was visible in the output. They are codified
* here so the failure cannot recur silently.
*
* ── GUARD 1: THE BOTH-SIDES TELL ─────────────────────────────────────────
* A prop population usually carries BOTH the over and the under. Their p_wins
* sum to ~1 and their outcomes are complementary, so ANY population-level
* calibration statistic over the raw set is pinned to 0.5 by construction — not
* by the model being calibrated.
*
* Measured: the raw population read +0.0002 bias on hits ("perfectly
* calibrated"); deduped to the model-picked side it read +0.0868. Same rows,
* opposite conclusion. The tell was mean p_win sitting at 0.4998 on all four
* stats at once, which is not something a real forecaster does.
*
* So picked-side dedup is MANDATORY preprocessing, and this asserts it.
*
* ── GUARD 2: A NULL THAT SCORES ITSELF ───────────────────────────────────
* `fitIsotonic` returns null below its minimum and `applyIsotonic` then returns
* null per row. In JavaScript `(null - 1) ** 2 === 1` and `(null - 0) ** 2 === 0`,
* so a Brier score computed over nulls silently equals the WIN RATE — a number
* in the right range, monotone in the data, and completely meaningless. It
* reported hits at 0.5567 against a 0.5684 win rate.
*
* This is the `Number(null) === 0` breach the TRUTH LAW names, wearing a metric.
* A null prediction must refuse, never score.
*/
const { knownNumber } = require('../../utils/known');
/** How close to 0.5 counts as the both-sides signature. */
const BALANCED_TOLERANCE = 0.02;
/**
* Does this population still contain both sides of the same prop?
*
* @param {Array} rows [{ p, side, propKey }]
* @returns {object} { violated, reason, ... } — never throws, so a caller can
* decide between refusing and hard-failing.
*/
function checkPickedSideDedup(rows) {
const usable = (rows || []).filter((r) => knownNumber(r && r.p) !== null);
if (usable.length < 2) return { violated: false, reason: 'too few rows to judge', n: usable.length };
const sidesByProp = new Map();
for (const r of usable) {
const k = r.propKey == null ? null : String(r.propKey);
if (k === null) continue;
if (!sidesByProp.has(k)) sidesByProp.set(k, new Set());
if (r.side) sidesByProp.get(k).add(String(r.side).toLowerCase());
}
let bothSides = 0;
for (const s of sidesByProp.values()) if (s.size > 1) bothSides += 1;
const propCount = sidesByProp.size;
const bothShare = propCount ? bothSides / propCount : 0;
const meanP = usable.reduce((s, r) => s + knownNumber(r.p), 0) / usable.length;
const balanced = Math.abs(meanP - 0.5) <= BALANCED_TOLERANCE;
// The violation is the CONJUNCTION: both sides present AND the mean pinned at
// 0.5. Either alone is unremarkable — a genuinely balanced book of one-sided
// picks is fine, and both sides present with a skewed mean means someone
// already deduped.
const violated = bothSides > 0 && balanced;
return {
violated,
n: usable.length,
props: propCount,
both_sides_props: bothSides,
both_sides_share: round4(bothShare),
mean_p: round4(meanP),
reason: violated
? `both sides present on ${bothSides}/${propCount} props while mean p_win is ${round4(meanP)}`
+ 'the population is balanced by construction and any calibration statistic over it is meaningless. '
+ 'Dedup to the model-picked side first.'
: null,
};
}
/** Same check, but refuses to continue. Use at the top of a measurement. */
function assertPickedSideDedup(rows) {
const r = checkPickedSideDedup(rows);
if (r.violated) throw new Error(`CALIBRATION GUARD: ${r.reason}`);
return r;
}
/**
* Brier score that refuses rather than scoring a null.
*
* @param {Array} preds
* @param {Array} outcomes
* @param {object} opts { onNull: 'throw' | 'refuse' } default 'refuse'
* @returns {number|null} null when any prediction is unreadable
*/
function safeBrier(preds, outcomes, opts = {}) {
const ps = preds || [];
const ys = outcomes || [];
if (ps.length === 0 || ps.length !== ys.length) return null;
let sum = 0;
for (let i = 0; i < ps.length; i += 1) {
const p = knownNumber(ps[i]);
const y = knownNumber(ys[i]);
if (p === null || y === null) {
// NEVER score it. (null - 1) ** 2 === 1 would pass silently.
if (opts.onNull === 'throw') {
throw new Error('CALIBRATION GUARD: a null prediction reached a Brier term');
}
return null;
}
sum += (p - y) ** 2;
}
return sum / ps.length;
}
/**
* Map a population through a calibration map, refusing unreadable rows rather
* than letting them through as nulls.
*/
function applyOrRefuse(map, rows, applyFn) {
if (!map) return { ok: false, reason: 'no calibration map could be fitted', rows: [] };
const out = [];
let dropped = 0;
for (const r of rows || []) {
const pc = applyFn(map, r.p);
if (knownNumber(pc) === null) { dropped += 1; continue; }
out.push({ ...r, pc });
}
return {
ok: out.length > 0,
rows: out,
dropped,
reason: out.length === 0 ? 'every row was unmappable' : null,
};
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = {
checkPickedSideDedup, assertPickedSideDedup, safeBrier, applyOrRefuse,
BALANCED_TOLERANCE,
};
+163
View File
@@ -0,0 +1,163 @@
'use strict';
/**
* calibrationRegistry — which stats are allowed to serve a calibrated number.
*
* ── WHY THIS IS NOT THE FACTOR REGISTRY ──────────────────────────────────
* A factor makes a CAUSAL claim, so it needs a Bonferroni slot, a cluster-robust
* interval, and a bar that rises with every hypothesis the programme tests.
* Calibration makes no causal claim: it is a monotone shrink toward what was
* actually observed, its failure mode is bounded (it can only over- or
* under-shrink), and it consumes no test slot.
*
* Applying the factor gate's >=40 date-cluster interval floor to it was the
* wrong instrument. The real risk for a calibration layer is that the correction
* is DATE-DRIVEN, and leave-one-date-out tests that directly — and harder.
*
* ── TWO TIERS, AND AUTO-DEMOTION IS WHAT MAKES PROVISIONAL HONEST ────────
* PROVISIONAL LODO passes AND the point-in-time held-out CI excludes zero.
* Serves, labelled, inside its certified band only.
* PROMOTED the original >=40 date-cluster bar, now correctly scoped as the
* PROMOTION bar rather than the deploy bar.
*
* A provisional deploy that cannot be taken away is just a deploy. `reverify`
* runs on every newly settled date and demotes on the first breach.
*/
const { knownNumber } = require('../../utils/known');
const STATUS = Object.freeze({ NONE: 'none', PROVISIONAL: 'provisional', PROMOTED: 'promoted' });
/**
* THE LODO TEST — a COHERENT pair, replacing the incoherent one at 1f40014.
*
* ── WHAT WAS WRONG ───────────────────────────────────────────────────────
* The previous gate paired a 1-SE per-drop informativeness bar with a
* ZERO-reversal decision rule. At exactly 1 SE a genuinely STABLE stat's drop
* reverses with probability Phi(-1) = 0.159, so on four informative drops the
* chance of at least one reversal is 1 - 0.841^4 = 0.50. The rule failed stable
* stats half the time BY CONSTRUCTION. And n* was pooled across four stats whose
* signed effects differ several-fold, so one number meant four different things:
* measured, the pooled 70 was too LOW for hits (77) and runs (81) and too HIGH
* for total_bases (60) and rbi (54).
*
* ── THE COHERENT PAIR ────────────────────────────────────────────────────
* The bar and the rule are chosen TOGETHER, per stat, for a stated error rate:
*
* informative bar n*_k = k^2 * (sigma_row / |g|)^2 per stat
* decision rule FAIL iff reversals > cutoff, where under stability
* R ~ Binomial(D, Phi(-k)) and cutoff is the smallest c
* with P(R > c) <= 0.05
*
* `g` is the mean SIGNED per-row improvement — the quantity whose sign a
* reversal flips. k = 1 is chosen because it maximises informative drops (D),
* which is the binding scarcity here, while the binomial cutoff holds the
* false-positive rate at 0.004-0.031 across the four stats.
*
* ── AND THE TEST STILL HAS ALMOST NO POWER ───────────────────────────────
* At the stated alternative (date-to-date SD of the effect equal to |g| — a
* strong instability), power is 0.093 / 0.093 / 0.045 / 0.014. The test would
* MISS a real date-driven failure more than nine times in ten. Across every k
* examined, the best any stat reaches is 0.337.
*
* So a PASS here means "no instability was detected", NOT "it is stable", and a
* gate that cannot fail is not a gate. LODO_POWER_FLOOR makes that structural: a
* stat whose test power falls below it is UNTESTABLE-BY-LODO and may not claim
* LODO stability at all, whatever its reversal count.
*
* Derived BLIND — the derivation script prints no reversal, no verdict and no
* reversing date. It ran, and these were committed, before any stat was re-read.
*/
const LODO_K = 1.0;
/** Below this power the test cannot fail, so it cannot pass either. */
const LODO_POWER_FLOOR = 0.50;
/** Per-stat, from (g, sigma_row) measured blind. */
const LODO_TEST = Object.freeze({
hits: { g: -0.01288, sigma_row: 0.11251, n_star: 77, informative_drops: 5, cutoff: 2, fp: 0.0310, power: 0.093 },
total_bases: { g: -0.01380, sigma_row: 0.10680, n_star: 60, informative_drops: 5, cutoff: 2, fp: 0.0310, power: 0.093 },
rbi: { g: -0.00884, sigma_row: 0.06459, n_star: 54, informative_drops: 4, cutoff: 2, fp: 0.0141, power: 0.045 },
runs: { g: -0.00902, sigma_row: 0.08080, n_star: 81, informative_drops: 3, cutoff: 2, fp: 0.0040, power: 0.014 },
});
/** Legacy name kept so nothing silently reads a stale pooled value. */
const LODO_MIN_HELD_ROWS = null;
/** The ORIGINAL floor, correctly scoped: promotion, not deploy. */
const PROMOTION_DATE_CLUSTERS = 40;
function createRegistry(initial = {}) {
const state = new Map(Object.entries(initial));
const log = [];
/** Deploy requires BOTH gates. Neither can be waived. */
function deploy(stat, evidence = {}) {
const lodo = evidence.lodo_pass === true;
const ci = Array.isArray(evidence.ci) && evidence.ci.length === 2 && evidence.ci[1] < 0;
if (!lodo || !ci) {
return {
ok: false,
status: STATUS.NONE,
reason: !lodo
? 'LODO did not pass — the correction may be date-driven'
: 'the point-in-time held-out interval does not exclude zero',
};
}
if (!evidence.map) return { ok: false, status: STATUS.NONE, reason: 'no calibration map supplied' };
state.set(stat, {
status: STATUS.PROVISIONAL,
map: evidence.map,
certified_bands: evidence.certified_bands || [],
date_clusters: knownNumber(evidence.date_clusters) ?? 0,
ci: evidence.ci,
deployed_at: evidence.at || null,
});
log.push({ stat, event: 'deployed_provisional', at: evidence.at || null });
return { ok: true, status: STATUS.PROVISIONAL };
}
/**
* Re-verify on a newly settled date. Demotes on the FIRST breach — either the
* interval ceasing to exclude zero, or the favourite over-prediction flipping
* sign (which would mean the correction is now pushing the wrong way).
*/
function reverify(stat, obs = {}) {
const cur = state.get(stat);
if (!cur || cur.status === STATUS.NONE) return { status: STATUS.NONE, changed: false };
const ciHolds = Array.isArray(obs.ci) && obs.ci.length === 2 && obs.ci[1] < 0;
const signHolds = obs.favourite_bias == null ? true : knownNumber(obs.favourite_bias) > 0;
if (!ciHolds || !signHolds) {
state.set(stat, { status: STATUS.NONE, demoted_at: obs.date || null, demoted_reason: !ciHolds ? 'ci_no_longer_excludes_zero' : 'favourite_bias_flipped' });
log.push({ stat, event: 'auto_demoted', at: obs.date || null, reason: !ciHolds ? 'ci_no_longer_excludes_zero' : 'favourite_bias_flipped' });
return { status: STATUS.NONE, changed: true, reason: !ciHolds ? 'ci_no_longer_excludes_zero' : 'favourite_bias_flipped', breaking_date: obs.date || null };
}
// PROMOTION uses the original floor, now correctly scoped.
const dc = knownNumber(obs.date_clusters) ?? cur.date_clusters;
if (cur.status === STATUS.PROVISIONAL && dc >= PROMOTION_DATE_CLUSTERS) {
state.set(stat, { ...cur, status: STATUS.PROMOTED, date_clusters: dc, promoted_at: obs.date || null });
log.push({ stat, event: 'promoted', at: obs.date || null, date_clusters: dc });
return { status: STATUS.PROMOTED, changed: true };
}
if (dc !== cur.date_clusters) state.set(stat, { ...cur, date_clusters: dc });
return { status: cur.status, changed: false };
}
/** Is this stat allowed to serve a calibrated number for THIS p_win? */
function serves(stat, p) {
const cur = state.get(stat);
if (!cur || cur.status === STATUS.NONE) return { serve: false, reason: 'not deployed' };
const x = knownNumber(p);
if (x === null) return { serve: false, reason: 'no p_win' };
const inBand = (cur.certified_bands || []).some((b) => x >= b[0] && x < b[1]);
if (!inBand) return { serve: false, reason: 'outside the certified band', status: cur.status };
return { serve: true, status: cur.status, provisional: cur.status === STATUS.PROVISIONAL };
}
const get = (stat) => state.get(stat) || { status: STATUS.NONE };
const all = () => Object.fromEntries([...state.entries()].map(([k, v]) => [k, { status: v.status, certified_bands: v.certified_bands, date_clusters: v.date_clusters }]));
return { deploy, reverify, serves, get, all, log: () => log.slice() };
}
module.exports = { createRegistry, STATUS, PROMOTION_DATE_CLUSTERS, LODO_K, LODO_TEST, LODO_POWER_FLOOR, LODO_MIN_HELD_ROWS };
+68 -4
View File
@@ -40,7 +40,24 @@ const { knownNumber } = require('../../utils/known');
/** Minimum mean |Δp| for a factor to count as having moved anything at all. */
const MIN_MOVEMENT = 0.01;
/** Observations needed for the effect ESTIMATE to be stable. */
const MIN_N = 500;
/**
* Clusters needed for the cluster-robust INTERVAL to be trustworthy.
*
* These two floors answer different questions and must not be collapsed. Rows
* govern whether the point estimate is stable; clusters govern whether the
* interval around it means anything. Transplanting the 500-row bar onto clusters
* refuses a factor measured over 1,059 rows and 85 games — which has ample
* observations AND ample clusters — while telling us nothing about either.
*
* 40 is the conventional floor below which cluster-robust inference is known to
* under-cover regardless of how many rows sit inside the clusters. It is a
* statement about when the bootstrap can be believed, not a bar tuned to let
* anything through: a venue-constant factor still caps at 30 ballparks and is
* still refused, permanently.
*/
const MIN_CLUSTERS = 40;
const brier = (ps, ys) => (ps.length
? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
@@ -92,11 +109,41 @@ function improvement(rows, iters = 4000, seed = 20260805, cumulativeTests = 1) {
if (usable.length < 30) return null;
const rnd = makeRnd(seed);
const diffs = [];
// ── PSEUDO-REPLICATION ────────────────────────────────────────────────────
// A factor that assigns ONE value per game (park, weather, opposing starter)
// gives every prop row in that game the identical treatment. Resampling ROWS
// then treats 18 hitters in one ballpark as 18 independent readings of that
// ballpark, and the interval collapses to a width the evidence never earned —
// so the gate PASSES a factor on sample it does not have. Measured here: 928
// total_bases rows carry only 53 distinct games.
//
// When rows carry a `cluster`, resample whole clusters. The interval then
// reflects the unit the treatment actually varies over. Rows without a
// cluster keep the original row-resampling path byte-for-byte.
const clustered = usable.some((r) => r.cluster != null);
const groups = new Map();
if (clustered) {
for (const r of usable) {
const k = String(r.cluster);
if (!groups.has(k)) groups.set(k, []);
groups.get(k).push(r);
}
}
const keys = clustered ? [...groups.keys()] : null;
for (let it = 0; it < iters; it += 1) {
const b = []; const c = []; const y = [];
for (let i = 0; i < usable.length; i += 1) {
const r = usable[Math.floor(rnd() * usable.length)];
b.push(r.baseline); c.push(r.conditioned); y.push(r.won > 0 ? 1 : 0);
if (clustered) {
for (let i = 0; i < keys.length; i += 1) {
const g = groups.get(keys[Math.floor(rnd() * keys.length)]);
for (const r of g) { b.push(r.baseline); c.push(r.conditioned); y.push(r.won > 0 ? 1 : 0); }
}
} else {
for (let i = 0; i < usable.length; i += 1) {
const r = usable[Math.floor(rnd() * usable.length)];
b.push(r.baseline); c.push(r.conditioned); y.push(r.won > 0 ? 1 : 0);
}
}
diffs.push(brier(c, y) - brier(b, y));
}
@@ -113,6 +160,9 @@ function improvement(rows, iters = 4000, seed = 20260805, cumulativeTests = 1) {
const ys = usable.map((r) => (r.won > 0 ? 1 : 0));
return {
n: usable.length,
// The number the gate must actually judge sample against.
effective_n: clustered ? keys.length : usable.length,
cluster_unit: clustered ? 'cluster' : 'row',
brier_baseline: round4(brier(usable.map((r) => r.baseline), ys)),
brier_conditioned: round4(brier(usable.map((r) => r.conditioned), ys)),
brier_delta: round4(brier(usable.map((r) => r.conditioned), ys) - brier(usable.map((r) => r.baseline), ys)),
@@ -138,9 +188,22 @@ function adjudicate(rows, opts = {}) {
const base = { factor: opts.factor || null, archetype: opts.archetype || null, stat: opts.stat || 'hits', movement: mv, improvement: imp };
// TWO FLOORS, because they answer different questions. Rows decide whether the
// point estimate is stable; clusters decide whether the interval around it can
// be believed. A factor needs both.
if (mv.n < minN) {
return { ...base, verdict: 'CANDIDATE_PENDING_SAMPLE', reason: `n ${mv.n} < ${minN}`, rows_needed: minN - mv.n };
}
const minClusters = opts.minClusters ?? MIN_CLUSTERS;
if (imp && imp.cluster_unit === 'cluster' && imp.effective_n < minClusters) {
return {
...base,
verdict: 'CANDIDATE_PENDING_SAMPLE',
reason: `${mv.n} rows but only ${imp.effective_n} independent clusters < ${minClusters}`
+ ' — the rows are not independent readings and the interval cannot be trusted at this cluster count',
clusters_needed: minClusters - imp.effective_n,
};
}
if (mv.mean_abs_shift === null || mv.mean_abs_shift < minMove) {
// It never moved the number, so it cannot be reading anything.
return { ...base, verdict: 'INERT', reason: `mean |shift| ${mv.mean_abs_shift} < ${minMove}` };
@@ -176,4 +239,5 @@ function adjudicate(rows, opts = {}) {
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = { movement, improvement, adjudicate, MIN_MOVEMENT, MIN_N };
module.exports = {
MIN_CLUSTERS, movement, improvement, adjudicate, MIN_MOVEMENT, MIN_N };
+235
View File
@@ -0,0 +1,235 @@
'use strict';
/**
* gradeBands — WHAT DOES A LETTER MEAN, PER ARCHETYPE, AND WHO SAYS SO?
*
* A grade is the product. So a letter has to be backed by a realized rate the
* ledger can stand behind — not by a threshold someone chose. This builds bands
* from settled outcomes, per archetype, and refuses to publish any band the
* evidence cannot support.
*
* ── LIFT, NOT RAW RATE ────────────────────────────────────────────────────
* An A must beat the archetype's OWN naive base rate. A power hitter and a
* contact hitter have different base rates for a hit, so the same 62% realized
* rate is a strong read for one and slightly below water for the other. Bands
* are therefore drawn where the realized rate SEPARATES from that archetype's
* base rate, and a band whose interval still contains the base rate has shown no
* lift — whatever its raw number looks like.
*
* ── THE TWO-BAR RULE IS STRUCTURAL, NOT A HABIT ───────────────────────────
* A band may be described as FACTOR-INFORMED only when that archetype's factors
* are both PROVEN and CALIBRATED. There is no argument that overrides it: the
* caller passes evidence, and without it every band comes back `basis:
* 'base_rate'` and says so. A test asserts that with nothing proven — which is
* the state today — no factor-informed band can be produced at all.
*
* This is the same shape as featureRegistry.liveFeatures(): the honest state is
* the DEFAULT, and the richer claim has to be earned past a gate, so it cannot
* be reached by forgetting.
*
* ── WHY A BAND CAN BE REFUSED ─────────────────────────────────────────────
* Thin bands are labelled PROVISIONAL rather than dropped, because "we are still
* counting" and "there is nothing here" are different claims. But a band whose
* corrected interval spans the base rate does NOT get a lift letter — it is
* reported as the base-rate read it is.
*/
const { knownNumber } = require('../../utils/known');
/** Below this a band is published but flagged — counting, not concluded. */
const PROVISIONAL_N = 30;
/** Below this a band is not published at all; there is nothing to stand behind. */
const MIN_BAND_N = 12;
/** Letters, richest first. Assigned by measured lift, never by curve. */
const LETTERS = ['A', 'B', 'C', 'D', 'F'];
/**
* Wilson score interval, widened by the cumulative correction.
*
* Wilson rather than normal-approximation because these bands are small and
* rates sit near the edges, where the normal interval runs past 0 and 1 and
* quietly implies impossible rates.
*/
function wilson(hits, n, cumulativeTests = 1) {
const k = knownNumber(hits); const N = knownNumber(n);
if (k === null || N === null || N <= 0) return null;
const tests = Math.max(1, Math.round(knownNumber(cumulativeTests) ?? 1));
const z = zFor(1 - (0.05 / tests) / 2);
const p = k / N;
const d = 1 + (z * z) / N;
const centre = (p + (z * z) / (2 * N)) / d;
const half = (z * Math.sqrt((p * (1 - p)) / N + (z * z) / (4 * N * N))) / d;
return [round4(Math.max(0, centre - half)), round4(Math.min(1, centre + half))];
}
/** Inverse normal CDF (Acklam) — good to ~1e-9, enough for an interval bound. */
function zFor(p) {
if (p <= 0 || p >= 1) return 0;
const a = [-3.969683028665376e+01, 2.209460984245205e+02, -2.759285104469687e+02,
1.383577518672690e+02, -3.066479806614716e+01, 2.506628277459239e+00];
const b = [-5.447609879822406e+01, 1.615858368580409e+02, -1.556989798598866e+02,
6.680131188771972e+01, -1.328068155288572e+01];
const c = [-7.784894002430293e-03, -3.223964580411365e-01, -2.400758277161838e+00,
-2.549732539343734e+00, 4.374664141464968e+00, 2.938163982698783e+00];
const d = [7.784695709041462e-03, 3.224671290700398e-01, 2.445134137142996e+00,
3.754408661907416e+00];
const pl = 0.02425;
let q; let r;
if (p < pl) {
q = Math.sqrt(-2 * Math.log(p));
return (((((c[0] * q + c[1]) * q + c[2]) * q + c[3]) * q + c[4]) * q + c[5])
/ ((((d[0] * q + d[1]) * q + d[2]) * q + d[3]) * q + 1);
}
if (p > 1 - pl) {
q = Math.sqrt(-2 * Math.log(1 - p));
return -(((((c[0] * q + c[1]) * q + c[2]) * q + c[3]) * q + c[4]) * q + c[5])
/ ((((d[0] * q + d[1]) * q + d[2]) * q + d[3]) * q + 1);
}
q = p - 0.5; r = q * q;
return (((((a[0] * r + a[1]) * r + a[2]) * r + a[3]) * r + a[4]) * r + a[5]) * q
/ (((((b[0] * r + b[1]) * r + b[2]) * r + b[3]) * r + b[4]) * r + 1);
}
/**
* Build bands for ONE archetype.
*
* @param {Array} rows [{ p, won }] — p is the model's probability, won 0/1
* @param {object} opts
* opts.archetype label
* opts.cumulativeTests Bonferroni denominator (applied to every interval)
* opts.proven this archetype's factors passed the gate
* opts.calibrated this archetype's probabilities are certified calibrated
* opts.targetBands how many bands to attempt (default 5)
*/
function buildBands(rows, opts = {}) {
const usable = (rows || []).filter((r) =>
knownNumber(r && r.p) !== null && knownNumber(r && r.won) !== null);
const n = usable.length;
const tests = opts.cumulativeTests ?? 1;
// THE TWO-BAR RULE. Both, or the bands are a base-rate read and say so.
const factorInformed = Boolean(opts.proven) && Boolean(opts.calibrated);
if (n < MIN_BAND_N) {
return {
archetype: opts.archetype || null,
basis: factorInformed ? 'factor_informed' : 'base_rate',
n,
base_rate: null,
bands: [],
refused: 'insufficient settled outcomes to stand behind any band',
};
}
const wins = usable.reduce((s, r) => s + (r.won > 0 ? 1 : 0), 0);
const baseRate = wins / n;
// Candidate cuts by quantile of the model's own probability, then merged so
// every published band is distinguishable from its neighbour.
const sorted = [...usable].sort((a, b) => b.p - a.p);
const target = Math.max(2, Math.min(LETTERS.length, opts.targetBands || 5));
const per = Math.max(MIN_BAND_N, Math.floor(n / target));
let raw = [];
for (let i = 0; i < sorted.length; i += per) {
const slice = sorted.slice(i, i + per);
// A trailing remainder too small to stand alone joins the previous band
// rather than being published as its own thin claim.
if (slice.length < MIN_BAND_N && raw.length) raw[raw.length - 1].push(...slice);
else raw.push(slice);
}
// Merge neighbours whose corrected intervals overlap — if we cannot tell two
// bands apart, publishing them as different letters is a distinction we have
// not measured.
let merged = true;
while (merged && raw.length > 1) {
merged = false;
for (let i = 0; i < raw.length - 1; i += 1) {
const a = stats(raw[i], tests); const b = stats(raw[i + 1], tests);
if (a.ci && b.ci && a.ci[0] <= b.ci[1] && b.ci[0] <= a.ci[1]) {
raw.splice(i, 2, raw[i].concat(raw[i + 1]));
merged = true;
break;
}
}
}
const bands = raw.map((slice, i) => {
const s = stats(slice, tests);
// LIFT: does this band's interval clear the archetype's own base rate?
const clearsAbove = s.ci !== null && s.ci[0] > baseRate;
const clearsBelow = s.ci !== null && s.ci[1] < baseRate;
return {
letter: LETTERS[Math.min(i, LETTERS.length - 1)],
n: slice.length,
p_range: [round4(Math.min(...slice.map((r) => r.p))), round4(Math.max(...slice.map((r) => r.p)))],
realized_rate: s.rate,
ci: s.ci,
lift_vs_archetype_base: round4(s.rate - baseRate),
// The claim the letter is allowed to make.
shows_lift: clearsAbove,
shows_deficit: clearsBelow,
separation: clearsAbove ? 'above_base_rate' : (clearsBelow ? 'below_base_rate' : 'indistinguishable_from_base_rate'),
provisional: slice.length < PROVISIONAL_N,
basis: factorInformed ? 'factor_informed' : 'base_rate',
};
});
return {
archetype: opts.archetype || null,
basis: factorInformed ? 'factor_informed' : 'base_rate',
// Stated explicitly so a consumer never has to infer why it is a base-rate read.
two_bar: {
proven: Boolean(opts.proven),
calibrated: Boolean(opts.calibrated),
factor_informed_allowed: factorInformed,
...(factorInformed ? {} : {
reason: !opts.proven
? 'no factor has passed the gate for this archetype'
: 'factors proved but this archetype\'s probabilities are not certified calibrated',
}),
},
n,
base_rate: round4(baseRate),
bands,
bonferroni_tests: Math.max(1, Math.round(knownNumber(tests) ?? 1)),
bands_showing_lift: bands.filter((b) => b.shows_lift).length,
};
}
function stats(slice, tests) {
const n = slice.length;
const w = slice.reduce((s, r) => s + (r.won > 0 ? 1 : 0), 0);
return { rate: round4(n ? w / n : null), ci: wilson(w, n, tests) };
}
/**
* The sentence attached to a grade. TRUE or absent — never a fluent fallback.
*
* A factor-informed reason names the factors that proved FOR THIS ARCHETYPE. A
* base-rate reason says plainly that it is a base-rate read, because a user who
* is told "favourable matchup" when nothing about the matchup was read has been
* given a fabricated reason, and that is worse than being given none.
*/
function reasoning(band, opts = {}) {
if (!band) return null;
const arch = opts.archetype ? String(opts.archetype).toUpperCase() : null;
if (band.basis === 'factor_informed') {
const proved = (opts.provenFactors || []).filter(Boolean);
if (!proved.length) return null; // cannot name what did not prove
return `${arch ? `${arch}: ` : ''}${proved.join(' + ')} — proved for this profile`;
}
const sep = band.separation === 'above_base_rate'
? 'above this profile\'s base rate'
: (band.separation === 'below_base_rate' ? 'below this profile\'s base rate' : 'at this profile\'s base rate');
return `${arch ? `${arch}: ` : ''}base-rate read, ${sep}`
+ ' — no matchup factor is proven for this profile yet';
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = {
buildBands, reasoning, wilson,
PROVISIONAL_N, MIN_BAND_N, LETTERS,
};
+144
View File
@@ -0,0 +1,144 @@
'use strict';
/**
* hitsFactorContext — load the proven factors' inputs ONCE per slate.
*
* The three hits factors each need a database read (batter spray, team
* positional defence, platoon splits, pitcher contact profile). Doing that per
* prop would put four queries inside a loop that runs across the whole board, so
* the tables are loaded once and indexed, and the per-prop lookup is a map hit.
*
* Loaded BEFORE grading, which is the whole point of this order — the same data
* was previously fetched after the grade it should have informed.
*
* Every load is best-effort: a missing table yields an empty index, the factor
* finds nothing readable, and the forecast is served unadjusted. A factor layer
* must never be able to break the pipeline it rides in.
*/
const { knownNumber } = require('../../utils/known');
const { nameKey } = require('../../utils/playerName');
/** Rows a factor table must have before we trust it at all. */
const MIN_ROWS = 1;
/**
* Page a factor table.
*
* ORDERING IS PER-TABLE. These tables have COMPOSITE primary keys
* (as_of_date, sport, player_key) and NO `id` column, so ordering by `id`
* errors — and an error here returns an empty index, which reads exactly like
* "this feed has no data". That is the third time in this codebase that a
* wiring fault has worn the costume of an honest absence, so the error is now
* surfaced rather than swallowed.
*/
async function page(sb, table, select, orderBy, apply) {
const out = [];
for (let from = 0; ; from += 1000) {
const q = apply ? apply(sb.from(table).select(select)) : sb.from(table).select(select);
const { data, error } = await q.order(orderBy, { ascending: true }).range(from, from + 999);
if (error) throw new Error(`${table}: ${error.message}`);
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < 1000) break;
}
return out;
}
/** Keep the most recent dated row per key. */
function latestBy(rows, keyFn, dateFn) {
const m = new Map();
for (const r of rows) {
const k = keyFn(r);
if (!k) continue;
const prev = m.get(k);
if (!prev || String(dateFn(r)) > String(dateFn(prev))) m.set(k, r);
}
return m;
}
/**
* @returns {function|null} a `(prop, sport) => context` resolver, or null when
* nothing loaded — null means "serve unadjusted", never a stub context.
*/
async function build(sb, opts = {}) {
if (!sb) return null;
let spray; let defense; let platoon; let statcast;
try {
[spray, defense, platoon, statcast] = await Promise.all([
page(sb, 'batter_spray', '*', 'player_key', (q) => q.eq('sport', 'mlb')),
page(sb, 'team_defense', '*', 'team', (q) => q.eq('sport', 'mlb')),
page(sb, 'platoon_splits', '*', 'player_key', (q) => q.eq('sport', 'mlb')),
page(sb, 'statcast_aggregates', 'player_key, source_id, role, bats, throws, hard_hit_pct', 'player_key', (q) => q.eq('sport', 'mlb')),
]);
} catch (e) {
// Surfaced, not silent: a load failure must be distinguishable from a feed
// that genuinely holds nothing.
console.warn('[factors] context load FAILED (not an empty feed):', e.message);
return null;
}
if (!spray.length && !defense.length && !platoon.length) return null;
const sprayBy = latestBy(spray, (r) => r.player_key, (r) => r.as_of_date);
const defBy = latestBy(defense, (r) => r.team, (r) => r.as_of_date);
const platBy = latestBy(platoon, (r) => r.player_key, (r) => r.as_of_date);
const batBy = new Map();
const pitBy = new Map();
for (const r of statcast) {
if (!r.player_key) continue;
if (r.role === 'pitcher') pitBy.set(r.player_key, r);
else batBy.set(r.player_key, r);
}
// statcast stores PERCENTAGES (0-100); the factor wants a fraction.
const asFraction = (v) => {
const n = knownNumber(v);
if (n === null) return null;
return n > 1 ? n / 100 : n;
};
const resolver = (prop) => {
const key = nameKey(prop && prop.player);
if (!key) return null;
const bat = batBy.get(key);
const bats = bat && bat.bats ? String(bat.bats)[0] : null;
// The opposing team and its starter, from whatever the prop carries.
const oppName = prop && (prop.opponent || prop.opp_team || null);
const def = oppName ? (defBy.get(oppName) || defBy.get(String(oppName).split(' ').pop())) : null;
const pitKey = prop && prop.opposing_pitcher ? nameKey(prop.opposing_pitcher) : null;
const pit = pitKey ? pitBy.get(pitKey) : null;
const sp = platBy.get(key);
const splits = sp ? {
vl: { pa: sp.vl_pa, atBats: sp.vl_ab, hits: sp.vl_hits },
vr: { pa: sp.vr_pa, atBats: sp.vr_ab, hits: sp.vr_hits },
} : null;
const ctx = {
spray: sprayBy.get(key) || null,
positionOaa: def && def.position_oaa ? def.position_oaa : null,
bats,
throws: pit && pit.throws ? String(pit.throws)[0] : null,
pitcherHardHit: pit ? asFraction(pit.hard_hit_pct) : null,
platoonSplits: splits,
};
// Nothing readable at all -> null, so the engine skips the factor block
// entirely rather than walking an empty context.
const anything = ctx.spray || ctx.pitcherHardHit !== null || ctx.platoonSplits;
return anything ? ctx : null;
};
resolver.__stats = {
spray_players: sprayBy.size,
defense_teams: defBy.size,
platoon_players: platBy.size,
pitcher_profiles: pitBy.size,
batter_profiles: batBy.size,
};
return resolver;
}
module.exports = { build, MIN_ROWS };
+127
View File
@@ -0,0 +1,127 @@
'use strict';
/**
* hitsFactors — the three PROVEN hits factors, applied to the forecast.
*
* These passed the two-part gate (they move the prediction AND improve
* out-of-sample Brier) and then sat unwired: `sprayDefense.js` and
* `platoonSeverity.js` were required by nothing in `src/`, and the served
* `p_win` read four inputs, none of them these. They were computed downstream of
* the grade they should inform.
*
* ── PIPELINE ORDER IS A CORRECTNESS PROPERTY ─────────────────────────────
* base rate -> FACTORS -> CALIBRATE -> GRADE
* Calibration must always correct the factor-adjusted number. Reversing it would
* calibrate a forecast that is not the one served.
*
* ── EVERY UNREADABLE GUARD FROM THE ORIGINAL PROOFS SURVIVES ─────────────
* A factor applies only where it proved. A switch hitter has no readable spray
* side; a thin platoon split is refused rather than shrunk to a league guess; a
* pitcher with no contact profile contributes nothing. In each case the factor
* returns NULL and the forecast is left alone — never nudged toward a default,
* which would be fabricating a read from an absence.
*/
const sd = require('./sprayDefense');
const pss = require('./platoonSeverity');
const { knownNumber, knownRate } = require('../../utils/known');
/** League mean hard-hit rate allowed; the pitcher factor is signed off this. */
const LEAGUE_HARD_HIT = 0.389;
/** Bound on the pitcher-contact adjustment, as proved. */
const PITCHER_MAX = 0.15;
/** Bound on the composed adjustment — no stack of three may run away. */
const COMBINED_MAX = 0.25;
/**
* A contact-allowing arm concedes better contact. Null without a profile.
*/
function pitcherContactMultiplier(hardHitAllowed) {
const h = knownRate(hardHitAllowed);
if (h === null) return null;
return 1 + Math.max(-PITCHER_MAX, Math.min(PITCHER_MAX, (h - LEAGUE_HARD_HIT) * 1.2));
}
/**
* Compose the three factors for one hits prop.
*
* @param {object} ctx
* spray batter_spray row (pull/straight/oppo x gb/air)
* positionOaa opposing team's per-position OAA
* bats 'R' | 'L' | 'S'
* throws opposing starter's hand
* pitcherHardHit opposing starter's hard-hit rate allowed
* platoonSplits { vl, vr } for this hitter
* @returns {object} { multiplier, applied[], skipped[] } — multiplier is 1 when
* nothing is readable, which is a no-op rather than a claim.
*/
function hitsFactorMultiplier(ctx = {}) {
const applied = [];
const skipped = [];
let mult = 1;
// ── defense_by_direction ──
const spray = ctx.spray;
const posOaa = ctx.positionOaa;
if (!spray || !posOaa || !ctx.bats) {
skipped.push({ factor: 'defense_by_direction', reason: 'no spray profile or positional defence' });
} else {
const out = sd.sprayDefenseMultiplier({ spray, bats: ctx.bats, positionOaa: posOaa });
if (!out || !Number.isFinite(out.multiplier)) {
// Switch hitters land here: he bats opposite by choice, so the SIDE of the
// field his contact goes to is not determined pre-game.
skipped.push({ factor: 'defense_by_direction', reason: 'unreadable (switch hitter or no covered zone)' });
} else {
mult *= out.multiplier;
applied.push({ factor: 'defense_by_direction', multiplier: round4(out.multiplier), coverage: out.coverage });
}
}
// ── pitcher_contact_profile ──
const pm = pitcherContactMultiplier(ctx.pitcherHardHit);
if (pm === null) {
skipped.push({ factor: 'pitcher_contact_profile', reason: 'no pitcher contact profile' });
} else {
mult *= pm;
applied.push({ factor: 'pitcher_contact_profile', multiplier: round4(pm) });
}
// ── platoon_severity ──
if (!ctx.platoonSplits || !ctx.bats || !ctx.throws) {
skipped.push({ factor: 'platoon_severity', reason: 'no splits or no pitcher hand' });
} else {
const out = pss.platoonRead({ splits: ctx.platoonSplits, bats: ctx.bats, throws: ctx.throws });
if (!out || !out.readable || !Number.isFinite(out.multiplier)) {
// A thin split is REFUSED, not shrunk — a heavily-shrunk severity is
// indistinguishable from a measured league-average one.
skipped.push({ factor: 'platoon_severity', reason: (out && out.reason) || 'unreadable split' });
} else {
mult *= out.multiplier;
applied.push({ factor: 'platoon_severity', multiplier: round4(out.multiplier), split: out.observed_split });
}
}
const bounded = Math.max(1 - COMBINED_MAX, Math.min(1 + COMBINED_MAX, mult));
return {
multiplier: round4(bounded),
unbounded: round4(mult),
applied,
skipped,
factors_fired: applied.length,
};
}
/** Apply to a probability, clamped into a usable range. Null in, null out. */
function adjustProbability(p, ctx = {}) {
const raw = knownNumber(p);
if (raw === null) return { p_adjusted: null, ...hitsFactorMultiplier(ctx) };
const f = hitsFactorMultiplier(ctx);
return { p_adjusted: round4(Math.max(0.01, Math.min(0.99, raw * f.multiplier))), p_base: round4(raw), ...f };
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = {
hitsFactorMultiplier, adjustProbability, pitcherContactMultiplier,
LEAGUE_HARD_HIT, PITCHER_MAX, COMBINED_MAX,
};
+127
View File
@@ -0,0 +1,127 @@
'use strict';
/**
* lowParamCalibrator — a two-parameter favourite-longshot correction.
*
* Isotonic has one free parameter per distinct prediction level, which on 19
* dates is far more freedom than the sample can discipline — it can and does
* chase a single night's structure. Platt scaling has exactly TWO parameters
* over the whole curve:
*
* p_cal = sigmoid(a * logit(p) + b)
*
* `a < 1` flattens an over-confident forecaster toward the base rate, which is
* precisely the favourite-longshot shape measured here (over-prediction rising
* monotonically from -0.008 at p~0.55 to +0.245 above 0.9). Two parameters
* cannot represent "this Tuesday was odd", which is the entire point.
*
* ── SHRINKAGE TOWARD IDENTITY ────────────────────────────────────────────
* Even two parameters are fitted on few dates, so the correction is blended
* back toward the raw forecast by a weight tied to how many dates were seen:
*
* w = D / (D + D0)
* p_final = w * p_platt + (1 - w) * p_raw
*
* At 5 fit dates w = 0.33 — the correction is applied at a third of its fitted
* strength. At 40 dates it is 0.80. A thin-sample fit therefore cannot
* over-correct, and the blend is monotone because both inputs are.
*
* The DIRECTION of this correction is bootstrap-robust; its MAGNITUDE is
* thin-sample. Shrinkage is how that distinction is expressed in the number
* rather than only in a label.
*/
const { knownNumber } = require('../../utils/known');
/** Dates at which the fit earns half its weight. */
const SHRINK_HALF_DATES = 10;
/** Probabilities are clamped off 0/1 before the logit. */
const EPS = 1e-6;
/** Below this many rows there is nothing to fit. */
const MIN_FIT_ROWS = 100;
/**
* A slope at or below this is not a correction. Ordering must be preserved and
* the curve must not collapse to a constant.
*/
const MIN_SLOPE = 0.05;
const clamp01 = (p) => Math.min(1 - EPS, Math.max(EPS, p));
const logit = (p) => Math.log(clamp01(p) / (1 - clamp01(p)));
const sigmoid = (z) => 1 / (1 + Math.exp(-z));
/**
* Fit `a` and `b` by NewtonRaphson on the log-likelihood. Two parameters, so
* this converges in a handful of steps and has no tuning of its own.
*/
function fitPlatt(rows, opts = {}) {
const pts = (rows || [])
.map((r) => ({ x: logit(knownNumber(r.p)), y: knownNumber(r.won) }))
.filter((r) => Number.isFinite(r.x) && (r.y === 0 || r.y === 1));
if (pts.length < (opts.minRows ?? MIN_FIT_ROWS)) return null;
let a = 1; let b = 0;
for (let it = 0; it < 100; it += 1) {
let g0 = 0; let g1 = 0; let h00 = 0; let h01 = 0; let h11 = 0;
for (const { x, y } of pts) {
const p = sigmoid(a * x + b);
const e = p - y;
const w = p * (1 - p);
g0 += e * x; g1 += e;
h00 += w * x * x; h01 += w * x; h11 += w;
}
const det = h00 * h11 - h01 * h01;
if (!Number.isFinite(det) || Math.abs(det) < 1e-12) break;
const da = (g0 * h11 - g1 * h01) / det;
const db = (g1 * h00 - g0 * h01) / det;
a -= da; b -= db;
if (Math.abs(da) < 1e-10 && Math.abs(db) < 1e-10) break;
}
if (!Number.isFinite(a) || !Number.isFinite(b)) return null;
// ── THE SLOPE MUST CORRECT, NOT ABANDON ────────────────────────────────
// `a` in (0, 1] is a flattening: ordering preserved, over-confidence pulled
// in. `a <= 0` INVERTS the forecast, and `a` near zero collapses the curve to
// a constant — the fit has decided p_win carries nothing and is predicting the
// base rate for everything. That lowers Brier (shrinking a miscalibrated
// forecaster toward its base rate always does) while destroying resolution,
// so it would score as a win while making the product worthless.
//
// Measured: runs fitted a = -0.032. Refused here rather than deployed.
if (a <= (opts.minSlope ?? MIN_SLOPE)) {
return { refused: true, a: round5(a), b: round5(b), reason: a <= 0
? 'fitted slope is not positive — the correction would invert the forecast'
: 'fitted slope is near zero — the fit collapses to a constant and abandons the forecast' };
}
const dates = new Set((rows || []).map((r) => r.date).filter(Boolean)).size;
const half = opts.shrinkHalfDates ?? SHRINK_HALF_DATES;
const shrink = dates > 0 ? dates / (dates + half) : 0;
return {
a: round5(a),
b: round5(b),
fit_rows: pts.length,
fit_dates: dates,
shrinkage: round4(shrink),
/** Flattening a forecaster means a < 1; reported so the shape is checkable. */
flattens: a < 1,
};
}
/**
* Apply the fitted correction, shrunk toward the raw forecast.
* Returns null when unreadable — never a silently uncorrected number.
*/
function applyPlatt(model, p) {
const x = knownNumber(p);
if (!model || model.refused || x === null) return null;
const raw = clamp01(x);
const corrected = sigmoid(model.a * logit(raw) + model.b);
const w = model.shrinkage;
return round5(w * corrected + (1 - w) * raw);
}
const round5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = { fitPlatt, applyPlatt, SHRINK_HALF_DATES, MIN_FIT_ROWS, MIN_SLOPE, logit, sigmoid };
+102
View File
@@ -0,0 +1,102 @@
'use strict';
/**
* lowParamService — the production side of the two-parameter correction.
*
* Mirrors calibrationService's interface so the serving path swaps cleanly, but
* fits a Platt curve instead of an isotonic map. The reason for the swap is
* capacity, not score: on 19 dates we cannot certify the stability of a map with
* one free parameter per prediction level, and LODO turned out to have 1.4-9.3%
* power to tell us otherwise. Two parameters cannot encode "this Tuesday was
* odd", which is exactly the failure we cannot rule out for isotonic.
*
* Stated plainly because it is a judgement rather than a measurement: on the
* held-out window isotonic scored BETTER than this on hits (+0.0028) and rbi
* (+0.0042) and tied on total_bases. That window spans 2-4 date blocks, so it is
* weak evidence either way, and it is consistent with a flexible map having
* captured structure shared by fit and evaluation periods.
*
* Same point-in-time cut as before: fitted ONLY on games that are already over.
*/
const lp = require('./lowParamCalibrator');
const cal = require('./calibration');
const { knownNumber } = require('../../utils/known');
const MIN_FIT = 200;
const HOLDOUT_FRACTION = 0.35;
/** Build from settled rows: fit on the older part, certify bands on the newer. */
function build(rows, opts = {}) {
const clean = (rows || [])
.map((r) => ({ p: knownNumber(r.p), won: knownNumber(r.won), date: String(r.date || '') }))
.filter((r) => r.p !== null && (r.won === 0 || r.won === 1))
.sort((a, b) => a.date.localeCompare(b.date));
if (clean.length < (opts.minFit ?? MIN_FIT)) return null;
const cut = Math.floor(clean.length * (1 - (opts.holdout ?? HOLDOUT_FRACTION)));
const fitRows = clean.slice(0, cut);
const certRows = clean.slice(cut);
if (fitRows.length < (opts.minFit ?? MIN_FIT) || certRows.length < 50) return null;
const model = lp.fitPlatt(fitRows, opts);
// A refused fit (inverting or collapsed slope) yields no calibrator at all.
if (!model || model.refused) return null;
const corrected = certRows
.map((r) => ({ ...r, p: lp.applyPlatt(model, r.p) }))
.filter((r) => knownNumber(r.p) !== null);
const bands = cal.certifyBands(corrected, {
tolerance: opts.tolerance ?? 0.05,
minBin: opts.minBin ?? 40,
});
return {
model,
bands,
fit_n: fitRows.length,
certify_n: certRows.length,
fitted_through: fitRows[fitRows.length - 1].date,
shrinkage: model.shrinkage,
calibrate(p) {
const raw = knownNumber(p);
if (raw === null) return { p_raw: null, p_calibrated: null, calibrated: false, reason: 'absent' };
const c = lp.applyPlatt(model, raw);
if (c === null) return { p_raw: raw, p_calibrated: null, calibrated: false, reason: 'no_model_value' };
const inBand = cal.inCertifiedBand(bands, c);
return {
p_raw: raw,
p_calibrated: Math.round(c * 1000) / 1000,
calibrated: inBand,
reason: inBand ? null : 'outside_certified_band',
};
},
};
}
/** Load settled history and build, POINT-IN-TIME (strictly before today). */
async function fromLedger(sb, { sport = 'mlb', stat = 'hits', before = null, ...opts } = {}) {
if (!sb) return null;
const cutoff = before || new Intl.DateTimeFormat('en-CA', {
timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit',
}).format(new Date());
const rows = [];
for (let from = 0; ; from += 1000) {
const { data, error } = await sb.from('ledger_entries')
.select('p_win, outcome, game_date, quarantine_reason')
.eq('sport', sport).is('user_id', null).eq('stat', stat)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null)
.lt('game_date', cutoff)
.range(from, from + 999);
if (error || !data || data.length === 0) break;
rows.push(...data);
if (data.length < 1000) break;
}
const clean = rows
.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'))
.map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, date: String(r.game_date) }));
const built = build(clean, opts);
return built ? { ...built, cutoff } : null;
}
module.exports = { build, fromLedger, MIN_FIT, HOLDOUT_FRACTION };
+179
View File
@@ -0,0 +1,179 @@
'use strict';
/**
* parkWeather — PARK GEOMETRY AND AIR, READ ONTO HIT TYPE.
*
* The crude park factor is a single number per stadium ("Coors inflates offence
* 1.15x") applied to every hitter and every outcome alike. It fails for the same
* reason team-average defence failed: it is not the unit the causal story runs
* through. A deep left-centre gap does not create hits, it converts fly balls
* that would have been caught into DOUBLES, and it converts home runs into
* outs. Those move total bases in opposite directions, and one multiplier
* cannot express both.
*
* So this atom does not touch P(hit). It reshapes the HIT-TYPE distribution —
* single / double / triple / home run — and lets the total-bases convolution
* carry the consequence.
*
* ── WIND IS REFUSED, AND THAT IS THE POINT ───────────────────────────────
* Wind is the largest weather effect on carry, and we have the wind: Open-Meteo
* returns speed and compass bearing for every one of these games. What we do NOT
* have is park ORIENTATION — which compass direction each stadium's centre field
* faces. Without it, a 15 mph wind from 220° is unresolvable: it is blowing out
* to right at one park and straight in at another, and those are opposite
* predictions.
*
* The tempting move is to use wind SPEED alone as a magnitude of disruption.
* That is fabrication with a plausible face — it asserts an effect while
* discarding the sign that determines what the effect IS. Wind stays unreadable
* and says so, until orientation is a real column. `wind_readable: false` is the
* honest carrier of that.
*
* ── WHAT IS ACTUALLY READ ────────────────────────────────────────────────
* AIR DENSITY temperature and elevation. Both have unambiguous sign — warmer
* and higher is thinner air is more carry — and neither needs
* orientation to interpret. Under a closed roof, temperature is
* the building's, not the sky's, so it is neutralised.
* GEOMETRY each park against the league, per direction. Short lines make
* home runs; deep gaps make doubles and triples out of the same
* batted ball.
*/
const { knownNumber } = require('../../utils/known');
/** Bound on how far this atom may reshape any single hit-type share. */
const MAX_EFFECT = 0.15;
/** Reference conditions — the shares are calibrated to a temperate sea-level park. */
const REF_TEMP_F = 72;
const REF_ELEVATION_FT = 500;
/** Per-degree and per-1000ft carry response, applied to the home-run share. */
const CARRY_PER_DEG_F = 0.004;
const CARRY_PER_KFT = 0.030;
const isClosed = (roof) => /dome|closed|retractable/i.test(String(roof || ''));
/**
* League geometry, computed from the parks actually held rather than hardcoded,
* so it cannot drift away from the data it is compared against.
*/
function leagueGeometry(parks) {
const keys = ['left_line', 'left_center', 'center', 'right_center', 'right_line'];
const out = {};
for (const k of keys) {
const vals = (parks || []).map((p) => knownNumber(p[k])).filter((v) => v !== null);
out[k] = vals.length ? vals.reduce((a, b) => a + b, 0) / vals.length : null;
}
return out;
}
/**
* The read for one game.
*
* @param {object} dims a park_dimensions row
* @param {object} wx { wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg }
* @param {object} league output of leagueGeometry
* @returns {object|null} null when there is nothing readable — never a 1.0 that
* looks measured.
*/
function parkWeatherRead({ dims, wx, league } = {}) {
if (!dims || !league) return null;
const closed = isClosed(dims.roof_type);
const temp = knownNumber(wx && wx.wx_temp_f);
const elev = knownNumber(dims.elevation);
// ── AIR ────────────────────────────────────────────────────────────────
// Under a closed roof the outside temperature is not the air the ball flies
// through, so it contributes nothing rather than contributing zero.
let carry = 0;
const airParts = [];
if (!closed && temp !== null) {
carry += (temp - REF_TEMP_F) * CARRY_PER_DEG_F;
airParts.push('temperature');
}
if (elev !== null) {
carry += ((elev - REF_ELEVATION_FT) / 1000) * CARRY_PER_KFT;
airParts.push('elevation');
}
// ── GEOMETRY ───────────────────────────────────────────────────────────
// Lines govern home runs; gaps and centre govern extra bases on balls that
// stay in the park. Deeper than league = fewer home runs, more doubles.
const rel = (k) => {
const v = knownNumber(dims[k]); const l = knownNumber(league[k]);
return v !== null && l !== null && l > 0 ? (v - l) / l : null;
};
const lines = [rel('left_line'), rel('right_line')].filter((v) => v !== null);
const gaps = [rel('left_center'), rel('right_center'), rel('center')].filter((v) => v !== null);
const lineDepth = lines.length ? lines.reduce((a, b) => a + b, 0) / lines.length : null;
const gapDepth = gaps.length ? gaps.reduce((a, b) => a + b, 0) / gaps.length : null;
if (lineDepth === null && gapDepth === null && !airParts.length) return null;
const clamp = (v) => Math.max(-MAX_EFFECT, Math.min(MAX_EFFECT, v));
// Deep lines suppress home runs; thin air and heat restore them.
const hr = clamp(carry - (lineDepth ?? 0) * 1.2);
// Deep gaps turn caught fly balls into doubles and the occasional triple.
const dbl = clamp((gapDepth ?? 0) * 0.8 - carry * 0.3);
const tpl = clamp((gapDepth ?? 0) * 1.5);
// Singles are the residual: what the ball did instead of clearing the fence.
const sgl = clamp(-(hr * 0.25 + dbl * 0.35));
return {
readable: true,
multipliers: {
single: round4(1 + sgl),
double: round4(1 + dbl),
triple: round4(1 + tpl),
home_run: round4(1 + hr),
},
carry: round4(carry),
line_depth_vs_league: round4(lineDepth),
gap_depth_vs_league: round4(gapDepth),
roof_closed: closed,
air_inputs: airParts,
// Stated on every read so a consumer cannot mistake silence for neutrality.
wind_readable: false,
wind_reason: 'park orientation unknown — a bearing cannot be resolved to out or in',
};
}
/** A checkable sentence, or nothing. */
function explain(read, parkName) {
if (!read || !read.readable) return null;
const m = read.multipliers;
const bits = [];
if (read.line_depth_vs_league !== null) {
bits.push(`lines ${read.line_depth_vs_league >= 0 ? 'deeper' : 'shorter'} than league`);
}
if (read.gap_depth_vs_league !== null) {
bits.push(`gaps ${read.gap_depth_vs_league >= 0 ? 'deeper' : 'shorter'}`);
}
if (read.air_inputs.length) bits.push(`air via ${read.air_inputs.join(' and ')}`);
return `${parkName || 'this park'}${bits.join(', ')}; home runs x${m.home_run}, doubles x${m.double}`
+ (read.roof_closed ? ' (roof closed, outside temperature not applied)' : '');
}
/** Reshape a hit-type share vector, renormalised so it stays a distribution. */
function applyToShares(shares, read) {
if (!shares || !read || !read.readable) return shares || null;
const m = read.multipliers;
const out = {
single: (knownNumber(shares.single) ?? 0) * m.single,
double: (knownNumber(shares.double) ?? 0) * m.double,
triple: (knownNumber(shares.triple) ?? 0) * m.triple,
home_run: (knownNumber(shares.home_run) ?? 0) * m.home_run,
};
const sum = out.single + out.double + out.triple + out.home_run;
if (!(sum > 0)) return shares;
for (const k of Object.keys(out)) out[k] = round4(out[k] / sum);
return out;
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = {
parkWeatherRead, leagueGeometry, applyToShares, explain,
MAX_EFFECT, REF_TEMP_F, REF_ELEVATION_FT, CARRY_PER_DEG_F, CARRY_PER_KFT,
};
+101
View File
@@ -0,0 +1,101 @@
'use strict';
/**
* penQuality — the PROVEN half of Link 2.
*
* Link 2 was asked twice. Naming the individual reliever failed on merit (17.2%
* accuracy — wrong five times in six), because managers mix and match and the
* individual genuinely is noise. Asked at the COARSE grain the chain actually
* needs, it proves: predicted pen quality separates a realized 2.70pp hit-rate
* difference between the pens we call best and worst.
*
* The archetype grain did NOT prove (0.567 vs a 0.531 modal-guess baseline,
* corrected interval spanning zero) and is deliberately absent from this module.
* Two grains were tested; one earned a place.
*
* ── POINT-IN-TIME ON BOTH SIDES ──────────────────────────────────────────
* An arm's quality is his allowed-hit-rate over appearances strictly BEFORE the
* game in question, and a club's pen forecast comes only from its prior games.
* The target is which KNOWN-quality arms appeared — never how they happened to
* pitch that night, which would be scoring against the answer.
*
* ── ABSTAIN, NEVER IMPUTE ────────────────────────────────────────────────
* An arm below the appearance floor has no readable quality, and a club without
* enough prior games has no readable pen. Both return null. A league-average
* stand-in would assert "this is an ordinary bullpen", which is a claim, and
* usually the wrong one for exactly the clubs whose pens have just turned over.
*/
const { knownNumber } = require('../../utils/known');
/** Appearances before an arm's quality is readable at all. */
const MIN_ARM_PA = 40;
/** Prior games before a club's pen is readable at all. */
const MIN_PRIOR_GAMES = 5;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/**
* One arm's quality from his prior line. Null below the floor — a 12-batter
* sample is not a scouting report.
*/
function armQuality(prior) {
const n = knownNumber(prior && prior.pa);
const h = knownNumber(prior && prior.hits);
if (n === null || h === null || n < MIN_ARM_PA) return null;
return h / n;
}
/**
* The pen a hitter's later plate appearances will face.
*
* @param {Array} priorGames [{ quality }] this club's prior relief outings
* @returns {object|null} null when unreadable — never a league-average guess.
*/
function projectPen(priorGames) {
const qs = (priorGames || []).map((g) => knownNumber(g && g.quality)).filter((v) => v !== null);
if (qs.length < MIN_PRIOR_GAMES) {
return null;
}
return {
readable: true,
quality: round4(mean(qs)),
games_read: qs.length,
// Stated so a consumer cannot mistake this for a reliever-identity claim.
grain: 'pen_quality',
individual_arm_refused: 'naming the specific reliever did not prove (17.2% accuracy) — managers mix and match',
archetype_refused: 'the archetype grain did not prove at the corrected bar',
};
}
/**
* The measured relationship between pen quality and hit rate, for a consumer
* that wants the consequence rather than the input. Anchored on the observed
* league mean; the slope is the measured tercile separation, not a fitted
* parameter, and the effect is bounded because it was measured over a range.
*/
const LEAGUE_PEN_QUALITY = 0.2261;
const HIT_RATE_PER_QUALITY = 0.87; // 2.70pp realized over a 0.031 quality gap
const MAX_SHIFT = 0.03;
function hitRateShift(penQuality) {
const q = knownNumber(penQuality);
if (q === null) return null; // absent stays absent
const raw = (q - LEAGUE_PEN_QUALITY) * HIT_RATE_PER_QUALITY;
return round4(Math.max(-MAX_SHIFT, Math.min(MAX_SHIFT, raw)));
}
/** A checkable sentence, or nothing. */
function explain(pen) {
if (!pen || !pen.readable) return null;
const d = pen.quality - LEAGUE_PEN_QUALITY;
if (Math.abs(d) < 0.005) return `bullpen reads league-average over ${pen.games_read} prior games`;
return `bullpen reads ${d > 0 ? 'weaker' : 'stronger'} than league over ${pen.games_read} prior games`;
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = {
armQuality, projectPen, hitRateShift, explain,
MIN_ARM_PA, MIN_PRIOR_GAMES, LEAGUE_PEN_QUALITY, MAX_SHIFT,
};
+160
View File
@@ -0,0 +1,160 @@
'use strict';
/**
* predictionGate — the two-part gate for a CONTINUOUS prediction.
*
* factorGate answers this for probabilities, where the loss is Brier. A link in
* the reliever chain predicts a quantity — how many batters the starter faces,
* which arm throws the fourth plate appearance — so the loss is squared error or
* a hit rate, and the binarisation factorGate applies to outcomes would silently
* destroy the target.
*
* The discipline is identical and deliberately so:
*
* (a) MOVEMENT the prediction differs from the naive baseline at all
* (b) IMPROVEMENT paired bootstrap on the loss difference, interval excluding
* zero at the cumulative-corrected level
*
* A link that predicts the league average very precisely has learned nothing, and
* without (a) it would pass (b) by tying. That is the same THEATER failure the
* probability gate exists to name, wearing different units.
*
* ── WHY CLUSTERING MATTERS HERE TOO ──────────────────────────────────────
* The same starter appears many times in a season, so his starts are not
* independent readings: a pitcher the model happens to fit well contributes a
* run of correlated wins. Clustering on the pitcher makes the interval reflect
* how many ARMS we have read, not how many starts we have counted.
*/
const { knownNumber } = require('../../utils/known');
const MIN_N = 500;
const MIN_CLUSTERS = 40;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Absolute error, the honest default for a quantity a user would read. */
const absLoss = (pred, actual) => Math.abs(pred - actual);
/** Squared error, when large misses should dominate. */
const sqLoss = (pred, actual) => (pred - actual) ** 2;
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
/**
* @param {Array} rows [{ baseline, prediction, actual, cluster? }]
* @param {object} opts { loss, iters, seed, cumulativeTests, minN, minClusters }
*/
function adjudicate(rows, opts = {}) {
const loss = opts.loss === 'squared' ? sqLoss : absLoss;
const usable = (rows || []).filter((r) =>
knownNumber(r && r.baseline) !== null
&& knownNumber(r && r.prediction) !== null
&& knownNumber(r && r.actual) !== null);
const n = usable.length;
const shifts = usable.map((r) => Math.abs(r.prediction - r.baseline));
const movement = {
n,
mean_abs_shift: n ? round4(mean(shifts)) : null,
max_abs_shift: n ? round4(Math.max(...shifts)) : null,
};
const minN = opts.minN ?? MIN_N;
const minClusters = opts.minClusters ?? MIN_CLUSTERS;
const base = { link: opts.link || null, movement };
if (n < minN) {
return { ...base, verdict: 'PENDING_SAMPLE', reason: `n ${n} < ${minN}`, rows_needed: minN - n };
}
const clustered = usable.some((r) => r.cluster != null);
const groups = new Map();
if (clustered) {
for (const r of usable) {
const k = String(r.cluster);
if (!groups.has(k)) groups.set(k, []);
groups.get(k).push(r);
}
}
const keys = clustered ? [...groups.keys()] : null;
if (clustered && keys.length < minClusters) {
return {
...base,
verdict: 'PENDING_SAMPLE',
reason: `${n} rows but only ${keys.length} independent clusters < ${minClusters}`,
clusters_needed: minClusters - keys.length,
};
}
const lossBase = mean(usable.map((r) => loss(r.baseline, r.actual)));
const lossPred = mean(usable.map((r) => loss(r.prediction, r.actual)));
const delta = lossPred - lossBase; // negative = the link is better
const rnd = makeRnd(opts.seed ?? 20260806);
const iters = opts.iters ?? 3000;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const b = []; const p = [];
if (clustered) {
for (let i = 0; i < keys.length; i += 1) {
for (const r of groups.get(keys[Math.floor(rnd() * keys.length)])) {
b.push(loss(r.baseline, r.actual)); p.push(loss(r.prediction, r.actual));
}
}
} else {
for (let i = 0; i < n; i += 1) {
const r = usable[Math.floor(rnd() * n)];
b.push(loss(r.baseline, r.actual)); p.push(loss(r.prediction, r.actual));
}
}
diffs.push(mean(p) - mean(b));
}
diffs.sort((x, y) => x - y);
const tests = Math.max(1, Math.round(knownNumber(opts.cumulativeTests) ?? 1));
const alpha = 0.05 / tests;
const q = (x) => round4(diffs[Math.floor(Math.min(diffs.length - 1, Math.max(0, x * (diffs.length - 1))))]);
const ci = [q(alpha / 2), q(1 - alpha / 2)];
const improvement = {
n,
effective_n: clustered ? keys.length : n,
cluster_unit: clustered ? 'cluster' : 'row',
loss_baseline: round4(lossBase),
loss_prediction: round4(lossPred),
loss_delta: round4(delta),
ci,
ci_level: round4(1 - alpha),
bonferroni_tests: tests,
improves: ci[1] < 0,
degrades: ci[0] > 0,
};
const out = { ...base, improvement };
if (movement.mean_abs_shift === null || movement.mean_abs_shift < (opts.minMovement ?? 0)) {
return { ...out, verdict: 'INERT', reason: 'the link never departs from the naive baseline' };
}
if (improvement.improves) {
return { ...out, verdict: 'PROVES', reason: `beats the naive baseline by ${-improvement.loss_delta} (CI ${JSON.stringify(ci)} at ${improvement.ci_level}, corrected for ${tests} tests)` };
}
if (delta < 0) {
return {
...out,
verdict: 'NOT_PROVEN_AT_CORRECTED_BAR',
reason: `point estimate improves by ${-improvement.loss_delta} but the corrected interval spans zero (${JSON.stringify(ci)})`,
note: 'a real candidate held to a rising bar — not theatre',
};
}
return {
...out,
verdict: 'THEATER',
reason: `moves ${movement.mean_abs_shift} off the baseline while accuracy does NOT improve (delta ${improvement.loss_delta})`,
consequence: 'wiring this would make the projection LOOK like it read the game script while reading nothing',
};
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
module.exports = { adjudicate, absLoss, sqLoss, MIN_N, MIN_CLUSTERS };
+63
View File
@@ -0,0 +1,63 @@
'use strict';
/**
* reAuditEligibility — which settled rows may be measured on.
*
* The champion was repaired on 2026-08-07: it had been reading ten games as its
* season rate. Everything measured before that ran against a forecaster that
* lost to a frequency table, so calibration maps fitted on those rows correct
* toward a bias the current forecast may not have, and every factor verdict was
* scored against a sub-trivial baseline.
*
* The temptation is to reconstruct the repaired forecast over old rows and
* measure on that. It is REFUSED: a reconstruction is not what was served, and
* scoring a served product against a simulation of itself is the same class of
* error as scoring a map on the window it was fitted to.
*
* So eligibility is mechanical — a row qualifies only if the snapshot that
* produced it carries the repaired champion's version marker.
*/
const { REPAIRED_CHAMPION_VERSION } = require('../retentionService');
/**
* Minimum settled DATES before each forward measurement is honest.
*
* Not row counts: the binding scarcity all session has been dates, and every
* interval that mattered was date-clustered. Stated here so the thresholds
* cannot drift toward whichever answer arrives first.
*/
const ACCRUAL = Object.freeze({
calibration_refit: 10, // isotonic/low-param need a fit AND a held-out window
hits_factor_lift: 10, // a date-block CI on a composed lift
prior_verdict_reaudit: 14, // re-running gates that previously returned nulls
rbi_lineup_slot_gate: 14, // a fresh two-part gate on a new factor
});
/** Was this row produced by the repaired champion? */
function isEligible(row) {
if (!row) return false;
return String(row.model_version || '') === REPAIRED_CHAMPION_VERSION;
}
/**
* @returns {object} { eligible, dates, ready:{...}, blocked_reason }
* `ready` is per-measurement, so one can unblock before another.
*/
function assess(rows) {
const eligible = (rows || []).filter(isEligible);
const dates = new Set(eligible.map((r) => String(r.game_date || ''))).size;
const ready = Object.fromEntries(Object.entries(ACCRUAL)
.map(([k, need]) => [k, { need, have: dates, ready: dates >= need }]));
return {
eligible_rows: eligible.length,
total_rows: (rows || []).length,
eligible_dates: dates,
ready,
blocked_reason: dates === 0
? 'no settled rows yet carry the repaired champion marker — nothing may be measured'
: null,
};
}
module.exports = { isEligible, assess, ACCRUAL, REPAIRED_CHAMPION_VERSION };
+142
View File
@@ -0,0 +1,142 @@
'use strict';
/**
* servedGrade — the letter a user sees, derived from the forecast.
*
* ── WHY THIS EXISTS ──────────────────────────────────────────────────────
* The letter being served came from engine1's additive factor index, computed
* independently of `p_win`. Measured on 3,417 settled props:
*
* grade n realized mean p_win
* A 8 0.500 0.647 <- the TOP grade did worst
* B 985 0.640 0.700
* C 1,695 0.602 0.676
* D 303 0.558 0.604
* F 426 0.535 0.588
*
* letter resolution 0.00116 (0.48% of variance)
* p_win resolution 0.00715 (2.98%)
*
* The letter carried ONE SIXTH the information of the number printed beside it,
* and its best grade hit worse than its worst. So the grade is derived from
* `p_win` here instead — not because p_win is good (2.98% is not good) but
* because shipping the weaker of two available signals as the headline is
* indefensible.
*
* ── NO MANUFACTURED A ────────────────────────────────────────────────────
* S91 established the honest ceiling: once the numbers are truthful this model
* has no 80%-plus reads. The realized rate PLATEAUS around 0.65-0.68 from p_win
* 0.70 upward — the 0.9+ bucket does no better than the 0.8 bucket.
*
* So A and A+ are NOT ISSUABLE. Not "rare" — structurally absent, because no
* band of this forecast has ever realized a rate that would justify one. A test
* asserts that no input produces an A. When resolution improves enough for a
* band to earn it, the ceiling is raised deliberately and visibly, not by a
* threshold quietly drifting.
*
* ── EVERY LETTER CARRIES ITS OWN MEANING ─────────────────────────────────
* The band's realized rate travels with the grade, so the surface can state what
* a B actually means rather than implying a spread the model does not have.
*/
const { knownNumber } = require('../../utils/known');
/**
* Bands over p_win, anchored on MEASURED realized rates (3,417 settled props,
* pooled across the four batter stats).
*
* The realized column is what this band has actually done — not a target, not a
* projection. `separates` says whether the band's realized rate is
* distinguishable from the pooled base rate; today only the extremes are.
*/
const BASE_RATE = 0.6005;
const BANDS = Object.freeze([
{ letter: 'B+', min: 0.780, realized: 0.663, separates: true,
meaning: 'the strongest read this model produces — realized about 66%' },
{ letter: 'B', min: 0.700, realized: 0.646, separates: true,
meaning: 'above this profile\'s base rate — realized about 65%' },
{ letter: 'C+', min: 0.640, realized: 0.615, separates: false,
meaning: 'slightly above base rate, not distinguishable from it' },
{ letter: 'C', min: 0.560, realized: 0.589, separates: false,
meaning: 'a base-rate read — the model sees nothing that separates this' },
{ letter: 'C-', min: 0.480, realized: 0.548, separates: false,
meaning: 'at or below base rate' },
{ letter: 'D', min: 0.350, realized: 0.512, separates: true,
meaning: 'below base rate — the model reads this as weak' },
{ letter: 'F', min: 0.000, realized: 0.447, separates: true,
meaning: 'well below base rate' },
]);
/** Letters this forecast cannot justify. Absent by construction, not by rarity. */
const UNISSUABLE = Object.freeze(['A+', 'A', 'A-']);
/**
* The grade for one served prop.
*
* @param {object} prop { p_win, refused, refusal_reason, factor_adjustment }
* @returns {object} always a renderable state — never null, never blank.
*/
function gradeFor(prop = {}) {
// ── REFUSAL IS A REAL STATE, NOT A BLANK ──
if (prop.refused || prop.insufficient_data) {
return {
letter: null,
state: 'refused',
label: 'NO READ',
meaning: prop.refusal_reason === 'juiced_no_edge'
? 'the book has priced the vig past any edge on this side'
: 'not enough history to call this one',
basis: 'refusal',
separates: false,
};
}
const p = knownNumber(prop.p_win);
if (p === null) {
return {
letter: null,
state: 'no_forecast',
label: 'NO READ',
meaning: 'no forecast could be produced for this prop',
basis: 'absent',
separates: false,
};
}
const band = BANDS.find((b) => p >= b.min) || BANDS[BANDS.length - 1];
const factored = Array.isArray(prop.factor_adjustment && prop.factor_adjustment.applied)
&& prop.factor_adjustment.applied.length > 0;
return {
letter: band.letter,
state: 'graded',
label: band.letter,
meaning: band.meaning,
band_realized_rate: band.realized,
separates_from_base_rate: band.separates,
base_rate: BASE_RATE,
// What the letter was computed from, stated so the surface cannot imply more.
basis: factored ? 'forecast_plus_matchup_factors' : 'forecast_only',
factors_applied: factored ? prop.factor_adjustment.applied.map((a) => a.factor) : [],
// Calibration is withdrawn; nothing here rides on a calibrated number.
calibrated: false,
};
}
/**
* The scale legend — the ceiling stated as a POSITION, not left as an absence.
*
* Without this a user is left to wonder why they never see an A, and the most
* natural guess ("the model is being coy" or "the slate is bad tonight") is
* wrong. The honest answer is that top grades are earned from realized outcomes
* and this forecast has not earned one.
*/
const SCALE_LEGEND = Object.freeze({
headline: 'Grades are earned from realized outcomes, not issued on confidence.',
ceiling: `Our honest ceiling right now is a strong B+ — those reads have landed about ${Math.round(BANDS[0].realized * 100)}% of the time against a ${Math.round(BASE_RATE * 100)}% baseline.`,
no_a: 'We do not issue A grades. No band of this model has hit at a rate that would justify one, and we would rather show you the ceiling than invent a letter above it.',
base_rate_note: 'Grades marked "base-rate read" are ones we cannot separate from the baseline. That is most of any slate, and saying so is the point.',
when_a_returns: 'If the model earns an A, this legend changes and we will say why.',
});
module.exports = { gradeFor, BANDS, UNISSUABLE, BASE_RATE, SCALE_LEGEND };
+169
View File
@@ -0,0 +1,169 @@
'use strict';
/**
* E10 — THE VYNDR REPORT, the designed issue template.
*
* The gap audit had this as PARTIAL, not absent: `newsletterService` already
* builds the daily report's CONTENT and lints its voice. What was missing is the
* designed HYBRID SHELL. So this is a template over that builder, not a second
* report — the same compose-don't-fork call made for the movement strip.
*
* ── THE SPEC'S STRUCTURAL LAW ────────────────────────────────────────────
* "Hybrid: dark billboard header that survives every client, light paper body
* Gmail can't wreck. 600px, stacked, no webfont dependence."
*
* That is an engineering constraint, not a look. Gmail strips `<style>` blocks,
* Outlook ignores flexbox, and a dark body renders as a black rectangle in
* several clients. Hence: tables, inline styles only, 600px fixed, system-font
* stacks, and no image required to read the issue.
*
* Its content law is equally short: *"One email per slate day. Top read, what
* changed, the record. Nothing else."*
*
* ── FACT-CONTRACTED ──────────────────────────────────────────────────────
* Same discipline as the content engine: a section whose data is absent is
* OMITTED, never filled. There is no code path producing a placeholder figure,
* and the designer's sample values (Nabers 1,120.5, Nº 128, 9-4) are a spec for
* what a live issue renders — never content to paste.
*/
const { knownNumber } = require('../../utils/known');
/** Design tokens, inlined because email clients drop stylesheets. */
const T = Object.freeze({
billboard: '#06060B', // dark header — survives every client
paper: '#F7F6F2', // light body — Gmail cannot wreck it
ink: '#1A1A22',
greenOnDark: '#00D4A0',
greenOnPaper: '#00A57D', // the green SHIFTS on paper for contrast
rule: '#D8D5CC',
dim: '#6B6B76',
mono: "'SFMono-Regular',Consolas,'Liberation Mono',Menlo,monospace",
sans: "-apple-system,BlinkMacSystemFont,'Segoe UI',Helvetica,Arial,sans-serif",
});
const WIDTH = 600;
const esc = (s) => String(s == null ? '' : s)
.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;').replace(/"/g, '&quot;');
/** Present means renderable. Zero is present; empty and null are not. */
const has = (v) => v !== null && v !== undefined && !(typeof v === 'string' && v.trim() === '')
&& !(Array.isArray(v) && v.length === 0) && !(typeof v === 'number' && !Number.isFinite(v));
/**
* The movement line — E1's law expressed in email-safe text.
*
* A `<canvas>` or SVG strip cannot be relied on in email, so the primitive's
* RULE travels even though its rendering cannot: green only when the move
* favours the read, and a flat market says FLAT rather than showing nothing.
*/
function movementText(m) {
if (!m || !has(m.from) || !has(m.to)) return null;
const from = knownNumber(m.from);
const to = knownNumber(m.to);
if (from === null || to === null) return null;
if (from === to) return { text: `FLAT${has(m.days) ? ` · ${m.days}D` : ''}`, colour: T.dim };
const favours = m.dir === 'toward';
return {
text: `${from}${to}`,
colour: favours ? T.greenOnPaper : '#B0762A', // amber-on-paper for against
};
}
const row = (inner) => `<tr><td style="padding:0 28px;">${inner}</td></tr>`;
/**
* @param {object} issue
* number, date_label, read_time
* top_read { grade, subject, line_text, movement:{from,to,dir,days}, model, best_book, note }
* changed [{ time, tag, text }]
* record { line, hit, miss, pct|null, note|null }
* honesty { graded, cleared_ceiling, ceiling_letter, ceiling_realized, base_rate, unissuable }
* @returns {object} { html, text, omitted[] } — omitted names what had no data.
*/
function renderIssue(issue = {}) {
const omitted = [];
const parts = [];
// ── DARK BILLBOARD HEADER ──
parts.push(`<tr><td style="background:${T.billboard};padding:26px 28px;">
<div style="font-family:${T.mono};font-size:20px;font-weight:800;letter-spacing:2px;color:#FFFFFF;">VYND<span style="color:${T.greenOnDark};">R</span></div>
<div style="font-family:${T.mono};font-size:11px;letter-spacing:2px;color:#8A8A96;padding-top:8px;">
THE REPORT${has(issue.number) ? ` &middot; N&ordm; ${esc(issue.number)}` : ''}${has(issue.date_label) ? ` &middot; ${esc(issue.date_label)}` : ''}${has(issue.read_time) ? ` &middot; READ TIME ${esc(issue.read_time)}` : ''}
</div></td></tr>`);
// ── TOP READ OF THE DAY ──
const tr = issue.top_read;
if (tr && has(tr.subject) && has(tr.grade)) {
const mv = movementText(tr.movement);
parts.push(row(`<div style="padding:22px 0 0;">
<div style="font-family:${T.mono};font-size:10px;letter-spacing:2px;color:${T.dim};">TOP READ OF THE DAY</div>
<div style="font-family:${T.sans};font-size:19px;font-weight:700;color:${T.ink};padding-top:10px;">
<span style="font-family:${T.mono};color:${T.greenOnPaper};font-weight:800;">${esc(tr.grade)}</span>
&nbsp;${esc(tr.subject)}${has(tr.line_text) ? ` <span style="font-weight:400;">${esc(tr.line_text)}</span>` : ''}
</div>
${mv || has(tr.model) || has(tr.best_book) ? `<div style="font-family:${T.mono};font-size:12px;color:${T.dim};padding-top:8px;">
${mv ? `<span style="color:${mv.colour};font-weight:700;">${esc(mv.text)}</span>` : ''}
${has(tr.model) ? ` &middot; VYNDR ${esc(tr.model)}` : ''}
${has(tr.best_book) ? ` &middot; BEST: ${esc(tr.best_book)}` : ''}
</div>` : ''}
${has(tr.note) ? `<p style="font-family:${T.sans};font-size:14px;line-height:1.6;color:${T.ink};padding-top:12px;margin:0;">${esc(tr.note)}</p>` : ''}
</div>`));
} else omitted.push('top_read');
// ── WHAT CHANGED ──
if (Array.isArray(issue.changed) && issue.changed.length) {
const rows = issue.changed.map((c) => `<div style="padding:7px 0;border-top:1px solid ${T.rule};">
<span style="font-family:${T.mono};font-size:11px;color:${T.dim};">${esc(c.time)}</span>
<span style="font-family:${T.mono};font-size:11px;font-weight:800;color:${T.ink};">&nbsp;${esc(String(c.tag).toUpperCase())}</span>
<span style="font-family:${T.sans};font-size:13px;color:${T.ink};">&nbsp;${esc(c.text)}</span></div>`).join('');
parts.push(row(`<div style="padding:26px 0 0;">
<div style="font-family:${T.mono};font-size:10px;letter-spacing:2px;color:${T.dim};padding-bottom:6px;">WHAT CHANGED</div>${rows}</div>`));
} else omitted.push('changed');
// ── THE RECORD ── (dark band: survives every client, per spec)
const rec = issue.record;
if (rec && has(rec.line)) {
parts.push(`<tr><td style="background:${T.billboard};padding:18px 28px;margin-top:20px;">
<div style="font-family:${T.mono};font-size:10px;letter-spacing:2px;color:#8A8A96;">THE RECORD</div>
<div style="font-family:${T.mono};font-size:16px;font-weight:800;color:#FFFFFF;padding-top:6px;">${esc(rec.line)}</div>
${has(rec.note) ? `<div style="font-family:${T.mono};font-size:11px;color:#8A8A96;padding-top:6px;">${esc(rec.note)}</div>` : ''}
</td></tr>`);
} else omitted.push('record');
// ── THE HONESTY BLOCK ── real figures or nothing.
const h = issue.honesty;
if (h && has(h.graded) && has(h.ceiling_letter)) {
parts.push(row(`<div style="padding:22px 0 0;">
<div style="font-family:${T.mono};font-size:10px;letter-spacing:2px;color:${T.dim};">HONESTLY</div>
<p style="font-family:${T.sans};font-size:13px;line-height:1.65;color:${T.ink};padding-top:8px;margin:0;">
We graded ${esc(h.graded)} props${has(h.cleared_ceiling) ? ` and ${esc(h.cleared_ceiling)} cleared ${esc(h.ceiling_letter)}` : ''}.
${has(h.ceiling_realized) && has(h.base_rate) ? `Those reads land about ${esc(h.ceiling_realized)}% against a ${esc(h.base_rate)}% baseline. ` : ''}
${has(h.unissuable) ? `We do not issue ${esc(h.unissuable)} — no band of this model has hit at a rate that would justify one.` : ''}
</p></div>`));
} else omitted.push('honesty');
const html = `<!doctype html><html><body style="margin:0;padding:0;background:${T.paper};">
<table role="presentation" width="100%" cellpadding="0" cellspacing="0" style="background:${T.paper};">
<tr><td align="center">
<table role="presentation" width="${WIDTH}" cellpadding="0" cellspacing="0" style="width:${WIDTH}px;max-width:${WIDTH}px;background:${T.paper};">
${parts.join('\n')}
<tr><td style="padding:22px 28px 30px;">
<div style="font-family:${T.mono};font-size:10px;color:${T.dim};line-height:1.7;">
One email per slate day. Top read, what changed, the record. Nothing else.<br>
No outcome is promised. 21+. <a href="{{ UnsubscribeURL }}" style="color:${T.dim};">Unsubscribe</a>.
</div></td></tr>
</table></td></tr></table></body></html>`;
const text = [
`THE REPORT${has(issue.number) ? ` No ${issue.number}` : ''}${has(issue.date_label) ? `${issue.date_label}` : ''}`,
tr && has(tr.subject) ? `\nTOP READ: ${tr.grade} ${tr.subject}${has(tr.line_text) ? ` ${tr.line_text}` : ''}` : '',
tr && has(tr.note) ? tr.note : '',
rec && has(rec.line) ? `\nRECORD: ${rec.line}` : '',
'\nOne email per slate day. Top read, what changed, the record. Nothing else.',
].filter(Boolean).join('\n');
return { html, text, omitted, width: WIDTH };
}
module.exports = { renderIssue, movementText, T, WIDTH, has };
+18 -1
View File
@@ -42,7 +42,23 @@ function etDateOf(iso) {
* Bump when the grading model changes in a way that makes rows non-comparable.
* This is the marker `ledger_entries` never had.
*/
const MODEL_VERSION = process.env.MODEL_VERSION || 'engine1@2026-07-20';
/**
* CHAMPION VERSION — the eligibility marker for every forward re-audit.
*
* Bumped when the forecaster itself changes, so a settled row is
* self-identifying: rows tagged `engine1@2026-08-07-fullwindow` were produced by
* the REPAIRED champion (full season log, recency weight 0.20); anything earlier
* came from the retired ten-game forecaster.
*
* This is what makes the re-audit rule mechanical rather than a promise.
* Calibration may only be re-fit, and factor verdicts may only be re-audited, on
* rows carrying the current marker — never on reconstructions of a retired
* forecast, and never on a mixture of the two, which is the trap that would
* otherwise be invisible once both generations sit in the same table.
*/
const MODEL_VERSION = process.env.MODEL_VERSION || 'engine1@2026-08-07-fullwindow';
/** Rows at or after this marker are eligible for forward re-audit. */
const REPAIRED_CHAMPION_VERSION = 'engine1@2026-08-07-fullwindow';
function codeSha() {
return process.env.SOURCE_COMMIT || process.env.GIT_SHA || process.env.COOLIFY_GIT_COMMIT_SHA || null;
@@ -229,6 +245,7 @@ function newSnapshotId() {
module.exports = {
MODEL_VERSION,
REPAIRED_CHAMPION_VERSION,
codeSha,
rowsFromSides,
createCollector,
+127 -24
View File
@@ -275,6 +275,54 @@ async function loadPitcherArsenals(sport) {
} catch { return out; }
}
/**
* Stats that may serve a calibrated number, and on what basis.
*
* ── THE BIAS IS ROBUST; THE MAP WAS NOT CERTIFIABLE ──────────────────────
* Tested model-free and map-free on the picked-side population: the model
* over-predicts its own favourites, and the sign survives 99.5% of date-block
* resamples pooled and replicates in 4 of 4 stats. Over-prediction rises
* monotonically from -0.008 near p=0.55 to +0.245 above 0.9.
*
* What could NOT be certified on 19 dates is the stability of a specific
* isotonic MAP — LODO has 1.4-9.3% power there. So isotonic is retired and the
* correction is a TWO-PARAMETER Platt curve, which has no capacity to encode a
* single odd day, shrunk toward the raw forecast by fit-date count.
*
* Validated as a NEW estimator against RAW, date-block bootstrap:
* hits a=0.406 shrink 0.565 0.2626 -> 0.2540 CI [-0.0112,-0.0069]
* total_bases a=0.472 shrink 0.333 0.2490 -> 0.2429 CI [-0.0062,-0.0059]
*
* rbi is WITHDRAWN (deployed last order on isotonic): the low-parameter fit does
* not beat raw, CI [-0.0007, 0] touching zero. runs is refused by the slope
* guard — it fitted a = -0.032, which would invert the forecast rather than
* flatten it. Both now serve RAW.
*/
const CALIBRATION_DEPLOYED = Object.freeze([]);
/**
* NOTHING IS SERVED CALIBRATED, AND THIS IS DELIBERATE.
*
* The low-parameter maps were fitted on the OLD forecast — the one whose base
* rate was a ten-game frequency. That distribution no longer exists: the
* champion now reads the full season log at a 0.20 recency weight, which tripled
* its resolution on hits (0.00251 -> 0.00817) and roughly doubled it on
* total_bases and runs.
*
* A calibration map applied to a forecast it was not fitted on is the stale-fit
* trap this session has already been caught by once, and it corrects toward a
* bias the new forecast may not have. `fromLedger` cannot rescue it either: the
* settled ledger rows still carry OLD p_win values, so refitting today would fit
* the retired forecast again.
*
* So calibration is OFF until enough dates settle under the repaired champion to
* refit honestly, and the favourite-longshot bias must be re-measured on the new
* forecast rather than assumed to have survived. Serving the raw repaired number
* is the honest state, not a regression.
*
* The shadow duel is likewise void — it accumulated against the old forecast.
*/
const CALIBRATION_BASIS = Object.freeze({});
async function runSnapshot(sport, opts = {}) {
const sp = String(sport || '').toLowerCase();
const deps = {
@@ -417,7 +465,26 @@ async function runSnapshot(sport, opts = {}) {
// Grade the slate via the existing service; capture the envelope instead of
// letting it write (we re-write an ENRICHED version below).
let envelope = null;
// ── PROVEN FACTORS, LOADED BEFORE THE GRADE ─────────────────────────────
// This ordering IS the fix. The same inputs were previously read at line 640+,
// downstream of the grade they should inform, so three proven factors never
// once moved a served number. Best-effort: a failed load means the slate is
// graded unadjusted, exactly as before.
let factorContext = null;
if (sp === 'mlb') {
try {
const ctxSvc = deps.hitsFactorContext || require('./model/hitsFactorContext');
const sbc = require('../utils/supabase').getSupabaseServiceClient();
factorContext = sbc ? await ctxSvc.build(sbc) : null;
if (factorContext) console.log(`[factors] ${sp} hits context loaded — ${JSON.stringify(factorContext.__stats)}`);
else console.log(`[factors] ${sp} — no factor context; grading unadjusted`);
} catch (e) {
console.warn('[factors] context skipped:', e.message);
}
}
await deps.gradeAndCacheSlate(sp, props, {
factorContext,
// Bisect hook (2026-08-01): lets the internal trigger run a bounded slate
// without a prod env change, so a cap regression can be isolated by
// measurement instead of guessed at. Omitted => gradeSlateService's own
@@ -700,37 +767,71 @@ async function runSnapshot(sport, opts = {}) {
console.warn(`[challenger] ${sp} skipped:`, e.message);
}
// ── FORWARD CALIBRATION (hits) ────────────────────────────────────────
// ── FORWARD CALIBRATION + THE SHADOW DUEL ─────────────────────────────
// Fitted on games that are OVER, applied to tonight's props. `p_win` is NOT
// touched — the counter stays byte-identical and the calibrated value rides
// touched — the counter stays byte-identical and the corrected value rides
// beside it, because a calibration map is a correction TO a forecast, not a
// different forecast.
//
// `calibrated` is true only inside a band certified out-of-sample, and it is
// what `chain.chainAcross` requires before it will compound anything. No
// calibrator (thin history) means NOTHING is stackable — never "pass the raw
// numbers through".
// WHAT IS SERVED, AND WHY IT IS A BET RATHER THAN A RESULT. On identical
// held-out rows the ISOTONIC map scored BETTER than the low-parameter one
// (hits +0.0028, rbi +0.0042, total_bases tied). We serve the low-parameter
// map anyway, on the argument that isotonic's in-window edge is daily
// structure shared between the fit and evaluation windows and will not
// transmit forward. At 19 dates no instrument here can test that argument —
// LODO has 1.4-9.3% power — so it is a BET, not evidence.
//
// So both are computed on every prop and the shadow is logged. Real
// out-of-window dates adjudicate it:
//
// PRE-REGISTERED: once >=10 forward dates have settled that NEITHER map was
// fitted on, if isotonic beats low-param with a date-block bootstrap CI
// excluding zero, the capacity argument is REFUTED and hits/TB revert to
// isotonic. If low-param wins or ties, the bet was right. The season
// decides, not the argument.
//
// Serving is unchanged until that bar is met. `calibrated` is true only inside
// a band certified out-of-sample, and it is what `chain.chainAcross` requires
// before it will compound anything.
if (sp === 'mlb') {
try {
const calSvc = deps.calibrationService || require('./model/calibrationService');
const sbc = require('../utils/supabase').getSupabaseServiceClient();
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat: 'hits' }) : null;
if (calibrator) {
let marked = 0;
for (const g of enriched) {
if (String(g.stat_type || g.stat || '').toLowerCase() !== 'hits') continue;
const out = calibrator.calibrate(g.p_win);
g.p_win_calibrated = out.p_calibrated;
g.calibrated = out.calibrated;
g.calibration_reason = out.reason;
if (out.calibrated) marked += 1;
for (const stat of CALIBRATION_DEPLOYED) {
try {
const calSvc = deps.calibrationService || require('./model/lowParamService');
const shadowSvc = deps.shadowCalibrationService || require('./model/calibrationService');
const sbc = require('../utils/supabase').getSupabaseServiceClient();
const calibrator = sbc ? await calSvc.fromLedger(sbc, { sport: 'mlb', stat }) : null;
// The shadow must never break serving: its own try, and a null shadow
// simply means the duel has no entry for tonight.
let shadow = null;
try { shadow = sbc ? await shadowSvc.fromLedger(sbc, { sport: 'mlb', stat }) : null; } catch { shadow = null; }
if (calibrator) {
let marked = 0; let shadowed = 0;
for (const g of enriched) {
if (String(g.stat_type || g.stat || '').toLowerCase() !== stat) continue;
const out = calibrator.calibrate(g.p_win);
g.p_win_calibrated = out.p_calibrated;
g.p_win_lowparam = out.p_calibrated; // named, so the duel is legible
g.calibrated = out.calibrated;
g.calibration_reason = out.reason;
g.calibration_status = 'provisional';
g.calibration_basis = CALIBRATION_BASIS[stat] || null;
if (out.calibrated) marked += 1;
if (shadow) {
const sh = shadow.calibrate(g.p_win);
// SHADOW ONLY. Never read by serving, never by chainAcross.
g.p_win_isotonic_shadow = sh.p_calibrated;
g.calibration_duel_fitted_through = shadow.fitted_through || null;
if (sh.p_calibrated != null) shadowed += 1;
}
}
console.log(`[calibration] ${sp} ${stat} (low-param, PROVISIONAL) — ${marked} stackable; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}, a=${calibrator.model.a} shrink=${calibrator.shrinkage}; shadow logged on ${shadowed}`);
} else {
console.log(`[calibration] ${sp} ${stat} — no calibrator (thin history); nothing is stackable`);
}
console.log(`[calibration] ${sp} hits — ${marked} stackable of ${enriched.filter((g) => String(g.stat_type || g.stat || '').toLowerCase() === 'hits').length}; fit n=${calibrator.fit_n} through ${calibrator.fitted_through}, bands ${JSON.stringify(calibrator.bands.map((b) => [b.lo, b.hi]))}`);
} else {
console.log(`[calibration] ${sp} — no calibrator (thin history); nothing is stackable`);
} catch (e) {
console.warn(`[calibration] ${stat} skipped:`, e.message);
}
} catch (e) {
console.warn('[calibration] skipped:', e.message);
}
}
@@ -843,5 +944,7 @@ module.exports = {
generateTickerEvents,
pushTickerItems,
ACTIVE_SPORTS,
CALIBRATION_DEPLOYED,
CALIBRATION_BASIS,
__internals: { propKey, gradedAtFor, indexOdds, lastName, isTopGrade, DELTA_NOISE, DELTA_MOVE, TICKER_CAP },
};
+21 -6
View File
@@ -1,3 +1,13 @@
/**
* NOTE ON `engine_grade` (2026-08-07 grade cutover).
*
* The user-facing `grade` now derives from `p_win`, not engine1's additive
* factor index -- that index carried 0.16x the information of p_win on 3,417
* settled props and its A hit worse than its F. The engine's OWN decision
* (graded vs suppressed) is preserved as `engine_grade`, so behaviour
* assertions read that; suppression assertions still read `grade`, since a
* suppressed prop has no letter either way.
*/
// Fix 2 (Session 7f) — verifies the end-to-end shape from
// computeFeaturesForProp → engine1 → adapter → concrete reasoning.
@@ -46,8 +56,13 @@ describe('analyzeViaEngine1 — happy path', () => {
});
// Adapter-collapsed grade.
expect(out.grade).toBe('A');
expect(out.confidence).toBe(78);
expect(out.engine_grade).toBe('A');
// Confidence now derives from p_win (0.95 -> 95), not from a grade-band
// midpoint of the old letter. The old 78 was the midpoint for engine1's "A"
// -- a number that carried no information beyond the letter it was looked up
// from, and which would now contradict the served B+.
expect(out.confidence).toBe(95);
expect(out.confidence_basis).toBe('p_win');
expect(out.player).toBe('Jalen Brunson');
expect(out.stat_type).toBe('points');
expect(out.line).toBe(25.5);
@@ -104,7 +119,7 @@ describe('analyzeViaEngine1 — happy path', () => {
player: 'P', stat_type: 'points', line: 24.5, direction: 'over', sport: 'nba',
});
expect(out.grade).toBe('D');
expect(out.engine_grade).toBe('D');
expect(out.reasoning.summary).toContain('Playing on the road');
expect(out.reasoning.summary).toContain('OKC');
expect(out.reasoning.summary).toContain('top-tier defense');
@@ -185,7 +200,7 @@ describe('analyzeViaEngine1 — graceful degradation', () => {
player: 'P', stat_type: 'points', line: 25, direction: 'over', sport: 'nba',
});
expect(out.grade).toBe('B');
expect(out.engine_grade).toBe('B');
expect(out.insufficient_data).toBeUndefined();
expect(out.projection).toBe(28.4); // the REAL model reference, never the line
});
@@ -205,7 +220,7 @@ describe('analyzeViaEngine1 — interface verifications', () => {
player: 'X', stat_type: 'points', line: 25, direction: 'over', sport: 'nba',
});
expect(out).toBeDefined();
expect(out.grade).toBeDefined();
expect(out.engine_grade).toBeDefined();
});
test('every legacy field DemoScan reads is present', async () => {
@@ -222,7 +237,7 @@ describe('analyzeViaEngine1 — interface verifications', () => {
});
// DemoScan reads: grade, confidence, reasoning.summary,
// kill_conditions_triggered[].code, edge_pct, line, player, stat_type.
expect(out.grade).toBeDefined();
expect(out.engine_grade).toBeDefined();
expect(typeof out.confidence).toBe('number');
expect(typeof out.reasoning.summary).toBe('string');
expect(Array.isArray(out.kill_conditions_triggered)).toBe(true);
+106
View File
@@ -0,0 +1,106 @@
'use strict';
/**
* THE WINDOW GUARD — a fixed N must never stand in for a season.
*
* This class has now cost six paths across two sports:
*
* MLB getStatRows res.last10 -> a 10-game "season rate"
* MLB mlbGameLogFeatures res.last10 -> l20_avg was 10 games
* NBA espnStatsAdapter rows.slice(0, 20) -> every downstream rate capped at 20
* NBA nbaGameLogFeatures vals.slice(0, 20) -> l20_avg was 20 games
* NBA getStatRows (python) getGameLogs(.., 20) -> a 20-game base rate
* NBA gameLogFeatures getGameLogs(.., 20) -> same
*
* Every one looked correct in isolation. `slice(0, 20)` is unremarkable code;
* what made it a defect was the QUESTION it was answering — "what is this
* player's season rate?" — and no test could see that mismatch, because the
* value it produced was always a plausible number.
*
* So this asserts the property directly on the source: the base-rate and
* season-reference paths must not narrow to a fixed window. Basketball is
* OFFLINE, which is exactly why it is guarded now — a regression there would
* otherwise surface only when it had already served grades.
*/
const fs = require('fs');
const path = require('path');
const ROOT = path.join(__dirname, '..', '..');
const read = (rel) => fs.readFileSync(path.join(ROOT, rel), 'utf8');
/** Body of a named function, for targeted assertions. */
function fnBody(src, name) {
const i = src.indexOf(`function ${name}(`);
if (i < 0) return null;
let depth = 0; let started = false;
for (let j = i; j < src.length; j += 1) {
if (src[j] === '{') { depth += 1; started = true; }
else if (src[j] === '}') { depth -= 1; if (started && depth === 0) return src.slice(i, j + 1); }
}
return null;
}
describe('no fixed window may stand in for a season', () => {
const featureCache = read('src/services/intelligence/featureCache.js');
it('the season log depth is a named constant past a full season', () => {
// A literal 20 buried in a call is how every instance of this bug looked.
expect(featureCache).toMatch(/const SEASON_LOG_DEPTH = (\d+);/);
const depth = Number(/const SEASON_LOG_DEPTH = (\d+);/.exec(featureCache)[1]);
expect(depth).toBeGreaterThanOrEqual(82); // a basketball season
});
it('no game-log request asks for a hardcoded small count', () => {
const calls = [...featureCache.matchAll(/getGameLogs\([^)]*\)/g)].map((m) => m[0]);
expect(calls.length).toBeGreaterThan(0);
for (const c of calls) {
const literal = /,\s*(\d+)\s*\)/.exec(c);
if (literal) expect(Number(literal[1])).toBeGreaterThanOrEqual(82);
}
});
it('the BASKETBALL season reference (l20_avg) is not a 20-game slice', () => {
const body = fnBody(featureCache, 'nbaGameLogFeatures');
expect(body).not.toBeNull();
// The exact pre-fix line. It read as ordinary code and was the defect.
expect(body).not.toMatch(/m20\s*=\s*avg\(vals\.slice\(0,\s*20\)\)/);
expect(body).toMatch(/m20\s*=\s*avg\(vals\)/);
});
it('the MLB season reference is a real season aggregate, not a slice', () => {
const body = fnBody(featureCache, 'mlbGameLogFeatures');
expect(body).not.toBeNull();
expect(body).toMatch(/l20_avg\s*=\s*seasonTotal\s*\/\s*games/);
});
it('the MLB base-rate path reads the full log, not last10', () => {
const body = fnBody(featureCache, 'getStatRows');
expect(body).toMatch(/res\.fullLog/);
});
it('the ESPN adapter does not cap the parsed game log', () => {
const adapter = read('src/services/adapters/espnStatsAdapter.js');
const body = fnBody(adapter, 'parseGameLog');
expect(body).not.toBeNull();
// `return rows.slice(0, 20)` capped every downstream basketball rate.
expect(body).not.toMatch(/return\s+rows\.slice\(/);
});
});
describe('the guard would have caught the real defects', () => {
it('flags the exact pre-fix basketball slice', () => {
const preFix = 'const m20 = avg(vals.slice(0, 20));';
expect(/m20\s*=\s*avg\(vals\.slice\(0,\s*20\)\)/.test(preFix)).toBe(true);
});
it('flags the exact pre-fix ESPN cap', () => {
expect(/return\s+rows\.slice\(/.test(' return rows.slice(0, 20);')).toBe(true);
});
it('flags a hardcoded small game-log request', () => {
const preFix = 'await gameLogs.getGameLogs(playerName, sp, 20);';
const literal = /,\s*(\d+)\s*\)/.exec(/getGameLogs\([^)]*\)/.exec(preFix)[0]);
expect(Number(literal[1])).toBeLessThan(82); // i.e. would fail the guard
});
});
+47
View File
@@ -0,0 +1,47 @@
'use strict';
/**
* Which stats serve a calibrated number, and on what basis.
*
* The claim that survived on this sample is the bias DIRECTION, tested model-free
* and map-free. What did not survive is any certification of a specific map's
* stability. These lock that distinction into the serving path.
*/
const snapshotService = require('../../src/services/snapshotService');
const lp = require('../../src/services/model/lowParamCalibrator');
describe('the deploy set rides a low-parameter correction, not isotonic', () => {
it('serves NOTHING while the maps are stale against the repaired champion', () => {
// hits and total_bases were deployed at 74cf1ce on maps fitted to the OLD
// forecast, whose base rate was a ten-game frequency. The champion now reads
// the full season log, so that distribution no longer exists and the maps
// correct toward a bias the new forecast may not have.
expect(snapshotService.CALIBRATION_DEPLOYED).toEqual([]);
});
it('WITHDRAWS rbi — the low-parameter fit does not beat raw', () => {
// rbi was deployed at ced4042 on isotonic. Its CI vs raw is [-0.0007, 0],
// which touches zero, so it serves raw again.
expect(snapshotService.CALIBRATION_DEPLOYED).not.toContain('rbi');
});
it('does NOT serve runs — its fitted slope would invert the forecast', () => {
expect(snapshotService.CALIBRATION_DEPLOYED).not.toContain('runs');
});
it('carries no basis claim while nothing is deployed', () => {
expect(snapshotService.CALIBRATION_BASIS).toEqual({});
});
it('the served correction cannot encode a single odd day', () => {
// Two parameters over the whole curve is the entire reason for the swap.
expect(typeof lp.fitPlatt).toBe('function');
expect(lp.MIN_SLOPE).toBeGreaterThan(0);
});
it('is frozen, so a stat cannot be added at runtime', () => {
expect(Object.isFrozen(snapshotService.CALIBRATION_DEPLOYED)).toBe(true);
expect(() => { snapshotService.CALIBRATION_DEPLOYED.push('hits'); }).toThrow();
});
});
+80
View File
@@ -0,0 +1,80 @@
'use strict';
/**
* The forward adjudication of a bet made against the measurement.
*
* We serve the map that scored WORSE in-window, on an argument the sample
* cannot test. These lock the rule that decides whether that argument survives —
* written before any forward date exists, so the bar cannot drift toward
* whichever answer arrives.
*/
const duel = require('../../src/services/model/calibrationDuel');
/** Forward rows where `edge` favours the shadow when positive. */
function rows(dates, perDate, edge, seed = 2) {
let s = seed;
const rnd = () => (s = (s * 1103515245 + 12345) % 2147483648) / 2147483648;
const out = [];
for (let d = 0; d < dates; d += 1) {
for (let i = 0; i < perDate; i += 1) {
const won = rnd() < 0.6 ? 1 : 0;
const err = 0.25 + rnd() * 0.1;
out.push({
date: `2026-09-${String(d + 1).padStart(2, '0')}`,
fitted_through: '2026-08-31',
won,
served: won ? 1 - err : err,
shadow: won ? 1 - err + edge : err - edge,
});
}
}
return out;
}
describe('the rule is pre-registered and cannot be met early', () => {
it('is PENDING below the forward-date bar, whatever the numbers say', () => {
// Even with the shadow winning decisively, 5 dates is not a verdict.
const v = duel.adjudicate(rows(5, 40, 0.15));
expect(v.verdict).toBe('PENDING');
expect(v.dates_needed).toBe(duel.MIN_FORWARD_DATES - 5);
expect(v.action).toMatch(/keep serving/);
});
it('REFUTES the bet when isotonic wins out-of-window at the bar', () => {
const v = duel.adjudicate(rows(12, 40, 0.15));
expect(v.verdict).toBe('REFUTED');
expect(v.ci[1]).toBeLessThan(0);
expect(v.action).toMatch(/REVERT/);
});
it('UPHOLDS the bet when the served map is not beaten', () => {
const v = duel.adjudicate(rows(12, 40, -0.15));
expect(v.verdict).toBe('UPHELD');
expect(v.action).toMatch(/keep serving/);
});
it('UPHOLDS on a tie — the burden is on refutation, not on us', () => {
const v = duel.adjudicate(rows(12, 40, 0));
expect(v.verdict).toBe('UPHELD');
});
});
describe('only genuinely out-of-window dates count', () => {
it('drops rows inside the fit window — that would score memorisation', () => {
const inWindow = rows(12, 40, 0.15).map((r) => ({ ...r, fitted_through: '2026-12-31' }));
const v = duel.adjudicate(inWindow);
expect(v.verdict).toBe('PENDING');
expect(v.forward_dates).toBe(0);
});
it('drops rows with no fit provenance rather than assuming they are forward', () => {
const noProv = rows(12, 40, 0.15).map(({ fitted_through, ...r }) => r);
expect(duel.adjudicate(noProv).forward_dates).toBe(0);
});
it('drops rows missing either map — a duel needs both entrants', () => {
const half = rows(12, 40, 0.15).map((r) => ({ ...r, shadow: null }));
expect(duel.adjudicate(half).forward_dates).toBe(0);
});
});
+112
View File
@@ -0,0 +1,112 @@
'use strict';
/**
* The two ways a calibration measurement lies.
*
* Both produced a confident, plausible, completely wrong number in the
* settlement session, and neither was visible in the output. These lock them out.
*/
const g = require('../../src/services/model/calibrationGuards');
const cal = require('../../src/services/model/calibration');
/** A population carrying BOTH sides of each prop, as the snapshot table does. */
function bothSides(n) {
const rows = [];
for (let i = 0; i < n; i += 1) {
const p = 0.55 + (i % 7) * 0.05;
rows.push({ propKey: `prop${i}`, side: 'over', p });
rows.push({ propKey: `prop${i}`, side: 'under', p: 1 - p });
}
return rows;
}
describe('GUARD 1 — the both-sides tell', () => {
it('catches the 0.4998 signature: both sides present AND mean pinned at 0.5', () => {
const rows = bothSides(200);
const r = g.checkPickedSideDedup(rows);
expect(r.violated).toBe(true);
expect(r.both_sides_share).toBe(1);
expect(Math.abs(r.mean_p - 0.5)).toBeLessThanOrEqual(g.BALANCED_TOLERANCE);
expect(r.reason).toMatch(/balanced by construction/);
});
it('assert form REFUSES rather than returning a number', () => {
expect(() => g.assertPickedSideDedup(bothSides(100))).toThrow(/CALIBRATION GUARD/);
});
it('passes once deduped to the model-picked side', () => {
// The picked side is the one the model favoured, so the mean sits well
// above 0.5 — which is what a real forecaster's book looks like.
const picked = bothSides(200).filter((r) => r.p > 0.5);
const r = g.checkPickedSideDedup(picked);
expect(r.violated).toBe(false);
expect(r.mean_p).toBeGreaterThan(0.5 + g.BALANCED_TOLERANCE);
});
it('does NOT fire on a genuinely balanced one-sided book', () => {
// Either condition alone is unremarkable. A book of one-sided picks that
// happens to average 0.5 is honest, and flagging it would be a false alarm.
const rows = Array.from({ length: 300 }, (_, i) => ({
propKey: `p${i}`, side: 'over', p: i % 2 ? 0.45 : 0.55,
}));
const r = g.checkPickedSideDedup(rows);
expect(r.both_sides_props).toBe(0);
expect(r.violated).toBe(false);
});
it('does NOT fire when both sides are present but the mean is skewed', () => {
const rows = bothSides(50).concat(
Array.from({ length: 400 }, (_, i) => ({ propKey: `x${i}`, side: 'over', p: 0.8 })));
const r = g.checkPickedSideDedup(rows);
expect(r.both_sides_props).toBeGreaterThan(0);
expect(r.violated).toBe(false); // already deduped elsewhere
});
});
describe('GUARD 2 — a null must never score itself', () => {
it('(null-1)**2 can no longer pass as a metric', () => {
// This is the exact breach: JS scores null as 1 against a win and 0 against
// a loss, so the "Brier" silently equals the win rate.
const outcomes = [1, 1, 0, 1, 0];
const naive = outcomes.reduce((s, y, i) => s + ((null - y) ** 2), 0) / outcomes.length;
const winRate = outcomes.reduce((a, b) => a + b, 0) / outcomes.length;
expect(naive).toBeCloseTo(winRate, 10); // the trap, demonstrated
expect(g.safeBrier([null, null, null, null, null], outcomes)).toBeNull();
});
it('refuses when ANY single prediction is null', () => {
expect(g.safeBrier([0.6, 0.4, null], [1, 0, 1])).toBeNull();
});
it('can be made to hard-fail instead of refusing', () => {
expect(() => g.safeBrier([0.6, null], [1, 0], { onNull: 'throw' }))
.toThrow(/null prediction reached a Brier term/);
});
it('scores normally when every prediction is real', () => {
expect(g.safeBrier([1, 0], [1, 0])).toBe(0);
expect(g.safeBrier([0.5, 0.5], [1, 0])).toBeCloseTo(0.25, 10);
});
it('an unfittable map refuses instead of producing null predictions', () => {
// fitIsotonic returns null below its minimum; this is what must happen next.
const map = cal.fitIsotonic([{ p: 0.6, won: 1 }, { p: 0.4, won: 0 }]);
expect(map).toBeNull();
const out = g.applyOrRefuse(map, [{ p: 0.6 }], cal.applyIsotonic);
expect(out.ok).toBe(false);
expect(out.reason).toMatch(/no calibration map/);
expect(out.rows).toEqual([]);
});
it('drops unmappable rows rather than passing nulls downstream', () => {
const fit = [];
for (let i = 0; i < 400; i += 1) fit.push({ p: 0.3 + (i % 60) / 100, won: i % 3 === 0 ? 1 : 0 });
const map = cal.fitIsotonic(fit);
expect(map).not.toBeNull();
const out = g.applyOrRefuse(map, [{ p: 0.5 }, { p: null }], cal.applyIsotonic);
expect(out.rows.length).toBe(1);
expect(out.dropped).toBe(1);
});
});
+176
View File
@@ -0,0 +1,176 @@
'use strict';
/**
* Which stats may serve a calibrated number.
*
* The thing these protect is the meaning of PROVISIONAL: a provisional deploy
* that cannot be taken away is just a deploy.
*/
const { createRegistry, STATUS, PROMOTION_DATE_CLUSTERS } = require('../../src/services/model/calibrationRegistry');
const MAP = [{ lo: 0.5, hi: 0.7, value: 0.55, n: 300 }];
const GOOD = { lodo_pass: true, ci: [-0.0061, -0.0045], map: MAP, certified_bands: [[0.6, 0.8]], date_clusters: 7, at: '2026-08-06' };
describe('deploy needs BOTH gates', () => {
it('deploys when LODO passes and the interval excludes zero', () => {
const r = createRegistry();
expect(r.deploy('total_bases', GOOD).status).toBe(STATUS.PROVISIONAL);
});
it('refuses on a LODO failure however good the interval', () => {
const r = createRegistry();
const out = r.deploy('runs', { ...GOOD, lodo_pass: false });
expect(out.ok).toBe(false);
expect(out.reason).toMatch(/date-driven/);
});
it('refuses when the interval spans zero however clean the LODO', () => {
const r = createRegistry();
const out = r.deploy('hits', { ...GOOD, ci: [-0.01, 0.002] });
expect(out.ok).toBe(false);
expect(out.reason).toMatch(/does not exclude zero/);
});
it('refuses without a map — there is nothing to serve', () => {
const r = createRegistry();
expect(r.deploy('hits', { ...GOOD, map: null }).ok).toBe(false);
});
});
describe('auto-demotion is what makes provisional honest', () => {
it('demotes on the first date where the interval stops excluding zero', () => {
const r = createRegistry();
r.deploy('total_bases', GOOD);
const out = r.reverify('total_bases', { ci: [-0.004, 0.001], date: '2026-08-07' });
expect(out.status).toBe(STATUS.NONE);
expect(out.reason).toBe('ci_no_longer_excludes_zero');
expect(out.breaking_date).toBe('2026-08-07');
expect(r.serves('total_bases', 0.65).serve).toBe(false);
});
it('demotes when the favourite over-prediction flips sign', () => {
// A flip means the correction is now pushing the wrong way.
const r = createRegistry();
r.deploy('total_bases', GOOD);
const out = r.reverify('total_bases', { ci: [-0.006, -0.004], favourite_bias: -0.03, date: '2026-08-08' });
expect(out.status).toBe(STATUS.NONE);
expect(out.reason).toBe('favourite_bias_flipped');
});
it('logs the demotion with its breaking date', () => {
const r = createRegistry();
r.deploy('total_bases', GOOD);
r.reverify('total_bases', { ci: [0.001, 0.004], date: '2026-08-09' });
const ev = r.log().find((e) => e.event === 'auto_demoted');
expect(ev).toMatchObject({ stat: 'total_bases', at: '2026-08-09' });
});
it('stays deployed while both conditions hold', () => {
const r = createRegistry();
r.deploy('total_bases', GOOD);
const out = r.reverify('total_bases', { ci: [-0.007, -0.003], favourite_bias: 0.17, date: '2026-08-07' });
expect(out.status).toBe(STATUS.PROVISIONAL);
expect(out.changed).toBe(false);
});
});
describe('the >=40 date-cluster bar is the PROMOTION bar, not the deploy bar', () => {
it('does not block deployment', () => {
const r = createRegistry();
expect(r.deploy('total_bases', { ...GOOD, date_clusters: 7 }).ok).toBe(true);
});
it('promotes out of provisional once it is met', () => {
const r = createRegistry();
r.deploy('total_bases', GOOD);
const out = r.reverify('total_bases', { ci: [-0.006, -0.004], date_clusters: PROMOTION_DATE_CLUSTERS, date: '2026-09-15' });
expect(out.status).toBe(STATUS.PROMOTED);
});
it('does not promote while the interval has stopped holding', () => {
const r = createRegistry();
r.deploy('total_bases', GOOD);
const out = r.reverify('total_bases', { ci: [-0.001, 0.003], date_clusters: 60, date: '2026-09-15' });
expect(out.status).toBe(STATUS.NONE);
});
});
describe('serving is band-limited', () => {
it('serves inside the certified band and refuses outside it', () => {
const r = createRegistry();
r.deploy('total_bases', GOOD);
expect(r.serves('total_bases', 0.65).serve).toBe(true);
expect(r.serves('total_bases', 0.65).provisional).toBe(true);
expect(r.serves('total_bases', 0.95).serve).toBe(false);
expect(r.serves('total_bases', 0.95).reason).toMatch(/outside the certified band/);
});
it('an undeployed stat never serves', () => {
const r = createRegistry();
expect(r.serves('hits', 0.6).serve).toBe(false);
expect(r.serves('hits', 0.6).reason).toBe('not deployed');
});
it('a missing p_win serves nothing', () => {
const r = createRegistry();
r.deploy('total_bases', GOOD);
expect(r.serves('total_bases', null).serve).toBe(false);
});
});
describe('the LODO test is a COHERENT pair, not a bar plus an unrelated rule', () => {
const { LODO_K, LODO_TEST, LODO_POWER_FLOOR } = require('../../src/services/model/calibrationRegistry');
const normCdf = (z) => {
const t = 1 / (1 + 0.2316419 * Math.abs(z));
const d = 0.3989422804014327 * Math.exp(-z * z / 2);
const p = d * t * (0.319381530 + t * (-0.356563782 + t * (1.781477937 + t * (-1.821255978 + t * 1.330274429))));
return z >= 0 ? 1 - p : p;
};
const binomPmf = (n, k, p) => {
let logC = 0;
for (let i = 0; i < k; i += 1) logC += Math.log(n - i) - Math.log(i + 1);
return Math.exp(logC + k * Math.log(p) + (n - k) * Math.log(1 - p));
};
const tail = (n, c, p) => { let s = 0; for (let k = c + 1; k <= n; k += 1) s += binomPmf(n, k, p); return s; };
it('each n* is what its own (sigma_row, g) produce — no pooled value', () => {
for (const [stat, t] of Object.entries(LODO_TEST)) {
expect(Math.ceil(LODO_K ** 2 * (t.sigma_row / Math.abs(t.g)) ** 2)).toBe(t.n_star);
}
// And the four differ, which is exactly why one pooled number mis-credited them.
const stars = Object.values(LODO_TEST).map((t) => t.n_star);
expect(new Set(stars).size).toBeGreaterThan(1);
});
it('each cutoff is the smallest one holding the false-positive rate at 0.05', () => {
const p = normCdf(-LODO_K);
for (const [stat, t] of Object.entries(LODO_TEST)) {
expect(tail(t.informative_drops, t.cutoff, p)).toBeLessThanOrEqual(0.05);
if (t.cutoff > 0) expect(tail(t.informative_drops, t.cutoff - 1, p)).toBeGreaterThan(0.05);
expect(t.fp).toBeCloseTo(tail(t.informative_drops, t.cutoff, p), 3);
}
});
it('the OLD rule is demonstrably incoherent — it failed stable stats ~half the time', () => {
// Zero-reversal rule at a 1-SE bar, on four informative drops.
const pNoise = normCdf(-1);
const falseFail = 1 - (1 - pNoise) ** 4;
expect(falseFail).toBeGreaterThan(0.45);
expect(falseFail).toBeLessThan(0.55);
});
it('every stat falls below the power floor, so none may claim LODO stability', () => {
// A gate that cannot fail is not a gate. This is the honest state at this
// date count, and the floor makes it structural rather than a footnote.
for (const [stat, t] of Object.entries(LODO_TEST)) {
expect(t.power).toBeLessThan(LODO_POWER_FLOOR);
}
});
it('the stale pooled threshold is nulled so nothing can read it', () => {
const { LODO_MIN_HELD_ROWS } = require('../../src/services/model/calibrationRegistry');
expect(LODO_MIN_HELD_ROWS).toBeNull();
});
});
+55
View File
@@ -0,0 +1,55 @@
'use strict';
/**
* The champion's forecast window, and what depends on it.
*
* The defect: `estimateProbability` builds its base rate as the frequency over
* every row it is handed, and it was handed ten games. So the "season rate" was
* a ten-game rate, and 0.4 of the forecast was the last five OF THOSE TEN.
* Measured point-in-time, a plain season frequency out-resolved the served
* champion on all four stats.
*/
const est = require('../../src/services/intelligence/probabilityEstimator');
const snapshotService = require('../../src/services/snapshotService');
/** n games where the player cleared the line at the given rate, most-recent-first. */
const logs = (n, rate, statType = 'hits') => Array.from({ length: n }, (_, i) => ({
date: `2026-06-${String((i % 28) + 1).padStart(2, '0')}`,
[statType]: (i % Math.round(1 / rate)) === 0 ? 2 : 0,
}));
describe('the forecast is no longer dominated by five games', () => {
it('a long cold streak inside a good season does not swing the forecast wildly', () => {
// Ten recent zeros on top of a strong season. At the old 0.40 weight this
// pulled the number a long way off a better one.
const season = logs(80, 0.6);
const cold = Array.from({ length: 5 }, (_, i) => ({ date: `2026-07-0${i + 1}`, hits: 0 }));
const withCold = [...cold, ...season];
const out = est.estimateProbability({ gameLogs: withCold, line: 0.5, statType: 'hits', features: {} });
const seasonOnly = est.estimateProbability({ gameLogs: season, line: 0.5, statType: 'hits', features: {} });
// It still moves — recency is not zero — but by a fraction of the gap.
expect(out.p_over).toBeLessThan(seasonOnly.p_over);
expect(seasonOnly.p_over - out.p_over).toBeLessThan(0.25);
});
it('more history produces a steadier forecast than ten games', () => {
const ten = logs(10, 0.6);
const many = logs(80, 0.6);
const a = est.estimateProbability({ gameLogs: ten, line: 0.5, statType: 'hits', features: {} });
const b = est.estimateProbability({ gameLogs: many, line: 0.5, statType: 'hits', features: {} });
expect(Number.isFinite(a.p_over)).toBe(true);
expect(Number.isFinite(b.p_over)).toBe(true);
});
});
describe('calibration is off while its maps are stale', () => {
it('serves nothing calibrated — the maps were fit on the retired forecast', () => {
expect(snapshotService.CALIBRATION_DEPLOYED).toEqual([]);
expect(snapshotService.CALIBRATION_BASIS).toEqual({});
});
it('the deploy list is still frozen, so nothing can re-enable it at runtime', () => {
expect(Object.isFrozen(snapshotService.CALIBRATION_DEPLOYED)).toBe(true);
});
});
+164
View File
@@ -0,0 +1,164 @@
'use strict';
/**
* THE TRUTH-LAW PROOF.
*
* The claim this engine makes is that a post cannot contain a number that was
* not pulled. These tests are that claim, made falsifiable. If any of them can
* be made to pass while a fabricated string escapes, the moat is decorative.
*/
const engine = require('../../src/services/content/contentEngine');
const { toSvg } = require('../../src/services/content/cardRenderer');
const base = {
id: 'test_tpl', sport: 'mlb', requires: ['n'],
pull: async () => ({ n: 3, name: 'Real Player' }),
copy: () => 'we graded {n} props',
card: () => ({ title: 'T', lines: [{ text: '{n} props', style: 'stat' }] }),
};
const reg = (over = {}) => engine.registerTemplate({ ...base, ...over, id: over.id || `t_${Math.random()}` });
describe('a template CANNOT render an unbacked claim', () => {
it('refuses a token with no pulled fact', async () => {
// The failure this prevents: a template author writes {edge} and the engine
// helpfully renders "undefined" or, worse, an empty string that reads fine.
const t = reg({ copy: () => 'edge is {edge_that_was_never_pulled}%' });
const out = await engine.generate(t.id, {});
expect(out.ok).toBe(false);
expect(out.skipped).toBe(true);
expect(out.reason).toMatch(/TRUTH LAW: unbacked token/);
});
it('refuses an unbacked token on the CARD too, not just the copy', async () => {
const t = reg({ card: () => ({ title: 'T', lines: [{ text: '{ghost_stat}', style: 'stat' }] }) });
const out = await engine.generate(t.id, {});
expect(out.ok).toBe(false);
expect(out.reason).toMatch(/unbacked token/);
});
it('render() throws directly — the gate is not bypassable by a caller', () => {
expect(() => engine.render('{nope}', { yes: 1 })).toThrow(/TRUTH LAW/);
});
it('a null fact is ABSENT, not rendered as "null"', async () => {
const t = reg({ pull: async () => ({ n: null }), copy: () => '{n} props' });
const out = await engine.generate(t.id, {});
expect(out.ok).toBe(false);
expect(String(out.copy || '')).not.toMatch(/null/);
});
it('an empty string is absent — a hole in a sentence is a lie by omission', () => {
expect(engine.isPresent('')).toBe(false);
expect(engine.isPresent(' ')).toBe(false);
});
it('ZERO is PRESENT — "0 cleared B+" is a real and important claim', () => {
// Treating 0 as missing is the Number(null) === 0 breach wearing its
// opposite coat, and it would silently delete our most honest post.
expect(engine.isPresent(0)).toBe(true);
expect(engine.render('{n} cleared', { n: 0 })).toBe('0 cleared');
});
it('NaN and Infinity are absent — they are arithmetic failures, not facts', () => {
expect(engine.isPresent(NaN)).toBe(false);
expect(engine.isPresent(Infinity)).toBe(false);
});
});
describe('the fact contract is checked BEFORE any string is built', () => {
it('a contract gap skips with the missing field named', async () => {
const t = reg({ requires: ['n', 'missing_field'] });
const out = await engine.generate(t.id, {});
expect(out.ok).toBe(false);
expect(out.reason).toMatch(/fact-contract gap: missing_field/);
});
it('a gap can emit an HONEST ABSENCE instead of nothing', async () => {
const t = reg({
requires: ['n', 'absent_thing'],
absent: () => ({ copy: 'nothing tonight, and we say so', card: { title: 'NONE' } }),
});
const out = await engine.generate(t.id, {});
expect(out.ok).toBe(true);
expect(out.honest_absence).toBe(true);
expect(out.copy).toMatch(/nothing tonight/);
});
it('a failing pull skips rather than rendering a half-post', async () => {
const t = reg({ pull: async () => { throw new Error('source down'); } });
const out = await engine.generate(t.id, {});
expect(out.ok).toBe(false);
expect(out.reason).toMatch(/pull failed: source down/);
});
});
describe('copy and card cannot disagree', () => {
it('both render from ONE fact object', async () => {
const t = reg({
pull: async () => ({ n: 7 }),
copy: () => 'we graded {n}',
card: () => ({ title: 'T', lines: [{ text: '{n} graded', style: 'stat' }] }),
});
const out = await engine.generate(t.id, {});
expect(out.copy).toMatch(/7/);
expect(out.card.lines[0].text).toMatch(/7/);
});
it('the card SVG contains the same pulled number', async () => {
const t = reg({ pull: async () => ({ n: 42 }), copy: () => '{n}', card: () => ({ title: 'T', lines: [{ text: '{n} graded', style: 'stat' }] }) });
const out = await engine.generate(t.id, {});
expect(toSvg(out.card)).toMatch(/42 graded/);
});
});
describe('the real templates obey the contract', () => {
const hot = require('../../src/services/content/templates/hotHitters');
const flex = require('../../src/services/content/templates/honestyFlex');
const streak = require('../../src/services/content/templates/streakList');
it.each([[hot], [flex], [streak]])('every token in %s is declarable', (t) => {
const tokens = new Set([
...engine.tokensIn(t.copy({})),
...engine.tokensIn(JSON.stringify(t.card({}))),
]);
expect(tokens.size).toBeGreaterThan(0);
// Each template must ship an absent-variant, or a thin night silently
// produces nothing and the flywheel stops without anyone noticing.
expect(typeof t.absent).toBe('function');
});
it('hot hitters refuses a hitter with too little history', async () => {
engine.registerTemplate(hot);
const out = await engine.generate('hot_hitters', {
date: '2026-08-07',
hitterForm: async () => ([{ name: 'Rookie', season_games: 4, recent_rate: 0.9, season_rate: 0.2 }]),
});
// 4 games is not a season, so nobody QUALIFIES -- which is a source problem,
// not a quiet night. It must SKIP rather than publish honest-absence copy.
expect(out.ok).toBe(false);
expect(out.reason).toMatch(/source problem, not a quiet night/);
});
it('hot hitters DOES emit honest absence when a real pool has nobody hot', async () => {
// The distinction that matters: candidates existed, were judged, none hot.
engine.registerTemplate(hot);
const out = await engine.generate('hot_hitters', {
date: '2026-08-07',
hitterForm: async () => Array.from({ length: 30 }, (_, i) => ({
name: `P${i}`, season_games: 90, recent_rate: 0.30, season_rate: 0.40,
})),
});
expect(out.honest_absence).toBe(true);
expect(out.copy).toMatch(/No hitter is meaningfully hot/);
});
it('streaks refuse a run not verified from settled results', async () => {
engine.registerTemplate(streak);
const out = await engine.generate('streak_list', {
date: '2026-08-07',
settledStreaks: async () => ([{ name: 'X', streak: 9, verified_from_settled: false }]),
});
expect(out.honest_absence).toBe(true);
});
});
+94
View File
@@ -153,3 +153,97 @@ describe('the measurements themselves', () => {
expect(fg.adjudicate(r, { factor: 'backwards' }).verdict).toBe('THEATER');
});
});
describe('pseudo-replication — sample counted in the unit the factor varies over', () => {
// A game-level factor (park, weather, opposing starter) hands every prop row
// in a game the identical treatment. Eighteen hitters in one ballpark are one
// reading of that ballpark, not eighteen.
const build = (games, perGame, seed = 1) => {
let s = seed;
const rnd = () => (s = (s * 1103515245 + 12345) % 2147483648) / 2147483648;
const rows = [];
for (let g = 0; g < games; g += 1) {
const shift = (rnd() - 0.5) * 0.06; // the game's treatment
// Outcomes are correlated WITHIN a game — a high-scoring night lifts every
// hitter in it. That shared component is exactly what row-resampling
// cannot see and what makes 18 rows worth far less than 18 readings.
const gameLevel = (rnd() - 0.5) * 0.5;
for (let i = 0; i < perGame; i += 1) {
const base = 0.3 + rnd() * 0.4;
const p = Math.max(0.02, Math.min(0.98, base + gameLevel));
rows.push({ cluster: `g${g}`, baseline: base, conditioned: base + shift, won: rnd() < p ? 1 : 0 });
}
}
return rows;
};
it('judges sample by CLUSTERS, so 900 rows over 50 games is 50 readings', () => {
const rows = build(50, 18);
// 900 rows clears the row floor; 50 clusters is what actually decides it.
const v = fg.adjudicate(rows, { factor: 'park', minN: 500, minClusters: 100 });
expect(rows.length).toBeGreaterThan(500); // looks like plenty
expect(v.verdict).toBe('CANDIDATE_PENDING_SAMPLE'); // and is not
expect(v.improvement.effective_n).toBe(50);
expect(v.improvement.cluster_unit).toBe('cluster');
expect(v.reason).toMatch(/not independent readings/);
});
it('the clustered interval is WIDER than the row interval on the same rows', () => {
// This is the whole hazard: resampling rows would have manufactured a
// confidence the evidence never supported.
const rows = build(40, 20, 7);
const clustered = fg.adjudicate(rows, { factor: 'park', minN: 10, minClusters: 5 });
const flat = fg.adjudicate(rows.map(({ cluster, ...r }) => r), { factor: 'park', minN: 10 });
const width = (v) => v.improvement.ci[1] - v.improvement.ci[0];
expect(width(clustered)).toBeGreaterThan(width(flat));
});
it('rows with no cluster keep the original row-resampling behaviour', () => {
const rows = build(40, 20, 3).map(({ cluster, ...r }) => r);
const v = fg.adjudicate(rows, { factor: 'x', minN: 10, minClusters: 5 });
expect(v.improvement.cluster_unit).toBe('row');
expect(v.improvement.effective_n).toBe(v.improvement.n);
});
});
describe('the two floors answer different questions', () => {
const rows = [];
let s2 = 11;
const rnd = () => (s2 = (s2 * 1103515245 + 12345) % 2147483648) / 2147483648;
for (let g = 0; g < 60; g += 1) {
for (let i = 0; i < 20; i += 1) {
const base = 0.35 + rnd() * 0.3;
rows.push({ cluster: `g${g}`, baseline: base, conditioned: base - 0.02, won: rnd() < base ? 1 : 0 });
}
}
it('ample rows with too FEW clusters is refused on the cluster floor', () => {
const v = fg.adjudicate(rows, { minN: 500, minClusters: 200 });
expect(v.verdict).toBe('CANDIDATE_PENDING_SAMPLE');
expect(v.reason).toMatch(/independent clusters/);
expect(v.clusters_needed).toBe(140);
});
it('too few ROWS is refused on the row floor even with many clusters', () => {
const thin = rows.filter((_, i) => i % 20 === 0); // 60 rows, 60 clusters
const v = fg.adjudicate(thin, { minN: 500, minClusters: 40 });
expect(v.verdict).toBe('CANDIDATE_PENDING_SAMPLE');
expect(v.reason).toMatch(/^n 60 < 500/);
});
it('a venue-constant factor stays refused however many rows accrue', () => {
// 30 ballparks is the whole universe; rows can grow forever and the
// interval never becomes trustworthy.
const venues = [];
for (let v = 0; v < 30; v += 1) {
for (let i = 0; i < 200; i += 1) {
const base = 0.4 + rnd() * 0.2;
venues.push({ cluster: `v${v}`, baseline: base, conditioned: base - 0.03, won: rnd() < base ? 1 : 0 });
}
}
const v = fg.adjudicate(venues, { minN: 500 });
expect(venues.length).toBe(6000);
expect(v.verdict).toBe('CANDIDATE_PENDING_SAMPLE');
expect(v.improvement.effective_n).toBe(30);
});
});
+153
View File
@@ -0,0 +1,153 @@
'use strict';
/**
* What a letter is allowed to mean.
*
* The failure these guard against is the one the whole programme keeps circling:
* a grade that LOOKS like it read tonight's matchup while reading nothing. Here
* that would be a band labelled factor-informed on an archetype where no factor
* ever proved — which is the state of every archetype today.
*/
const gb = require('../../src/services/model/gradeBands');
/** n rows whose outcome rate genuinely tracks p, interleaved (never front-loaded). */
function rows(specs) {
const out = [];
for (const [p, n, rate] of specs) {
for (let i = 0; i < n; i += 1) {
const won = Math.floor((i + 1) * rate) > Math.floor(i * rate) ? 1 : 0;
out.push({ p, won });
}
}
return out;
}
describe('the two-bar rule is structural', () => {
const data = rows([[0.75, 120, 0.75], [0.6, 120, 0.6], [0.45, 120, 0.45]]);
it('with NOTHING proven — the state today — no band can be factor-informed', () => {
const out = gb.buildBands(data, { archetype: 'GHOST' });
expect(out.basis).toBe('base_rate');
expect(out.two_bar.factor_informed_allowed).toBe(false);
expect(out.two_bar.reason).toMatch(/no factor has passed the gate/);
expect(out.bands.every((b) => b.basis === 'base_rate')).toBe(true);
});
it('PROVEN but not calibrated is still a base-rate read, and names why', () => {
// Both bars, or neither claim. A proven factor whose numbers are not
// certified honest cannot carry a letter that asserts a rate.
const out = gb.buildBands(data, { archetype: 'GHOST', proven: true });
expect(out.basis).toBe('base_rate');
expect(out.two_bar.reason).toMatch(/not certified calibrated/);
});
it('CALIBRATED but not proven is still a base-rate read', () => {
const out = gb.buildBands(data, { archetype: 'GHOST', calibrated: true });
expect(out.basis).toBe('base_rate');
expect(out.two_bar.reason).toMatch(/no factor has passed the gate/);
});
it('only BOTH unlocks factor-informed', () => {
const out = gb.buildBands(data, { archetype: 'GHOST', proven: true, calibrated: true });
expect(out.basis).toBe('factor_informed');
expect(out.bands.every((b) => b.basis === 'factor_informed')).toBe(true);
});
});
describe('lift is measured against the archetype\'s OWN base rate', () => {
it('a band is only credited when its interval clears the base rate', () => {
const out = gb.buildBands(rows([[0.8, 200, 0.80], [0.5, 200, 0.50], [0.3, 200, 0.30]]), {
archetype: 'GHOST', targetBands: 3,
});
const top = out.bands[0];
expect(top.shows_lift).toBe(true);
expect(top.separation).toBe('above_base_rate');
expect(top.lift_vs_archetype_base).toBeGreaterThan(0);
const bottom = out.bands[out.bands.length - 1];
expect(bottom.shows_deficit).toBe(true);
});
it('the same realized rate is lift for one archetype and not for another', () => {
// 62% is a real read for a profile that hits 45%, and slightly under water
// for one that hits 68%. A raw-rate band would call both the same letter.
const lowBase = gb.buildBands(rows([[0.62, 300, 0.62], [0.4, 300, 0.40]]), { archetype: 'LOW', targetBands: 2 });
const highBase = gb.buildBands(rows([[0.62, 300, 0.62], [0.75, 300, 0.75]]), { archetype: 'HIGH', targetBands: 2 });
expect(lowBase.bands[0].shows_lift).toBe(true);
expect(highBase.bands.find((b) => b.p_range[0] === 0.62).shows_deficit).toBe(true);
});
it('a band indistinguishable from the base rate claims NO lift', () => {
const flat = gb.buildBands(rows([[0.7, 150, 0.55], [0.5, 150, 0.55], [0.3, 150, 0.55]]), { archetype: 'FLAT' });
for (const b of flat.bands) {
expect(b.shows_lift).toBe(false);
expect(b.separation).toBe('indistinguishable_from_base_rate');
}
expect(flat.bands_showing_lift).toBe(0);
});
});
describe('bands the ledger cannot stand behind', () => {
it('indistinguishable neighbours are MERGED, not published as different letters', () => {
// If we cannot tell two bands apart, calling one A and one B is a
// distinction we have not measured.
const out = gb.buildBands(rows([[0.7, 200, 0.60], [0.65, 200, 0.60], [0.6, 200, 0.60]]), {
archetype: 'SAME', targetBands: 5,
});
expect(out.bands.length).toBe(1);
});
it('thin bands are PROVISIONAL, not silently dropped', () => {
// "Still counting" and "nothing here" are different claims.
const out = gb.buildBands(rows([[0.8, 20, 0.8], [0.3, 20, 0.3]]), { archetype: 'THIN', targetBands: 2 });
expect(out.bands.some((b) => b.provisional)).toBe(true);
});
it('too few outcomes → no bands at all, with the refusal stated', () => {
const out = gb.buildBands(rows([[0.6, 5, 0.6]]), { archetype: 'TINY' });
expect(out.bands).toEqual([]);
expect(out.refused).toMatch(/insufficient settled outcomes/);
});
it('the cumulative correction WIDENS every interval', () => {
const data = rows([[0.75, 150, 0.75], [0.45, 150, 0.45]]);
const one = gb.buildBands(data, { archetype: 'X', cumulativeTests: 1, targetBands: 2 });
const many = gb.buildBands(data, { archetype: 'X', cumulativeTests: 99, targetBands: 2 });
const width = (o) => o.bands[0].ci[1] - o.bands[0].ci[0];
expect(width(many)).toBeGreaterThan(width(one));
});
it('the interval never runs past 0 or 1', () => {
const out = gb.buildBands(rows([[0.99, 60, 1.0], [0.01, 60, 0.0]]), { archetype: 'EDGE', targetBands: 2 });
for (const b of out.bands) {
expect(b.ci[0]).toBeGreaterThanOrEqual(0);
expect(b.ci[1]).toBeLessThanOrEqual(1);
}
});
});
describe('the reason attached to a grade is true or absent', () => {
it('a base-rate band SAYS it is a base-rate read', () => {
const out = gb.buildBands(rows([[0.75, 120, 0.75], [0.45, 120, 0.45]]), { archetype: 'BOMBER', targetBands: 2 });
const text = gb.reasoning(out.bands[0], { archetype: 'BOMBER' });
expect(text).toMatch(/base-rate read/);
expect(text).toMatch(/no matchup factor is proven/);
// The thing it must never do is imply a matchup was read.
expect(text).not.toMatch(/matchup edge|favourable matchup/i);
});
it('a factor-informed band with NO named proven factors returns nothing', () => {
// Better to say nothing than to invent the why.
const out = gb.buildBands(rows([[0.75, 120, 0.75], [0.45, 120, 0.45]]),
{ archetype: 'GHOST', proven: true, calibrated: true, targetBands: 2 });
expect(gb.reasoning(out.bands[0], { archetype: 'GHOST', provenFactors: [] })).toBeNull();
});
it('a factor-informed band names the factors that actually proved', () => {
const out = gb.buildBands(rows([[0.75, 120, 0.75], [0.45, 120, 0.45]]),
{ archetype: 'GHOST', proven: true, calibrated: true, targetBands: 2 });
const text = gb.reasoning(out.bands[0], { archetype: 'GHOST', provenFactors: ['defense_by_direction'] });
expect(text).toMatch(/GHOST: defense_by_direction/);
expect(text).toMatch(/proved for this profile/);
});
});
+88
View File
@@ -0,0 +1,88 @@
'use strict';
/**
* The three proven hits factors, on the serving path at last.
*
* What these lock: the SIGN each factor moves the forecast, and that every
* unreadable case leaves it completely alone. A factor that nudges toward a
* default on missing input is fabricating a read from an absence.
*/
const hf = require('../../src/services/model/hitsFactors');
const pullGround = { pull_gb: 0.45, straight_gb: 0.10, oppo_gb: 0.05, pull_air: 0.20, straight_air: 0.12, oppo_air: 0.08 };
const pos = (o) => Object.fromEntries(Object.entries(o).map(([k, v]) => [k, { oaa: v, fielders: 2 }]));
const side = (avg, pa) => ({ pa, atBats: Math.round(pa * 0.9), hits: Math.round(pa * 0.9 * avg) });
const bigSplit = { vl: side(0.284, 183), vr: side(0.221, 291) };
describe('each factor moves the forecast in the direction it proved', () => {
it('TOUGH spray defence lowers the hit forecast', () => {
const eliteLeft = pos({ '3B': 12, SS: 10, '1B': 0, '2B': 0, LF: 0, CF: 0, RF: 0 });
const out = hf.adjustProbability(0.6, { spray: pullGround, bats: 'R', positionOaa: eliteLeft });
expect(out.p_adjusted).toBeLessThan(0.6);
expect(out.applied.map((a) => a.factor)).toContain('defense_by_direction');
});
it('a CONTACT-ALLOWING pitcher raises it; a bat-misser lowers it', () => {
const soft = hf.adjustProbability(0.6, { pitcherHardHit: 0.46 });
const tough = hf.adjustProbability(0.6, { pitcherHardHit: 0.31 });
expect(soft.p_adjusted).toBeGreaterThan(0.6);
expect(tough.p_adjusted).toBeLessThan(0.6);
});
it('a PLATOON disadvantage lowers it, the edge raises it', () => {
const edge = hf.adjustProbability(0.6, { platoonSplits: bigSplit, bats: 'R', throws: 'L' });
const wrongSide = hf.adjustProbability(0.6, { platoonSplits: bigSplit, bats: 'R', throws: 'R' });
expect(edge.p_adjusted).toBeGreaterThan(0.6);
expect(wrongSide.p_adjusted).toBeLessThan(0.6);
});
});
describe('unreadable means UNTOUCHED, never nudged to a default', () => {
it('a SWITCH hitter gets no spray adjustment', () => {
const out = hf.adjustProbability(0.6, { spray: pullGround, bats: 'S', positionOaa: pos({ '3B': 12, SS: 10 }) });
expect(out.applied.find((a) => a.factor === 'defense_by_direction')).toBeUndefined();
expect(out.skipped.map((s) => s.factor)).toContain('defense_by_direction');
});
it('a THIN platoon split is refused, not shrunk toward league', () => {
const thin = { vl: side(0.350, 25), vr: side(0.250, 400) };
const out = hf.adjustProbability(0.6, { platoonSplits: thin, bats: 'R', throws: 'L' });
expect(out.applied.find((a) => a.factor === 'platoon_severity')).toBeUndefined();
});
it('no pitcher profile contributes nothing at all', () => {
expect(hf.pitcherContactMultiplier(null)).toBeNull();
const out = hf.adjustProbability(0.6, { pitcherHardHit: null });
expect(out.p_adjusted).toBe(0.6);
});
it('with NOTHING readable the forecast is returned exactly', () => {
const out = hf.adjustProbability(0.6, {});
expect(out.p_adjusted).toBe(0.6);
expect(out.multiplier).toBe(1);
expect(out.factors_fired).toBe(0);
});
it('a null forecast stays null — no factor invents one', () => {
expect(hf.adjustProbability(null, { pitcherHardHit: 0.46 }).p_adjusted).toBeNull();
});
});
describe('composition', () => {
it('three factors compound, and the stack is bounded', () => {
const out = hf.adjustProbability(0.6, {
spray: pullGround, bats: 'R', positionOaa: pos({ '3B': -12, SS: -12, '1B': -12, '2B': -12, LF: -12, CF: -12, RF: -12 }),
pitcherHardHit: 0.46, platoonSplits: bigSplit, throws: 'L',
});
expect(out.factors_fired).toBe(3);
expect(out.multiplier).toBeLessThanOrEqual(1 + hf.COMBINED_MAX);
expect(out.multiplier).toBeGreaterThanOrEqual(1 - hf.COMBINED_MAX);
});
it('partial readability applies only what is readable', () => {
const out = hf.adjustProbability(0.6, { pitcherHardHit: 0.46, bats: 'S', spray: pullGround, positionOaa: pos({ '3B': 5 }) });
expect(out.factors_fired).toBe(1);
expect(out.applied[0].factor).toBe('pitcher_contact_profile');
});
});
+148
View File
@@ -0,0 +1,148 @@
'use strict';
/**
* The two-parameter correction that replaces isotonic on a thin sample.
*
* What these protect: that it CANNOT chase day-structure (two parameters over
* the whole curve), and that a thin fit is applied at reduced strength rather
* than at face value.
*/
const lp = require('../../src/services/model/lowParamCalibrator');
/** An over-confident forecaster: predicts p, actually hits closer to the mean. */
function overConfident(n, dates = 10, seed = 3) {
let s = seed;
const rnd = () => (s = (s * 1103515245 + 12345) % 2147483648) / 2147483648;
const rows = [];
for (let i = 0; i < n; i += 1) {
const p = 0.5 + rnd() * 0.45;
const truth = 0.5 + (p - 0.5) * 0.4; // real skill is 40% of claimed
rows.push({ date: `d${i % dates}`, p, won: rnd() < truth ? 1 : 0 });
}
return rows;
}
describe('it fits the favourite-longshot shape', () => {
it('FLATTENS an over-confident forecaster (a < 1)', () => {
const m = lp.fitPlatt(overConfident(2000));
expect(m).not.toBeNull();
expect(m.flattens).toBe(true);
expect(m.a).toBeLessThan(1);
});
it('pulls high predictions down and leaves the middle nearly alone', () => {
const m = lp.fitPlatt(overConfident(2000));
const hi = lp.applyPlatt(m, 0.92);
const mid = lp.applyPlatt(m, 0.55);
expect(hi).toBeLessThan(0.92);
expect(Math.abs(mid - 0.55)).toBeLessThan(Math.abs(hi - 0.92));
});
it('stays monotone — ordering is never disturbed', () => {
const m = lp.fitPlatt(overConfident(2000));
let prev = -1;
for (let p = 0.05; p <= 0.95; p += 0.05) {
const v = lp.applyPlatt(m, p);
expect(v).toBeGreaterThan(prev);
prev = v;
}
});
it('leaves an already-honest forecaster essentially alone', () => {
let s = 11;
const rnd = () => (s = (s * 1103515245 + 12345) % 2147483648) / 2147483648;
const rows = [];
for (let i = 0; i < 2000; i += 1) {
const p = 0.3 + rnd() * 0.6;
rows.push({ date: `d${i % 12}`, p, won: rnd() < p ? 1 : 0 });
}
const m = lp.fitPlatt(rows);
expect(Math.abs(lp.applyPlatt(m, 0.8) - 0.8)).toBeLessThan(0.06);
});
});
describe('shrinkage — a thin fit is applied at reduced strength', () => {
it('scales with the number of fit DATES, not rows', () => {
const few = lp.fitPlatt(overConfident(2000, 5));
const many = lp.fitPlatt(overConfident(2000, 40));
expect(few.fit_rows).toBe(many.fit_rows); // same rows
expect(few.shrinkage).toBeLessThan(many.shrinkage); // different dates
expect(few.shrinkage).toBeCloseTo(5 / 15, 3);
expect(many.shrinkage).toBeCloseTo(40 / 50, 3);
});
it('a 5-date fit corrects less than a 40-date fit on the same input', () => {
const few = lp.fitPlatt(overConfident(2000, 5));
const many = lp.fitPlatt(overConfident(2000, 40));
// Both flatten; the thin one is held closer to the raw number.
expect(Math.abs(lp.applyPlatt(few, 0.92) - 0.92))
.toBeLessThan(Math.abs(lp.applyPlatt(many, 0.92) - 0.92));
});
});
describe('honesty', () => {
it('refuses to fit below the row floor', () => {
expect(lp.fitPlatt(overConfident(40))).toBeNull();
expect(lp.fitPlatt([])).toBeNull();
expect(lp.fitPlatt(null)).toBeNull();
});
it('an unreadable input returns null, never an uncorrected number', () => {
const m = lp.fitPlatt(overConfident(2000));
expect(lp.applyPlatt(m, null)).toBeNull();
expect(lp.applyPlatt(null, 0.7)).toBeNull();
});
it('has exactly two parameters — it CANNOT encode a single odd day', () => {
// This is the whole reason it replaces isotonic here.
const m = lp.fitPlatt(overConfident(2000));
const shape = Object.keys(m).filter((k) => k === 'a' || k === 'b');
expect(shape.sort()).toEqual(['a', 'b']);
});
});
describe('the slope must CORRECT, not abandon the forecast', () => {
/** A forecaster whose p_win carries no information at all. */
function uninformative(n, seed = 7) {
let s = seed;
const rnd = () => (s = (s * 1103515245 + 12345) % 2147483648) / 2147483648;
const rows = [];
for (let i = 0; i < n; i += 1) rows.push({ date: `d${i % 8}`, p: 0.4 + rnd() * 0.5, won: rnd() < 0.55 ? 1 : 0 });
return rows;
}
it('REFUSES a fit whose slope collapses to a constant', () => {
// Shrinking a miscalibrated forecaster toward its base rate always lowers
// Brier, so this would score as a win while destroying all resolution.
const m = lp.fitPlatt(uninformative(3000));
expect(m.refused).toBe(true);
expect(m.reason).toMatch(/collapses to a constant|invert/);
});
it('a refused fit produces no calibrated number at all', () => {
const m = lp.fitPlatt(uninformative(3000));
expect(lp.applyPlatt(m, 0.8)).toBeNull();
});
it('an INVERTING slope is refused by name', () => {
// Real case: runs fitted a = -0.032, which would reverse every ordering.
let s = 5;
const rnd = () => (s = (s * 1103515245 + 12345) % 2147483648) / 2147483648;
const rows = [];
for (let i = 0; i < 3000; i += 1) {
const p = 0.4 + rnd() * 0.5;
rows.push({ date: `d${i % 8}`, p, won: rnd() < (0.9 - p) ? 1 : 0 }); // backwards
}
const m = lp.fitPlatt(rows);
expect(m.refused).toBe(true);
expect(m.a).toBeLessThanOrEqual(lp.MIN_SLOPE);
});
it('still accepts a genuine flattening', () => {
const m = lp.fitPlatt(overConfident(2000));
expect(m.refused).toBeUndefined();
expect(m.a).toBeGreaterThan(lp.MIN_SLOPE);
expect(m.a).toBeLessThan(1);
});
});

Some files were not shown because too many files have changed in this diff Show More