main
272 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
71d3b7b786 |
E10 Report issue template + E12 /report archive, to spec
PHASE 0 — the spec, read not recalled. E10: "Hybrid: dark billboard header that survives every client, light paper body Gmail can't wreck. 600px, stacked, no webfont dependence." Content law: "One email per slate day. Top read, what changed, the record. Nothing else." E12: "EVERY ISSUE SHOWS ITS OWN DAY RECORD -- THE ARCHIVE IS A LEDGER TOO." COMPOSED, NOT FORKED. The audit had E10 as PARTIAL, not absent: newsletterService already builds the daily report's CONTENT and lints its voice. What was missing is the designed hybrid SHELL, so reportTemplate.js is a template over that builder rather than a second report -- the same call made for the movement strip, and for the same reason. PHASE 1 — the hybrid shell is an ENGINEERING constraint, not a look, and the tests say so: Gmail strips style blocks, Outlook ignores flexbox, and a dark body renders as a black rectangle in several clients. Hence tables, inline styles, 600px fixed, system fonts, no image required to read, and the green SHIFTS from #00D4A0 to #00A57D on paper because the dark-mode green is unreadable there. FACT-CONTRACTED: a section whose data is absent is OMITTED and NAMED in `omitted`, never filled. There is no code path producing a placeholder figure. The honesty block carries the real numbers -- graded count, cleared-ceiling count, the realized rate against baseline, and that we do not issue A grades. E1'S LAW TRAVELS EVEN THOUGH ITS RENDERING CANNOT. An SVG strip is not reliable in email, so movementText carries the RULE: green only when the move favours the read, and a flat market says FLAT · [N]D rather than showing nothing. NO DESIGNER SAMPLE DATA. Nabers 1,120.5, No 128, DAY RECORD 9-4 are a spec for what a live issue renders; pasting them in would be fabrication carrying a designer's authority and would look entirely correct. Tested. PHASE 2 — /report is now the real archive, REPLACING the S41 redirect to /blog. That redirect existed because the surface did not; E12 built it, so the placeholder is correctly gone and the S41 test is updated rather than worked around. Every row carries its own day record, and an unknown record says UNSETTLED -- never a dash that reads as zero. Empty archive is an honest state. Backend: public read-only /api/report over Redis issues, plus the Next proxy. Both surfaces registered under the reachability guard. A test bug I made twice now: my check for forbidden sample values matched the template's own doc block, which NAMES those values as things never to paste. Documentation worth keeping, so both suites strip comments before matching -- a guard that reads its own warning is not reading the code. WAVE-2 STATUS: E1, F9-F11, E10, E12 done. Still gated -- F5 article media and E16/F8 on the card-system reconciliation; the in-season hub IA on the social chat's formula; E9/E15 on model; E2/E6 on licensing. Read-only throughout; serving fingerprint unchanged including newsletterService; accrual clock unchanged at 0 eligible dates. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
49e76068da |
Doctrine + E1 movement strip + F9-F11 offseason hub shell, to spec
PHASE 0 — specs/ARCHETYPE-TAXONOMY-DOCTRINE.md records the ruling as
shared law: 83 designed glyphs are the full four-sport taxonomy; a glyph
renders ONLY where its archetype is modeled and proven. 39 of 83 map to a
real archetype and are wired; the 44 unmapped are DORMANT SLOTS for
WNBA/NBA/Soccer, not a wiring gap. Wiring them would mean inventing 44
archetypes to consume artwork -- decoration presented as classification,
which is forbidden. DUAL THREAT and PAINT BOSS are modeled archetypes with
no mark: the mirror gap, flagged to the design side. When a sport's
archetypes ship, activation is a MANIFEST lookup, not new art.
PHASE 1 — E1 movement strip. The spec's own line is "the movement strip is
defined once here and reused everywhere a line has a past", so it is a
primitive, not a fourth chart.
RECONCILED RATHER THAN FORKED: lib/gradeShift.js ALREADY implements E1's
colour law -- toward/against/flat, including the direction flip that makes
an UNDER's favourable move the opposite sign of an OVER's. MovementStrip
CONSUMES buildGradeTimeline instead of reimplementing it, and a test
asserts it never redefines isUnder. GradeShift stays the grade-history
view; this is the reusable strip. That is the card-fork lesson applied
before it could happen again.
Spec laws honoured: STEPS NOT CURVES (H then V, no smoothing -- a curve
invents prices that never traded, and a test rejects any C/S/Q/T command);
green only when the move FAVOURS the read; FLAT renders as a hairline plus
FLAT · [N]D because a flat market is a finding; and too little history
says NO MOVEMENT HISTORY rather than rendering blank.
PHASE 2 — F9-F11 offseason hub shell, built from Vyndr Offseason.dc.html.
The spec's load-bearing words are used verbatim: "OUTLOOKS REPRICE ON NEWS
· NOT GAME ODDS" (an offseason number is not a game line), the QUIET WIRE
empty state ("No outlook-moving news since X. We don't manufacture
movement."), WHAT CHANGED TODAY as the hero with the countdown ambient and
top-right, the tag-colour-is-meaning row anatomy, the open -> NOW -> FAIR
triplet with the movement strip embedded, and the OUTLOOK ONLY block where
every row carries NOT GRADED.
THE DESIGN FILE'S SAMPLE DATA IS NOT IN THE COMPONENT. Wembanyama +420 ->
+330, Nabers cleared 11:42 AM, the Summer League names -- all of it is a
SPEC for what a live feed renders, and copying it in would be fabrication
carrying a designer's authority. A test asserts none of those strings
appear.
The IN-SEASON information architecture is NOT invented here. The spec
covers an offseason hub; nothing specifies how content, articles, wire and
the live slate share year-round navigation. That remains the open design
gap, and the route notes it.
Two test bugs caught and fixed: my first assertions matched my own doc
comments -- the ordering check found "WHAT CHANGED TODAY" in the header
block and the no-curves check caught the word "curve" in the sentence
explaining why curves are wrong. A guard that reads its own explanation is
not reading the render; both now strip comments first.
PHASE 3 — both surfaces registered under the reachability guard. Read-only
throughout, serving fingerprint unchanged, accrual clock unchanged at 0
eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
60469422af |
Wave D1 primitives — and the audit says E1/E9 were not Wave 1
PHASE 0 — the order proposed E1 + E9 as Wave-1 and told me to follow the
audit if it disagreed. It disagrees: E9 calibration curve is WAVE D3,
gated on MODEL work ("resolve n>=20 vs N30, accrue buckets"), and E1
movement strip is WAVE D6, a large surface build. E9's gate is live right
now -- calibration is WITHDRAWN at 0 eligible dates, so the curve could
only render its empty state today. Building it would ship a component
whose entire purpose is unavailable.
PHASE 1 — three of the five real Wave D1 items were ALREADY DONE, and the
2026-07-31 audit has aged:
D1 glyph library audit: 38/83 wired (46%)
now: COMPLETE for everything wireable -- 39 of 83
designed glyphs map to a real archetype, all 39
are wired, colours match the registry exactly
(0 disagreements).
A1 card token audit: BUILT-BUT-DRIFTED, "in only 1 file"
now: BUILT-TO-SPEC -- it IS the --bg-1 token,
consumed by 32 files. The audit counted literal
hex, which is what a correctly tokenised value
looks like.
B1 boundary blue audit: PARTIAL, hex in 2 files
now: BUILT -- --priced-out/#8fb2de is a token with a
documented colour law, 4 consumers.
The 44 unwired glyphs are NOT a wiring gap: they have no backend
archetype, so wiring them means inventing 44 archetypes to consume
artwork -- the fabrication this programme refuses. That is the 41-vs-74
scope question and it is Kev's call. Separately, 2 registry archetypes
have NO designed glyph (DUAL THREAT, PAINT BOSS) -- a design gap.
PHASE 2 — what was genuinely absent is now built. web/src/lib/motion.js:
nudge() capped at 180ms so it reads as acknowledgement rather than
latency; bootStagger capped at 240ms because uncapped, row 40 waits 1.1s
and the stagger BECOMES the latency it exists to disguise; rowHover
returns handlers not CSS so touch cannot stick a hover state; and
revealOnIntersect returns an unobserve in every path and reveals
IMMEDIATELY when there is no IntersectionObserver or motion is reduced --
content is never hidden behind a capability check.
Reduced motion is honoured, not softened. The sharpest of the 10 tests:
bootStagger under reduced motion returns opacity 1, not merely delay 0 --
if the CSS animation supplies the opacity, skipping it leaves the row
invisible forever.
PHASE 3 — the reachability guard gains a PRIMITIVES section: a module
built to be embedded must declare its exports AND name its intended
consumers, because a primitive imported by nothing is the same
built-but-unread class as an unmounted component.
WAVE-2 UNGATED: F9-F11 offseason hub, F5 article media, E10/E12 Report,
E1 movement strip. GATED: E9 + E15 on model, E16/F8 on the resolution tail
and the card-system reconciliation, E2/E6 on licensing, E13 on another
order.
Read-only throughout; serving fingerprint unchanged; accrual clock
unchanged at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
c575a708c7 |
Content studio API + preview page; widen the reachability guard; correct
two inventory errors INVENTORY CORRECTION, and it was mine. Phase 2's two "orphans" are NOT orphans -- my board grepped only web/src/app and missed component-level mounting. The transitive check says both are already mounted: BookComparisonPanel -> GradeResultCard -> app/scan/page.tsx NewsWire -> ExploreHub -> app/explore/page.tsx So book comparison is DONE (wired to /api/books, rendering on the grade card) and THE WIRE is DONE-BY-DESIGN, mounted in ExploreHub. Its header names an "Offseason Hub" as its home, and that hub genuinely does not exist -- but that is board item #8, not a mounting bug, and inventing a surface to satisfy a comment would be the wrong fix. The lesson is the same one this session keeps teaching: I checked one directory and reported a conclusion the check could not support. ALSO CAUGHT: I overwrote src/routes/content.js, which was the Session-29 content-templates route, by picking a filename without looking. Restored from git with no work lost; the new surface lives at /api/content-studio and both now coexist. PHASE 0/1 — /api/content-studio serves finished posts (copy, branded card, card_svg, the fact_contract each was REQUIRED to have, and the facts that actually backed it) plus a POST for editorial status in Redis. Private via internal key; the Next proxy holds the key server-side so the browser never does. /studio renders it as a thin client -- copy and card side by side with the fact contract visible, because reviewing copy by reading it is exactly how a wrong number ships. Never-blank: a night with nothing generated says so. API-FIRST is the point: the endpoint an autonomous poster will call is the one the page already renders, so the agent handoff is a pointer change, not a rebuild. Contract documented at docs/CONTENT-STUDIO-API.md. EXPRESS 5 BROKE 23 SUITES at first: `router.get('/:date?')` throws at mount time in Express 5, taking down everything that imports app.js. Two explicit routes instead. PHASE 3 — the reachability guard is widened from grade-fields-only to a general built-but-unread check. Book comparison, THE WIRE and the content studio are now registered surfaces; a page counts as its own entry point (Next mounts it by convention) while everything else must trace to one. 22 checks green; a registered-but-unimported surface still goes red. FULLY ISOLATED: read-only on model/slate/ledger, serving fingerprint verified unchanged, accrual clock unchanged at 0 eligible dates. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
74aa75945e |
Content engine: posts that structurally cannot lie
PHASE 0 — contentEngine makes Truth Law structural, not careful. Copy is
token-substituted and an unbacked {token} REFUSES to render -- there is no
code path that produces a plausible default. The fact contract is asserted
before any string is built. Card and copy render from ONE fact object, so
a caption and a card cannot disagree. No live model writes factual claims:
the voice is in the template, the facts are pulled, and the voice-polish
port is deliberately unwired, because an LLM that can rewrite a sentence
can rewrite a number.
18 tests carry the proof. The one that matters most: ZERO IS PRESENT.
"0 cleared B+" is our most honest possible post, and treating 0 as missing
would be the Number(null)===0 breach wearing its opposite coat -- it would
silently delete exactly the post the brand is built on.
PHASE 1 — three templates, generating real posts from tonight's data:
hot hitters off the repaired full-season log, the honesty flex off the
real servedGrade distribution (2,140 graded / 70 cleared B+ / 42% not
separable / A unissuable), and streaks verified from settled outcomes only.
THE ENGINE CAUGHT A BUG IN ITSELF, and it is the sharpest lesson here. The
first run published "No hitter is meaningfully hot tonight -- we could
dress up a middling week as a streak. We don't." That was FALSE: the
box-score cache spans only the settled window, every player had under 20
games, and the pool was empty. A broken pull was publishing as considered
editorial judgement -- the fourth appearance of this class tonight and the
first where our OWN HONESTY COPY was the disguise.
Fixed structurally rather than by patching the number: an absent() variant
may now DECLINE to speak, and the template separates "no candidates at
all" (SKIP with a reason) from "candidates judged, none hot" (honest
absence). Both locked by test. Source corrected to mlbStatsAdapter.fullLog,
the same log the repaired champion reads.
PHASE 2 — cardRenderer emits SVG rather than canvas: it is text, so it
diffs in review and its numbers are greppable, which matters when the
whole claim is that the numbers are real. VYND white + R green, slashed-Y,
scanlines, mono. The card never formats its own facts -- every string
arrives pre-rendered and gate-checked.
PHASE 3 — scripts/generate-content.js writes copy + card per template to
.content-out/<date>/. Template N+1 is a registry entry: requires, pull,
copy, card, absent. Queued as stubs, not built: hot takes, daily reads,
"grades we DIDN'T give", cross-sport streak variants (the streak template
is already sport-agnostic -- settled outcomes and a noun).
FULLY ISOLATED: read-only on every source, zero writes to serving, model
or ledger tables. Serving fingerprint verified unchanged. The accrual clock
is untouched at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
55b210cb95 |
Fix the dormant basketball window-bug before it ships; guard the class
PHASE 0 — audit. espnStatsAdapter's slice(0,20) was already fixed at |
||
|
|
981a05cbd6 |
Render-reachability guard: make built-but-unread a CI failure
Three consecutive orders shipped a backend-correct field that never
reached a screen, all on a green suite: gradeBands (required by no
serving code), served_grade (dropped at the adapter boundary),
GradeScaleLegend (imported by nothing). Each was caught by luck on a later
re-check, and in two of the three I had already reported the wiring done.
WHY GREEN TESTS COULD NOT SEE IT: backend tests stop at the API payload.
They prove a field is PRODUCED and say nothing about whether it is
CONSUMED. Invisible by construction, not an oversight in any one test.
THE TRAP, NAMED: the difficulty pools in the backend, so by the time a
field exists on the payload it feels finished. What remains is a
three-line adapter change nobody considers worth verifying, so it gets
claimed rather than traced. The last inch is the one with no friction,
which is exactly why it gets skipped. "I added the field" and "a user can
see it" are different claims and only the first is fun.
THE GUARD traces each promised field the whole way: payload -> adapter
consumes -> component renders -> component is MOUNTED. Mounted is
transitive to a Next entry point (page/layout/template), the only thing
that puts a pixel on screen, depth-limited so an import cycle cannot hang
the suite.
Container rows are exempted EXPLICITLY, not silently: served_grade carries
container:true plus a rendersVia list, and a separate assertion checks
every named part actually renders. The exemption is auditable and cannot
hide an unrendered field.
The guard tests itself -- an orphan component must report unmounted, and
the contract must be non-empty, since an empty contract passing vacuously
is how this would most plausibly rot.
RETRO-PROOF: run unchanged against
|
||
|
|
3591c7626e |
Total grade cutover + the ceiling stated as a position
PHASE 0 caught my own repeat of the failure I diagnosed one order ago.
|
||
|
|
91927a4a8a |
Serve an honest grade: the letter was carrying 1/6 the information of the
number beside it PHASE 0 corrects the order's premise. A grade letter has been served all along -- engine1.gradeProp builds it from an additive factor index, computed INDEPENDENTLY of p_win. gradeBands is orphaned for a different reason than assumed: it defines what a letter MEANS from realized outcomes, and every band collapses to base-rate at current resolution. The measurement that changed this order, on 3,417 settled props: grade n realized mean p_win A 8 0.500 0.647 <- the TOP grade did WORST B 985 0.640 0.700 C 1,695 0.602 0.676 D 303 0.558 0.604 F 426 0.535 0.588 letter resolution 0.00116 (0.48% of variance) p_win resolution 0.00715 (2.98%) -> the letter carried 0.16x the information of the number beside it Concretely, from the hand-verify: Christian Encarnacion's 0.95 over graded C and his 0.05 under ALSO graded C -- same hitter, opposite forecasts, same letter. The gap was never that grades don't ship; it is that the weaker of two available signals shipped as the headline. PHASE 1 — model/servedGrade.js derives the letter from p_win with bands anchored on MEASURED realized rates (B+ 0.663 / B 0.646 / C+ 0.615 / C 0.589 / C- 0.548 / D 0.512 / F 0.447, base 0.6005). NO MANUFACTURED A, structurally: A+/A/A- are UNISSUABLE, not rare. The realized rate plateaus at 0.65-0.68 above p_win 0.70, so no band has earned a top letter; a test sweeps every p_win 0..1 and asserts none produces one. Even 0.99 tops out at B+ with its realized 0.663 attached. Raising that ceiling later is a deliberate, visible act. Bands that cannot separate SAY so -- C+/C/C- carry separates_from_base_rate false and copy naming it, which is the honest description of a forecast explaining 3% of variance. Every grade states its basis (forecast_only vs forecast_plus_matchup_factors, naming which factors fired) and calibrated:false. engine1.grade is preserved as engine_grade so nothing downstream breaks. PHASE 2 — refusals render real states: insufficient_data -> "not enough history to call this one"; juiced_no_edge -> "the book has priced the vig past any edge on this side". 1,870 refused snapshots carry exactly those two reasons and both now surface. PHASE 3 — hand-verified on 12 real served props. Freeman/Rice/Encarnacion 0.95 overs now B+ (was B, C, B); the 0.05 unders now F (was C). Refused doubles render NO READ with their reason. never-blank PASS, no-manufactured-A PASS. Serving change; nine frozen model modules unchanged including engine1; p_win never mutated; no calibrated number leaks (deployed set empty); no Bonferroni slot. STILL TRUE: the forecast explains ~3% of outcome variance. This order did not make the model better. It made the letter stop overstating it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
494c83cf76 |
Hunt the window-bug class: three more paths, and the forward re-audit rule
in code PHASE 0 — getStatRows is the single base-rate path, so every branch is audited, plus the feature builders since l20_avg is the season reference projectionFor reads: getStatRows MLB -> estimator base fullLog CORRECT ( |
||
|
|
929fd81940 |
Repair the champion: it was reading ten games, not a season
PHASE 0 — the defect is real past the peek. Against a FAIR point-in-time baseline (each player's rate over games strictly before that date, >=10 prior games, box scores back to 05-01), the served champion LOSES on all four stats, three of four CIs excluding zero: hits 0.00251 vs 0.00774 CI [-0.0074,-0.0011] TB 0.00393 vs 0.00619 CI [-0.0055,-0.0003] rbi 0.02481 vs 0.03133 CI [-0.0153,-0.0005] runs 0.00181 vs 0.00683 CI [-0.0114,+0.0008] PHASE 1 — the cause is the WINDOW, not the weights. estimateProbability builds its base rate as the frequency over every row it is handed, and featureCache.getStatRows handed it res.last10. So the "season rate" was a TEN-GAME rate, and 0.4 of the forecast was the last five OF THOSE TEN. The 0.40 recency weight costs resolution on all four stats (-0.00086, -0.00107, -0.00562, -0.00365). Nudges are mixed and small -- harmful on hits and rbi, marginally helpful on TB and runs -- so they are left alone. PHASE 2 — two lines, no new data, no extra API call, because fullLog was already fetched by the same adapter call that produced last10: getStatRows now reads fullLog, and RECENCY_WEIGHT goes 0.40 -> 0.20. hits 0.00251 -> 0.00817 (tripled; now above the fair baseline) TB 0.00393 -> 0.00734 (above baseline; vs old CI [0.0020,0.0067]) rbi 0.02481 -> 0.02727 (still below baseline, CI includes zero) runs 0.00181 -> 0.00436 (still below baseline, CI includes zero) Gate stated exactly: hits and TB now exceed the fair baseline on the point estimate; rbi and runs remain below but EVERY CI now includes zero, so no stat reliably loses to a frequency table. That is a tie on rbi/runs, not a win, and it is reported as one. Only TB's improvement over the old champion is CI-confirmed; the rest are directional. STALE-FIT GATE: CALIBRATION_DEPLOYED is now EMPTY. The low-param maps were fitted on the retired forecast and fromLedger cannot rescue them -- settled ledger rows still carry OLD p_win, so refitting today would refit the retired forecast. Nothing is served calibrated until dates settle under the repaired champion, and the favourite-longshot bias must be re-measured rather than assumed to survive. The shadow duel is void. PHASE 3 — the hits factor lift is NOT re-measured, and cannot be yet: it needs settled rows produced BY the repaired champion, which ships in this commit. Replaying would score the factors against a reconstruction rather than the served forecast. Deferred, explicitly. The factors remain wired and transmitting; only their lift is unquantified on the new baseline. PHASE 4 — standing flag, and it is large: EVERY factor verdict in this programme, every null and every THEATER, was measured against a champion worse than a frequency table. Signal added to noise reads as noise. Prior verdicts may deserve re-audit. Logged, not re-run. Re-queued not built: rbi lineup-slot / RISP opportunity through the two-part gate, now landing on a repaired champion. Serving-path change by design; the byte-identical invariant inverted and all four stats move. Nine frozen model modules verified unchanged. No Bonferroni slot -- resolution accounting on the champion's own knobs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
43f65d30cb |
Wire the three proven hits factors pre-grade: transmission proven, gain
inconclusive THE BUG THIS NEARLY SHIPPED AS A FINDING. The first audit reported 0 factors fired on all 1,140 rows. Not a result -- my paging helper ordered by `id`, and batter_spray, team_defense, platoon_splits and statcast_aggregates have composite primary keys with NO id column. The query errored, the loop broke on error, and four fully-populated tables read as empty. hitsFactorContext.js -- the PRODUCTION loader -- had the identical defect, so live wiring would have loaded nothing and served unadjusted while logging success. Third occurrence of this class in one session. Both loaders now order by a real column and THROW rather than degrade. The Phase 2 gate is what caught it: no resolution number was quoted until transmission was proved. PHASE 1 — pipeline is now base -> FACTORS -> CALIBRATE -> GRADE. Context built in snapshotService BEFORE gradeAndCacheSlate (was line 640+, grade at 454), threaded per prop, applied to p_over before p_win is set with p_win_prefactor and a full trace retained. Hits only. Coverage 859/1140 rows (75%): 474 with all three factors, 256 two, 129 one, 281 none. PHASE 2 — TRANSMISSION PROVEN, 12/12 sign-correct, 4/4 per factor, each applied IN ISOLATION. My first table compared each factor's expected sign against the COMPOSITE change and showed 3 false failures -- with three factors firing the net can oppose any single member; that was a flaw in the test, not the wiring. Two under-side rows confirm the flip is handled: a factor raising p(over) correctly lowers p_win. Switch hitters (Bailey, Bell, Rocchio) took no spray adjustment while their other factors fired normally -- the refusal is selective, not a blanket skip. PHASE 3/4 — both maps refit on the factor-adjusted forecast; the shadow-duel baseline is VOID and restarts, since it accumulated against a different forecast. Point-in-time, 765 held-out rows: reliability 0.00795 -> 0.00828 RESOLUTION 0.00229 -> 0.00345 (variance explained 0.93% -> 1.39%) Brier 0.25398 -> 0.25305 delta -0.00093 CI [-0.00225,+0.00002] Resolution rose 51% relative. The CI TOUCHES ZERO on 4 eval dates, so the composition does NOT earn a proven keep -- three isolated passes did not grant a composed pass. INCONCLUSIVE, reported as such. The gain is far below the sum of the isolated effects, which is expected: all three run through the same pitcher-batter confrontation and share signal. PHASE 5 — 1.39% of variance is still far below what band separation needs. The pivot was correct and incomplete: the plumbing defect was real and is fixed, three proven factors reach the served number for the first time, and transmission alone did not buy grade separation. Next arc is factor STRENGTH and BREADTH, not more plumbing. PHASE 6 — rbi anomaly logged, not chased: 14.51% variance explained vs hits 1.03%, on the stat we do not serve corrected and which has no proven factors. Either the biggest lever on the board or a mirage; it deserves its own order. The byte-identical invariant INVERTED for hits by design. All 13 frozen non-hits modules verified unchanged, probabilityEstimator included -- the factors ride outside it. No new Bonferroni slot; the composed OOS claim is reported with its CI and not claimed as a pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
e872eff4ce |
Instrument the calibration duel forward; diagnose the resolution ceiling
— the proven factors were never wired in
PHASE 0 — two truths recorded. The swap is a BET, not an OOS win:
isotonic beat low-param on identical held-out rows (hits +0.0028, rbi
+0.0042, TB tied) and we serve low-param anyway on an untestable prior
about shared daily structure. At 19 dates nothing here can test it. And
the MIN_SLOPE catch is preserved as standing rationale: a near-zero or
negative slope collapses toward base-rate-for-everything, which LOWERS
Brier while destroying all resolution -- a metric win that guts the
product.
PHASE 1 — the duel is now falsifiable. Both corrections computed on every
hits/TB prop; p_win_lowparam served, p_win_isotonic_shadow logged in its
own try so it can never break serving. calibrationDuel.adjudicate encodes
the rule IN CODE before any forward date exists: >=10 forward dates and
isotonic winning with a date-block CI excluding zero => REFUTED, revert;
otherwise UPHELD; under 10 dates PENDING regardless of the numbers. A
date counts as forward only if NEITHER map was fitted on it -- otherwise
we would be scoring which map memorised better. Nothing swaps now.
PHASE 2 — the ceiling, quantified via Murphy decomposition:
stat reliability RESOLUTION uncertainty variance explained
hits 0.01353 0.00252 0.24532 1.03%
TB 0.01419 0.00442 0.24329 1.82%
rbi 0.00654 0.03268 0.22531 14.51%
runs 0.00788 0.00130 0.23182 0.56%
Calibration did exactly what theory says and nothing more: hits
reliability 0.01353 -> 0.00233 (-0.0112, 83% of the error removed) while
resolution moved -0.0002. Unexpected: rbi has 13x the resolution of hits
and is the one stat we do NOT serve corrected -- it needs calibration
least and discriminates most.
PHASE 2 DIAGNOSIS — NOT-TRANSMITTED, and not weak, ABSENT. Traced in code:
sprayDefense.js and platoonSeverity.js are required by NOTHING in src/,
only by analysis scripts and their own tests. The served p_win
(intelligence/probabilityEstimator.js:54) reads exactly four inputs --
game-log frequency, opp_rank_stat +/-0.03, home_away +/-0.015, and a cv
pull -- with zero occurrences of spray, platoon, hard-hit or
contact-profile. And snapshotService grades at line 454 while computing
challenger/context at 640+, so everything proven is computed DOWNSTREAM of
the grade it would inform. The three proven hits factors have never once
moved a served number.
That reframes the recent nulls: "calibrated p_win does not separate within
archetype" was never a statement about factors. The factors were not in
the forecast.
PHASE 3 — bands rebuilt on SERVED values (hits/TB low-param, rbi/runs
raw): 28 archetype slots across four stats, ZERO show lift. No longer an
open shrug -- it is the arithmetic of resolution 0.0013-0.0327 against
uncertainty ~0.23. A forecast explaining 1% of variance cannot produce
separating bands, and no correction to its numbers will change that.
HEADLINE: calibration is complete, delivered honest numbers on two stats
and zero grade separation, because the counter has no resolution -- and
the proven factors are not wired into the forecast at all. The second is
the reason for the first, and it is plumbing rather than a modelling wall.
Per-archetype grades need proven factors that actually reach p_win. Last
calibration order.
Serving unchanged from
|
||
|
|
74cf1ce974 |
Robust bias established; low-parameter correction replaces isotonic
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.
PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.
PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:
hits a=0.406 shrink 0.565 0.2626 -> 0.2540 CI [-0.0112,-0.0069] DEPLOY
total_bases a=0.472 shrink 0.333 0.2490 -> 0.2429 CI [-0.0062,-0.0059] DEPLOY
rbi a=0.775 shrink 0.231 0.2011 -> 0.2007 CI [-0.0007, 0] REFUSE
runs a=-0.032 REFUSE
A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.
STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.
PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at
|
||
|
|
ced40421ed |
Audit the LODO instrument: it cannot evaluate any stat, and both prior
FAILs were false
PHASE 0 — the gate at
|
||
|
|
1f40014256 |
Power-derive the LODO threshold: hits restored through the gate, rbi/runs
routed as date-driven PHASE 0 — threshold derived BLIND, before any stat was re-read. A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count. Per-row Brier difference d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single stat's verdict could shape the threshold deciding it: pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175 n* = (0.09816/0.01175)^2 = 69.8 -> 70 The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect this removes, and why the previous verdict moved with the number. Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it equals the constant, so it cannot drift from its own justification. The derivation script prints no stat verdict, no date and no reversal. PHASE 1 — LODO at n*, applied cold: hits 5 informative drops, 0 reversals PASS total_bases 4 informative drops, 0 reversals PASS rbi reverses 2026-08-01 (n=99) FAIL runs reverses 08-01 (n=86), 08-05 (244) FAIL hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across 123-272 row dates, favourite sign holding on every testable drop. THIS IS THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the withdrawal at |
||
|
|
6ae11f1193 |
LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration layer without challenging the binding. That floor is a cluster-robust interval bar for a CAUSAL claim. Calibration makes no causal claim, has a bounded failure mode (it can only over- or under-shrink) and consumes no Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and leave-one-date-out tests that directly -- a STRICTER bar, since a cluster count cannot detect a single day carrying the effect. The >=40 floor is retained, correctly scoped as the PROMOTION bar. PHASE 1 — both guards codified, 11 tests, green before Phase 2. Demonstrated on live data: raw population violated=true, mean_p 0.4962, both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The null guard's test demonstrates the trap explicitly, since (null-1)**2 is 1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate. PHASE 2 — LODO: hits n=1140 dates=17 2 reversals (07-22 n=20, 07-26 n=25) FAIL total_bases n=1050 dates=7 0 reversals, 0 sign flips PASS rbi n= 630 dates=5 1 reversal (08-01 n=99) FAIL runs n= 597 dates=5 2 reversals (08-01 n=86, 08-05 n=244) FAIL Threshold sensitivity reported because the verdict moves: total_bases passes at every held-size threshold, runs fails at every one, and hits fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20 before seeing which stats passed and did not move it afterwards to preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a standard error several times the effect, so the instrument is underpowered per-drop -- an argument for pre-registering a higher threshold, which is a Roundtable call, not one to make while holding the results. PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and runs REFUSE. HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService hardcoded it since S91; it fails LODO, so it is out. A stat that cannot survive dropping one day was never calibrated, it was fitted to that day. The consequence is real -- hits props become unstackable for chain.chainAcross -- and it errs toward withdrawing a claim rather than preserving one on a fragile verdict. Deployment is now driven by a frozen, tested CALIBRATION_DEPLOYED set, not a hardcoded stat name. PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither waivable. reverify auto-demotes on the first breach (CI stops excluding zero, or the favourite bias flips sign) and logs the breaking date. Promotion needs the original >=40 bar. A provisional deploy that cannot be taken away is just a deploy. PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The two-bar rule still bites: calibrated YES, proven NO, so they stay a base-rate read, now honestly numbered. Every archetype still collapses to one band -- calibrated p_win separates within archetype no better than raw. PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI, and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did not survive); the refused set is a map of missing inputs; a low-parameter calibrator is queued unbuilt. p_win never mutated; calibration rides as p_win_calibrated with calibration_status provisional. No Bonferroni slot consumed. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
b2e4c6c4fb |
Link 2 at the coarse grain: pen QUALITY proves, archetype does not
The refinement was right. Naming the individual reliever failed; the same
question at the grain the chain needs passes, and it transmits more than
anything else measured in this chain.
WHY IT WAS WORTH RE-ASKING: last session's null (the pen is on average no
softer, +0.0010 on 35,760 PAs) does NOT rule this out, and treating it as
though it did would have been the error. An average washing out is fully
consistent with quality VARIATION mattering. It does -- actual arm quality
moves the hit rate monotonically across quartiles, 0.2244 / 0.2293 /
0.2410 / 0.2501, a 2.57pp spread, larger than the whole times-through-
the-order effect.
CLUSTER UNIT CORRECTED, THEN CHECKED RATHER THAN ARGUED. Last session
refused Link 2 partly as team-borne (30 bullpens, the park ceiling). My
first re-check was that 76% of pen-quality variance is within-team -- but
that is a statement about TREATMENT variance, not about where errors
correlate, and stopping there would have been picking the convenient
answer. Measured the actual thing: ICC of prediction error by team =
0.0261, design effect 1.41, SEs inflated ~19%. So the verdict was run
three ways:
unclustered CI [-0.0067,-0.0010] excludes zero
team-clustered (30) CI [-0.0086,-0.0003] excludes zero (below the
40-cluster floor -- indicative, not a pass)
design-effect adjusted CI [-0.0072,-0.0005] excludes zero
QUALITY GRAIN PROVES on the concentrated elevated-early-exit subset:
n=501 team-games, 426 clusters, MAE 0.0294 -> 0.0260, delta -0.0034, CI
[-0.0063,-0.0005] at 110 cumulative tests. Pooled also proves, so it is
not a subset artefact.
ARCHETYPE GRAIN DOES NOT: 0.5669 vs a 0.5309 modal-guess baseline,
corrected interval [-0.1073,+0.0268] spans zero. Two grains tested, one
earned a place -- penQuality.js exposes no archetype and a test asserts
it.
WHAT LINK 3 RECEIVES, which is the number that actually matters -- not
the MAE gain but realized outcome separation, prediction strictly
point-in-time:
predicted BEST pen 167 games 2,044 PAs hit rate 0.2231 +/-0.0180
predicted WORST pen 167 games 1,799 PAs hit rate 0.2501 +/-0.0200
2.70pp separated, intervals non-overlapping, capturing nearly all the
2.57pp available at the quartile grain. Caveat stated not buried: the
tercile cut is chosen in-sample; the prediction driving it is not.
BUILT: penQuality.js + 9 tests. Abstains below 5 prior club games and 40
arm appearances -- a league-average stand-in would assert "this is an
ordinary bullpen", which is a claim, and usually the wrong one for exactly
the clubs whose pens just turned over.
Link 3 is unblocked on a proven Link 2 at the quality grain only. Not run
here; this order scopes to building and gating Link 2.
Parallel track logged unchanged: TB n=948 pooled, BOMBER x TB 340, short
by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
e4dae0e6b0 |
Reliever chain: Link 1 proves, Link 2 does not, and the premise inverts
The causal insight is right -- the game is a sequence and the matchup does shift mid-game. The direction is backwards, measured on 93,663 plate appearances from 1,238 games pulled free from statsapi. LINK 1 PROVES. Starter batters-faced, point-in-time from his own prior starts only, clustered on the pitcher: MAE 3.2226 -> 2.7990, delta -0.4236, CI [-0.6006,-0.2731] at 0.9995 corrected for 107 tests, 1,706 starts across 204 pitchers. It finds the tail the chain needed -- early exits are a 23.2% base rate, model-flagged starts are 34.0% early, lift +10.8pp. Scope correction inside Link 1: the order specifies fatigue x GAME SCRIPT, but game script is not available at grade time -- whether he gets hit tonight is the thing being projected, not an input to it. Only the workload half is measured; the in-game half is recorded as a live feature, out of scope, rather than quietly folded in. LINK 2 DOES NOT PROVE, twice over. Model accuracy 17.2% vs an 8.6% baseline -- doubling it sounds good and is not, since naming a specific arm is wrong five times in six. And structurally the entity is the BULLPEN: 39,629 post-starter plate appearances across 30 clubs is 30 readings, below the 40-cluster floor, the same permanent ceiling as park geometry and team defence. LINK 3 NOT RUN, per the order's own rule. THE PREMISE IS REFUTED, and this chains on nothing so it was safe to measure: vs STARTER n=48,492 hit rate 0.2444 +/-0.0038 vs BULLPEN n=35,760 hit rate 0.2373 +/-0.0044 The pen is 0.7pp HARDER. The specific effect the chain exists to exploit -- early exit making later at-bats softer -- is +0.0010 on 35,760 PAs. A well-powered null, not a sample problem. What IS real is times through the order: TTO1 0.2351 -> TTO2 0.2515 -> TTO3 0.2518. A starter does decay as the lineup sees him again, but that advantage is SURRENDERED when he leaves, not extended -- the pen is harder than his second and third time through. A modern bullpen is a queue of fresh specialists throwing one inning each; there is no tiring arm to punish. So the insight survives inverted, and Link 1 stays valuable for the opposite reason it was built: a likely early hook predicts the hitter LOSES his third-time-through look (0.2518 -> 0.2373 on that PA). The mispricing is on hitters who get an EXTRA look at a starter going deep. BUILT: predictionGate.js + tests -- the two-part gate for a continuous prediction. factorGate binarises outcomes for Brier, which would destroy a target like batters faced. Same discipline, same THEATER verdict, real scale. PRE-REGISTERED NOT RUN: Link 2' using a PA-weighted bullpen AGGREGATE rather than a named arm. Recorded rather than substituted in -- running Link 3 on a swapped-in Link 2 is the assumed-link failure the order forbids. Given the premise result its expected value is now low. PARALLEL TRACK logged: total_bases n=948 pooled, BOMBER x TB 340, short by 160. Sample-readiness only, not a verdict. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
3081c92e00 |
Per-archetype grade bands: built, gated, and the rescale blocked twice
The premise does not hold. proven-status.js run fresh: PROVEN_SET is EMPTY, no archetype x stat reaches the gate. pitcher_contact_profile has a CI upper bound of exactly 0.0000 and platoon_severity is held on 4.5%-contaminated splits, so the proven set is one factor, pooled, not three archetype-conditioned ones. The specific pattern the order names -- defense strong for GHOST/BRUSH, null for BOMBER -- is the one I measured running the OTHER WAY yesterday, both noise-dominated. But the second blocker is new and matters more, because it would stop the rescale even if the factors had proved: the grade does not separate within any archetype. Every archetype collapses to ONE band at the corrected bar, because bands merge when their intervals overlap and publishing two letters we cannot tell apart is a distinction we have not measured. Uncorrected, so the ranking is visible rather than hidden by the bar, this INVERTS the order's design. The order gives contact types the factor-rich treatment and power types honest base-rate, reasoning that single-game hits are variance for a power profile. Measured: BOMBER n=466 corr(p_win,outcome) +0.207 quintiles 0.75 0.62 0.60 0.48 0.48 GHOST n=192 corr(p_win,outcome) -0.007 quintiles 0.47 0.63 0.74 0.58 0.45 BOMBER is the one archetype the model ranks, and it splits into a real A 0.660 / B 0.481 at 95%. GHOST is flat, and non-monotone -- its most confident reads hit 47% while its middle reads hit 74%. Shipping as specified would have given the factor-rich treatment to the archetype the model reads worst and left base-rate on the one it reads best. That is mechanically sensible in hindsight: a power hitter's hit tracks whether he can damage the arm, a contact hitter's depends on balls finding holes. BOMBER's split does not survive the cumulative correction at 106 tests. Exposing it by loosening the correction is the curve-to-make-A's the order forbids, so it stays one band. BUILT: gradeBands.js -- lift against the archetype's OWN base rate (the same 62% is lift for a 45% profile and a deficit for a 68% one), indistinguishable neighbours merged, thin bands PROVISIONAL not dropped, Wilson intervals widened by the cumulative correction. The two-bar rule is structural: proven-alone, calibrated-alone and neither all return base_rate with the reason stated, so with nothing proven no factor-informed band can be produced at all. reasoning() is built and tested but NOT wired to the card -- there is no per-archetype band being served, so attaching the copy now would ship product language for a rescale that does not exist. NOT BUILT: the specified power-type reason "the matchup edge is in total_bases". total_bases is recorded INCONCLUSIVE (+0.0038, CI [-0.068,+0.075]). Wiring it would assert an edge measured as indistinguishable from zero -- the exact fabricated-reason failure this module exists to prevent. BOMBER x hits is 29 rows short of the gate and is the archetype the model actually reads. That is the first slot to test, not GHOST. Counter and frozen clusters byte-identical. No letter was moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
6b17f79367 |
Per-archetype re-audit: no slot reaches 500, and the replication unit
decided everything The premise does not hold. prove-hit-factors.js has no date filter anywhere in it and pages the full table -- there was never a window to widen. Full clean history is 1,266 rows, not 2,715. platoon was not "proved" last session, it was explicitly held on 4.5%-median-contaminated season-to-date splits, and pitcher_contact_profile was demoted. The proven set going in was one factor, not three. STEP 1: no archetype slot reaches n>=500 on full history. Best is BOMBER at 408, and BOMBER is the most common archetype on the board. GHOST 173, BRUSH 64, DRIVER 43, CATALYST 16. These are confirmed genuinely short, not artifacts. STEP 2 is where the real finding is. park_hits initially PROVED at 619 rows across 45 games -- but those games only ever visited 14 distinct park values. A park effect is replicated across parks, and unmodelled park heterogeneity is confounded with the thing being estimated. Each factor is now clustered on the coarser of the game and the entity its treatment rides on. That flipped two verdicts and confirms Kev's causal-correctness thesis from a new direction: defense_by_direction has 442 hitter-team units of replication where crude team defense has 26. The correct atom is not just more accurate, it is the only one measurable at all. park_hits (14) and defense (26) can never be validated however long the ledger runs -- the same ceiling as park dimensions, reached independently. Also fixed a bar I got wrong last session: I transplanted the 500-row floor onto clusters, which refused a factor with 1,059 rows over 85 games while answering neither question. Two floors now -- rows>=500 for a stable estimate, clusters>=40 for a trustworthy interval. Not a lowered bar: park_hits and defense are still refused. PROVEN: defense_by_direction only, pooled, [-0.0054,-0.0012] at 99 tests. It stays POOLED-ONLY -- no per-archetype reasoning wired, nothing grandfathered. The card must not say "GHOST: defence matchup strong" because we have not earned that sentence. The predicted fingerprint did not appear either: BOMBER -0.0036 vs GHOST -0.0024, the opposite direction, both noise-dominated. Recorded so it is not claimed later. RESCALE: NOT READY. One proven factor worth -0.0031 Brier. Rescaling on that is relabelling. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
7b85934dc3 |
Under-querying vs out of data: the answer depends on the unit
The platoon test's n=452 described how much of the JOIN survived, not how much data exists. There are 1,266 clean settled hits rows and zero quarantined ones. platoon_splits had been ingested from tonight's lineups only (315 players), so any hitter who settled a prop without appearing in an ingest-day lineup was silently absent from every test. Backfilled all 380 hitters (81 fetched, 0 unresolved). Re-ran on 1,059 rows, up from 452. THE DEMOTION IS THE HEADLINE. pitcher_contact_profile, the strongest proven factor in the programme (-0.0064, CI [-0.0113,-0.0014]), roughly halved to -0.0034 on more than double the sample and its corrected interval now spans zero. The Bonferroni denominator also rose to 55, which widens every interval -- but a denominator cannot move a point estimate, and that halved on its own. platoon and platoon_severity now clear the bar and are NOT promoted. Upper bound -0.0001, on season-to-date splits that contain the games they predict: measured contamination is 4.5% median, 12.4% at p90, 137% worst. I had assumed ~1%. They stay CANDIDATE pending point-in-time splits. GAME-LEVEL IS A DIFFERENT PROBLEM. game_context held zero weather rows ever -- not because the fetcher was wrong (it correctly targets Open-Meteo's archive) but because ledger_entries keys a game as mlb:2026-08-03:Away@Home and game_context keys it as mlb:823437. Every lookup missed and NULL columns read as honest absence. Third occurrence of that class. Fixed the join: 96/101 settled games now carry actual archived weather, park dimensions backfilled 15 -> 30 venues. But 928 total_bases rows sit on 47 games at 17.6 rows per game. Park and weather assign one value per game, so resampling rows would have manufactured a pass. factorGate now resamples clusters when rows carry one and judges sample against effective_n; unclustered rows keep the original path byte-for-byte. Verdict: 47 clusters < 500, and the point estimate is +0.0011 -- worse, not merely unproven. Weather needs ~57 more days. Park dimensions need never: there are 30 ballparks in MLB, so a venue-constant factor can never reach 500 independent units. That bar was built for player-level factors and does not transfer. Wind is refused. We have speed and bearing for all 96 games; we lack park orientation, and 220 degrees is blowing out at one park and in at another. Using speed alone would assert an effect while discarding the sign that decides what it is. Counter and frozen clusters untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
de0077f6f9 |
Causally-correct platoon + park-dimensions ingest
Applying the method that worked for defence to the two factors the code flagged as still crude. PLATOON. The flat version is 'lefty versus righty, add a boost', and it failed the two-part gate for the same reason team-average defence did: it is not the unit the causal story runs through. The advantage is only worth what THIS hitter's split is actually worth -- measured on a real hitter, .284 against left-handed pitching versus .221 against right-handed, a 63-point split, where the flat factor applied the same six percent to him and to a hitter with none. Most of the work is sample discipline, and the second rule matters more than the first. Severity shrinks toward the league split weighted by the SMALLER side's plate appearances, because a 500-against-40 split is a 40-PA read. And below a floor it REFUSES outright rather than shrinking, because a heavily-shrunk severity is indistinguishable from a measured league-average one and those are different claims -- without the refusal the atom would quietly assert a league-typical split about every September call-up in the league. Switch hitters turn out to be the easy case misread as the hard one. He bats opposite by choice so the direction is never in doubt, but the per-side value of his swing is a different question and one this sample cannot answer, so he is unreadable rather than credited with an automatic edge. PARK DIMENSIONS. Free from statsapi's venue endpoint, which carries fence distances, roof, turf and elevation outright -- Wrigley returns 355 down the left line, 400 to centre, 353 to right, at 595 feet. parkFactors holds run COEFFICIENTS, which structurally cannot express a park that turns outs into hits without scoring, and that is why the crude park factor failed. The park join is by the venue the game is ACTUALLY at, carried from the schedule feed, never inferred from the home team -- neutral-site and international games break that assumption and they break it silently. A venue with no geometry at all is absent rather than a park with zero dimensions. Both tables dated in the primary key. Venue geometry changes rarely but it does change, and by now that is the default rather than a lesson. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
405180e791 |
Build the causally-correct defence atom: spray x positional OAA
Team-average defence failed the two-part gate for hits, and the reason was the unit rather than the signal. A left-handed pull-ground hitter meets the first baseman and the second baseman and almost nobody else, so a team total averages in five fielders who will never touch his ball. Both halves were already free on the host we pull from. Statcast publishes spray x trajectory per hitter -- pull/straight/oppo crossed with ground/air, 608 hitters -- and the OAA feed already carries each fielder's position, so per-position defence is a regrouping of data ingested last week rather than a new source. Zero new sourcing, as the order expected. Handedness is what joins them and getting it backwards would be invisible: pull for a right-handed hitter is the left side, pull for a left-handed hitter is the right side, so a model ignoring bats would send half the league's grounders to the wrong infielders and still look like it was reading defence. A switch hitter bats opposite the pitcher, which this does not resolve, so he is unreadable rather than guessed. Two properties the crude version could not express, both locked by test: two teams with the SAME total defence read differently for a pull hitter, and a ground-ball hitter and an air hitter read the same team in opposite directions. Unmeasured zones are renormalised away rather than contributing a zero, which would assert an exactly-average fielder standing there, and states honestly what share of a hitter's contact we could actually read. Nothing readable at all returns null, so the caller falls back to the base rate instead of to an invented 1.0 that looks measured. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
a9ee55550b |
Build the two-part factor gate: one factor proves, and zero are theatre
The question was whether the hit grade reads tonight's game or just says he is due. Answering it needed a gate that correlation cannot provide, because correlation cannot separate the two ways a factor looks alive: it reads the game, or it moves the number and reads nothing. The second is what a product ships by accident -- arch-v1 moved 76% of rows by 2.5 points, changed resolution by 0.0000, and was live for months, and no user could have told. So a factor must now clear both conditions: move the prediction off the player's own leave-one-out base rate, AND improve out-of-sample Brier. Brier rather than correlation, because correlation asks whether the ordering improved and this asks whether the NUMBER got closer to what happened -- and for a graded probability the number is the product. The correction applies to the interval itself, which turned out to matter more than expected. A plain 95% CI is the right bar for one test; at fifty cumulative tests roughly two or three intervals exclude zero by chance alone. Widening to 1 - 0.05/tests, currently 99.9%, flipped both defence and platoon out of "proves". A 95% interval would have shipped two unproven factors into the grade, with reasoning text explaining them to users. That forced a distinction I had initially collapsed. Defence and platoon have FAVOURABLE point estimates whose corrected intervals merely span zero, and calling that THEATER would repeat the error this codebase keeps correcting: insufficient evidence is not evidence of absence. THEATER is now reserved for its one real meaning -- moves the number, reads nothing -- and NOT_PROVEN_AT_CORRECTED_BAR names a real candidate held to a bar that rises with every hypothesis the programme tests. Result on 741 settled hits rows: pitcher_contact_profile PROVES, improving Brier by 0.0066 with a 99.9% interval of [-0.0114, -0.0016]. Defence (-0.0043) and platoon (-0.0039) are not proven at the corrected bar. Park is sample-blocked at n=405. Zero factors are theatre, which is the genuinely good news: nothing decorative is being wired. Per-archetype every slot is sample-blocked (BOMBER 252-294, GHOST 67-125). Two spec gaps worth recording. The approach identities the order names -- SPRAY, DAMAGE-DEALER, COUNT-WORKER -- do not exist in the registry; the MLB batter archetypes are BOMBER, GHOST, TORCH, BRUSH, DRIVER, FLEX, ALPHA, HYBRID and CATALYST. And parkFactors maps hits to run_base, so there is no hits-specific park factor at all: a park that turns outs into hits without producing runs is invisible to the input we have. The grade rescale is NOT run. It was explicitly gated on the factor proving, and one pooled factor worth 0.0066 of Brier is not a factor-informed distribution -- rescaling on it would dress a base-rate model as a matchup model, which is the exact thing this gate was built to prevent. 4,286 tests green (340 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
4d1803f6d7 |
Calibrate hits point-in-time: partial pass, and an honest ceiling of 0.667
Fitted the isotonic map on game_date < 2026-08-02 (n=589) and evaluated it on everything from that date forward (n=383). The map never saw the evaluation rows, which is the only thing that makes the result mean anything -- fitting and evaluating on the same rows always looks perfectly calibrated, because the map is reciting the answers it was built from. It works, on most of the distribution. Held-out after correction: 0.477 comes back 0.506, 0.587 comes back 0.580, 0.667 comes back 0.603 -- against raw errors of +0.191, +0.279 and +0.246 in the same bins. Ordering survived, and that was verified pairwise rather than assumed, because a broken map would silently destroy the one thing this model does well. Two findings matter more than the pass. First, the honest ceiling is 0.667. Once the numbers are truthful this model has no 80%-plus hit reads at all -- the top of its range was miscalibration, not confidence. A four-leg ticket at the ceiling is 0.198, where the raw numbers implied 0.686. The high-floor parlay is a two-thirds-per-leg proposition, and that is the number to say out loud. Second, calibration is certified BY BAND rather than by a blanket flag. Held-out error was -0.029 and +0.007 through the middle but -0.167 at the bottom and +0.063 at the top: the model is trustworthy over most of its mass and untrustworthy at both edges. A single true/false would either throw away the 72% that works or ship the edges that do not. Only a probability inside a certified band is marked stackable, and that flag is what chainAcross requires before it will compound anything. The certified band is 0.40 to 0.60, n=276. A methodological catch on the way: my first pass condition demanded honest bins at 0.70 and above -- but honest calibration REMOVES those bins, since the ceiling drops to 0.667. The gate would have failed the repair for succeeding. It now tests the highest remaining band instead of a fixed threshold. Wired forward with the same discipline: calibrationService fits strictly before today, splits by time rather than at random, and returns null on thin history so that "no calibrator" means nothing is stackable rather than "trust the raw numbers". p_win is never mutated -- the calibrated value rides beside it as p_win_calibrated, because a calibration map is a correction to a forecast, not a different forecast, and the counter stays byte-identical. 4,275 tests green (339 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
9c5b968351 |
chaining-v1: the portable chain, and the gate that blocks the parlay surface
The order's own prerequisite for the hit-parlay surface was to verify the hit probability is calibrated. It is not, and the failure is exactly the shape that destroys a parlay. Measured on 972 settled hits props: the model is monotonically over-confident at the top and flat above 0.70. Predicted 0.911 comes back 0.630. Predicted 0.844 comes back 0.630. Predicted 0.747 comes back 0.605. There is no discrimination at all in the range a parlay is built from, and the error runs in the flattering direction. Four "91%" legs are 0.686 by the model and 0.157 in fact -- a 4.4x overstatement that compounds with every leg added. Single props survive a calibration error of that size. A parlay multiplies it. So chainAcross REFUSES to compound atoms not marked calibrated, and refusing is the feature rather than a limitation: a ticket built on these numbers would be confidently wrong in the direction the user pays for. calibration.js provides the reliability table, the gate (tolerance 0.05, weighted to the high end because that is where tickets live) and an isotonic fit. Isotonic is the honest repair here because it is monotone: the model's ordering survives untouched while the numbers move to what actually happened. The fitted map says 0.65 -> 0.594, 0.85 -> 0.639, 0.91 -> 0.639. chain.js is the portable core -- base events plus context, through a chain function, into a PLUGGABLE aggregator: across players for a compound ticket, up to the team for expected scoring. The sport-specific parts are inputs rather than code paths, so basketball plugs in as content. The archetype redistribution hook is there now, dormant in baseball because a nine-run lead does not change who bats next, and live in basketball where a blowout fades the star and feeds the bench. Two judgement calls worth naming. Treating same-game legs as independent errs in the FLATTERING direction, since they share pitcher, park and weather -- so correlation shifts the compound toward the weakest leg, bounded, and is labelled an approximation rather than a joint distribution. And market divergence does NOT downgrade confidence: it flags a contested script whose props are either the best or the worst on the board, and which one is unknown until settled. Internal inconsistency does downgrade it, because per-entity reads failing to sum to the team read means one of them is wrong and we do not know which. Not built: the independent game-script projection. It needs proven team-level atoms and out-of-sample validation against actual margins, and no atom has passed the gate yet. Building it now would produce something plausible rather than something proven, which is the failure mode this whole programme exists to avoid. 4,269 tests green (339 suites); web build exit 0; counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
08276c0880 |
Ingest lineup + baserunner context: the input RBI and runs always needed
RBI is power TIMES opportunity. The same swing drives in one run or three depending on who is on base, and a hitter batting with the bases empty cannot drive anyone in however hard he hits it. Every context-free model of RBI here has failed, and the failure kept being read as 'skill inputs don't work for RBI' when the truth was that we were modelling half the stat. Both halves are free from statsapi.mlb.com, which we already call for game logs, schedules and probable pitchers. No new provider, no key, no quota. RUNG 1, batting order: schedule?hydrate=lineups returns homePlayers and awayPlayers as ORDERED arrays of nine, and the order IS the batting order -- index 0 is the leadoff hitter. That single fact gives CATALYST its identity and supplies lineup-position context for every context-dependent stat. RUNG 2 turned out cheap, which the cheapest-first rule did not expect. It looked like it would need play-by-play reconstruction across a season; statsapi serves situational splits directly, so 'how often does this hitter bat with runners to drive in' is ONE call per player rather than one per game. Measured on a real hitter: 87 plate appearances with runners in scoring position producing 25 RBI, against 302 with the bases empty producing 17. That ratio is the opportunity half of the stat and it is the thing no amount of exit velocity can tell you. Both tables are dated in the primary key. statcast_aggregates was built upsert-in-place and that silently made every backtest leak the games it was predicting; a lineup is worse still, because it is a PRE-GAME fact that changes by the hour, so an in-place table would overwrite what we knew at grade time with what turned out to be true. Absent stays absent throughout: no lineup posted is an empty slate rather than a guessed order, a short lineup records fewer slots rather than padding to nine, and a hitter with no splits is null rather than a zero RISP share -- which would assert he never bats with runners on, a strong claim and usually a false one. Wired into the snapshot best-effort, so a context failure can never break the pipeline it rides in. The three pre-registered theories are now marked input-ready rather than input-blocked: DRIVER's power x runners-on and power x lineup-position, and CATALYST's speed x on-base x power-behind. They are sample-blocked from here, and the proofs run under native cumulative correction as sample accumulates -- ingesting is not proving. Counter and frozen clusters byte-identical. 4,250 tests green (338 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
ff037e40c2 |
Re-adjudicate: nothing to demote, and close the hole that would have mattered
There is nothing to re-adjudicate. The proven set is empty and always has
been -- verified three ways: proven-status reports EMPTY, validatedSkills()
returns {} for every archetype, and zero conditioning entries have ever
reached PROVEN. The one PROVEN feature is recent_frequency_prior, which is the
incumbent counter itself, proven by the S78 ablation as ~100% of the
champion's resolution. It is the baseline every challenger is measured
against, not a conditioning interaction, and demoting it would leave the model
with nothing to grade from.
A correction to the premise: the cumulative gate did NOT catch a false
positive last session. It caught nothing, because there was nothing in the
proven set to catch. What it did was tighten alpha from 0.0026 to 0.0013
within one session, which demonstrated the mechanism working rather than a
demotion. So steps 3 and 4 -- demote, recalibrate -- are vacuous here, and
readjudicateAll says so plainly rather than glossing a no-op.
But the worry behind the order was well founded, and the audit found the real
exposure: promote() did not require the cumulative denominator. It checked n,
lift and CI, and nothing stopped a future session from testing eight
hypotheses, correcting by eight, and promoting on a p-value that would not
survive the programme's real denominator. That is precisely the hole that
makes a retroactive re-adjudication pass necessary later, so it is closed at
promotion time instead. isSufficient now refuses evidence carrying no
correction, evidence corrected against fewer tests than the cumulative count,
and any p-value that does not clear 0.05 over its own test count. The same
rule guards a PROVEN conditioning entry.
The second audit found two of four analysis scripts still correcting
per-session; pitcher-prove-k and tb-solo-and-interactions now use the
cumulative ledger, so the correction is native on every path.
reAblation.js is the standing second line: pure and injectable, so the
decision rule cannot drift from the gate's, and every verdict records both
p-values and both test counts so a demotion is re-derivable by anyone. A
feature promoted at alpha 0.05/20 can demote on the same p-value once the bar
is 0.05/60 -- correct, because the bar rose only after the programme had more
chances to get lucky. No fresh measurement is PENDING_RETEST and never a
demotion: absence of a re-test is not evidence, and demoting on it would
punish whichever stat happens to be off-season.
Net effect on the proven set is zero. No demotions, no recalibrations, and no
public ledger event -- announcing "recalibrated after re-adjudication" when
nothing changed would itself be a false signal of rigour.
4,238 tests green (337 suites); web build exit 0; counter byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
ece2b9f5f9 |
Ingest defence, and make Bonferroni cumulative across the programme
Two things shipped that stand regardless of sample. DEFENCE. Statcast Outs Above Average is free on the host we already pull six feeds from, so there was nothing to decide. 514 fielders, aggregated to team level -- the unit a batter's prop actually needs, the defence behind the pitcher he faces -- and persisted as 31 team rows. Verified in production. Cubs +56 best, Mariners -29 worst. Unknown is not zero, and it bites unusually hard here: an OAA of 0 is a REAL reading meaning exactly average, so coercing absence to 0 would assert that every unmeasured fielder is league-average, which is the commonest defensive profile there is. team_defense also carries as_of_date in its primary key from the first row -- statcast_aggregates was built upsert-in-place and that silently made every backtest leak the games it predicted, so point-in-time is available here before it is needed rather than after a wrong answer. A bug worth recording as a class: BASE already ends in /leaderboard, so the new feed built a doubled path and 404'd. Because a failing feed degrades to an empty index by design -- correct, so one broken source cannot fail the whole pull -- it surfaced as "fielding_oaa: 0 rows", which reads exactly like "Statcast has no fielding data". Graceful degradation makes a wiring bug look like an honest absence. CUMULATIVE CORRECTION. Bonferroni had been applied per session throughout: a run testing eight features corrected by eight. Across a programme's lifetime that is wrong in the dangerous direction, because every order gets a fresh generous alpha and the false-positive rate compounds quietly. Correcting by 8 when sixty have been tried is how a noise result eventually gets recorded as PROVEN with a p-value to point at. The denominator is now distinct hypotheses ever tested, persisted, and it moved 19 -> 38 within this session alone, alpha 0.0026 -> 0.0013. Re-tests deliberately do not inflate it: re-asking the same question on more data is not a new shot on goal, and counting it would punish the discipline of waiting for sample. THE MEASUREMENT. The differential the theory predicted is present: defence correlates with the counter's residual at +0.130 for GHOST, the contact and speed archetype, and -0.018 for BOMBER, the power archetype. A GHOST's hits depend on whether anyone can range to the ball; a BOMBER's barrels clear the defence entirely. So a flat BOMBER result is the theory working rather than the test failing. It is not a result. GHOST is n=104 against a 500 bar, with p=0.188 against a corrected alpha of 0.0013 -- three orders of magnitude short. Both are recorded as CANDIDATE with their measured lift, tagged contact-skill, so the re-run at full sample compares against a recorded baseline. Nothing proved, so nothing was recalibrated and nothing shipped. 4,228 tests green (336 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
843c8c6d4b |
Build the pitcher engine, and find the cap was eating the whole board
Strikeouts are NOT proven -- n=57 against a bar of 500. But the finding that matters is not a correlation. THE CAP. Measured on the live slate via the refusal diagnostic: 1,244 unique gradeable props exist, the 500 cap graded about 334, and because dedupeProps takes first-row-wins in FEED ORDER, what survives is decided by feed position rather than value. Pitchers are 2.6% of a batter-dominated feed, so we were grading SIX strikeout props a slate against 32 available -- putting n>=500 three months away for every pitcher stat. Pitcher props were never being refused (graded 5, refused 0, suppressed 0); it was truncation. Raised 500 -> 1500 on measured cost: 721ms per prop at concurrency 5 is about 179 seconds for the full board, against a cron that runs five times a day and a fire-and-forget caller that never holds an HTTP response. statsapi is free and unlimited. Concurrency stays at 5 -- one variable at a time. This unblocks every n-blocked stat in the programme, not just pitchers. THE ENGINE. pitcherEngine.js is its own engine, not the batter engine pointed at pitchers: the batter model asks whether contact becomes a hit and reads contact quality, the pitcher model asks whether the plate appearance ends without contact at all and reads stuff. Archetypes are FLAME (whiff-led), SCALPEL (chase-led), SINKER (pitches to contact) and DEFAULT, and a test asserts the weight keys are not the batter engine's. The projection is K% by log5 against THIS lineup, times batters faced, through a binomial. An unclassifiable arm gets the balanced map, never a guessed archetype. THE MEASUREMENT, at n=57 and contaminated. Four solo features clear the 0.15 effect bar and fail only on sample: arm angle at -0.250 -- the largest correlation measured anywhere in this programme -- then whiff +0.213, k rate +0.206, chase +0.195. The batter cluster's best was 0.135. Head to head, pitch-v1 resolves 0.1285 against the counter's -0.0639, delta +0.192 with a CI spanning zero. That negative is the interesting number. The counter is ANTI-PREDICTIVE on strikeouts: counting a pitcher's recent Ks is worse than useless, because his recent totals track which lineups he drew and how long he was left in rather than his skill. It is the one stat where the incumbent has no defensible edge. A bug caught on the way. resolveTeam wants an abbreviation and the game log supplies full team names, so the roster join silently resolved nothing and the first run reported 0% lineup coverage -- the theorized stuff x lineup carrier was never being tested, not failing. Fixed; coverage is now 94.7%. The carrier still shows no incremental signal over whiff alone, and adding the lineup term lowered head-to-head resolution, which is recorded rather than dropped. Calibration was not reached: nothing passed the first bar. The batter model and the counter are byte-identical, verified by diff. 4,221 tests green (335 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
4aab18096f |
Prove both on total bases -- and find that my own fix destroyed the backtest
Nothing passed. Nothing promoted. Counter byte-identical. THE BLOCKER, which is the real finding. statcast_aggregates is upserted in place and holds exactly one as-of date. Yesterday's skill backtest was honest only by accident: the nightly refresh was unreachable code, so the profiles sat frozen at 2026-07-21 -- before the settled window. Repairing that cron was right for production and it refreshed them to today, destroying every prior version. Scoring a 2026-07-25 game now uses a season aggregate that contains that game. Point-in-time validation is structurally impossible from that table, so every number in this run is contaminated and directional, and none of it is a gate verdict. Fixed forward: statcast_history retains a dated snapshot on every refresh, so point-in-time becomes "as_of_date < game_date, most recent". Retention is best-effort and cannot fail the refresh; both properties are unit-tested. It has one day of data, which is not yet a window. SOLO BASELINE, n=383, Bonferroni across 12 tests (alpha 0.00417): nothing passes. hard_hit_pct is closest at marginal r 0.135 with p 0.0080, failing both the 0.15 effect bar and the corrected alpha. And it drifted DOWN from 0.153 at n=295 -- an estimate regressing as noise averages out, not an effect firming up. I called that number encouraging yesterday; on 88 more rows it is fading, and it should not keep being quoted at its best value. INTERACTIONS, each scored by partial correlation against the counter residual controlling for both of its own components: none pass. Only barrel x power archetype has an incremental exceeding its parts (-0.101 against 0.019) at n=260 -- the shape Discipline 2 predicts, but a lead, not a finding. A methodological catch worth keeping. The archetype conditioner was first built as barrel_pct over league barrel -- a monotone transform of one of its own components -- so the "interaction" was barrel squared, measuring nonlinearity in barrel rate rather than any archetype effect, and it produced this run's only positive result. A Gauss-Jordan pivot test does not catch that, because the two columns differ by a scale factor. Fixed with a scale-free collinearity check plus real archetype labels joined from model_snapshots. Without it this document would have reported a fabricated interaction as the session's finding. COMBINED vs COUNTER on total bases: 0.2718 against 0.2647, delta +0.0071, CI [-0.065, +0.079] -- inconclusive, and the first time a challenger has not lost. The same engine on hits was -0.116 with a CI excluding zero. That contrast is the whole argument for total bases, and it is what the physics said: contact quality governs extra bases, not whether a grounder finds a hole. Also built: the compound TB projection. skillProjection no longer refuses total bases -- a deterministic bases-per-hit multiplier had made P(TB>=2) exactly P(hits>=1), a relabelled hits curve. It is now a convolution over per-PA base outcomes with hit-type shares shifted by skill. Non-degeneracy is locked by test. 4,204 tests green (334 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
c7cc8f5e52 |
Build the gate, run it, and find we were proving things on the wrong stat
PREMISE CORRECTION FIRST. statModel.js and correlateValidator.js do not exist in this repository. The validation spec's only prior form is src/services/python/blueprints/unconventional.py -- a Flask blueprint in the Python service that is offline in production, scoring NBA factors against a warehouse that was never populated -- and tests/unit/supplementSystems.test.js requires only fs and path while defining its own validateFactor inline at line 368. Those tests assert a re-implementation of the thresholds, not an implementation, which is exactly why they passed for months while nothing was connected. The diagnosis behind the order is right -- every challenger was measured without a gate -- but the cause is that there was no gate on the Node side to import. So it is built, to the exact spec. correlateValidator: n>=500, |r|>=0.15, p<0.05, Bonferroni across the sweep. The p-value is exact rather than approximated (t-transform through a regularized incomplete beta) and is verified in the suite against known values, because scipy is not available here. Pairs with an unknown side are dropped, never zero-filled -- a zero-fill inside a correlation does not add noise, it invents a point at the origin. THE RUN, hits, n=570, Bonferroni-8: every skill feature fails, and not narrowly. The strongest marginal correlation against the counter's residual is 0.062 against a 0.15 bar. That is an effect-size failure at a sample that would have found a real effect comfortably -- a clean, well-powered negative. The head-to-head agrees: value engine 0.0499 against the counter's 0.166, delta -0.116 with CI [-0.189, -0.043]. Not promoted. THE RUN, total bases, n=295: cannot be tested, and that is the finding. hard_hit_pct shows a marginal r of 0.153 -- above the threshold -- and exit velo 0.124, refused solely because n is 205 short of 500. It is the most encouraging number this work has produced, and it is what the physics predicts: contact quality governs extra bases, not whether a grounder finds a hole. We have been testing skill inputs on the one stat where they should not matter much. Two things the run forced. Feature verdicts are now PER STAT, because marking these DEAD sport-wide on hits evidence would have killed, for total bases, the features that look most alive there -- per-sport doctrine one level deeper. And the gate now reports r and p even when underpowered, because "not enough data yet" and "nothing here" demand opposite decisions and a bare refusal was hiding the best signal on the board. Next: build the compound TB projection (skillProjection still refuses total bases by design, since a deterministic bases-per-hit made P(TB>=2) identical to P(hits>=1)), accrue to n>=500, re-run this gate. Leave hits alone. 4,200 tests green (334 suites); web build exit 0; counter byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
258d8a6655 |
The skill engine: built, gated by construction, and Stage A honestly lost
Built src/services/model/ -- the forward, archetype-selected, skill-based projection, as a challenger. The champion is untouched. featureRegistry makes "earn its place or it's out" structural rather than aspirational: CANDIDATE / PROVEN / DEAD per feature per sport, liveFeatures() returns PROVEN only, promotion requires n>=200 with positive lift and a CI excluding zero, and there is deliberately no override argument. It ships with exactly ONE proven feature -- the incumbent counter, because it is the only one with a measurement. A test asserts that with only PROVEN features allowed the projection returns null, so an unproven model cannot reach a user by accident. The three champion adjustment layers are registered DEAD with their reasons so they cannot be silently rebuilt. skillProjection is a PA outcome tree: K and BB combined by log5 odds-ratio against league (both identities unit-tested), then archetype-weighted contact quality against contact allowed, then Binomial(PA, p_hit) mixed over a PA distribution. Archetype is a FEATURE SELECTOR, not a nudge -- BOMBER reads barrels at 0.50 and ground-ball speed at 0.00, GHOST inverts it -- and a test locks that the same hitter read two ways moves more than 0.15. STAGE A: IT LOSES. Out-of-sample on 570 settled hits props with 91.9% opposing-pitcher coverage, resolution 0.0499 against the champion's 0.166, delta -0.116 with CI [-0.189, -0.043]. It is not selective either: its eight most confident picks hit 50%, a lift of -0.065. Not promoted. The gate did its job on its first real test, which is the point of having built it that way. Two false starts, both recorded because they nearly produced a wrong verdict: statcast_aggregates stores PERCENTAGES, so raw rows made bip = 1-29.6-17.1 and refused 568 of 576 -- the honest-absent guards made a units bug loud instead of silent, and the conversion now lives at one chokepoint. And the first run resolved an opposing pitcher for 1 of 570 rows, because ledger team/opponent are NULL, so it would have reported "skill-v1 loses" while measuring a batter-only model with no matchup in it at all. The verdict above is from the corrected run. The loss is real but partial: park was passed as 1.0, handedness and opportunity_drift never fired, PA is season-PA over a constant, and the skill profiles carry no recency at all while the champion has a last-5 term. Also fixed: the Statcast nightly refresh was unreachable code. It sat inside tick() below "if (!HOURS_UTC.includes(h)) return" while testing h === 11, so it had never run once; the aggregates were 13 days stale and both of its alerts were in the same dead branch. It now runs on its own tick, and the test that passed happily throughout -- it only checked the string existed -- is replaced by one that asserts it is not behind the guard. 4,182 tests green (333 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
b06a84af80 |
Settlement has been dead since 2026-08-01: a 500-id filter overflowed the URL
The self-learning loop stopped two days ago and reported success the whole
time. 1,444 ledger rows from 2026-08-01 sit unsettled with settle_attempts=0
-- never even attempted -- and every accruing challenger has been starved of
settled sample as a result.
ROOT CAUSE. settleLedger fetched open ids, then REFETCHED the full rows with
.in('id', ids). PostgREST puts filters in the URL, so 500 UUIDs became an
18,499-character request that the fetch layer rejects with "TypeError: fetch
failed". The result was destructured as `const { data: rows } = ...` with NO
error binding, so rows came back null, the loop body never executed, and the
function returned {settled:0, voided:0, unrecoverable:0, pending:0} --
byte-identical to a clean "nothing to settle". Reproduced against prod before
changing anything.
WHY IT HID FOR TWO DAYS. It is volume-triggered. Daily volume ran 20-260 rows
and settled perfectly for weeks; 2026-08-01 was the first day past the 500-row
fetch limit. And the zero-settle ops alarm reads these very return values, so
pending:0 told the watchdog the backlog was empty -- the alarm built to catch
exactly this could not see it.
THE FIX. The refetch existed only to add game_date/settle_attempts/
dclv_computed_at. Selecting them in the first query removes the id list
entirely, so there is no URL to overflow at any volume. A failed fetch now
surfaces its error instead of being reported as an empty backlog.
captureClosing carried the same shape one level down -- .in('id', g.ids) on an
UPDATE, which fails identically once a single line|odds group gets large on a
big slate. Its id filters are now chunked at 100 (~3.7 KB).
Tests: the regression is locked by asserting settlement issues NO id-list
filter at 500 rows, and that a failed fetch is never reported as an empty
backlog -- the two properties that would have caught this. Two existing
suites asserted the old two-query shape and were updated to the real one.
4,159 tests green (332 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
07626de3de |
hits-v1: built on the right structure, measured honestly, REFUTED
Hits was diagnosed as a family mismatch: 84% of hits rows trade at 0.5, so the stat rides on P(0), and a negative binomial has unbounded support and no notion of opportunity at all. hits-v1 models it as the bounded conversion it is -- N ~ the player's empirical at-bat distribution, hits|N ~ Binomial(N,q), with the multiplier scaling q (conversion) and never N (opportunity). STEP 0 confirmed the inputs before the model existed: 30/30 real ledger players, 100% combined-input coverage. Every read goes through knownRate -- a row with no atBats is dropped, never counted as a 0-at-bat game. It FIRES: 158/159 hits props (99.4%) on the live production snapshot, through the real attachProjection path. Scoping by book IDENTITY rather than price shape kept 94 out-of-promotion-band props on the board, 93 of them modelled -- 59% that a price rule would have deleted. And it LOST. Point-in-time replay (game log truncated strictly before each row's game_date, real grade-time multiplier), hits-only, direction-aligned, n=242: resolution champion 0.195 / ladder 0.048 / hits-v1 0.026. Paired bootstrap on the same rows: hits-v1 - ladder = -0.022, CI95 excluding zero. Not promoted. The value is in what it eliminates. The family was wrong AND the mean was not the constraint -- hits-v1 moved the line-0.5 mean 0.554 -> 0.581 toward a 0.598 base rate while resolution fell. What is left is per-prop discrimination: the ladder's inputs, not its distribution. The pre-registered fallback is recorded as WRONG rather than deleted. It said hits might be genuinely low-resolution for anyone; the champion scores 0.276 on the identical 189 rows, so there is real signal and the ceiling claim was the comfortable reading, not the honest one. Its own control refuted it, and that control was already in hand when the branch was written. hits-v1 stays wired as a challenger writing its own ledger columns so the forward accrual can confirm the backtest. Champion, ladder, ranking, calibration, reference ruler and the four accruing verdicts are byte-identical -- the diff has zero deleted lines. Tests 4,156 green (332 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
d103ecf4c3 |
Disambiguate takeable: THREE questions shared one word, now three names
BYTE-IDENTICAL. The audit found no consumer getting the wrong axis, so this is a disambiguation, not a bug fix. 4,131 tests / 331 suites green. STEP 1 AUDIT -- and the order's premise was wrong in a useful way: the four accruing challengers read the flag ZERO times (not four) the ranking gate wants PROMOTION, gets promotion [correct] the ledger column holds the LEDGER band, consumed as such the UI (LiveHeroProp) TYPES a `takeable` field it never renders THERE ARE THREE DEFINITIONS, NOT TWO -- and I only found the third by tracing the ranking gate: 1. IDENTITY can it be bet? book identity (takeability) 2. LEDGER BAND worth recording? odds >= -160, UNCAPPED plus 3. PROMOTION worth crowning? -160..+200, i.e. band PLUS a ceiling (2) and (3) genuinely disagree, and I measured it rather than asserting it: 439 rows -- 28.2% of all takeable=true ledger rows -- carry prices above +200, up to +1300. A +1300 longshot is a real bet worth RECORDING and not one worth CROWNING. Both are correct for their own purpose. THE DANGER WAS NEVER THE LOGIC. It was that three questions shared one word, so a reader could not tell which answer they held -- and hits, which must model thin/juiced/one-sided REAL markets, would have been the next reader to guess wrong. RESOLUTION: all three now have distinct names in config/takeability.js; gradeRanking calls isWithinPromotionBand so its intent is self-evident (a test pins it byte-identical to the old valueEngine call across the whole price range); the ledger dual-writes within_price_band with `takeable` kept as a documented DEPRECATED MIRROR so nothing breaks. Column comments in the database now say what each column actually holds. I did NOT redefine `takeable` in place. Four readers and a ranking gate sit on it, and silently changing its meaning under cover of a naming change is exactly the class of move this session keeps removing. Gates: 4,131 tests / 331 suites green; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
8c764c22a4 |
Structural hardening: unknown-is-not-zero + takeability-is-book-identity
Both guards are ADDITIVE. The full suite (4,111 -> 4,126 tests, 331 suites) passes unchanged through the migration, which is the evidence that no currently-correct output moved: served path, champion, reference ruler and the four accruing challengers are byte-identical. GUARD 1 -- src/utils/known.js. Number(null)===0 has produced at least SIX separate defects here, including one in a module written the same week its author documented the trap. Per-module vigilance has demonstrably failed, so the rule lives in one place and SEVEN sites now delegate: platoonSplits, projectionChallenger, challengerProjection, contactChallenger, statcastAggregateService, consensusRuler, gradeRanking -- plus compoundTotalBases moved onto knownRate. Two functions, deliberately: knownNumber (any finite number -- a REAL 0 is a fact and must survive) and knownRate (non-negative, rejects booleans -- for counts/rates where `true` or -1 is broken, not thin). Collapsing them is how the next variant gets in. firstKnown() exists because `a || b` discards a measured 0 and `a ?? b` does not. MY OWN GUARD HAD THE BUG IT EXISTS TO PREVENT, and its own test caught it: Number([]) === 0, so an empty array coerced to a measured ZERO. Same trap wearing a different type. Both helpers now reject objects outright. GUARD 2 -- src/config/takeability.js. Takeability is BOOK IDENTITY and never price shape. Baseball prop markets are genuinely thin, juiced and one-sided, and all three are NORMAL structure: betrivers and hardrockbet legitimately quote one side only (5 such rows surfaced in yesterday's re-stamp), and a hits-over at -300 is a real placeable bet. A rule that inferred un-takeability from price extremity or one-sidedness would throw those away while still admitting a DFS book at an ordinary -119 -- exactly backwards, because the -119 is the fake one. THE DISTINCTION THAT MUST NOT COLLAPSE, now enforced by test: isTakeableMarket(book) -- CAN it be bet? (identity) isWithinPriceBand(odds) -- SHOULD we promote? (policy band, floor -160) A -300 DraftKings prop is takeable AND out of band; a PrizePicks -119 is in band AND not takeable. Independent axes. FLAGGED, NOT SILENTLY CHANGED: the ledger's `takeable` column is the PRICE-BAND answer, and its name predates this distinction. Four challengers and the ranking gate read it, so renaming or redefining it is its own order -- doing it here would have changed correct current behaviour under cover of a hardening change. Fixtures are REAL prod rows from the 2026-08-02 re-stamp, not invented. Gates: 4,126 tests / 331 suites green; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
5de464330c |
URGENT: anchor the ledger price/book/takeable to TAKEABLE books
Ships before tonight's settle. Served path, champion, ranking and the
reference ruler are untouched.
TWO leaks, not one. The audit found ledgerService.indexProps; tracing the
lock price found that snapshotService.indexOdds has the SAME defect -- it
also indexed the full props list, so gradedAt.odds (the price a grade is
locked at) could itself be a DFS or exchange price. Fixing only the ledger
would have left the contamination flowing in through the lock.
Both now gate on TAKEABLE_BOOKS -- deliberately NOT MODEL_BOOKS. pinnacle
is model-eligible and correctly not takeable, so a MODEL gate would
re-break this the moment pinnacle's feed recovers. A test asserts pinnacle
cannot anchor a price.
TWO INDEXES, TWO ROLES, because the row needs two different things from a
prop and they have different correctness rules:
PRICE / BOOK / TAKEABLE -- takeable books only.
GAME FACTS (game_time, game_date, team/opponent) -- book-INDEPENDENT.
First pitch is first pitch whichever book listed it, so these still
come from any book. Gating them too would drop otherwise-valid rows
for no gain.
Collapsing those roles into one index is precisely the bug.
No takeable quote leaves the key ABSENT and the price null. An honest
missing price beats a price from a book you cannot bet -- and it keeps the
takeable flag from being computed off a DFS number, which is what made it
wrong on its own terms rather than merely mislabelled.
Gates: 4,111 tests / 330 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
|
||
|
|
eabf3b5bcf |
tb-v1: model total_bases as a compound outcome (challenger)
Current ladder (proj_p_over_line) and champion p_win are BYTE-IDENTICAL. tb-v1 writes alongside them, on total_bases props only. STEP 0 -- components confirmed on real data, not assumed. statsapi has no singles field, but hits - doubles - triples - homeRuns reproduces stored totalBases EXACTLY on a real 10-game log. So the decomposition is exact, not an approximation. THE MODEL. Each component gets its own per-game Poisson rate; TB is their weighted sum, and the PMF is built by exact convolution rather than simulated (TB support is small). It inherits the SAME combined multiplier proj-v1.1 computes, so the two models differ only in STRUCTURE. Why this is the fix: with identical mean TB of 1.0, a pure-HR hitter and a pure-singles hitter get P(TB>=4) of 0.221 vs 0.019 -- a 12x difference an NB on TB alone cannot express, because it treats one home run as four events. A test asserts that separation, and asserts P(TB>=4) for a pure-HR hitter equals P(at least one HR) exactly. INDEPENDENCE IS AN APPROXIMATION AND IS LABELLED AS ONE: a plate appearance that becomes a double cannot also become a single, so the components are weakly negatively correlated and independent Poissons slightly overstate the tail. Closer to the truth than what it replaces; not a solved problem. HONEST-ABSENT throughout: fewer than 3 usable games, or no derivable component, returns null and the prop keeps the current ladder value. An inconsistent row (hits < extra-base hits) is SKIPPED rather than clamped to zero -- clamping would invent a plausible line out of a broken one. I HIT THE Number(null)===0 TRAP IN MY OWN CODE and a test caught it: a null rate passed a naive finite check and was treated as a measured zero, which is the difference between "this player never triples" and "we do not know his triple rate". Both tbPmf and tbMean now reject null/''/boolean strictly. Holdout committed: TB ROWS ONLY (49 of 437 settled -- averaging into other stats would hide the effect) and DIRECTION-ALIGNED, since the unaligned comparison is the artifact that accounted for 41% of the ladder's apparent loss. If tb-v1 does NOT improve, the family-mismatch hypothesis is wrong and the mean/similarity branch reopens -- recorded in the query header. Migration applied: proj_tb_p_over + proj_tb_meta, NULL-meaningful. Gates: 4,104 tests / 329 suites green; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
9ebd77b68e |
Build the matchup/platoon axis: three joins fixed, axis now FIRES
The axis was already wired and firing on 0/634 prod rows. Three separate
absences kept it silent, and all three are now joined:
1. oppPitcherByTeam 0 -> the self-origin /api/schedule/mlb/pitchers route
returned nothing in prod. Added the statsapi probable-pitcher hydrate as
a fallback, mirroring the one the schedule step already uses. 29/30
team-sides, one free request.
2. handById 0 -> follows from (1); the batched people call now has ids.
3. bats 0/120 -> batter hand rode ONLY on statcast aggregate rows, which do
not cover the slate. The season player list we ALREADY fetch and cache
carries batSide on 1342/1342, so this is a join, not a fetch.
Switch-hitters ('S') are preserved as-is; platoonSplits decides what to
do with them, not the map.
Verified end-to-end against the live API: opp_declared 29,
pitchers_with_hand 29, batters_with_hand 1342, and a real read --
multiplier 0.966, L vs R, 287 observed PA, weight 0.324 -- composing
alongside environment in one challenger.
FALLBACK LADDER, and a deliberate deviation from the order. Shipped tier:
`batter_own_split` (the hitter's OWN vs-L/vs-R line, regressed toward HIS
OWN overall rate), labelled on every adjustment.
`league_generic` is deliberately NOT implemented. platoonSplits already
handles thin evidence by regressing toward the hitter's own rate, which
covers the thin case per-player; its own doc-comment argues a hitter with
no split evidence should get NO adjustment. A league split applied to such
a hitter models the LEAGUE, not the player -- the doctrine breach the order
itself names in the same step. Adding it would have produced more firing
rows and a weaker signal.
`archetype_x_archetype` is scoped, not built: it needs the opposing
starter classified per game, which is real work and a separate order. The
tier vocabulary is in place for it.
Honest-absent on every join: no starter, no pitcher hand, or no batter hand
-> NO matchup adjustment, never a fabricated neutral. A neutral multiplier
produces no adjustment row at all.
Holdout committed (scripts/matchup-axis-holdout.sql), filtered to
matchup-carrying rows, and it keeps MATCHUP'S OWN nudge visible rather than
only the combined challenger -- arch-v1 composes four axes into one
p_win_challenger, so a combined-only view could not tell which axis earned
the movement, or which one is dragging.
Champion p_win, ranking, calibration, the armed invariant and the two
accruing verdicts are untouched.
Gates: 4,093 tests / 328 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
|
||
|
|
9fc17a4689 |
Fix the team resolve properly: backfill the name BEFORE confirmation
My first attempt did not work in prod -- team stayed 0/323 after deploy.
I resolved the team name AFTER the hint-confirmation check, but the check
itself reads hit.currentTeam.name, which is undefined because
/sports/1/players returns { id, link }. With a FULL-NAME hint (what
snapshotService passes) neither branch of teamRecordMatchesHint could
match: the name branch had no name, and the abbr branch cannot resolve a
full name to an abbr. Confirmation failed, the team was nulled, and my
later backfill ran on an already-null value.
withTeamName() now backfills the name from the cached /teams list BEFORE
any comparison, and is used at all three confirmation sites plus the
return. Verified against the live API on all four cases: no hint, FULL-NAME
hint, abbr hint -> "Philadelphia Phillies"; WRONG hint -> null.
That last case matters most: a wrong hint must still REFUSE. The
confirmation exists so a namesake collision cannot tag a player to a team
he is not on, which would fabricate opponents downstream. Making the match
succeed must not make it succeed wrongly, and a test locks it.
Gates: 4,087 tests / 327 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
|
||
|
|
03efdda33c |
Arm the S59 invariant by fixing its input; matchup sourcing = BUILDABLE
PART 1 -- PREMISE CORRECTION, then the real fix.
The order said the invariant's blocker was removed because "team is now
populated 416/416". It is not: what became 416/416 is home_team/away_team.
`team` (the PLAYER'S roster team) is still 0/416. Arming the guard off
home_team would compare the prop's game to itself -- always a match, a
permanent no-op that LOOKS armed. That would be worse than leaving it
disarmed, because it would read as a working guard.
The guard is also ALREADY fail-safe by construction (`if (knownTeam &&
gameTeams && ...)`), so Part 1's requirement was met in code all along.
What was missing was the data.
ROOT CAUSE: /sports/1/players returns currentTeam as { id, link } with NO
name, so searchPlayer's `hit.currentTeam?.name` was ALWAYS undefined and
every resolve returned team: null. The id is present on 1342/1342 and the
/teams list (already cached 24h) maps id -> name, so resolving it costs no
new request. Verified: Schwarber -> Philadelphia Phillies, Ohtani -> Los
Angeles Dodgers, Judge -> New York Yankees.
Five tests lock the fail-safe: drops only on a positive not-in-game;
abstains on unknown player team; abstains on unknown game participants;
and a row carrying only home_team/away_team does NOT satisfy the guard --
so the tautology can never be reintroduced.
PART 2 -- MATCHUP SOURCING: BUILDABLE. Measured on tonight's real board
against the free feeds, by VALUE not endpoint presence (the environment
trap: wired and null 634/634):
opposing starter 29/30 team-sides (home 14/15, away 15/15)
pitcher hand 1342/1342 (pitchHand.code)
batter hand 1342/1342 (batSide.code; L 416 / R 848 / S 78)
SHARED DEPENDENCY, and it is the finding: /sports/1/players -- a list we
ALREADY fetch and cache -- carries currentTeam.id, batSide AND pitchHand.
One join unlocks the invariant's input and two of the three matchup inputs
at once. The third (probable starter) comes from the schedule hydrate that
already exists.
So matchup is BUILDABLE and is the next order; SOURCE-LINEUPS-first is NOT
needed. Archetype-level reach on the opposing starter is available too
(the SP resolves to a player id, so the existing classifier applies) --
noted, not built.
Champion p_win, ranking, calibration and both accruing verdicts untouched.
Gates: 4,082 tests / 327 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
|
||
|
|
4435856f46 |
Audit finds env/matchup axes DEAD in prod; fix the environment join
STEP 0 AUDIT -- the "already partly live" premise was half true: the CODE is wired, the axes are NOT firing. Across 634 graded prod rows the environment and matchup axes fired on ZERO rows, while 13 archetype axes fired normally (power 80, swing_miss 69, contact 56, launch 51, line_drive 43, ...) plus opportunity 142. Ledger confirms it from the other side: env_multiplier, env_park_base, env_weather_mod, wx_forecast and env_weather_state are ALL null on 634/634. ROOT CAUSE, located rather than inferred. A drop-off audit against the live snapshot: with_team_field 0/120, with_bats 0/120, with_playerId 120/120, oppPitcherByTeam 0, handById 0. `team` is a KEY on every stored grade and NULL on 416/416 -- so an environment resolver keyed off the player's roster team could never find a venue, while buildContext sat there with all 30 teams mapped and 14 weather forecasts resolved and unused. Coors composes to 1.241 the moment it gets a key. FIX -- and it is the more correct join, not just a workaround. The park and the weather belong to the GAME, not to the player's roster team, and the game rides on the prop from the odds feed. gradeBestSide now carries home_team/away_team onto the graded row (the legacy grade shape dropped them), and contextFor joins on the game first, keeping the roster team as a fallback. This no longer depends on a stats-resolve that can legitimately fail. MATCHUP/PLATOON IS NOT FIXED HERE and is not claimed as fixed: it needs the opposing starter and both hands, and the audit shows oppPitcherByTeam=0, handById=0 and bats=0 on the slate -- three separate absences. Per "one axis at a time" that is its own order with its own diagnosis, not a second fix smuggled into this one. Champion p_win, ranking, calibration and opportunity_drift's accruing verdict are all untouched. Gates: 4,077 tests / 326 suites green; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
092f8f09cd |
Build opportunity_drift axis on challengerProjection (arch-v1)
Champion p_win and the live grade path are BYTE-IDENTICAL: the axis writes only to p_win_challenger / challenger_adjustments in the ledger. STEP 1 -- MAP THE INPUT. MLB_LOG_FIELD now maps at_bats -> 'atBats'. Deliberately NOT added to outcomeService's map or liveTracking's LIVE_BOX_FIELD: those exist to SETTLE and TRACK graded props, and nothing grades at-bats, so adding it there would imply a settlement path for a market we do not carry. A test asserts the settle map still lacks it. STEP 2 -- DRIFT, NOT LEVEL. opportunity_drift = mean(last-5 atBats) / (season atBats / games). The LEVEL is collinear with l20_avg (same games denominator; hits/game ~= (hits/AB) x (AB/game)), so the projection already embeds it multiplicatively and adding it would double-count. A deviation from the player's own baseline is the part the projection does not contain. HONEST ABSENCE throughout: fewer than 3 at-bat rows, no at-bats in the logs, or no season baseline all leave drift UNDEFINED -- never 1.0 by default and never 0. Number(null) === 0 here would read as "zero at-bats", the strongest possible fade, invented from missing data. Four tests cover the absent paths. STEP 3 -- THE AXIS. opportunityNudge composes in the same log-odds space as park and platoon (log of a ratio), with two guards the measured axes do not need: a +/-10% DEADBAND (a rest day or a blowout can move a 5-game window without any role change) and a tighter cap (0.15 vs the environment's 0.30) so a noisy PROXY cannot outvote measured signals. Every adjustment carries is_proxy: true and proxy_for: 'confirmed_batting_order' so nothing downstream can mistake it for a lineup feed. The axis can stand ALONE -- without it the early return would gate opportunity off on exactly the thin-classification rows it is most likely to help. Zero extra I/O: analyzeViaEngine1 attaches drift from the feature vector it has already built, and attachChallenger reads it off the grade. Nothing re-fetches in a loop that runs over hundreds of props. COLLINEARITY GUARD added to the coverage probe: Pearson r of drift against l20_avg / l5_avg / ab_per_game, returning null under n=8 rather than reporting a correlation on a handful of rows. If drift just re-encodes the projection, the axis is dead signal and gets shelved. Gates: 4,073 tests / 326 suites green; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
8a02c75aec |
Step 0 input check: stop before wiring opportunity, and why
READ-ONLY. Live grade path byte-identical -- no layer wired, no threshold
moved, no challenger added, no holdout run.
INPUTS ARE 100% POPULATED (n=80 real MLB props, through the grader's own
path): ab_per_game, rest_days, l5_avg, l20_avg, l10_stddev and
game_count_in_7d all 100%; opp_rank_stat 65% overall and 0% on
stolen_bases. So there is no honest-degradation problem to solve.
FOUR FINDINGS THAT STOP THE WIRING, three of which would have made the
work unmeasurable or wrong:
1. THE PREMISE IS WRONG. There is no built opportunity layer to connect.
ab_per_game is consumed in exactly one place -- analyzeViaEngine1:379,
which renders "4.3 AB/G" on the grade card. engine1 has NO opportunity
or usage factor at all. A projected opportunity was never built;
building one is construction, not connection.
2. THE INPUT IS THE WRONG SHAPE. ab_per_game = season atBats/games. It is
a per-player CONSTANT (measured: varies for 3 of 20 players, and those
cannot be legitimate since the value can't depend on stat_type), so it
can only move all of a player's props together, never separate them.
And it is collinear with the projection: l20_avg = seasonTotal/games,
the SAME denominator, so l20_avg already embeds opportunity
multiplicatively. Adding it additively double-counts.
3. THE REAL INPUT DOES NOT EXIST. depthChartService returns battingOrder:
null for MLB ("the one lineup slot the free schedule feed exposes") and
PropLine /context carries lineup_confirmed as a BOOLEAN, not the order.
4. ARCHITECTURE: wiring it into engine1 would be unmeasurable BY THIS
ORDER'S OWN TEST. Step 2 proves reliability and resolution, both
measured on p_win. engine1 factors move the grade LETTER and never
touch p_win. The layer belongs in probabilityEstimator, which already
adjusts on opp_rank_stat, home_away and a consistency pull.
SEQUENCING IS ALSO STALE: challengerProjection (arch-v1) is already live
with archetype, matchup (platoon) and environment (park) axes, writing
p_win_challenger to the ledger. Step 2 of the order's sequence is partly
done -- and the harness this order needed already exists.
RECOMMENDED INSTEAD, as its own order: an `opportunity` axis on that
harness driven by DRIFT, not level -- recent AB/G (last 5) over season
AB/G. A deviation is not collinear the way the level is. Per-game atBats
is present in the statsapi log rows but MLB_LOG_FIELD never maps it, so it
is a small contained BUILD, which is why it gets its own order. Honest
caveat carried forward: it is still a proxy, not tonight's opportunity.
PROBE BUG RECORDED: the first run reported 0% for every feature including
l5_avg, on a pipeline that had just graded 365 props -- impossible, so the
probe was wrong. getFeatures takes camelCase and returns { features: {} };
I passed snake_case and read the top level. Fixed to call
computeFeaturesForProp. Same class as the earlier silent-false harness: a
measurement that makes working code look broken invites you to "fix"
something that was never broken.
Gates: 4,059 tests / 325 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
|
||
|
|
a7d6cf8e36 |
Raise the grade cap 25 -> 500 on measured cost; refusals are correct
PART 1 (read-only, measured on a live prod slate, n=80) OVERTURNS THE
PREMISE. The refusal rate is not a data problem -- it is 98% correct
behaviour. The cap is the entire problem, and it is worse than "25 of 546".
Composition: GRADED 44 (55.0%) | POLICY-SUPPRESSION 35 (43.8%) |
FETCHABLE-GAP 1 (1.3%) | FALSE-THRESHOLD 0 | ARCHETYPE-GAP 0 |
GENUINE-ABSENCE 0.
THE FIFTH BUCKET the order did not anticipate: all 35 "refusals" are
rare_event_over_below_line -- the 2026-07-19 betting-logic audit
deliberately refusing 0.5-line rare events, setting the SAME
insufficient_data flag as a real data gap, which is why they read as one.
They are entirely doubles (18) and stolen_bases (17), while hits (19/19),
rbi (19/19) and total_bases (5/5) grade at ~100%. Had we "fixed" this we
would have re-introduced exactly the bets a previous audit removed, and the
count would have looked like progress.
THE CAP: 585 unique gradeable props, cap 25 -> 560 discarded (95.7%).
Traced to Session 32 (
|
||
|
|
6c97f59546 |
WNBA truth correction + THE p_win FLIP (live, rollback armed)
PART A -- WNBA TRUTH CORRECTION (no behaviour change).
WNBA does not "abstain" and is not "anti-predictive". The -0.12 that
produced those words was NBA-template machinery run on WNBA data -- WNBA
has never had its own archetypes, variables, conditions or calibration,
which is precisely the "sport stubbed in on another sport's template"
CLAUDE.md forbids. That is an UNBUILT MODEL'S EXPECTED FAILURE, not a
verdict on the sport; reading it as a verdict would quietly retire a sport
we never actually attempted. Its own build is QUEUED, after MLB.
The guard CODE is unchanged -- FORECAST_RANKED_SPORTS = {'mlb'} and the
inheritance test are correct live safety either way. Only the meaning is
corrected, and generalised into the doctrine-as-a-gate: a sport ranks on
p_win ONLY once its OWN model is built and shown to predict (calibration
AND resolution on its own holdout). Others are held out as NOT-BUILT,
never as failed. Re-labelled across gradeRanking, snapshot route, tests,
MASTER-PLAN and the challenger report.
PART B -- THE FLIP, gated on a full-slate re-run.
The re-run found something better than a bigger sample. An induced
snapshot graded 7 props: gradeAndCacheSlate runs with DEFAULT_LIMIT = 25
and ~72% of those refuse for insufficient_data, while 546 props are
gradeable. So 8 props IS the board, structurally -- not a small sample of
it. Logged as its own finding; the cap is a separate order.
For a statistically meaningful delta I used 11 real historical boards
(n=328, board sizes 14-57): 79.9% of rows move, mean 5.16 places per
board, TOP READ CHANGES ON 9 OF 11 BOARDS. The re-ordering holds at real
board size. Query committed.
FLIPPED:
- rankGrades drops its edge key (safe for every sport: removes a
non-predictive tiebreak without putting p_win in front).
- selectTopGrades leads on forecast_rank, edge key removed.
- flattenToEdgeBoard sorts on forecastRank, not edge -- this board had
edge as its PRIMARY key, so the whole mobile board was ordered by a
quantity measured not to predict.
- forecast_rank threaded onto strip props.
Sports whose model is not built supply no forecast_rank, so their boards
fall through to the unchanged grade chain -- the fallback is the guard.
ROLLBACK ARMED: boards sort by forecast_rank WHEN PRESENT, so
FORECAST_RANK=0 reverts every surface on the next response -- no deploy,
no client release.
Edge is still computed, stored, carried and displayed as a labelled
diagnostic. Retired from ranking, not deleted.
Eight superseded tests updated to strictly stronger INVERSE properties --
they now fail if edge is ever re-introduced as a ranking key, which the
originals could not detect.
Gates: 4,045 tests / 323 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
|
||
|
|
ef4ac60b81 |
Per-sport rank guard + edge diagnostic-only display + delta report
DELTA MEASURED on live prod grades (live ordering unchanged): MLB 7/8 props move (87.5%), mean 2.5 places, TOP READ CHANGES (corey seager hits 1.5 under -> jake burger hits 0.5 over). WNBA 25/25 move, mean 4.1, max 12. This is a large re-ordering, not a tweak. Caveat recorded rather than buried: MLB had only 8 graded props at measurement time. The percentages are real; the sample is one small slate. Re-run before the flip -- it is one call. PER-SPORT DOCTRINE ENFORCED IN CODE. WNBA moves the most and must NOT adopt this: its p_win is anti-predictive, so ranking that board by p_win would sort it by a signal measured to point the WRONG WAY -- worse than the incumbent, not better. A comment would not have stopped a future flip from going global, so FORECAST_RANKED_SPORTS = Set(['mlb']) gates the forecast_rank stamp, with tests asserting no sport inherits MLB's result. A sport joins only by passing its own holdout. EDGE IS NOW DIAGNOSTIC-ONLY IN DISPLAY. MobileEdgeBoard.EdgeCell rendered green (--g-a) for positive edge and red (--miss) for negative. Two things were wrong: green/red IS a quality claim on a quantity that does not predict, and ROW-GRAMMAR reserves red for settled-negative ONLY -- a negative diagnostic is not a settled loss. Now neutral mono with a diagnostic tooltip; header reads "MKT GAP · DIAGNOSTIC". The number is still shown -- no display went blank. DeskShowcase neutralised likewise. PINNACLE LOGGED, NOT ENSHRINED. Per the order, "market-not-sharp" is PENDING-RECOVERY rather than a confirmed permanent limitation. The single question for PropLine is in BLOCKERS.md with its evidence, and MASTER-PLAN now carries the pending status instead of the permanent claim. Live sorts remain byte-identical: selectTopGrades, flattenToEdgeBoard and topGradedService all still call the incumbent. Gates: 4,041 tests / 323 suites green; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
86d123945c |
Rank on p_win: challenger instrument + retire edge from decisions
MEASURED BASIS (n=200 settled MLB rows): corr(p_win, outcome) = +0.26; corr(edge, outcome) = -0.010 incumbent ruler / -0.022 consensus ruler. Subtracting the market destroys the signal under BOTH rulers, so a quantity that does not predict must not rank, gate or decide. CHALLENGER-FIRST -- live ordering is byte-identical. rankGrades (the incumbent, grade-first with edge as its 4th key) is untouched and tested as untouched. NEW: rankByForecast -- takeable-gated p_win -> grade -> confidence -> stable order, with NO edge term anywhere. p_win LEADS and the letter follows, deliberately: the letter measured r ~ 0.005 and is inverted (B 52.4% < C 56.9%) while p_win measures +0.26, so leading with the letter would sort by the weaker signal and use the stronger one only to break ties. Recorded in the code: isotonic calibration is a MONOTONE transform, so ranking on raw vs calibrated p_win gives the SAME ORDER. Calibration matters when p_win is displayed or thresholded; it cannot change a ranking. Nothing here needs the calibrated value. rankingDelta + GET /api/internal/ranking-delta measure how far the board would move before any flip. The endpoint reports p_win coverage alongside the delta -- if p_win is absent the challenger degrades to grade order and the delta UNDERSTATES, which is worth saying rather than reporting a clean zero. forecast_rank is stamped on snapshot grades BEFORE stripModelPrice, so every tier gets the correct order without the paid values (the topGradedService precedent -- an ordinal can travel where the magnitude cannot). Additive only: nothing sorts by it yet. RETIRED AS DECISIONS (not rankings, so done now): - altLineScanner.compareToBookImplied no longer returns value_detected: edge > 0. Edge is still COMPUTED and returned -- losing the record would be worse than mis-using it -- but the verdict is an honest null with value_basis: 'retired:edge_does_not_predict'. - scanAltLines no longer filters to edge>0 or calls the survivor "optimal". The whole ladder is returned ranked and labelled 'price_gap_diagnostic_unvalidated'. The module has ZERO callers (verified) -- unwired like mlbGrader.js, left in place and made honest. An honest asymmetry recorded there: ranking props AGAINST EACH OTHER must not use edge, but choosing between RUNGS OF THE SAME PROP is inherently price-relative -- ranking rungs by model probability alone would always pick the lowest line, since P(over 0.5) > P(over 2.5) by construction. So the gap stays the rung key, explicitly labelled unvalidated. Two superseded tests updated to stronger properties. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |