11277a1b99359c8f6dc6ab10c5e0df90e6f9062a
511 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
11277a1b99 |
Materialization truth: a completed write is not a complete cohort
Transport truth says every intended write succeeded. Materialization truth says
every identity that should exist actually exists. The retention writer could
only report the first, and the gap is not theoretical.
THE CONFLICT IDENTITY, traced to the real index:
model_snapshots_cycle_prop_uniq UNIQUE (snapshot_id, player_key, stat, line, side)
written by `upsert(..., { onConflict: same five columns, ignoreDuplicates: true })`.
Measured in production: no key column is ever NULL (0 of 328,262 rows), so
NULLS DISTINCT never applies and the identity is plain column equality. `line`
is an unconstrained numeric, so identity normalises it — 0.5 in and "0.50" out
must not read as two identities for one stored row.
PRIOR-CYCLE COLLISION IS IMPOSSIBLE. `snapshot_id` is in the identity and is a
fresh UUID per cycle, so no row can be suppressed by an earlier cycle.
Append-only chronology across cycles is safe, and `captured_at` is not in the
identity, so a cohort cannot be split by timing.
INTRA-CYCLE COLLISION IS REAL, AND WE CAUSED IT. `canonical_event_id` is NOT in
the identity. A doubleheader — same hitter, same stat, same line, two genuinely
different games — is ONE identity. Demonstrated through the real collector: 4
outbound rows, 2 distinct identities, 2 rows discarded by ignoreDuplicates with
no error, `written` counting all 4 and the terminal status reading COMPLETE.
Before event-aware dedupe the second game was dropped before grading, so the
collision could not arise; that fix moved the loss downstream into retention.
The conflict identity is NOT changed here — that is a separate decision with its
own before/after. This makes the loss visible instead of silent.
EXPECTED vs ACTUAL. `expectedMaterialization(rows)` derives the identity set
from the FINAL outbound payload using the exact database identity — never from
`attempted`, which counts rows sent, not identities that can exist.
`reconcileMaterialization` compares SETS, not counts: two sets of equal size can
still differ, and a cohort that swapped one identity for another passes every
count test ever written. A collision passes set equality by construction (the
discarded row was never in the expected set) while real rows were lost, so
collision_count > 0 fails the cohort on its own.
A cohort is evidence-complete only when transport is COMPLETE, missing = 0,
extra = 0, and collisions = 0.
OBSERVABILITY stayed minimal. `last_retention` was already PER SPORT (a Map
keyed by sport), so no fix was needed there and the route is UNCHANGED — the new
fields ride the existing entry: outbound_rows, expected_materialized_count,
outbound_collision_count, expected_identity_digest. Counts and a digest only,
never the identities, which carry player names. The expected set is the one
materialization fact unrecoverable from the database afterwards, which is why it
is the only thing recorded at runtime.
A collision leaves transport COMPLETE, so the existing failure alert could never
see it. It now has its own high-severity alert naming the counts, the cycle and
the build, and says the cohort is not evidence-complete.
Seven teeth, each injection verified present, against a GREEN baseline of 63:
1 attempted===written as evidence completeness -> 1 fail
2 COUNT(*) equality instead of set equality -> 1 fail
3 snapshot_id dropped from expected identity -> 4 fail
4 unexpected collision allowed to qualify -> 1 fail
5 single global last_retention slot -> 2 fail
6 partial chunk failure treated as usable -> 3 fail
7 collision loses its announcement -> 1 fail
Restored byte-identically (retention 742f116473d97f49, snapshot 81129facbabeb280).
Three brittle assertions repaired, with the reason recorded: two windowed on a
byte count that a neighbouring block outgrew — a test failing because of its
neighbour, not its subject — now windowed to syntactic landmarks; and one
counted TERMINAL.COMPLETE occurrences, which a legitimate comparison
incremented. It now asserts one DECISION and one READ.
persist() and createCollector are BYTE-IDENTICAL. onConflict and
ignoreDuplicates appear in the diff only as prose. Model, event, ledger,
calibration, chain, lineage config, and the status route: UNCHANGED. Zero
lineage/publication files, zero cacheSet changes, zero web paths. Lineage OFF.
Schema contract unchanged: release 64, prod 67, prod-only 3 (debt, not
authorized), missing in prod 0.
384 suites / 5,144 tests pass. web tsc exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
9809626c99 |
Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.
So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.
TERMINAL RETENTION STATES (retentionService.classifyPersist):
NOTHING_TO_PERSIST attempted 0 — a refusal-only slate is still a cycle
SKIPPED_NO_DATABASE no database configured; not a failure
COMPLETE attempted > 0, written === attempted, no error
FAILED_ZERO_WRITE written === 0 — first chunk failed
FAILED_PARTIAL 0 < written < attempted — a later chunk failed
FAILED_UNRESOLVED_ERROR counts look complete but an error is unresolved;
unreachable through today's loop, and kept because
the alternative is reporting COMPLETE holding an error
The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.
classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to
|
||
|
|
35da190f2c |
Retention hotfix: drop published_side, derive the schema contract, break the silence
`createCollector.onPublished` set `published_side` beside `published`.
`published_side` is not a model_snapshots column. supabase-js declares the
UNION of row keys in the `columns=` parameter, so one invalid key made
PostgREST reject the ENTIRE batch with a 400 — every sport, every cycle.
Retention is best-effort, so nothing surfaced. Confirmed in edge logs.
The field was redundant as well as invalid: `side` is already on the row.
Deleted rather than added to the schema — a column would preserve an
accidental artifact.
Three things missed it, and each is now closed:
1. WRONG SHAPE INSPECTED. The manual check sampled the collector after
onGraded only and never called onPublished, so the offending key was
not yet on the row. It read a pre-publication shape and reported the
final outbound shape as clean. The new test captures the array actually
handed to .upsert(), after the full production call order.
2. NO CONTRACT. Every retention test injects a permissive fake client that
accepts any column set, so 381 suites proved the logic and never once
compared a row against the database. The contract is now DERIVED — the
migration chain applied to a disposable postgres, read out of
information_schema (scripts/generate-schema-contract.js). A
hand-maintained list would be a second opinion about the schema, and a
second opinion is what let this through. scripts/verify-schema-contract.js
checks the contract still describes a live database.
3. SILENT FAILURE. A failed batch reached one console.log. It now emits a
high-severity structured event carrying sport, snapshot id, stage,
error, code_sha and timestamp. Best-effort semantics are unchanged —
the product continues and says so — but the failure is observable.
`skipped` (no database configured) is not a failure and does not alert.
Teeth, each with the injection verified present before the run:
- published_side back into the final payload -> 4 tests fail; restored
byte-identically (sha 6a0ced7c52134135 both sides)
- settled_at (a REAL contract column) -> accepted, so the guard
discriminates by contract membership, not by novelty
- alert block deleted -> 3 tests fail; restored byte-identically
Model and product behaviour untouched: analyzeViaEngine1,
probabilityEstimator, gradeSlateService, lineageCanaryConfig all unchanged.
Lineage stays OFF. Net source change is one behavioural line plus the alert.
382 suites / 5,094 tests pass. web tsc exit 0 (zero web paths touched).
Measurement blackout recorded, NOT backfilled: last good retention write
2026-08-27T19:08:32Z;
|
||
|
|
ceaa896f77 |
Runtime observability: report the build and canary state the system acts on
The rollout stalled at RUNTIME_UNVERIFIED because two facts were answerable only
as a side effect of a scheduled snapshot writing a row: which build is running,
and whether MLB lineage is effectively enabled. Every state transition therefore
waited on cron rather than on asking the service.
- src/services/lineageCanaryConfig.js — THE canary resolver. Parsed once at
module load (unchanged semantics), normalised sorted/deduped/trimmed, frozen.
snapshotService's write gate now delegates to it, and the status probe reads
the SAME state. A route that parsed the environment itself would be a second
version of the truth, free to drift from the gate it claims to report.
- GET /api/internal/snapshot/status gains runtime.code_sha (the production
codeSha resolver — never git, never gitea/main; null when unavailable),
runtime.started_at (computed ONCE at module load, so it marks a boundary
rather than reading as now; deliberately not called deployed_at), and
lineage_canary {enabled, sports, configuration_source}.
- No raw environment value is returned; sports is the normalised set and
configuration_source says only ENVIRONMENT vs DEFAULT. Router-wide
requireInternalAuth is unchanged: 200 with key, 401 without.
- Effective lineage config is fixed for the process lifetime, so
runtime.started_at is a defensible lower bound for how long that state held.
Strictly observational — the handler still only reads Redis.
Model and decision code byte-identical to
|
||
|
|
8c6aef1e12 |
Event admission gate: an unresolved or contradicted game does not earn a Read
A fractured identity keeps two bad records from merging. It does not make an unknown game true. Until now a prop with an unresolved or verified-impossible event still continued into grading under that synthetic key; it no longer does. - event_binding_status contract: RESOLVED / UNRESOLVED / AMBIGUOUS / CONTRADICTED / UNSUPPORTED. CONTRADICTED and UNRESOLVED stay distinct — one means we know the association is wrong, the other that we do not know. - ONE admission gate (gradeSlateService.admitForGrading), before dedupe and before grading. Admission requires RESOLVED *and* a canonical_event_id; a legacy derived game_id can never satisfy it. Rejected props are returned, not discarded, so retention keeps them as evidence. - Roster evidence is now DATE-SCOPED. statsapi honours ?date= and it changes the answer (Joe Mack is on the 2026-08-26 Marlins roster, absent on 2026-04-15). Evidence that does not describe the slate's date can only yield UNRESOLVED, never CONTRADICTED — uncertainty must not become an accusation. - A mis-nested market is REFUSED, never re-bound. Knowing Joe Mack is a Marlin does not license moving a provider record into the Marlins game; that would invent provenance. Measured on the real 2026-08-26 slate: 2,878 props -> 2,874 admitted, 4 rejected (EVENT_PLAYER_TEAM_CONTRADICTION), each player's correct game still resolving. Doubleheader 824514/824478 both remain independently RESOLVED and admitted. No change to probability, projection, side, grade, confidence, ranking, normalization or calibration. Lineage remains disabled. Suite 380/5,061/0 from the release worktree; web tsc exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
352016790a |
MLB canonical event identity, impossible-binding refusal, event-aware dedupe, publication commit
Release-isolated slice built from
|
||
|
|
41ba38ed4e |
Master state board, repo-verified — and a live retention regression
Captures the complete position before the build-process pivot, derived from the repo and live database checks rather than from any summary. Every claim is marked VERIFIED / INFERRED / UNVERIFIED with the check shown. THE HEADLINE IS A PRODUCTION REGRESSION FOUND WHILE VERIFYING. model_snapshots stopped being written on ~2026-08-15: 14,718 rows on 08-12 -> 7,080 -> 378 -> 0 on 08-16 and 08-17. ledger_entries and closing_captures are still writing normally, so this is one write path failing, not a dead cron. Cause, verified three ways: retentionService.js:165 declares chain_shadow on every row; migration 038 was never applied so the column does not exist; and persist() catches the error because retention is best-effort by design. The chain-v1 report flagged this exact sequence as a blocking precondition. The commit was pushed and deployed; the migration was not applied. Consequence: both pending verdicts are frozen. A8 sits at 623 complete pairs (MIN_N 500 met) but only 20 of 40 distinct settled games, and has gained nothing since 08-15. The chain-vs-counter head-to-head has never accrued a row, because its column does not exist either. Four premise corrections the verification forced: - WNBA is BUILT, not ingested. All three WNBA tables are absent; the source is ESPN, not nba_api (WAF-blocked, package not installed); Basketball-Reference cannot serve as-of-date aggregates; wehoop has no lineup/on-off data. - Lineup reconstruction is not pending. It works, 40/40 games to a complete 5-a-side. What is missing is a column to store it in. - WNBA footprint measures 42.6 MB/season, not ~475 MB. The Pro conclusion may still hold, but not from that number. - The chain diagnostic is BOTH a difficulty-correlated component and an unexplained uniform offset, and the correlation is near-mechanical -- it shows the wiring reaches the number, not that the adjustment is right. Also records the fix options for the regression without taking either, since this order is documentation-only. Reverting retentionService.js:165 restores retention immediately and needs no credentials; applying 038 is the complete fix and needs DDL access this environment does not have. Documentation only. No code, schema, served path or accrual changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xcpku8k719mtCWqFKXg287 |
||
|
|
6c34af3414 |
checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion recon (Tony resurrection, STEP 0). This work existed only on the laptop disk. - chain shadow accrual + probe script (038_chain_shadow.sql) - WNBA possession feed: ESPN adapter, usage service, verify script (039_wnba_player_game.sql) - baseball chain - retention/snapshot service updates, tableKeys, matchupKeys - specs: chain-v1, wnba-possession-feed, wnba-source-survey - unit tests for the above Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn |
||
|
|
0657b71d18 |
Drop lock_lines and retire its writer (dead table, dead writer)
lock_lines was built in Session 64 for a staleness audit - join lock-time
per-book lines to closing_captures and ask whether our locked line was
stale-high vs consensus. That audit was never written. ruler-comparison.sql
records why: only 43 settled rows ever joined it with >=2 two-sided books.
Measured before removal: 367,595 rows, 104 MB, zero readers in src/, scripts/
or web/src/ - the only from('lock_lines') was an upsert, every other mention a
comment. Zero dependents: no FK, no view, no trigger. The newest pg_dump held
all 367,595 rows, pg_restore-verified before the drop.
The write is off too, because a dead table that keeps refilling is only half
solved: it was accruing 36,440 rows/day, 10.3 MB/day, 23% of all database
growth, for a question nobody was asking. buildLockRows is kept and still
tested - the logic was never what was wrong, and re-arming is one flag plus
re-creating the table.
DB 510 MB -> 406 MB: 81% of the 500 MB cap, +94 MB headroom, under it for the
first time in months.
Moat and grade untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fb0010222d |
Stop persisting missed_window refusals (halt the bleed at source)
missed_window means 'this capture pass ran after first pitch'. It is a fact about our cron cadence, not the market: once a game starts the same prop emits a fresh refusal every ~20 minutes for the rest of the night, per book, per side. Measured over 7 days of production that is 332,608 rows/day - 83.1% of all closing_captures writes, ~66 MB/day - and since B1 filtered both readers, nothing reads them. The filter lives in persist(), not buildCaptureRows(), and that is the whole trick: the caller computes the capture-rate alarm from the full in-memory array, so filtering at build time would have blinded the ops alarm to the exact condition it exists to catch. captureRateAlarm is pure; a test asserts the caller still passes the full array, and that a 10-priced/90-late pass still fires at 0.10. Narrow by design: one_sided_price still persists (liquidity signal), priced captures unchanged, and the rare fault refusals still persist because each names a pipeline fault worth seeing. A missed_window row carrying a price is kept. Growth drops 400,469 -> 67,862 rows/day. The 4.2M historical rows are now static, so the cleanup is a calm decision rather than a race. Nothing deleted. Grade untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2271f46ab2 |
dCLV close leg: read priced captures only (the unfixed twin)
computeDirectionalForRow selected the latest closing_captures row with no missed_reason filter, while its sibling attachClosingProb has had one since the CLV instrument repair. closing_captures records a refusal for every prop x book x side on every cycle after first pitch, so a refusal always carries a later captured_at than the last real price - latest-first returned a refusal on 26,448 of 26,448 identity groups, and computeDirectionalClv refuses on a missedReason. That is why dclv_state has been 'unknown' on 100% of rows since Session 64. Measured on 600 real settled rows: 100% unknown becomes flat 45.8%, negative 23.0%, positive 22.8%, unknown 8.3%. ClvBadge will render MOVED TOWARD US 114 and MOVED AWAY 117 per 600 - near-symmetric, which is the honest shape. Existing rows do not recompute (first-computation-wins). The re-stamp is described in BUILD-STATE, not run. The grade is untouched. No rows deleted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7c8ef8b12e |
Deploy verified: A1-A7 live, served grade unchanged (0 leaks of 425)
First post-deploy slot (2026-08-12T03:10:52Z,
|
||
|
|
f61ec6b391 |
Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move. A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS: 410-617 of 2,490 duplicated with an equal number never returned, while rows.length matched the server exactly. safePaginate orders on a real unique key, verifies the tuple at runtime, and THROWS on a query error instead of treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn through that reader, and defense_by_direction's distinct-n was likely below the gate floor all along. A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from pg_index (the context tables are dated-composite and had no single unique column). The unordered helper is deleted, not parked. A3 — ledgerService and retentionService defaulted the SAME env var to DIFFERENT versions, so no ledger row ever carried the marker eligibility requires. One source now. model_snapshots settlement moved onto the cron: 15,484 -> 28,894 settled, repaired-champion 0 -> 7,556. A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no row at-or-before the date means the factor does not apply, never the nearest row. Live path unchanged, proven 400/400 on real rows. A5 — factor_inputs freezes what the factor READ, never the multiplier, so an audit can recompute and check. It also recorded the finding: the three hits factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by the resolver and written by nothing. A6/A7 — matchupKeys resolves those keys from the posted lineup plus the schedule's probable pitchers, and fires the factors into a SHADOW freeze: 248 fires on 308 props, 245 of which would move the grade. The served forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test that decides whether they ever go live. Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity harness 34/34. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
387ae4d54e |
Handoff doc: ground-truth build state at 71d3b7b
Written from the repo, not summary -- every claim grep- or run-verified,
UNKNOWN where the repo cannot say.
State: HEAD and gitea main in sync at
|
||
|
|
71d3b7b786 |
E10 Report issue template + E12 /report archive, to spec
PHASE 0 — the spec, read not recalled. E10: "Hybrid: dark billboard header that survives every client, light paper body Gmail can't wreck. 600px, stacked, no webfont dependence." Content law: "One email per slate day. Top read, what changed, the record. Nothing else." E12: "EVERY ISSUE SHOWS ITS OWN DAY RECORD -- THE ARCHIVE IS A LEDGER TOO." COMPOSED, NOT FORKED. The audit had E10 as PARTIAL, not absent: newsletterService already builds the daily report's CONTENT and lints its voice. What was missing is the designed hybrid SHELL, so reportTemplate.js is a template over that builder rather than a second report -- the same call made for the movement strip, and for the same reason. PHASE 1 — the hybrid shell is an ENGINEERING constraint, not a look, and the tests say so: Gmail strips style blocks, Outlook ignores flexbox, and a dark body renders as a black rectangle in several clients. Hence tables, inline styles, 600px fixed, system fonts, no image required to read, and the green SHIFTS from #00D4A0 to #00A57D on paper because the dark-mode green is unreadable there. FACT-CONTRACTED: a section whose data is absent is OMITTED and NAMED in `omitted`, never filled. There is no code path producing a placeholder figure. The honesty block carries the real numbers -- graded count, cleared-ceiling count, the realized rate against baseline, and that we do not issue A grades. E1'S LAW TRAVELS EVEN THOUGH ITS RENDERING CANNOT. An SVG strip is not reliable in email, so movementText carries the RULE: green only when the move favours the read, and a flat market says FLAT · [N]D rather than showing nothing. NO DESIGNER SAMPLE DATA. Nabers 1,120.5, No 128, DAY RECORD 9-4 are a spec for what a live issue renders; pasting them in would be fabrication carrying a designer's authority and would look entirely correct. Tested. PHASE 2 — /report is now the real archive, REPLACING the S41 redirect to /blog. That redirect existed because the surface did not; E12 built it, so the placeholder is correctly gone and the S41 test is updated rather than worked around. Every row carries its own day record, and an unknown record says UNSETTLED -- never a dash that reads as zero. Empty archive is an honest state. Backend: public read-only /api/report over Redis issues, plus the Next proxy. Both surfaces registered under the reachability guard. A test bug I made twice now: my check for forbidden sample values matched the template's own doc block, which NAMES those values as things never to paste. Documentation worth keeping, so both suites strip comments before matching -- a guard that reads its own warning is not reading the code. WAVE-2 STATUS: E1, F9-F11, E10, E12 done. Still gated -- F5 article media and E16/F8 on the card-system reconciliation; the in-season hub IA on the social chat's formula; E9/E15 on model; E2/E6 on licensing. Read-only throughout; serving fingerprint unchanged including newsletterService; accrual clock unchanged at 0 eligible dates. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
49e76068da |
Doctrine + E1 movement strip + F9-F11 offseason hub shell, to spec
PHASE 0 — specs/ARCHETYPE-TAXONOMY-DOCTRINE.md records the ruling as
shared law: 83 designed glyphs are the full four-sport taxonomy; a glyph
renders ONLY where its archetype is modeled and proven. 39 of 83 map to a
real archetype and are wired; the 44 unmapped are DORMANT SLOTS for
WNBA/NBA/Soccer, not a wiring gap. Wiring them would mean inventing 44
archetypes to consume artwork -- decoration presented as classification,
which is forbidden. DUAL THREAT and PAINT BOSS are modeled archetypes with
no mark: the mirror gap, flagged to the design side. When a sport's
archetypes ship, activation is a MANIFEST lookup, not new art.
PHASE 1 — E1 movement strip. The spec's own line is "the movement strip is
defined once here and reused everywhere a line has a past", so it is a
primitive, not a fourth chart.
RECONCILED RATHER THAN FORKED: lib/gradeShift.js ALREADY implements E1's
colour law -- toward/against/flat, including the direction flip that makes
an UNDER's favourable move the opposite sign of an OVER's. MovementStrip
CONSUMES buildGradeTimeline instead of reimplementing it, and a test
asserts it never redefines isUnder. GradeShift stays the grade-history
view; this is the reusable strip. That is the card-fork lesson applied
before it could happen again.
Spec laws honoured: STEPS NOT CURVES (H then V, no smoothing -- a curve
invents prices that never traded, and a test rejects any C/S/Q/T command);
green only when the move FAVOURS the read; FLAT renders as a hairline plus
FLAT · [N]D because a flat market is a finding; and too little history
says NO MOVEMENT HISTORY rather than rendering blank.
PHASE 2 — F9-F11 offseason hub shell, built from Vyndr Offseason.dc.html.
The spec's load-bearing words are used verbatim: "OUTLOOKS REPRICE ON NEWS
· NOT GAME ODDS" (an offseason number is not a game line), the QUIET WIRE
empty state ("No outlook-moving news since X. We don't manufacture
movement."), WHAT CHANGED TODAY as the hero with the countdown ambient and
top-right, the tag-colour-is-meaning row anatomy, the open -> NOW -> FAIR
triplet with the movement strip embedded, and the OUTLOOK ONLY block where
every row carries NOT GRADED.
THE DESIGN FILE'S SAMPLE DATA IS NOT IN THE COMPONENT. Wembanyama +420 ->
+330, Nabers cleared 11:42 AM, the Summer League names -- all of it is a
SPEC for what a live feed renders, and copying it in would be fabrication
carrying a designer's authority. A test asserts none of those strings
appear.
The IN-SEASON information architecture is NOT invented here. The spec
covers an offseason hub; nothing specifies how content, articles, wire and
the live slate share year-round navigation. That remains the open design
gap, and the route notes it.
Two test bugs caught and fixed: my first assertions matched my own doc
comments -- the ordering check found "WHAT CHANGED TODAY" in the header
block and the no-curves check caught the word "curve" in the sentence
explaining why curves are wrong. A guard that reads its own explanation is
not reading the render; both now strip comments first.
PHASE 3 — both surfaces registered under the reachability guard. Read-only
throughout, serving fingerprint unchanged, accrual clock unchanged at 0
eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
60469422af |
Wave D1 primitives — and the audit says E1/E9 were not Wave 1
PHASE 0 — the order proposed E1 + E9 as Wave-1 and told me to follow the
audit if it disagreed. It disagrees: E9 calibration curve is WAVE D3,
gated on MODEL work ("resolve n>=20 vs N30, accrue buckets"), and E1
movement strip is WAVE D6, a large surface build. E9's gate is live right
now -- calibration is WITHDRAWN at 0 eligible dates, so the curve could
only render its empty state today. Building it would ship a component
whose entire purpose is unavailable.
PHASE 1 — three of the five real Wave D1 items were ALREADY DONE, and the
2026-07-31 audit has aged:
D1 glyph library audit: 38/83 wired (46%)
now: COMPLETE for everything wireable -- 39 of 83
designed glyphs map to a real archetype, all 39
are wired, colours match the registry exactly
(0 disagreements).
A1 card token audit: BUILT-BUT-DRIFTED, "in only 1 file"
now: BUILT-TO-SPEC -- it IS the --bg-1 token,
consumed by 32 files. The audit counted literal
hex, which is what a correctly tokenised value
looks like.
B1 boundary blue audit: PARTIAL, hex in 2 files
now: BUILT -- --priced-out/#8fb2de is a token with a
documented colour law, 4 consumers.
The 44 unwired glyphs are NOT a wiring gap: they have no backend
archetype, so wiring them means inventing 44 archetypes to consume
artwork -- the fabrication this programme refuses. That is the 41-vs-74
scope question and it is Kev's call. Separately, 2 registry archetypes
have NO designed glyph (DUAL THREAT, PAINT BOSS) -- a design gap.
PHASE 2 — what was genuinely absent is now built. web/src/lib/motion.js:
nudge() capped at 180ms so it reads as acknowledgement rather than
latency; bootStagger capped at 240ms because uncapped, row 40 waits 1.1s
and the stagger BECOMES the latency it exists to disguise; rowHover
returns handlers not CSS so touch cannot stick a hover state; and
revealOnIntersect returns an unobserve in every path and reveals
IMMEDIATELY when there is no IntersectionObserver or motion is reduced --
content is never hidden behind a capability check.
Reduced motion is honoured, not softened. The sharpest of the 10 tests:
bootStagger under reduced motion returns opacity 1, not merely delay 0 --
if the CSS animation supplies the opacity, skipping it leaves the row
invisible forever.
PHASE 3 — the reachability guard gains a PRIMITIVES section: a module
built to be embedded must declare its exports AND name its intended
consumers, because a primitive imported by nothing is the same
built-but-unread class as an unmounted component.
WAVE-2 UNGATED: F9-F11 offseason hub, F5 article media, E10/E12 Report,
E1 movement strip. GATED: E9 + E15 on model, E16/F8 on the resolution tail
and the card-system reconciliation, E2/E6 on licensing, E13 on another
order.
Read-only throughout; serving fingerprint unchanged; accrual clock
unchanged at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
7a318ccce0 |
Media/hub design inventory: the specs already exist, and so does a gap audit
READ-ONLY. Nothing designed, built or mounted.
THE HEADLINE: this inventory was largely already done. specs/
design-vs-build-gap-audit.md is a 61-item design-vs-build audit from
2026-07-31, every claim grep-verified, with a wave-ordered build plan. The
useful work is reconciling it, not redoing it.
WHAT ALREADY HAS A FULL DESIGN SPEC (and is NOT built):
- Offseason hub -- Vyndr Offseason.dc.html, 119KB: hub home desktop+390, an
NBA Summer-League variant with an OUTLOOK ONLY / NOT GRADED honesty
block, season-long board, season-read reveal with WHAT WOULD CHANGE THIS
READ, news/outlook feed row anatomy, quiet-wire empty state.
- Article media (S3) -- hero template plus four hero graphic archetypes,
all GENERATED data visuals never stock, inline figures with a one-pull-
quote max and a caption law that every figure names its data, article
card, OG 1200x630 with the master already rasterised.
- Movement strip (E1) -- steps not curves, green only when the move favours
the read, FLAT as hairline. THIS CORRECTS MY 2026-08-07 BOARD, which
listed it UNKNOWN and possibly satisfied by GradeShift. It is a specified
primitive that does not exist.
THE WIRE is BUILT (vyndr/Ticker on real ticker exhaust) and is a DIFFERENT
thing from NewsWire.tsx, which is the offseason news/outlook feed.
THE ONE REAL DESIGN GAP: the hub's information architecture across seasons.
The Offseason file specifies an OFFSEASON hub; nothing specifies what that
surface is IN-season, or how content, articles, wire and the live slate
share one navigation. Every component exists on paper; their composition
into a year-round media surface does not.
A FINDING AGAINST MY OWN RECENT WORK: I built a card renderer at 1080x1350
without checking whether a designed card system existed. It does -- five
master sizes with defined layouts and rasterised PNGs in exports/. The
content engine's cards are a parallel invention. Not wrong, but they should
conform to E16/F8 rather than diverge, and that is a design decision.
CORRECTION: the order lists content API, preview page, book-comparison
mount and guard widening as still pending. All four shipped at
|
||
|
|
c575a708c7 |
Content studio API + preview page; widen the reachability guard; correct
two inventory errors INVENTORY CORRECTION, and it was mine. Phase 2's two "orphans" are NOT orphans -- my board grepped only web/src/app and missed component-level mounting. The transitive check says both are already mounted: BookComparisonPanel -> GradeResultCard -> app/scan/page.tsx NewsWire -> ExploreHub -> app/explore/page.tsx So book comparison is DONE (wired to /api/books, rendering on the grade card) and THE WIRE is DONE-BY-DESIGN, mounted in ExploreHub. Its header names an "Offseason Hub" as its home, and that hub genuinely does not exist -- but that is board item #8, not a mounting bug, and inventing a surface to satisfy a comment would be the wrong fix. The lesson is the same one this session keeps teaching: I checked one directory and reported a conclusion the check could not support. ALSO CAUGHT: I overwrote src/routes/content.js, which was the Session-29 content-templates route, by picking a filename without looking. Restored from git with no work lost; the new surface lives at /api/content-studio and both now coexist. PHASE 0/1 — /api/content-studio serves finished posts (copy, branded card, card_svg, the fact_contract each was REQUIRED to have, and the facts that actually backed it) plus a POST for editorial status in Redis. Private via internal key; the Next proxy holds the key server-side so the browser never does. /studio renders it as a thin client -- copy and card side by side with the fact contract visible, because reviewing copy by reading it is exactly how a wrong number ships. Never-blank: a night with nothing generated says so. API-FIRST is the point: the endpoint an autonomous poster will call is the one the page already renders, so the agent handoff is a pointer change, not a rebuild. Contract documented at docs/CONTENT-STUDIO-API.md. EXPRESS 5 BROKE 23 SUITES at first: `router.get('/:date?')` throws at mount time in Express 5, taking down everything that imports app.js. Two explicit routes instead. PHASE 3 — the reachability guard is widened from grade-fields-only to a general built-but-unread check. Book comparison, THE WIRE and the content studio are now registered surfaces; a page counts as its own entry point (Next mounts it by convention) while everything else must trace to one. 22 checks green; a registered-but-unimported surface still goes red. FULLY ISOLATED: read-only on model/slate/ledger, serving fingerprint verified unchanged, accrual clock unchanged at 0 eligible dates. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
74aa75945e |
Content engine: posts that structurally cannot lie
PHASE 0 — contentEngine makes Truth Law structural, not careful. Copy is
token-substituted and an unbacked {token} REFUSES to render -- there is no
code path that produces a plausible default. The fact contract is asserted
before any string is built. Card and copy render from ONE fact object, so
a caption and a card cannot disagree. No live model writes factual claims:
the voice is in the template, the facts are pulled, and the voice-polish
port is deliberately unwired, because an LLM that can rewrite a sentence
can rewrite a number.
18 tests carry the proof. The one that matters most: ZERO IS PRESENT.
"0 cleared B+" is our most honest possible post, and treating 0 as missing
would be the Number(null)===0 breach wearing its opposite coat -- it would
silently delete exactly the post the brand is built on.
PHASE 1 — three templates, generating real posts from tonight's data:
hot hitters off the repaired full-season log, the honesty flex off the
real servedGrade distribution (2,140 graded / 70 cleared B+ / 42% not
separable / A unissuable), and streaks verified from settled outcomes only.
THE ENGINE CAUGHT A BUG IN ITSELF, and it is the sharpest lesson here. The
first run published "No hitter is meaningfully hot tonight -- we could
dress up a middling week as a streak. We don't." That was FALSE: the
box-score cache spans only the settled window, every player had under 20
games, and the pool was empty. A broken pull was publishing as considered
editorial judgement -- the fourth appearance of this class tonight and the
first where our OWN HONESTY COPY was the disguise.
Fixed structurally rather than by patching the number: an absent() variant
may now DECLINE to speak, and the template separates "no candidates at
all" (SKIP with a reason) from "candidates judged, none hot" (honest
absence). Both locked by test. Source corrected to mlbStatsAdapter.fullLog,
the same log the repaired champion reads.
PHASE 2 — cardRenderer emits SVG rather than canvas: it is text, so it
diffs in review and its numbers are greppable, which matters when the
whole claim is that the numbers are real. VYND white + R green, slashed-Y,
scanlines, mono. The card never formats its own facts -- every string
arrives pre-rendered and gate-checked.
PHASE 3 — scripts/generate-content.js writes copy + card per template to
.content-out/<date>/. Template N+1 is a registry entry: requires, pull,
copy, card, absent. Queued as stubs, not built: hot takes, daily reads,
"grades we DIDN'T give", cross-sport streak variants (the streak template
is already sport-agnostic -- settled outcomes and a noun).
FULLY ISOLATED: read-only on every source, zero writes to serving, model
or ledger tables. Serving fingerprint verified unchanged. The accrual clock
is untouched at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
08791520fc |
Board: mechanical inventory of every outstanding item, read from the repo
Inventory only, no build. Every line cites a file or a query; unverifiable items are marked UNKNOWN rather than asserted. Corrections to the assumed board: - Scanner is DONE on the blue-boundary format (scan/page.tsx:715-717), not still amber. - Screen conversion is DONE except ONE file -- RouteStub survives only in app/notifications/page.tsx across 47 route dirs. - Terminal is deliberately retired (S57 redirect), not debt. - ev_pct DOES render, in 5 frontend files. That flag was stale. - edge_pct is NOT a scale bug: (model-line)/line is correct arithmetic that explodes on small lines (a 5.5 projection on a 0.5 line is a legitimate 1000%). Live top values 900/860/700 on 68,364 rows. It is a display defect, not a broken computation, and it reaches a user only via alt_lines typing -- the grade card computes edge independently. - All FOUR sports have real archetype registries (nba 15, mlb 15, soccer 6, wnba 5) and ACTIVE_SPORTS runs all four. Only MLB BATTERS is MODELED; everything else is scaffolding with a sound base rate and no proven factors. MLB pitchers: pitcherEngine.js exists at 261 lines and is read by ZERO serving code. - Book comparison is built and mounted NOWHERE -- the same built-but-unread class the reachability guard exists for, and NOT covered by it, since that contract is grade fields only. UNKNOWN and marked as such: opp_rank_stat prod population, internal-key rotation, the ~9% void/DNP rate, and whether MovementStrip is a distinct spec item or satisfied by GradeShift/MarketBreadth. The ordered board puts five PARTIAL items first and flags accrual sensitivity: seven items are isolated design/guard work buildable today with zero clock impact; five touch the serving path and would muddy the accruing repaired-champion dates. edge_pct is the sharpest tension -- a real honesty defect a user can see, whose fix touches serving mid-accrual. Clock today: 0 eligible dates, all four re-audit items WAITING. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
55b210cb95 |
Fix the dormant basketball window-bug before it ships; guard the class
PHASE 0 — audit. espnStatsAdapter's slice(0,20) was already fixed at |
||
|
|
981a05cbd6 |
Render-reachability guard: make built-but-unread a CI failure
Three consecutive orders shipped a backend-correct field that never
reached a screen, all on a green suite: gradeBands (required by no
serving code), served_grade (dropped at the adapter boundary),
GradeScaleLegend (imported by nothing). Each was caught by luck on a later
re-check, and in two of the three I had already reported the wiring done.
WHY GREEN TESTS COULD NOT SEE IT: backend tests stop at the API payload.
They prove a field is PRODUCED and say nothing about whether it is
CONSUMED. Invisible by construction, not an oversight in any one test.
THE TRAP, NAMED: the difficulty pools in the backend, so by the time a
field exists on the payload it feels finished. What remains is a
three-line adapter change nobody considers worth verifying, so it gets
claimed rather than traced. The last inch is the one with no friction,
which is exactly why it gets skipped. "I added the field" and "a user can
see it" are different claims and only the first is fun.
THE GUARD traces each promised field the whole way: payload -> adapter
consumes -> component renders -> component is MOUNTED. Mounted is
transitive to a Next entry point (page/layout/template), the only thing
that puts a pixel on screen, depth-limited so an import cycle cannot hang
the suite.
Container rows are exempted EXPLICITLY, not silently: served_grade carries
container:true plus a rendersVia list, and a separate assertion checks
every named part actually renders. The exemption is auditable and cannot
hide an unrendered field.
The guard tests itself -- an orphan component must report unmounted, and
the contract must be non-empty, since an empty contract passing vacuously
is how this would most plausibly rot.
RETRO-PROOF: run unchanged against
|
||
|
|
09186ea609 |
Wire the composed surface: the honest fields now reach a user
PHASE 0 re-check caught the same failure a THIRD time, mine again. Last
turn I created GradeScaleLegend.tsx and never mounted it, and I reported
that the separation flag "renders per band" -- it did not. grep: zero
frontend references to separates_from_base_rate, served_grade or
factor_adjustment. The honest grade was reaching the API payload and dying
at the adapter boundary.
That is three occurrences in three orders of the same shape: built,
correct, unread. gradeBands, then served_grade, then the legend.
PHASE 2 — the composed surface is wired end to end:
analyzeViaEngine1 -> served_grade + factor_adjustment on the payload
scan/page.tsx -> ScanResponse types them and forwards them
gradeAdapter -> gradeMeaning, separatesFromBaseRate,
bandRealizedRate, factorsApplied, refusalReason
GradeResultCard -> renders "WHAT THIS GRADE MEANS"
What a user now sees that they could not before: what the band has
actually realized, an amber note when the read CANNOT be separated from
the baseline, and -- only where a factor actually fired, with its proven
sign -- what moved the read, in plain language rather than feature names
("where he hits it vs who is standing there").
No narrative on props where nothing fired: factorsApplied is empty and the
block self-hides. Only the three PROVEN hits factors have labels, so an
unproven factor cannot acquire prose by being added to the map.
PHASE 1 — GradeScaleLegend is now MOUNTED on the grade card (compact). The
ceiling is a stated position where the grade is, not a page a user would
have to find.
PHASE 3 hand-verified through the real adapter across twelve states: B+
with 3/0 factors, B with 1, C+/C/C- all flagged not-separable, D, F, a
switch-hitter case where 2 of 3 factors fire, two refusals and a
no-forecast. never-blank PASS, no-manufactured-A PASS, no-narrative-when-
nothing-fired PASS, separation-flag-reaches-card PASS.
No A-threshold loosening. No calibrated number leaks. p_win never mutated.
engine_grade still read by zero serving code. No Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
3591c7626e |
Total grade cutover + the ceiling stated as a position
PHASE 0 caught my own repeat of the failure I diagnosed one order ago.
|
||
|
|
91927a4a8a |
Serve an honest grade: the letter was carrying 1/6 the information of the
number beside it PHASE 0 corrects the order's premise. A grade letter has been served all along -- engine1.gradeProp builds it from an additive factor index, computed INDEPENDENTLY of p_win. gradeBands is orphaned for a different reason than assumed: it defines what a letter MEANS from realized outcomes, and every band collapses to base-rate at current resolution. The measurement that changed this order, on 3,417 settled props: grade n realized mean p_win A 8 0.500 0.647 <- the TOP grade did WORST B 985 0.640 0.700 C 1,695 0.602 0.676 D 303 0.558 0.604 F 426 0.535 0.588 letter resolution 0.00116 (0.48% of variance) p_win resolution 0.00715 (2.98%) -> the letter carried 0.16x the information of the number beside it Concretely, from the hand-verify: Christian Encarnacion's 0.95 over graded C and his 0.05 under ALSO graded C -- same hitter, opposite forecasts, same letter. The gap was never that grades don't ship; it is that the weaker of two available signals shipped as the headline. PHASE 1 — model/servedGrade.js derives the letter from p_win with bands anchored on MEASURED realized rates (B+ 0.663 / B 0.646 / C+ 0.615 / C 0.589 / C- 0.548 / D 0.512 / F 0.447, base 0.6005). NO MANUFACTURED A, structurally: A+/A/A- are UNISSUABLE, not rare. The realized rate plateaus at 0.65-0.68 above p_win 0.70, so no band has earned a top letter; a test sweeps every p_win 0..1 and asserts none produces one. Even 0.99 tops out at B+ with its realized 0.663 attached. Raising that ceiling later is a deliberate, visible act. Bands that cannot separate SAY so -- C+/C/C- carry separates_from_base_rate false and copy naming it, which is the honest description of a forecast explaining 3% of variance. Every grade states its basis (forecast_only vs forecast_plus_matchup_factors, naming which factors fired) and calibrated:false. engine1.grade is preserved as engine_grade so nothing downstream breaks. PHASE 2 — refusals render real states: insufficient_data -> "not enough history to call this one"; juiced_no_edge -> "the book has priced the vig past any edge on this side". 1,870 refused snapshots carry exactly those two reasons and both now surface. PHASE 3 — hand-verified on 12 real served props. Freeman/Rice/Encarnacion 0.95 overs now B+ (was B, C, B); the 0.05 unders now F (was C). Refused doubles render NO READ with their reason. never-blank PASS, no-manufactured-A PASS. Serving change; nine frozen model modules unchanged including engine1; p_win never mutated; no calibrated number leaks (deployed set empty); no Bonferroni slot. STILL TRUE: the forecast explains ~3% of outcome variance. This order did not make the model better. It made the letter stop overstating it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
9159b7e1b9 |
Diagnose the push block: it was the wrong remote, not the firewall
The GATE-0 theory is REJECTED for VYNDR, tested rather than assumed. Both remotes are ALREADY HTTPS -- there is no git@ SSH URL in this repo to be blocked -- and both hosts answer on 443: git.builtbykev.com HTTP 200 in 0.64s, github.com HTTP 200 in 0.17s, Gitea's git endpoint 200, GitHub's 401 (auth required, reachable). No connectivity failure of any kind. COLYRA's HTTPS-remote fix was right for COLYRA; VYNDR was already in the state that fix produces, so applying it would have minted a new token to solve a problem that did not exist -- and the pre-existing credential would have made the new token look like the cure. The real cause is in the error text, which named it exactly and was misread all session: "could not read Username for 'https://github.com'". Every attempt used `git push origin main`, and origin is GitHub (kev3109/betonblk) with no stored credential. The WORKING remote is `gitea` (git.builtbykev.com/builtbykev/vyndr.git), and a valid credential for it sat in ~/.git-credentials the entire time. The habit of typing `origin` is what kept twenty commits local. Nothing was blocked. FIX: git push gitea main. Pushed 6452926..ecf78b9, 21 commits, verified by git ls-remote matching local HEAD exactly. No firewall rule touched, no token created, no code or model change -- infra only. The bundle and patch series stay as belt-and-braces; they are no longer the only copy. docs/GIT-PUSH-DIAGNOSIS.md records it so no future session re-reaches for an infrastructure theory when the error message names a credential and a specific host. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
ecf78b911c |
Accrual watch + pre-registered resumption; the program is idle on modeling
PHASE 1 — the existence risk is resolved. Nineteen commits were local-only
with no push credentials. Three artifacts now exist off the working tree:
~/vyndr-full-history-2026-08-07.bundle 6.8M SELF-CONTAINED, clone it
~/vyndr-session-2026-08-07.bundle 244K needs existing history
~/vyndr-session-patches/ (20 patches) 1.5M
Use the full-history bundle -- `git bundle verify` shows the session bundle
requires ref
|
||
|
|
2391574f00 |
Live-surface integrity on the repaired champion; fix the label my own
repair falsified
PHASE 0 — four checks PASS, one defect found and fixed.
PASS gradeBands is required by NO serving code -- built across several
orders, never wired. No stale band derived from the retired
ten-game champion can reach a user, because none reaches a user
at all.
PASS CALIBRATION_DEPLOYED is [] and the calibrate loop iterates it, so
calibrate() is never called and p_win_calibrated is never set. The
only assignment site sits inside that empty loop. No withdrawn map
can leak.
PASS projectionFor reads l20_avg, which mlbGameLogFeatures now builds
from fullLog -- so refusals are computed on the repaired
full-window reference, not the retired ten-game one.
PASS factors still fire with correct sign across the repaired base
range (0.35/0.50/0.65/0.80): defense lowers, pitcher-contact
raises, platoon raises at every point. Mechanical firing check
only -- NOT a lift re-measurement, which waits for accrual.
DEFECT FIXED — my own repair falsified a user-facing sentence. The grade
card rendered "Last 20 games average: X" from l20_avg, and l20_avg is now
a FULL SEASON average. The number changed and the label did not, so the
surface was stating something the data no longer supported. Copy now reads
"Season average"; trapDetection's L20 explanations likewise. The field
name is kept -- it is read in many places -- but no rendered sentence
claims a window that isn't there.
That is the same class as everything else tonight, one layer out: a
correct-looking string describing data that moved underneath it.
Serving change; frozen model modules unchanged; p_win never mutated; no
Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
494c83cf76 |
Hunt the window-bug class: three more paths, and the forward re-audit rule
in code PHASE 0 — getStatRows is the single base-rate path, so every branch is audited, plus the feature builders since l20_avg is the season reference projectionFor reads: getStatRows MLB -> estimator base fullLog CORRECT ( |
||
|
|
929fd81940 |
Repair the champion: it was reading ten games, not a season
PHASE 0 — the defect is real past the peek. Against a FAIR point-in-time baseline (each player's rate over games strictly before that date, >=10 prior games, box scores back to 05-01), the served champion LOSES on all four stats, three of four CIs excluding zero: hits 0.00251 vs 0.00774 CI [-0.0074,-0.0011] TB 0.00393 vs 0.00619 CI [-0.0055,-0.0003] rbi 0.02481 vs 0.03133 CI [-0.0153,-0.0005] runs 0.00181 vs 0.00683 CI [-0.0114,+0.0008] PHASE 1 — the cause is the WINDOW, not the weights. estimateProbability builds its base rate as the frequency over every row it is handed, and featureCache.getStatRows handed it res.last10. So the "season rate" was a TEN-GAME rate, and 0.4 of the forecast was the last five OF THOSE TEN. The 0.40 recency weight costs resolution on all four stats (-0.00086, -0.00107, -0.00562, -0.00365). Nudges are mixed and small -- harmful on hits and rbi, marginally helpful on TB and runs -- so they are left alone. PHASE 2 — two lines, no new data, no extra API call, because fullLog was already fetched by the same adapter call that produced last10: getStatRows now reads fullLog, and RECENCY_WEIGHT goes 0.40 -> 0.20. hits 0.00251 -> 0.00817 (tripled; now above the fair baseline) TB 0.00393 -> 0.00734 (above baseline; vs old CI [0.0020,0.0067]) rbi 0.02481 -> 0.02727 (still below baseline, CI includes zero) runs 0.00181 -> 0.00436 (still below baseline, CI includes zero) Gate stated exactly: hits and TB now exceed the fair baseline on the point estimate; rbi and runs remain below but EVERY CI now includes zero, so no stat reliably loses to a frequency table. That is a tie on rbi/runs, not a win, and it is reported as one. Only TB's improvement over the old champion is CI-confirmed; the rest are directional. STALE-FIT GATE: CALIBRATION_DEPLOYED is now EMPTY. The low-param maps were fitted on the retired forecast and fromLedger cannot rescue them -- settled ledger rows still carry OLD p_win, so refitting today would refit the retired forecast. Nothing is served calibrated until dates settle under the repaired champion, and the favourite-longshot bias must be re-measured rather than assumed to survive. The shadow duel is void. PHASE 3 — the hits factor lift is NOT re-measured, and cannot be yet: it needs settled rows produced BY the repaired champion, which ships in this commit. Replaying would score the factors against a reconstruction rather than the served forecast. Deferred, explicitly. The factors remain wired and transmitting; only their lift is unquantified on the new baseline. PHASE 4 — standing flag, and it is large: EVERY factor verdict in this programme, every null and every THEATER, was measured against a champion worse than a frequency table. Signal added to noise reads as noise. Prior verdicts may deserve re-audit. Logged, not re-run. Re-queued not built: rbi lineup-slot / RISP opportunity through the two-part gate, now landing on a repaired champion. Serving-path change by design; the byte-identical invariant inverted and all four stats move. Nine frozen model modules verified unchanged. No Bonferroni slot -- resolution accounting on the champion's own knobs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
65ca6493db |
Decompose the rbi anomaly: it is lineup ROLE, and the counter is
out-resolved by a frequency table on three of four stats
PHASE 0 — the 14.51% is REAL. Re-derived with a paged pull asserted
against an exact count (rbi 7,930 == 7,930; hits 11,690; TB 12,086; runs
6,440), since this harness produced a false null three times tonight. rbi
resolution 0.03268 reproduces, deciles are monotone through the middle,
and 20 raw rows are in the artifact for hand audit.
CAVEAT GOVERNING EVERYTHING BELOW: the naive forecasts are leave-one-out
ON THE EVALUATION WINDOW, so they see the rows they are scored on while
the model is strictly point-in-time. They are upper bounds on available
resolution, not fair competitors, and every comparison is read that way.
PHASE 1 — the split:
stat MODEL (a)player-base (b)lineup-slot (c)within-stratum
rbi 0.03268 0.01167 0.03608 0.01908
hits 0.00252 0.00446 0.00100 0.00473
TB 0.00442 0.01331 0.03448 0.00607
runs 0.00130 0.00262 0.01156 0.01170
FINDING 1 — rbi's resolution is LINEUP ROLE almost exactly. Batting-order
slot alone resolves 0.03608 against the model's 0.03268. A single integer
accounts for the whole anomaly and slightly more. That is opportunity, not
skill -- the cleanup hitter bats with runners on. 36% is matched by player
identity alone. Within similar-base-rate strata the model still resolves
0.01908, 58% of its total and higher than any other stat's ENTIRE model
resolution, so genuine within-role discrimination exists on top.
FINDING 2 — on three of four stats the model is beaten by "he's a .270
hitter". Base-rate-only out-resolves the model 1.8x on hits, 3.0x on TB,
2.0x on runs. Even allowing for the window-peeking advantage, a 1.8-3.0x
gap is not explained by that alone: the served counter appears to DESTROY
discrimination relative to the player's own rate. rbi is the one stat
where the model beats the naive baseline.
FINDING 3 — lineup slot out-resolves the MODEL on three stats: TB 7.8x,
runs 8.9x, rbi 1.1x. Hits is the only stat where batting order carries
less, which is mechanically right -- a hit is a hit wherever you bat, but
runs, RBI and total bases all scale with opportunity.
PHASE 2 — all three worlds are partly true, in measured proportions.
World A ~90% true (slot covers rbi's entire resolution). World B ~36% true
for rbi, but the WHOLE story for hits/TB/runs where base rate alone wins.
World C true with a low ceiling: hits' total available spread resolution
is 0.00446, i.e. 1.8% of variance from a forecast that has seen the
answers.
PHASE 3 — the next arc is NOT "strengthen hits factors". Hits has the
lowest available resolution on the board and last order's wiring already
took it to 1.39% of a ~1.8% ceiling. Named first factor order for next
session: LINEUP SLOT / RISP OPPORTUNITY on rbi through the two-part gate --
input already ingested and prod-verified (S89), resolution measured not
hypothesised, causally-correct unit is plate appearances with runners on.
Measured availability is not a pass; it still faces the gate.
And higher-value than either: the counter being out-resolved by a
frequency table on three of four stats is a defect in the CHAMPION, not a
factor problem, and it costs nothing to test -- the recency blend and the
+/-0.03 / +/-0.015 nudges are three lines in probabilityEstimator.
The hits transmission win from
|
||
|
|
43f65d30cb |
Wire the three proven hits factors pre-grade: transmission proven, gain
inconclusive THE BUG THIS NEARLY SHIPPED AS A FINDING. The first audit reported 0 factors fired on all 1,140 rows. Not a result -- my paging helper ordered by `id`, and batter_spray, team_defense, platoon_splits and statcast_aggregates have composite primary keys with NO id column. The query errored, the loop broke on error, and four fully-populated tables read as empty. hitsFactorContext.js -- the PRODUCTION loader -- had the identical defect, so live wiring would have loaded nothing and served unadjusted while logging success. Third occurrence of this class in one session. Both loaders now order by a real column and THROW rather than degrade. The Phase 2 gate is what caught it: no resolution number was quoted until transmission was proved. PHASE 1 — pipeline is now base -> FACTORS -> CALIBRATE -> GRADE. Context built in snapshotService BEFORE gradeAndCacheSlate (was line 640+, grade at 454), threaded per prop, applied to p_over before p_win is set with p_win_prefactor and a full trace retained. Hits only. Coverage 859/1140 rows (75%): 474 with all three factors, 256 two, 129 one, 281 none. PHASE 2 — TRANSMISSION PROVEN, 12/12 sign-correct, 4/4 per factor, each applied IN ISOLATION. My first table compared each factor's expected sign against the COMPOSITE change and showed 3 false failures -- with three factors firing the net can oppose any single member; that was a flaw in the test, not the wiring. Two under-side rows confirm the flip is handled: a factor raising p(over) correctly lowers p_win. Switch hitters (Bailey, Bell, Rocchio) took no spray adjustment while their other factors fired normally -- the refusal is selective, not a blanket skip. PHASE 3/4 — both maps refit on the factor-adjusted forecast; the shadow-duel baseline is VOID and restarts, since it accumulated against a different forecast. Point-in-time, 765 held-out rows: reliability 0.00795 -> 0.00828 RESOLUTION 0.00229 -> 0.00345 (variance explained 0.93% -> 1.39%) Brier 0.25398 -> 0.25305 delta -0.00093 CI [-0.00225,+0.00002] Resolution rose 51% relative. The CI TOUCHES ZERO on 4 eval dates, so the composition does NOT earn a proven keep -- three isolated passes did not grant a composed pass. INCONCLUSIVE, reported as such. The gain is far below the sum of the isolated effects, which is expected: all three run through the same pitcher-batter confrontation and share signal. PHASE 5 — 1.39% of variance is still far below what band separation needs. The pivot was correct and incomplete: the plumbing defect was real and is fixed, three proven factors reach the served number for the first time, and transmission alone did not buy grade separation. Next arc is factor STRENGTH and BREADTH, not more plumbing. PHASE 6 — rbi anomaly logged, not chased: 14.51% variance explained vs hits 1.03%, on the stat we do not serve corrected and which has no proven factors. Either the biggest lever on the board or a mirage; it deserves its own order. The byte-identical invariant INVERTED for hits by design. All 13 frozen non-hits modules verified unchanged, probabilityEstimator included -- the factors ride outside it. No new Bonferroni slot; the composed OOS claim is reported with its CI and not claimed as a pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
e872eff4ce |
Instrument the calibration duel forward; diagnose the resolution ceiling
— the proven factors were never wired in
PHASE 0 — two truths recorded. The swap is a BET, not an OOS win:
isotonic beat low-param on identical held-out rows (hits +0.0028, rbi
+0.0042, TB tied) and we serve low-param anyway on an untestable prior
about shared daily structure. At 19 dates nothing here can test it. And
the MIN_SLOPE catch is preserved as standing rationale: a near-zero or
negative slope collapses toward base-rate-for-everything, which LOWERS
Brier while destroying all resolution -- a metric win that guts the
product.
PHASE 1 — the duel is now falsifiable. Both corrections computed on every
hits/TB prop; p_win_lowparam served, p_win_isotonic_shadow logged in its
own try so it can never break serving. calibrationDuel.adjudicate encodes
the rule IN CODE before any forward date exists: >=10 forward dates and
isotonic winning with a date-block CI excluding zero => REFUTED, revert;
otherwise UPHELD; under 10 dates PENDING regardless of the numbers. A
date counts as forward only if NEITHER map was fitted on it -- otherwise
we would be scoring which map memorised better. Nothing swaps now.
PHASE 2 — the ceiling, quantified via Murphy decomposition:
stat reliability RESOLUTION uncertainty variance explained
hits 0.01353 0.00252 0.24532 1.03%
TB 0.01419 0.00442 0.24329 1.82%
rbi 0.00654 0.03268 0.22531 14.51%
runs 0.00788 0.00130 0.23182 0.56%
Calibration did exactly what theory says and nothing more: hits
reliability 0.01353 -> 0.00233 (-0.0112, 83% of the error removed) while
resolution moved -0.0002. Unexpected: rbi has 13x the resolution of hits
and is the one stat we do NOT serve corrected -- it needs calibration
least and discriminates most.
PHASE 2 DIAGNOSIS — NOT-TRANSMITTED, and not weak, ABSENT. Traced in code:
sprayDefense.js and platoonSeverity.js are required by NOTHING in src/,
only by analysis scripts and their own tests. The served p_win
(intelligence/probabilityEstimator.js:54) reads exactly four inputs --
game-log frequency, opp_rank_stat +/-0.03, home_away +/-0.015, and a cv
pull -- with zero occurrences of spray, platoon, hard-hit or
contact-profile. And snapshotService grades at line 454 while computing
challenger/context at 640+, so everything proven is computed DOWNSTREAM of
the grade it would inform. The three proven hits factors have never once
moved a served number.
That reframes the recent nulls: "calibrated p_win does not separate within
archetype" was never a statement about factors. The factors were not in
the forecast.
PHASE 3 — bands rebuilt on SERVED values (hits/TB low-param, rbi/runs
raw): 28 archetype slots across four stats, ZERO show lift. No longer an
open shrug -- it is the arithmetic of resolution 0.0013-0.0327 against
uncertainty ~0.23. A forecast explaining 1% of variance cannot produce
separating bands, and no correction to its numbers will change that.
HEADLINE: calibration is complete, delivered honest numbers on two stats
and zero grade separation, because the counter has no resolution -- and
the proven factors are not wired into the forecast at all. The second is
the reason for the first, and it is plumbing rather than a modelling wall.
Per-archetype grades need proven factors that actually reach p_win. Last
calibration order.
Serving unchanged from
|
||
|
|
74cf1ce974 |
Robust bias established; low-parameter correction replaces isotonic
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.
PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.
PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:
hits a=0.406 shrink 0.565 0.2626 -> 0.2540 CI [-0.0112,-0.0069] DEPLOY
total_bases a=0.472 shrink 0.333 0.2490 -> 0.2429 CI [-0.0062,-0.0059] DEPLOY
rbi a=0.775 shrink 0.231 0.2011 -> 0.2007 CI [-0.0007, 0] REFUSE
runs a=-0.032 REFUSE
A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.
STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.
PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at
|
||
|
|
ced40421ed |
Audit the LODO instrument: it cannot evaluate any stat, and both prior
FAILs were false
PHASE 0 — the gate at
|
||
|
|
1f40014256 |
Power-derive the LODO threshold: hits restored through the gate, rbi/runs
routed as date-driven PHASE 0 — threshold derived BLIND, before any stat was re-read. A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count. Per-row Brier difference d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single stat's verdict could shape the threshold deciding it: pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175 n* = (0.09816/0.01175)^2 = 69.8 -> 70 The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect this removes, and why the previous verdict moved with the number. Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it equals the constant, so it cannot drift from its own justification. The derivation script prints no stat verdict, no date and no reversal. PHASE 1 — LODO at n*, applied cold: hits 5 informative drops, 0 reversals PASS total_bases 4 informative drops, 0 reversals PASS rbi reverses 2026-08-01 (n=99) FAIL runs reverses 08-01 (n=86), 08-05 (244) FAIL hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across 123-272 row dates, favourite sign holding on every testable drop. THIS IS THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the withdrawal at |
||
|
|
6ae11f1193 |
LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration layer without challenging the binding. That floor is a cluster-robust interval bar for a CAUSAL claim. Calibration makes no causal claim, has a bounded failure mode (it can only over- or under-shrink) and consumes no Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and leave-one-date-out tests that directly -- a STRICTER bar, since a cluster count cannot detect a single day carrying the effect. The >=40 floor is retained, correctly scoped as the PROMOTION bar. PHASE 1 — both guards codified, 11 tests, green before Phase 2. Demonstrated on live data: raw population violated=true, mean_p 0.4962, both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The null guard's test demonstrates the trap explicitly, since (null-1)**2 is 1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate. PHASE 2 — LODO: hits n=1140 dates=17 2 reversals (07-22 n=20, 07-26 n=25) FAIL total_bases n=1050 dates=7 0 reversals, 0 sign flips PASS rbi n= 630 dates=5 1 reversal (08-01 n=99) FAIL runs n= 597 dates=5 2 reversals (08-01 n=86, 08-05 n=244) FAIL Threshold sensitivity reported because the verdict moves: total_bases passes at every held-size threshold, runs fails at every one, and hits fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20 before seeing which stats passed and did not move it afterwards to preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a standard error several times the effect, so the instrument is underpowered per-drop -- an argument for pre-registering a higher threshold, which is a Roundtable call, not one to make while holding the results. PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and runs REFUSE. HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService hardcoded it since S91; it fails LODO, so it is out. A stat that cannot survive dropping one day was never calibrated, it was fitted to that day. The consequence is real -- hits props become unstackable for chain.chainAcross -- and it errs toward withdrawing a claim rather than preserving one on a fragile verdict. Deployment is now driven by a frozen, tested CALIBRATION_DEPLOYED set, not a hardcoded stat name. PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither waivable. reverify auto-demotes on the first breach (CI stops excluding zero, or the favourite bias flips sign) and logs the breaking date. Promotion needs the original >=40 bar. A provisional deploy that cannot be taken away is just a deploy. PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The two-bar rule still bites: calibrated YES, proven NO, so they stay a base-rate read, now honestly numbered. Every archetype still collapses to one band -- calibrated p_win separates within archetype no better than raw. PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI, and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did not survive); the refused set is a map of missing inputs; a low-parameter calibrator is queued unbuilt. p_win never mutated; calibration rides as p_win_calibrated with calibration_status provisional. No Bonferroni slot consumed. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
f976df47b8 |
Settle model_snapshots + four-stat calibration: works, deploys nowhere
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.
PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.
PHASE 1, two integrity findings both caught by the gate:
1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').
2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.
THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.
THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.
PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):
hits n=1140 bias +0.0868 brier 0.2626 -> 0.2511 d -0.0115 CI [-0.0139,-0.0097]
TB n=1050 bias +0.0834 brier 0.2490 -> 0.2438 d -0.0052 CI [-0.0061,-0.0045]
rbi n= 630 bias +0.0164 brier 0.2011 -> 0.1965 d -0.0046 CI [-0.0092,-0.0010]
runs n= 597 bias +0.0410 no map fittable (173 fit rows < 200)
ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.
A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.
PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.
PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.
PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.
p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
23d1b13176 |
runs + RBI: mostly-base-rate confirmed, and one level deeper than expected
Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest grade is the POOLED base rate. PREMISE NOTE: the order's closing line says the batter board is per-archetype-graded after this. Nothing has been rescaled for hits or total_bases either -- no archetype slot has ever reached sample and gradeBands remains built, gated and unwired. This is the fourth stat measured, not the completion of three. AUDIT: RBI 935 clean / 43 games; RUNS 617 clean / 33 games. Zero quarantined. No archetype slot reaches 500 -- and the signature archetypes the order names are the two SMALLEST slots on the board, RBI->DRIVER at n=24 and runs->CATALYST at n=9. RUNS is refused structurally before any factor is tested: 33 game clusters against a 40 floor. INPUTS RECONSTRUCTED rather than declared missing. lineup_context only covers 08-04 onward while settled rows start 07-31, so 187/617 runs rows joined. But the play-by-play cache runs from 05-01 and the batting order IS the order batters first appear -- slot, power-behind and reach-base all rebuilt point-in-time, coverage 187 -> 574. RBI, all THEATER: risp_opportunity +0.0047, extra_base_skill +0.0010, risp x extra_base +0.0056. RUNS, all refused on clusters and all pointing the wrong way: +0.0043 / +0.0008 / +0.0054. THE COMPOUND IS THE WORST VERSION IN BOTH STATS. The causally-correct compound was the most promising factor on the sheet and is the most harmful in each. Two multipliers that individually carry nothing do not cancel -- they compound each other's noise. Distinct from the collapsed-sequence lesson: there the product of two REAL effects was too small to use; here the product of two NULL effects is worse than either. THE ARCHETYPE DOES NOT RESCUE IT, and this is where the session nearly went wrong. The base rates look strongly differentiated (RBI DRIVER 0.609 vs BOMBER 0.413; runs GHOST 0.716 vs BOMBER 0.460). Gated directly against the pooled base rate: RBI +0.0010 CI [-0.0034,+0.0050] THEATER; runs -0.0028 CI [-0.0147,+0.0108] candidate at k=33. DRIVER's 0.609 is n=23 -- small-slot noise wearing a decimal point. Read off the table instead of gated, this would have shipped as "archetype differentiation is real and large". It is not. THE CROSS-STAT PATTERN THAT IS REAL -- the counter over-predicts every batter counting stat measured: total_bases p_win 0.5698 vs actual 0.5074 bias +0.0624 rbi p_win 0.4860 vs actual 0.4313 bias +0.0547 runs p_win 0.5949 vs actual 0.5749 bias +0.0200 Across four stats and three sessions, calibration is the systematic defect and factor scarcity is not. TB's held-out isotonic fix (-0.0039) still outperforms every factor tried on any stat, all null or theatre. NO RESCALE. Nothing proved, nothing certified calibrated, no slot at sample, and for RBI the archetype split is itself theatre -- so the honest band is the pooled base rate, which gradeBands returns by construction. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
6a327d9114 |
total_bases: every power factor is THEATER, and a units bug nearly hid it
PREMISE CORRECTION: the per-archetype rescale is not "proven and live on hits". gradeBands was built, gated and explicitly NOT wired two orders ago -- no hits archetype slot reached sample, every band came back base-rate, and only defense_by_direction proved pooled. This applies an unvalidated-at-archetype-level method to a second stat. FULL-HISTORY AUDIT: 988 clean settled TB rows (101 quarantined, 948 with p_win), 341 players, and only 9 DISTINCT GAME DATES. No archetype slot reaches 500 -- BOMBER 340, GHOST 147, BRUSH 55. Confirmed short on full history, not a windowed artifact. The 9-date figure matters more than the row count: ~49 games means any game- or venue-borne factor has almost no replication here. THE BASELINE HAD TO CHANGE, to a harder null. TB lines vary (1.5 on 559 rows, 0.5 on 345), so a per-line personal base rate would rest on ~2 rows per player-line and would have to be invented. The null is the counter's own p_win, which already prices the line -- beating the champion, not beating "he's due". THE UNITS BUG, caught, and it had produced the best result in the programme. The first run reported barrel_rate at Brier -0.0095, the largest improvement ever measured here. fromStatcastRow returns barrel_pct as a FRACTION (0.06) while the raw table stores 0-100, so (0.06 - 7.8) * 0.018 clamped EVERY row to the maximum negative shift. That uniform downward push "improved" Brier purely by leaning on the counter's over-prediction and contained no barrel information at all. Same family as the S80 trap, inverted. exit_velo was a second bug -- the column is avg_exit_velo, so it read null on every row and reported n=0. A zero is a wiring bug until proven an honest absence. GATE with units fixed, 138 cumulative tests: barrel_rate n=707 shift 0.0364 brier +0.0036 THEATER exit_velo n=707 shift 0.0229 brier +0.0022 THEATER hard_contact_allowed n=707 shift 0.0260 brier +0.0033 THEATER park_weather_hit_type n=651 36 entities PENDING (k<40) platoon_severity n=481 PENDING (n<500) THE PREDICTED INVERSION WENT THE OTHER WAY. BOMBER x barrel_rate is +0.0114, the single most harmful cell in the table, exactly where the strongest proof was predicted. GHOST +0.0012. All sample-blocked so not a verdict, but recorded so it is not claimed later. AND IT IS NOT DOUBLE-COUNTING -- tested and refuted: corr(barrel, p_win) = -0.061, the counter is not pricing barrel at all. The duller answer is corr(barrel, counter RESIDUAL) = -0.012. Barrel is a real skill that carries no information about what the counter gets wrong at this line. That also closes the S81 lead: hard_hit r=0.153 at n=295 drifted to 0.135 at n=383 and is THEATER at n=707. THE REAL FINDING: TB is miscalibrated, not under-factored. mean p_win 0.5698 vs actual 0.5074, bias +0.0624. Held out on a strict time split (fit < 2026-08-02, eval 651 unseen rows): raw 0.25007, constant de-bias 0.24740 (-0.00267), isotonic 0.24621 (-0.00386). Worth more than any factor tested and the only intervention pointing the right way -- and still refused at the corrected bar on 32 clusters. A CANDIDATE, not a result. It also explains the units bug's fake success exactly: a blanket downward shift is a crude de-bias. NO RESCALE. Nothing proved, nothing certified calibrated, no slot at sample -- every band would be the honest base-rate band gradeBands already returns by construction. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
8ab6557faa |
Collapsed sequence edge: two proven links whose product is too small to use
Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.
FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.
TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.
GATE on the concentrated subset, 114 cumulative tests:
early-exit x WEAK pen n=1931 brier -0.0001 CI [-0.0014,+0.0010] NOT_PROVEN
early-exit x STRONG pen n=2574 brier 0.0000 CI [-0.0011,+0.0010] THEATER
all early-exit later ABs n=6869 brier -0.0001 CI [-0.0007,+0.0005] NOT_PROVEN
pooled all later ABs n=17891 brier 0.0000 CI [-0.0004,+0.0003] THEATER
Not pooled-diluted -- the concentrated subset was gated alone and is no
better.
THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925
hit-rate swing across pen quality 0.0394
MAX JUSTIFIABLE ADJUSTMENT 0.00365
actually applied 0.01930 -> 5.3x over-movement
A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.
A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.
NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.
THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.
Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
b2e4c6c4fb |
Link 2 at the coarse grain: pen QUALITY proves, archetype does not
The refinement was right. Naming the individual reliever failed; the same
question at the grain the chain needs passes, and it transmits more than
anything else measured in this chain.
WHY IT WAS WORTH RE-ASKING: last session's null (the pen is on average no
softer, +0.0010 on 35,760 PAs) does NOT rule this out, and treating it as
though it did would have been the error. An average washing out is fully
consistent with quality VARIATION mattering. It does -- actual arm quality
moves the hit rate monotonically across quartiles, 0.2244 / 0.2293 /
0.2410 / 0.2501, a 2.57pp spread, larger than the whole times-through-
the-order effect.
CLUSTER UNIT CORRECTED, THEN CHECKED RATHER THAN ARGUED. Last session
refused Link 2 partly as team-borne (30 bullpens, the park ceiling). My
first re-check was that 76% of pen-quality variance is within-team -- but
that is a statement about TREATMENT variance, not about where errors
correlate, and stopping there would have been picking the convenient
answer. Measured the actual thing: ICC of prediction error by team =
0.0261, design effect 1.41, SEs inflated ~19%. So the verdict was run
three ways:
unclustered CI [-0.0067,-0.0010] excludes zero
team-clustered (30) CI [-0.0086,-0.0003] excludes zero (below the
40-cluster floor -- indicative, not a pass)
design-effect adjusted CI [-0.0072,-0.0005] excludes zero
QUALITY GRAIN PROVES on the concentrated elevated-early-exit subset:
n=501 team-games, 426 clusters, MAE 0.0294 -> 0.0260, delta -0.0034, CI
[-0.0063,-0.0005] at 110 cumulative tests. Pooled also proves, so it is
not a subset artefact.
ARCHETYPE GRAIN DOES NOT: 0.5669 vs a 0.5309 modal-guess baseline,
corrected interval [-0.1073,+0.0268] spans zero. Two grains tested, one
earned a place -- penQuality.js exposes no archetype and a test asserts
it.
WHAT LINK 3 RECEIVES, which is the number that actually matters -- not
the MAE gain but realized outcome separation, prediction strictly
point-in-time:
predicted BEST pen 167 games 2,044 PAs hit rate 0.2231 +/-0.0180
predicted WORST pen 167 games 1,799 PAs hit rate 0.2501 +/-0.0200
2.70pp separated, intervals non-overlapping, capturing nearly all the
2.57pp available at the quartile grain. Caveat stated not buried: the
tercile cut is chosen in-sample; the prediction driving it is not.
BUILT: penQuality.js + 9 tests. Abstains below 5 prior club games and 40
arm appearances -- a league-average stand-in would assert "this is an
ordinary bullpen", which is a claim, and usually the wrong one for exactly
the clubs whose pens just turned over.
Link 3 is unblocked on a proven Link 2 at the quality grain only. Not run
here; this order scopes to building and gating Link 2.
Parallel track logged unchanged: TB n=948 pooled, BOMBER x TB 340, short
by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
e4dae0e6b0 |
Reliever chain: Link 1 proves, Link 2 does not, and the premise inverts
The causal insight is right -- the game is a sequence and the matchup does shift mid-game. The direction is backwards, measured on 93,663 plate appearances from 1,238 games pulled free from statsapi. LINK 1 PROVES. Starter batters-faced, point-in-time from his own prior starts only, clustered on the pitcher: MAE 3.2226 -> 2.7990, delta -0.4236, CI [-0.6006,-0.2731] at 0.9995 corrected for 107 tests, 1,706 starts across 204 pitchers. It finds the tail the chain needed -- early exits are a 23.2% base rate, model-flagged starts are 34.0% early, lift +10.8pp. Scope correction inside Link 1: the order specifies fatigue x GAME SCRIPT, but game script is not available at grade time -- whether he gets hit tonight is the thing being projected, not an input to it. Only the workload half is measured; the in-game half is recorded as a live feature, out of scope, rather than quietly folded in. LINK 2 DOES NOT PROVE, twice over. Model accuracy 17.2% vs an 8.6% baseline -- doubling it sounds good and is not, since naming a specific arm is wrong five times in six. And structurally the entity is the BULLPEN: 39,629 post-starter plate appearances across 30 clubs is 30 readings, below the 40-cluster floor, the same permanent ceiling as park geometry and team defence. LINK 3 NOT RUN, per the order's own rule. THE PREMISE IS REFUTED, and this chains on nothing so it was safe to measure: vs STARTER n=48,492 hit rate 0.2444 +/-0.0038 vs BULLPEN n=35,760 hit rate 0.2373 +/-0.0044 The pen is 0.7pp HARDER. The specific effect the chain exists to exploit -- early exit making later at-bats softer -- is +0.0010 on 35,760 PAs. A well-powered null, not a sample problem. What IS real is times through the order: TTO1 0.2351 -> TTO2 0.2515 -> TTO3 0.2518. A starter does decay as the lineup sees him again, but that advantage is SURRENDERED when he leaves, not extended -- the pen is harder than his second and third time through. A modern bullpen is a queue of fresh specialists throwing one inning each; there is no tiring arm to punish. So the insight survives inverted, and Link 1 stays valuable for the opposite reason it was built: a likely early hook predicts the hitter LOSES his third-time-through look (0.2518 -> 0.2373 on that PA). The mispricing is on hitters who get an EXTRA look at a starter going deep. BUILT: predictionGate.js + tests -- the two-part gate for a continuous prediction. factorGate binarises outcomes for Brier, which would destroy a target like batters faced. Same discipline, same THEATER verdict, real scale. PRE-REGISTERED NOT RUN: Link 2' using a PA-weighted bullpen AGGREGATE rather than a named arm. Recorded rather than substituted in -- running Link 3 on a swapped-in Link 2 is the assumed-link failure the order forbids. Given the premise result its expected value is now low. PARALLEL TRACK logged: total_bases n=948 pooled, BOMBER x TB 340, short by 160. Sample-readiness only, not a verdict. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
3081c92e00 |
Per-archetype grade bands: built, gated, and the rescale blocked twice
The premise does not hold. proven-status.js run fresh: PROVEN_SET is EMPTY, no archetype x stat reaches the gate. pitcher_contact_profile has a CI upper bound of exactly 0.0000 and platoon_severity is held on 4.5%-contaminated splits, so the proven set is one factor, pooled, not three archetype-conditioned ones. The specific pattern the order names -- defense strong for GHOST/BRUSH, null for BOMBER -- is the one I measured running the OTHER WAY yesterday, both noise-dominated. But the second blocker is new and matters more, because it would stop the rescale even if the factors had proved: the grade does not separate within any archetype. Every archetype collapses to ONE band at the corrected bar, because bands merge when their intervals overlap and publishing two letters we cannot tell apart is a distinction we have not measured. Uncorrected, so the ranking is visible rather than hidden by the bar, this INVERTS the order's design. The order gives contact types the factor-rich treatment and power types honest base-rate, reasoning that single-game hits are variance for a power profile. Measured: BOMBER n=466 corr(p_win,outcome) +0.207 quintiles 0.75 0.62 0.60 0.48 0.48 GHOST n=192 corr(p_win,outcome) -0.007 quintiles 0.47 0.63 0.74 0.58 0.45 BOMBER is the one archetype the model ranks, and it splits into a real A 0.660 / B 0.481 at 95%. GHOST is flat, and non-monotone -- its most confident reads hit 47% while its middle reads hit 74%. Shipping as specified would have given the factor-rich treatment to the archetype the model reads worst and left base-rate on the one it reads best. That is mechanically sensible in hindsight: a power hitter's hit tracks whether he can damage the arm, a contact hitter's depends on balls finding holes. BOMBER's split does not survive the cumulative correction at 106 tests. Exposing it by loosening the correction is the curve-to-make-A's the order forbids, so it stays one band. BUILT: gradeBands.js -- lift against the archetype's OWN base rate (the same 62% is lift for a 45% profile and a deficit for a 68% one), indistinguishable neighbours merged, thin bands PROVISIONAL not dropped, Wilson intervals widened by the cumulative correction. The two-bar rule is structural: proven-alone, calibrated-alone and neither all return base_rate with the reason stated, so with nothing proven no factor-informed band can be produced at all. reasoning() is built and tested but NOT wired to the card -- there is no per-archetype band being served, so attaching the copy now would ship product language for a rescale that does not exist. NOT BUILT: the specified power-type reason "the matchup edge is in total_bases". total_bases is recorded INCONCLUSIVE (+0.0038, CI [-0.068,+0.075]). Wiring it would assert an edge measured as indistinguishable from zero -- the exact fabricated-reason failure this module exists to prevent. BOMBER x hits is 29 rows short of the gate and is the archetype the model actually reads. That is the first slot to test, not GHOST. Counter and frozen clusters byte-identical. No letter was moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
6b17f79367 |
Per-archetype re-audit: no slot reaches 500, and the replication unit
decided everything The premise does not hold. prove-hit-factors.js has no date filter anywhere in it and pages the full table -- there was never a window to widen. Full clean history is 1,266 rows, not 2,715. platoon was not "proved" last session, it was explicitly held on 4.5%-median-contaminated season-to-date splits, and pitcher_contact_profile was demoted. The proven set going in was one factor, not three. STEP 1: no archetype slot reaches n>=500 on full history. Best is BOMBER at 408, and BOMBER is the most common archetype on the board. GHOST 173, BRUSH 64, DRIVER 43, CATALYST 16. These are confirmed genuinely short, not artifacts. STEP 2 is where the real finding is. park_hits initially PROVED at 619 rows across 45 games -- but those games only ever visited 14 distinct park values. A park effect is replicated across parks, and unmodelled park heterogeneity is confounded with the thing being estimated. Each factor is now clustered on the coarser of the game and the entity its treatment rides on. That flipped two verdicts and confirms Kev's causal-correctness thesis from a new direction: defense_by_direction has 442 hitter-team units of replication where crude team defense has 26. The correct atom is not just more accurate, it is the only one measurable at all. park_hits (14) and defense (26) can never be validated however long the ledger runs -- the same ceiling as park dimensions, reached independently. Also fixed a bar I got wrong last session: I transplanted the 500-row floor onto clusters, which refused a factor with 1,059 rows over 85 games while answering neither question. Two floors now -- rows>=500 for a stable estimate, clusters>=40 for a trustworthy interval. Not a lowered bar: park_hits and defense are still refused. PROVEN: defense_by_direction only, pooled, [-0.0054,-0.0012] at 99 tests. It stays POOLED-ONLY -- no per-archetype reasoning wired, nothing grandfathered. The card must not say "GHOST: defence matchup strong" because we have not earned that sentence. The predicted fingerprint did not appear either: BOMBER -0.0036 vs GHOST -0.0024, the opposite direction, both noise-dominated. Recorded so it is not claimed later. RESCALE: NOT READY. One proven factor worth -0.0031 Brier. Rescaling on that is relabelling. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
b818626870 |
BUILD-STATE: under-querying vs out-of-data session
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
7b85934dc3 |
Under-querying vs out of data: the answer depends on the unit
The platoon test's n=452 described how much of the JOIN survived, not how much data exists. There are 1,266 clean settled hits rows and zero quarantined ones. platoon_splits had been ingested from tonight's lineups only (315 players), so any hitter who settled a prop without appearing in an ingest-day lineup was silently absent from every test. Backfilled all 380 hitters (81 fetched, 0 unresolved). Re-ran on 1,059 rows, up from 452. THE DEMOTION IS THE HEADLINE. pitcher_contact_profile, the strongest proven factor in the programme (-0.0064, CI [-0.0113,-0.0014]), roughly halved to -0.0034 on more than double the sample and its corrected interval now spans zero. The Bonferroni denominator also rose to 55, which widens every interval -- but a denominator cannot move a point estimate, and that halved on its own. platoon and platoon_severity now clear the bar and are NOT promoted. Upper bound -0.0001, on season-to-date splits that contain the games they predict: measured contamination is 4.5% median, 12.4% at p90, 137% worst. I had assumed ~1%. They stay CANDIDATE pending point-in-time splits. GAME-LEVEL IS A DIFFERENT PROBLEM. game_context held zero weather rows ever -- not because the fetcher was wrong (it correctly targets Open-Meteo's archive) but because ledger_entries keys a game as mlb:2026-08-03:Away@Home and game_context keys it as mlb:823437. Every lookup missed and NULL columns read as honest absence. Third occurrence of that class. Fixed the join: 96/101 settled games now carry actual archived weather, park dimensions backfilled 15 -> 30 venues. But 928 total_bases rows sit on 47 games at 17.6 rows per game. Park and weather assign one value per game, so resampling rows would have manufactured a pass. factorGate now resamples clusters when rows carry one and judges sample against effective_n; unclustered rows keep the original path byte-for-byte. Verdict: 47 clusters < 500, and the point estimate is +0.0011 -- worse, not merely unproven. Weather needs ~57 more days. Park dimensions need never: there are 30 ballparks in MLB, so a venue-constant factor can never reach 500 independent units. That bar was built for player-level factors and does not transfer. Wind is refused. We have speed and bearing for all 96 games; we lack park orientation, and 220 degrees is blowing out at one park and in at another. Using speed alone would assert an effect while discarding the sign that decides what it is. Counter and frozen clusters untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
6452926732 |
Retain raw weather, and record platoon severity 48 rows short
Two fixes in the weather path, and the second was hiding behind the first. The scalar weather_mod cannot express a hit-TYPE conversion at all -- wind out and warm turning fly balls into extra bases, and cold heavy air turning them into outs, collapse to the same number once multiplied -- so the raw temperature, wind speed and wind direction are now retained alongside it. And the old guard only kept the environment when the multiplier was not 1, which silently discarded the forecast for every ordinary night. That is the majority of games, and precisely the rows a hit-type model would need in order to learn what ordinary looks like. Platoon severity is built and measured at n=452, which is 48 rows short of the gate: CANDIDATE_PENDING, neither proven nor theatre. It moves less than flat platoon (0.021 against 0.026), consistent with the pattern, and its Brier point estimate is favourable but the corrected interval still spans zero. Worth naming: the refusal costs sample, and that is the design working. Flat platoon scores 741 rows because it will happily apply a boost to anyone; severity scores 452 because the other 289 are hitters whose split we cannot actually read at 60 plate appearances on the short side. Buying those rows back by shrinking instead of refusing would have produced a number indistinguishable from a measured league-average split, which is a different claim from the one the data supports. Park dimensions are ingested and verified in production across fifteen venues, joined by the venue the game is actually at rather than inferred from the home team -- neutral-site and international games break that assumption without surfacing an error. The park-and-weather-to-hit-type atom is NOT built. Its inputs landed this session and carry a single as_of date, so testing it on total_bases would be scoring games with inputs that postdate them. Building it now would produce something plausible rather than something proven. Proven factors for hits remain pitcher_contact_profile and defense_by_direction. 4,307 tests green (344 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
de0077f6f9 |
Causally-correct platoon + park-dimensions ingest
Applying the method that worked for defence to the two factors the code flagged as still crude. PLATOON. The flat version is 'lefty versus righty, add a boost', and it failed the two-part gate for the same reason team-average defence did: it is not the unit the causal story runs through. The advantage is only worth what THIS hitter's split is actually worth -- measured on a real hitter, .284 against left-handed pitching versus .221 against right-handed, a 63-point split, where the flat factor applied the same six percent to him and to a hitter with none. Most of the work is sample discipline, and the second rule matters more than the first. Severity shrinks toward the league split weighted by the SMALLER side's plate appearances, because a 500-against-40 split is a 40-PA read. And below a floor it REFUSES outright rather than shrinking, because a heavily-shrunk severity is indistinguishable from a measured league-average one and those are different claims -- without the refusal the atom would quietly assert a league-typical split about every September call-up in the league. Switch hitters turn out to be the easy case misread as the hard one. He bats opposite by choice so the direction is never in doubt, but the per-side value of his swing is a different question and one this sample cannot answer, so he is unreadable rather than credited with an automatic edge. PARK DIMENSIONS. Free from statsapi's venue endpoint, which carries fence distances, roof, turf and elevation outright -- Wrigley returns 355 down the left line, 400 to centre, 353 to right, at 595 feet. parkFactors holds run COEFFICIENTS, which structurally cannot express a park that turns outs into hits without scoring, and that is why the crude park factor failed. The park join is by the venue the game is ACTUALLY at, carried from the schedule feed, never inferred from the home team -- neutral-site and international games break that assumption and they break it silently. A venue with no geometry at all is absent rather than a park with zero dimensions. Both tables dated in the primary key. Venue geometry changes rarely but it does change, and by now that is the default rather than a lesson. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |