5cad851922f31ff844ef009e80f4fa90b5129686
78 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5cad851922 |
A digest names the curve; it does not vouch for the procedure that makes the next one
Artifact identity was the previous tranche's answer. It is not certification. `knot_digest` says WHICH mapping ran and nothing about whether tomorrow's refit deserves the same trust — that one would get its own digest and be equally "identified" while fitted on anything at all. So the object being certified is named: VYNDR certifies a PROCEDURE, not a frozen curve. A frozen curve goes stale against a live model and has to be replaced by hand on no schedule; "fit past, apply forward" is procedural by construction. `fitPolicy.POLICY_V1` declares it — data selection, horizon, algorithm and version, minimum rows, model-version restriction, sport, stat, and a support contract that a refit may NOT widen. Each artifact still carries its own digest. `fitPolicy.validate` is the Step-22 gate: an artifact does not become servable because the algorithm ran. It refuses a widened support, a wrong era, a wrong estimator, a thin fit, a missing identity or a missing training cutoff, and `servable` is false whenever any violation stands, with no override argument. The statistical bars stay where they already live in calibrationRegistry — this is not a second governance system. MEASURED, AND THE REASON LIVE SERVING STAYS BLOCKED: production does not match the declaration. `loadSettledRows` applies no model_version filter, so at fit_as_of 2026-09-02 the fit drew 6,084 rows from 9,361 settled — all 3,292 from the superseded engine1@2026-07-20 plus 2,792 current-era. 54.1% of the served map is fitted on a forecaster it was never certified for, while the artifact declares the current era. That is a provenance contradiction, not a performance claim: era-filtered scores 0.24065 against pooled 0.24068 on 1,120 out-of-sample rows and both intervals span zero. It is blocked because nothing prevents the next era change from repeating it, and because the freshness lag grows. `era_restricted` is answered STRUCTURALLY, not by an extra read — the query is in this service and applies no filter, so the artifact records ERA_NOT_RESTRICTED rather than claiming a restriction that did not hold. An unverified restriction is recorded as a violation, because "we did not check" is exactly the state production is in. The violation does NOT distort the shadow. Flipping every row to UNCERTIFIED would make the shadow measure the violation instead of the contract, so the policy state rides beside the resolution and a test asserts the shadow still reads CERTIFIED_CALIBRATED at 0.65 and UNCERTIFIED at 0.91. Nothing serves. CALIBRATION_DEPLOYED still []. Shadow still OFF (the production variable remains absent — the probe reads configuration_source "default"). Suite 402/402, 5,607 passed, 4 skipped. Teeth 10/10 + 23/23. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
556d186ff1 |
One evaluator for the shadow, so the probe cannot report a mode the pipeline is not in
"Is the shadow effective?" was only answerable by waiting for a snapshot to write a row. That leaves a blind spot with real cost: a variable SET IN COOLIFY BUT NOT YET APPLIED to the running process is indistinguishable from an unset one, and the runtime probe already proves the distinction matters — code_sha |
||
|
|
22cf51c4b0 |
The mapping that ran now has a name, and the row carries it
Runtime probes say the fleet is on
|
||
|
|
a8de676756 |
A probability is served because evidence supports it, not because nothing else answered
The band gate was blocked for its `else` branch. It read:
candidate = F(raw)
served = inCertifiedBand(candidate) ? candidate : RAW
and above raw 0.60 the model is measured overconfident — holdout raw 0.80-0.90
predicts 0.843 and realizes 0.639. So "the calibrator is not supported here" was
being answered with a number already proven wrong. Unsupported calibration does
not make raw true.
Four candidates were adjudicated on ONE split — fit on the earliest 60% of
train, decide support on the last 40%, evaluate on a holdout that saw neither:
A low-param 80.2% coverage 0.24374 REFUTED — its extra region
(raw 0.80-0.90) certified on cert (err +0.040, n=55) and
refuted on holdout (served 0.754 vs observed 0.639), and it
leaves a hole at 0.70-0.80 while serving the island above it
B isotonic 91.3% coverage 0.24337 CERTIFIED, contiguous raw [0.50,0.80)
C empirical band 91.3% coverage 0.24335 REFUTED — refitted point-in-time on
current-model hits the realized rates INVERT in grade order
(B+ 0.593 < B 0.614 < C+ 0.623), so the served function steps
down at raw 0.78. Its shipped constants come from 3,417 props
pooled across four batter stats and do not reproduce here
D raw identity 43.1% coverage 0.24866 certifies raw 0.50-0.60 and only there
Raw is candidate D, not a fallback. It earns exactly one region (holdout error
+0.010 on n=1,316), which is why the law is "raw must earn its region" rather
than "raw is never true". B already covers that region, so no hybrid is built.
Above raw 0.80 nothing is certified and nothing is served. That is the region
where raw is most wrong, isotonic over-corrects (cert err -0.093) and its LODO
mapping at 0.95 has spread 0.180. 8.7% of holdout rows land there.
The registry did not need changing. `serves(stat, p)` already tested certified
bands against the RAW p_win — support in the input domain, the correct question —
and returned {serve:false, reason}. It never said "serve raw". The output-space
gate and the raw fallback were both invented downstream in calibrationService.
ACTIVATION IS OFF. PROBABILITY_CONTRACT_SHADOW defaults to 0, CALIBRATION_DEPLOYED
stays frozen empty, and every served field is byte-identical. This releases the
support first, which is the required order. The shadow records raw belief, the
candidate served value, the state, the estimator identity, and what EV/Kelly/VALUE
would be under the actionability law — into its own column, read by nothing.
Migration 051 was applied to production BEFORE retentionService named the column.
PostgREST builds a bulk insert from the first row's shape, so a key whose column
does not exist 400s the whole batch silently — that is how migration 038 took
retention down for three days.
The user-facing contradiction is NOT fixed here. A B+ still says "realized about
66%" beside a confidence of 84. Fixing that is activation, and activation costs
32% of VALUE flags and 46% of Kelly recommendations on the holdout.
Suite 401/401, 5,580 passed, 4 skipped, deterministic across three runs.
Teeth 23/23, each independently injected and restored byte-identically.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
9dda9df132 |
Read History: the graph becomes something a person can read
The ancestry contract existed and nothing rendered it. This turns it into a product surface, and stops there. WHERE IT GOES. `LedgerCard` is the terminal surface — there is no Ledger detail view — so the history expands in place inside the card, matching the board's existing "ALL N READS" affordance rather than adding a page, a modal or a navigation category. It loads on first open, not on render. WHAT IT IS CALLED. "Read History". Lineage stays engineering vocabulary; a test asserts no rendered string contains lineage, natural key, ordinal, digest, graph, origin or recapture. ORIGIN reads "First published", REVISION reads "Updated", and the persisted `change_type` supplies "The price moved" / "The read changed" / "The read and the price changed". `change_type` is stored and trustworthy, so naming it is reporting; no field-level diff is persisted, so none is invented. THE SEPARATION, WHICH IS THE LOAD-BEARING PART. The grade strike means the LETTER changed. A history entry means the published CLAIM changed — often the price, sometimes the read, frequently with no letter change at all. The history uses no strike-through, shares no styling, and a tooth fails if it ever does. The ledger card's own strike is untouched. RECAPTURES ARE SUMMARISED, NEVER DESTROYED. A republishing board can produce hundreds of "unchanged" entries that bury the two that matter, so the UI collapses them to a count. The API still returns every one. WHAT EACH ENTRY SHOWS came from the acceptance run: without the published grade the history can say a Read changed but never what it changed to, which answers none of the questions someone opens a history to ask. `published_grade` / `published_p_win` / `published_line` are read off the SAME retained row the lineage action sits on — the authoritative record of the published claim, with lineage only the pointer to it. A tooth fails if they are ever synthesised from lineage metadata. TWO DEFECTS THE PRODUCTION ACCEPTANCE FOUND, NEITHER OF WHICH A TEST HAD. `ledger_entries.id` is a UUID and the route parsed it with Number.parseInt, so every real row would have 400'd. The unauthenticated probe that "proved the route was live" returns 401 from requireAuth before the handler runs, so it could never have seen this. And `chase burns / hits_allowed / under / 4.5` has SIX published captures and four lineage actions — two were published while the writer was off. The response said `chronology_complete: true` while showing four of six. Completeness now counts published-but-never-recorded states as well as attempted-and-failed ones, kept as separate numbers because the causes differ and the copy says which. Suite 399/5,544/0 · tsc 0 · web build 0 · teeth 15/15. Lint is not runnable in this repository (`next lint` removed in Next 16, no eslint.config.*) — pre-existing, untouched here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
e5b20b0509 |
Software watches coverage now, and ancestry becomes a product contract
Three closeout items, traced before building.
THE OBSERVER WAS DIAGNOSTICS, NOT MONITORING. Traced from the deployed tree:
exactly two callsites, both manual internal routes, and nothing in the scheduler
or ops path consumed it. A persistent lineage failure could have sat unnoticed
until a human asked.
The monitor now runs on the scheduler's per-minute tick, throttled to 30
minutes. It is placed there rather than after a snapshot on purpose: an
in-snapshot audit structurally cannot report that no snapshot ran, which is the
failure mode that matters most, and it would run under peak write contention
where the audit already demonstrably times out. It reads only the durable
observer and never `attachLineage`'s counters, and every failure path is
swallowed — a monitor that can take down the pipeline it watches is worse than
no monitor.
HEALTHY IS SILENCE; EVERYTHING ELSE SPEAKS. `coverageAlarm` is pure, so the
policy is testable and cannot drift into the scheduler. AUDIT_UNAVAILABLE says
"could not be measured — the audit did not run", deliberately worded so it can
never be read as "coverage is zero": those are different claims and collapsing
them is how a monitor starts lying in the reassuring direction. Alerts dedupe on
(health, cohort) so a standing fault states itself once and a NEW cohort with
the same fault speaks again.
ANCESTRY BECOMES A PRODUCT CONTRACT. It was internal-only. `GET
/api/ancestry/ledger/:id` (requireAuth, rate-limited) plus the Next proxy that
makes it browser-reachable, keyed on the LEDGER ROW ID — a stable identifier the
ledger API already returns — rather than a raw natural key exposed because it
was convenient. Its own router, so `routes/ledger.js` stays free of lineage
entirely and the grade-badge guard keeps its teeth. Every response declares
`authority: LOCAL, authority_scope: ANCESTRY_ONLY`.
THREE TEETH CAME BACK GREEN AND ALL THREE WERE MY TESTS, NOT SAFE DEFECTS.
The badge guard was CASE-SENSITIVE, so `LINEAGE_ANCESTRY` and `readAncestry`
walked straight past it. `try/finally` is valid JavaScript, so removing the
monitor's catch produced no load error and nothing asserted the containment.
And the multi-date cohort check was a grep for `.gte('game_date'` that the
head query satisfied on its own. All three replaced with behavioural tests,
including a scheduler double whose fake client HONOURS its filters — a
pass-through would have made a narrowed cohort walk look correct.
Suite 398/5,522/0 · tsc 0 · web build 0 · teeth 14/14 and 15/15.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
930d526b01 |
Lineage productization: a live graph, an observer that doesn't trust it, and a surface that owns only what it knows
The authority review found there was nothing to switch: lineage is a
write-only graph with no living permanent writer and no product consumer.
Authority theater would have been a switch on a consumer that does not exist,
reading a store that is not being written. So: make the graph live, measure it
from outside, and expose the one category it alone owns.
WRITE-PATH FAILURE SEMANTICS, TRACED FIRST. persist(:857) -> the authoritative
cacheSet snapshot:latest(:1404) -> the lineage gate(:1426) -> ledger(:1473).
Lineage failure cannot fail base retention (earlier, separate upsert), cannot
fail publication (the product write precedes the gate), cannot fail the Ledger
(lineageIndex defaults null; the ledger has its own guard), and cannot create a
new partial row (blankLineage() first, atomicity sweep after the catch). The
isolation this tranche needed already existed; only a mode was missing.
THREE MODES, ONE EVALUATOR. `lineageWriteMode` distinguishes OFF /
CANARY_LEASED / PERSISTENT_SHADOW, and the write gate and the status surface
both read it, so they cannot disagree. The canary is CONSULTED, never
converted: its <=4h absolute expiry, its dynamic evaluation and its
fail-closed parse are untouched.
AN AMBIGUOUS CONFIGURATION FAILS CLOSED. If a sport is named by both persistent
mode and an active lease, the two instructions disagree about WHEN WRITING
STOPS — the lease says 22:45, persistent says never. The dangerous reading is
the quiet one: an operator sets a bounded lease believing writing will stop
while persistent keeps it going. We cannot know which they meant, so that sport
writes nothing until the configuration says one thing. The sport allowlist is
the canary's own, so persistent mode can never widen past it.
THE OBSERVER MAY NOT ASK THE WRITER HOW IT DID. settleLedger returned
{settled:0,pending:0} — byte-identical to a healthy "nothing to settle" — while
1,444 rows sat unprocessed, and the watchdog believed it. So `lineageCoverage`
reads durable retained state only, and THE DENOMINATOR MAY NOT CONSULT
lineage_action: eligibility is "the row was published AND a natural key is
derivable from its own identity columns", neither of which the lineage path
writes. If expectation were derived from whether lineage exists, coverage would
be 100% by construction and the metric would be decoration. Zero-expected and
zero-written are kept as different answers.
THE LEGACY BOUNDARY IS OBSERVED, NOT DECLARED. `publication_id` is stamped only
by commitPublication, so the row itself says whether lineage ran. Verified on
production: 5,353 rows carry it — 4,234 complete actions plus exactly the 1,119
historical partial rows — and zero actions exist without one. No epoch constant
is invented; a date would have been a guess about when the writer was on.
publication_id NULL -> LEGACY_UNVERIFIED. Stamped but incomplete ->
LINEAGE_UNAVAILABLE, which is the honest answer for the 1,119 and is never
quietly rewritten as legacy.
THE GRADE-SHIFT BADGE IS UNTOUCHED. revised_from_grade answers "did the letter
change"; lineage answers "which published claim superseded which". Different
questions, and a test now fails if either route learns the word lineage.
Suite 397/5,496/0 · tsc 0 · 15/15 teeth.
TWO OF MY OWN TESTS WERE VACUOUS AND A TOOTH FOUND IT. Tooth 2 came back green
because the isolation tests asserted the slate was published — true whether or
not the exception propagated — while never reaching the lineage gate at all:
the fake grader never fired `onGraded`, so the collector stayed empty and
persistedRows stayed null. Fixed by firing the hook and counting the commit.
A green teeth run means the test is missing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
4aca33deb6 |
One human, one semantic identity — MLB participant convergence
The collision autopsy left two unrepaired defects, running in OPPOSITE
directions, and `outbound_collision_count` can only ever see one of them.
UNDER-COLLAPSE. Dedupe keys on `mlb:<personId>` when the participant is
proven and on the RAW PROVIDER SPELLING when it is not. Mickey Gasper
(681508) is on Boston's 40-man and not on its active roster, so an
active-only index could not identify him and every book's spelling of him
survived dedupe as its own proposition — retention was the first layer to
notice, far too late, and could only discard the loser.
SPLIT. The mirror image, and invisible to the collision metric because it
makes MORE identities, not fewer: Leo Jiménez (677870) is published as both
"Leo Jiménez" and "Leonardo Jimenez", so one human became two semantic
players in one game. Measured across the 15 MLB cohort slices since the
canonical-participant repair, this is a recurring class, not one case:
cam/cameron smith (5 slices), mitch/mitchell bratt, zac/zachary thornton,
leo/leonardo jimenez.
THE REPAIR READS MLB'S OWN RECORD. `hydrate=person` on the roster call the
pipeline already makes returns firstName / useName / useLastName, so the
legitimate name forms for a human come from the league rather than from an
alias table. An alias table is a list of the mistakes we happened to notice.
`nickName` is DELIBERATELY EXCLUDED: over 821 people it produced 14
ambiguous keys, because MLB's nickname field carries bare surnames and
shared clubhouse names — `nameKey('Smitty Smith')` is one string for both
Burch Smith and Will Smith. The four forms kept produce ZERO ambiguity.
Canonical participant reach widens to the 40-man; TEAM EVIDENCE still reads
the ACTIVE roster alone, so event admission and the impossible-binding
refusal are unchanged. Identity still fails closed: a name matching more
than one person in the event resolves to nobody.
CONTINUITY, MEASURED BEFORE WRITING ANY CODE. Over the real 19:00 cohort,
208 of 209 player_keys are unchanged and the one that moves is the defect —
`leonardo jimenez` converging onto `leo jimenez`, a key that already exists.
No new lineage family. The natural key contains game_date, so chains never
span dates and a forward change cannot fork a closed one.
DETERMINISTIC REPRESENTATIVE. Which book's payload survives was decided by
position. It is now decided by the existing MODEL_BOOKS declaration order —
reused, not authored; inventing a sportsbook ranking to settle a tiebreak
would be a market judgement smuggled in as a bug fix — with book name and a
content tiebreak. Stable under every input permutation.
TWO GUARDS, BOTH DIRECTIONS. split (one person, many identities) and merge
(one identity, many people). A merge is refused at the same single admission
seam event identity already uses; a split is counted and alerted but does not
cut the board, because it duplicates an identity rather than asserting a
falsehood.
RETENTION REMAINS AN INDEPENDENT CHECK. The old assertion grepped the source
for `player_key: nameKey(player)`. That expression stood in for a PROPERTY,
and a grep verifies a spelling. Replaced with the property itself, asserted
in both modes: when the producer emits two rows for one human, retention
still files them under one identity and still reports the collision.
Replay of the real cohort through the repair: 3,129 offerings, 100%
participants resolved, every one of 207 participants on exactly ONE semantic
key, collision 0, split 0, merge 0.
Suite 396/5,455/0 · tsc 0 · 15/15 teeth. Tooth 12 came back green first
time and that was a coverage hole, not a safe defect: nothing asserted
retention's append-only upsert. It does now.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
9809626c99 |
Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.
So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.
TERMINAL RETENTION STATES (retentionService.classifyPersist):
NOTHING_TO_PERSIST attempted 0 — a refusal-only slate is still a cycle
SKIPPED_NO_DATABASE no database configured; not a failure
COMPLETE attempted > 0, written === attempted, no error
FAILED_ZERO_WRITE written === 0 — first chunk failed
FAILED_PARTIAL 0 < written < attempted — a later chunk failed
FAILED_UNRESOLVED_ERROR counts look complete but an error is unresolved;
unreachable through today's loop, and kept because
the alternative is reporting COMPLETE holding an error
The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.
classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to
|
||
|
|
35da190f2c |
Retention hotfix: drop published_side, derive the schema contract, break the silence
`createCollector.onPublished` set `published_side` beside `published`.
`published_side` is not a model_snapshots column. supabase-js declares the
UNION of row keys in the `columns=` parameter, so one invalid key made
PostgREST reject the ENTIRE batch with a 400 — every sport, every cycle.
Retention is best-effort, so nothing surfaced. Confirmed in edge logs.
The field was redundant as well as invalid: `side` is already on the row.
Deleted rather than added to the schema — a column would preserve an
accidental artifact.
Three things missed it, and each is now closed:
1. WRONG SHAPE INSPECTED. The manual check sampled the collector after
onGraded only and never called onPublished, so the offending key was
not yet on the row. It read a pre-publication shape and reported the
final outbound shape as clean. The new test captures the array actually
handed to .upsert(), after the full production call order.
2. NO CONTRACT. Every retention test injects a permissive fake client that
accepts any column set, so 381 suites proved the logic and never once
compared a row against the database. The contract is now DERIVED — the
migration chain applied to a disposable postgres, read out of
information_schema (scripts/generate-schema-contract.js). A
hand-maintained list would be a second opinion about the schema, and a
second opinion is what let this through. scripts/verify-schema-contract.js
checks the contract still describes a live database.
3. SILENT FAILURE. A failed batch reached one console.log. It now emits a
high-severity structured event carrying sport, snapshot id, stage,
error, code_sha and timestamp. Best-effort semantics are unchanged —
the product continues and says so — but the failure is observable.
`skipped` (no database configured) is not a failure and does not alert.
Teeth, each with the injection verified present before the run:
- published_side back into the final payload -> 4 tests fail; restored
byte-identically (sha 6a0ced7c52134135 both sides)
- settled_at (a REAL contract column) -> accepted, so the guard
discriminates by contract membership, not by novelty
- alert block deleted -> 3 tests fail; restored byte-identically
Model and product behaviour untouched: analyzeViaEngine1,
probabilityEstimator, gradeSlateService, lineageCanaryConfig all unchanged.
Lineage stays OFF. Net source change is one behavioural line plus the alert.
382 suites / 5,094 tests pass. web tsc exit 0 (zero web paths touched).
Measurement blackout recorded, NOT backfilled: last good retention write
2026-08-27T19:08:32Z;
|
||
|
|
8c6aef1e12 |
Event admission gate: an unresolved or contradicted game does not earn a Read
A fractured identity keeps two bad records from merging. It does not make an unknown game true. Until now a prop with an unresolved or verified-impossible event still continued into grading under that synthetic key; it no longer does. - event_binding_status contract: RESOLVED / UNRESOLVED / AMBIGUOUS / CONTRADICTED / UNSUPPORTED. CONTRADICTED and UNRESOLVED stay distinct — one means we know the association is wrong, the other that we do not know. - ONE admission gate (gradeSlateService.admitForGrading), before dedupe and before grading. Admission requires RESOLVED *and* a canonical_event_id; a legacy derived game_id can never satisfy it. Rejected props are returned, not discarded, so retention keeps them as evidence. - Roster evidence is now DATE-SCOPED. statsapi honours ?date= and it changes the answer (Joe Mack is on the 2026-08-26 Marlins roster, absent on 2026-04-15). Evidence that does not describe the slate's date can only yield UNRESOLVED, never CONTRADICTED — uncertainty must not become an accusation. - A mis-nested market is REFUSED, never re-bound. Knowing Joe Mack is a Marlin does not license moving a provider record into the Marlins game; that would invent provenance. Measured on the real 2026-08-26 slate: 2,878 props -> 2,874 admitted, 4 rejected (EVENT_PLAYER_TEAM_CONTRADICTION), each player's correct game still resolving. Doubleheader 824514/824478 both remain independently RESOLVED and admitted. No change to probability, projection, side, grade, confidence, ranking, normalization or calibration. Lineage remains disabled. Suite 380/5,061/0 from the release worktree; web tsc exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
352016790a |
MLB canonical event identity, impossible-binding refusal, event-aware dedupe, publication commit
Release-isolated slice built from
|
||
|
|
6c34af3414 |
checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion recon (Tony resurrection, STEP 0). This work existed only on the laptop disk. - chain shadow accrual + probe script (038_chain_shadow.sql) - WNBA possession feed: ESPN adapter, usage service, verify script (039_wnba_player_game.sql) - baseball chain - retention/snapshot service updates, tableKeys, matchupKeys - specs: chain-v1, wnba-possession-feed, wnba-source-survey - unit tests for the above Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn |
||
|
|
f61ec6b391 |
Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move. A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS: 410-617 of 2,490 duplicated with an equal number never returned, while rows.length matched the server exactly. safePaginate orders on a real unique key, verifies the tuple at runtime, and THROWS on a query error instead of treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn through that reader, and defense_by_direction's distinct-n was likely below the gate floor all along. A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from pg_index (the context tables are dated-composite and had no single unique column). The unordered helper is deleted, not parked. A3 — ledgerService and retentionService defaulted the SAME env var to DIFFERENT versions, so no ledger row ever carried the marker eligibility requires. One source now. model_snapshots settlement moved onto the cron: 15,484 -> 28,894 settled, repaired-champion 0 -> 7,556. A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no row at-or-before the date means the factor does not apply, never the nearest row. Live path unchanged, proven 400/400 on real rows. A5 — factor_inputs freezes what the factor READ, never the multiplier, so an audit can recompute and check. It also recorded the finding: the three hits factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by the resolver and written by nothing. A6/A7 — matchupKeys resolves those keys from the posted lineup plus the schedule's probable pitchers, and fires the factors into a SHADOW freeze: 248 fires on 308 props, 245 of which would move the grade. The served forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test that decides whether they ever go live. Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity harness 34/34. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74aa75945e |
Content engine: posts that structurally cannot lie
PHASE 0 — contentEngine makes Truth Law structural, not careful. Copy is
token-substituted and an unbacked {token} REFUSES to render -- there is no
code path that produces a plausible default. The fact contract is asserted
before any string is built. Card and copy render from ONE fact object, so
a caption and a card cannot disagree. No live model writes factual claims:
the voice is in the template, the facts are pulled, and the voice-polish
port is deliberately unwired, because an LLM that can rewrite a sentence
can rewrite a number.
18 tests carry the proof. The one that matters most: ZERO IS PRESENT.
"0 cleared B+" is our most honest possible post, and treating 0 as missing
would be the Number(null)===0 breach wearing its opposite coat -- it would
silently delete exactly the post the brand is built on.
PHASE 1 — three templates, generating real posts from tonight's data:
hot hitters off the repaired full-season log, the honesty flex off the
real servedGrade distribution (2,140 graded / 70 cleared B+ / 42% not
separable / A unissuable), and streaks verified from settled outcomes only.
THE ENGINE CAUGHT A BUG IN ITSELF, and it is the sharpest lesson here. The
first run published "No hitter is meaningfully hot tonight -- we could
dress up a middling week as a streak. We don't." That was FALSE: the
box-score cache spans only the settled window, every player had under 20
games, and the pool was empty. A broken pull was publishing as considered
editorial judgement -- the fourth appearance of this class tonight and the
first where our OWN HONESTY COPY was the disguise.
Fixed structurally rather than by patching the number: an absent() variant
may now DECLINE to speak, and the template separates "no candidates at
all" (SKIP with a reason) from "candidates judged, none hot" (honest
absence). Both locked by test. Source corrected to mlbStatsAdapter.fullLog,
the same log the repaired champion reads.
PHASE 2 — cardRenderer emits SVG rather than canvas: it is text, so it
diffs in review and its numbers are greppable, which matters when the
whole claim is that the numbers are real. VYND white + R green, slashed-Y,
scanlines, mono. The card never formats its own facts -- every string
arrives pre-rendered and gate-checked.
PHASE 3 — scripts/generate-content.js writes copy + card per template to
.content-out/<date>/. Template N+1 is a registry entry: requires, pull,
copy, card, absent. Queued as stubs, not built: hot takes, daily reads,
"grades we DIDN'T give", cross-sport streak variants (the streak template
is already sport-agnostic -- settled outcomes and a noun).
FULLY ISOLATED: read-only on every source, zero writes to serving, model
or ledger tables. Serving fingerprint verified unchanged. The accrual clock
is untouched at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
929fd81940 |
Repair the champion: it was reading ten games, not a season
PHASE 0 — the defect is real past the peek. Against a FAIR point-in-time baseline (each player's rate over games strictly before that date, >=10 prior games, box scores back to 05-01), the served champion LOSES on all four stats, three of four CIs excluding zero: hits 0.00251 vs 0.00774 CI [-0.0074,-0.0011] TB 0.00393 vs 0.00619 CI [-0.0055,-0.0003] rbi 0.02481 vs 0.03133 CI [-0.0153,-0.0005] runs 0.00181 vs 0.00683 CI [-0.0114,+0.0008] PHASE 1 — the cause is the WINDOW, not the weights. estimateProbability builds its base rate as the frequency over every row it is handed, and featureCache.getStatRows handed it res.last10. So the "season rate" was a TEN-GAME rate, and 0.4 of the forecast was the last five OF THOSE TEN. The 0.40 recency weight costs resolution on all four stats (-0.00086, -0.00107, -0.00562, -0.00365). Nudges are mixed and small -- harmful on hits and rbi, marginally helpful on TB and runs -- so they are left alone. PHASE 2 — two lines, no new data, no extra API call, because fullLog was already fetched by the same adapter call that produced last10: getStatRows now reads fullLog, and RECENCY_WEIGHT goes 0.40 -> 0.20. hits 0.00251 -> 0.00817 (tripled; now above the fair baseline) TB 0.00393 -> 0.00734 (above baseline; vs old CI [0.0020,0.0067]) rbi 0.02481 -> 0.02727 (still below baseline, CI includes zero) runs 0.00181 -> 0.00436 (still below baseline, CI includes zero) Gate stated exactly: hits and TB now exceed the fair baseline on the point estimate; rbi and runs remain below but EVERY CI now includes zero, so no stat reliably loses to a frequency table. That is a tie on rbi/runs, not a win, and it is reported as one. Only TB's improvement over the old champion is CI-confirmed; the rest are directional. STALE-FIT GATE: CALIBRATION_DEPLOYED is now EMPTY. The low-param maps were fitted on the retired forecast and fromLedger cannot rescue them -- settled ledger rows still carry OLD p_win, so refitting today would refit the retired forecast. Nothing is served calibrated until dates settle under the repaired champion, and the favourite-longshot bias must be re-measured rather than assumed to survive. The shadow duel is void. PHASE 3 — the hits factor lift is NOT re-measured, and cannot be yet: it needs settled rows produced BY the repaired champion, which ships in this commit. Replaying would score the factors against a reconstruction rather than the served forecast. Deferred, explicitly. The factors remain wired and transmitting; only their lift is unquantified on the new baseline. PHASE 4 — standing flag, and it is large: EVERY factor verdict in this programme, every null and every THEATER, was measured against a champion worse than a frequency table. Signal added to noise reads as noise. Prior verdicts may deserve re-audit. Logged, not re-run. Re-queued not built: rbi lineup-slot / RISP opportunity through the two-part gate, now landing on a repaired champion. Serving-path change by design; the byte-identical invariant inverted and all four stats move. Nine frozen model modules verified unchanged. No Bonferroni slot -- resolution accounting on the champion's own knobs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
65ca6493db |
Decompose the rbi anomaly: it is lineup ROLE, and the counter is
out-resolved by a frequency table on three of four stats
PHASE 0 — the 14.51% is REAL. Re-derived with a paged pull asserted
against an exact count (rbi 7,930 == 7,930; hits 11,690; TB 12,086; runs
6,440), since this harness produced a false null three times tonight. rbi
resolution 0.03268 reproduces, deciles are monotone through the middle,
and 20 raw rows are in the artifact for hand audit.
CAVEAT GOVERNING EVERYTHING BELOW: the naive forecasts are leave-one-out
ON THE EVALUATION WINDOW, so they see the rows they are scored on while
the model is strictly point-in-time. They are upper bounds on available
resolution, not fair competitors, and every comparison is read that way.
PHASE 1 — the split:
stat MODEL (a)player-base (b)lineup-slot (c)within-stratum
rbi 0.03268 0.01167 0.03608 0.01908
hits 0.00252 0.00446 0.00100 0.00473
TB 0.00442 0.01331 0.03448 0.00607
runs 0.00130 0.00262 0.01156 0.01170
FINDING 1 — rbi's resolution is LINEUP ROLE almost exactly. Batting-order
slot alone resolves 0.03608 against the model's 0.03268. A single integer
accounts for the whole anomaly and slightly more. That is opportunity, not
skill -- the cleanup hitter bats with runners on. 36% is matched by player
identity alone. Within similar-base-rate strata the model still resolves
0.01908, 58% of its total and higher than any other stat's ENTIRE model
resolution, so genuine within-role discrimination exists on top.
FINDING 2 — on three of four stats the model is beaten by "he's a .270
hitter". Base-rate-only out-resolves the model 1.8x on hits, 3.0x on TB,
2.0x on runs. Even allowing for the window-peeking advantage, a 1.8-3.0x
gap is not explained by that alone: the served counter appears to DESTROY
discrimination relative to the player's own rate. rbi is the one stat
where the model beats the naive baseline.
FINDING 3 — lineup slot out-resolves the MODEL on three stats: TB 7.8x,
runs 8.9x, rbi 1.1x. Hits is the only stat where batting order carries
less, which is mechanically right -- a hit is a hit wherever you bat, but
runs, RBI and total bases all scale with opportunity.
PHASE 2 — all three worlds are partly true, in measured proportions.
World A ~90% true (slot covers rbi's entire resolution). World B ~36% true
for rbi, but the WHOLE story for hits/TB/runs where base rate alone wins.
World C true with a low ceiling: hits' total available spread resolution
is 0.00446, i.e. 1.8% of variance from a forecast that has seen the
answers.
PHASE 3 — the next arc is NOT "strengthen hits factors". Hits has the
lowest available resolution on the board and last order's wiring already
took it to 1.39% of a ~1.8% ceiling. Named first factor order for next
session: LINEUP SLOT / RISP OPPORTUNITY on rbi through the two-part gate --
input already ingested and prod-verified (S89), resolution measured not
hypothesised, causally-correct unit is plate appearances with runners on.
Measured availability is not a pass; it still faces the gate.
And higher-value than either: the counter being out-resolved by a
frequency table on three of four stats is a defect in the CHAMPION, not a
factor problem, and it costs nothing to test -- the recency blend and the
+/-0.03 / +/-0.015 nudges are three lines in probabilityEstimator.
The hits transmission win from
|
||
|
|
43f65d30cb |
Wire the three proven hits factors pre-grade: transmission proven, gain
inconclusive THE BUG THIS NEARLY SHIPPED AS A FINDING. The first audit reported 0 factors fired on all 1,140 rows. Not a result -- my paging helper ordered by `id`, and batter_spray, team_defense, platoon_splits and statcast_aggregates have composite primary keys with NO id column. The query errored, the loop broke on error, and four fully-populated tables read as empty. hitsFactorContext.js -- the PRODUCTION loader -- had the identical defect, so live wiring would have loaded nothing and served unadjusted while logging success. Third occurrence of this class in one session. Both loaders now order by a real column and THROW rather than degrade. The Phase 2 gate is what caught it: no resolution number was quoted until transmission was proved. PHASE 1 — pipeline is now base -> FACTORS -> CALIBRATE -> GRADE. Context built in snapshotService BEFORE gradeAndCacheSlate (was line 640+, grade at 454), threaded per prop, applied to p_over before p_win is set with p_win_prefactor and a full trace retained. Hits only. Coverage 859/1140 rows (75%): 474 with all three factors, 256 two, 129 one, 281 none. PHASE 2 — TRANSMISSION PROVEN, 12/12 sign-correct, 4/4 per factor, each applied IN ISOLATION. My first table compared each factor's expected sign against the COMPOSITE change and showed 3 false failures -- with three factors firing the net can oppose any single member; that was a flaw in the test, not the wiring. Two under-side rows confirm the flip is handled: a factor raising p(over) correctly lowers p_win. Switch hitters (Bailey, Bell, Rocchio) took no spray adjustment while their other factors fired normally -- the refusal is selective, not a blanket skip. PHASE 3/4 — both maps refit on the factor-adjusted forecast; the shadow-duel baseline is VOID and restarts, since it accumulated against a different forecast. Point-in-time, 765 held-out rows: reliability 0.00795 -> 0.00828 RESOLUTION 0.00229 -> 0.00345 (variance explained 0.93% -> 1.39%) Brier 0.25398 -> 0.25305 delta -0.00093 CI [-0.00225,+0.00002] Resolution rose 51% relative. The CI TOUCHES ZERO on 4 eval dates, so the composition does NOT earn a proven keep -- three isolated passes did not grant a composed pass. INCONCLUSIVE, reported as such. The gain is far below the sum of the isolated effects, which is expected: all three run through the same pitcher-batter confrontation and share signal. PHASE 5 — 1.39% of variance is still far below what band separation needs. The pivot was correct and incomplete: the plumbing defect was real and is fixed, three proven factors reach the served number for the first time, and transmission alone did not buy grade separation. Next arc is factor STRENGTH and BREADTH, not more plumbing. PHASE 6 — rbi anomaly logged, not chased: 14.51% variance explained vs hits 1.03%, on the stat we do not serve corrected and which has no proven factors. Either the biggest lever on the board or a mirage; it deserves its own order. The byte-identical invariant INVERTED for hits by design. All 13 frozen non-hits modules verified unchanged, probabilityEstimator included -- the factors ride outside it. No new Bonferroni slot; the composed OOS claim is reported with its CI and not claimed as a pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
e872eff4ce |
Instrument the calibration duel forward; diagnose the resolution ceiling
— the proven factors were never wired in
PHASE 0 — two truths recorded. The swap is a BET, not an OOS win:
isotonic beat low-param on identical held-out rows (hits +0.0028, rbi
+0.0042, TB tied) and we serve low-param anyway on an untestable prior
about shared daily structure. At 19 dates nothing here can test it. And
the MIN_SLOPE catch is preserved as standing rationale: a near-zero or
negative slope collapses toward base-rate-for-everything, which LOWERS
Brier while destroying all resolution -- a metric win that guts the
product.
PHASE 1 — the duel is now falsifiable. Both corrections computed on every
hits/TB prop; p_win_lowparam served, p_win_isotonic_shadow logged in its
own try so it can never break serving. calibrationDuel.adjudicate encodes
the rule IN CODE before any forward date exists: >=10 forward dates and
isotonic winning with a date-block CI excluding zero => REFUTED, revert;
otherwise UPHELD; under 10 dates PENDING regardless of the numbers. A
date counts as forward only if NEITHER map was fitted on it -- otherwise
we would be scoring which map memorised better. Nothing swaps now.
PHASE 2 — the ceiling, quantified via Murphy decomposition:
stat reliability RESOLUTION uncertainty variance explained
hits 0.01353 0.00252 0.24532 1.03%
TB 0.01419 0.00442 0.24329 1.82%
rbi 0.00654 0.03268 0.22531 14.51%
runs 0.00788 0.00130 0.23182 0.56%
Calibration did exactly what theory says and nothing more: hits
reliability 0.01353 -> 0.00233 (-0.0112, 83% of the error removed) while
resolution moved -0.0002. Unexpected: rbi has 13x the resolution of hits
and is the one stat we do NOT serve corrected -- it needs calibration
least and discriminates most.
PHASE 2 DIAGNOSIS — NOT-TRANSMITTED, and not weak, ABSENT. Traced in code:
sprayDefense.js and platoonSeverity.js are required by NOTHING in src/,
only by analysis scripts and their own tests. The served p_win
(intelligence/probabilityEstimator.js:54) reads exactly four inputs --
game-log frequency, opp_rank_stat +/-0.03, home_away +/-0.015, and a cv
pull -- with zero occurrences of spray, platoon, hard-hit or
contact-profile. And snapshotService grades at line 454 while computing
challenger/context at 640+, so everything proven is computed DOWNSTREAM of
the grade it would inform. The three proven hits factors have never once
moved a served number.
That reframes the recent nulls: "calibrated p_win does not separate within
archetype" was never a statement about factors. The factors were not in
the forecast.
PHASE 3 — bands rebuilt on SERVED values (hits/TB low-param, rbi/runs
raw): 28 archetype slots across four stats, ZERO show lift. No longer an
open shrug -- it is the arithmetic of resolution 0.0013-0.0327 against
uncertainty ~0.23. A forecast explaining 1% of variance cannot produce
separating bands, and no correction to its numbers will change that.
HEADLINE: calibration is complete, delivered honest numbers on two stats
and zero grade separation, because the counter has no resolution -- and
the proven factors are not wired into the forecast at all. The second is
the reason for the first, and it is plumbing rather than a modelling wall.
Per-archetype grades need proven factors that actually reach p_win. Last
calibration order.
Serving unchanged from
|
||
|
|
74cf1ce974 |
Robust bias established; low-parameter correction replaces isotonic
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.
PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.
PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:
hits a=0.406 shrink 0.565 0.2626 -> 0.2540 CI [-0.0112,-0.0069] DEPLOY
total_bases a=0.472 shrink 0.333 0.2490 -> 0.2429 CI [-0.0062,-0.0059] DEPLOY
rbi a=0.775 shrink 0.231 0.2011 -> 0.2007 CI [-0.0007, 0] REFUSE
runs a=-0.032 REFUSE
A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.
STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.
PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at
|
||
|
|
ced40421ed |
Audit the LODO instrument: it cannot evaluate any stat, and both prior
FAILs were false
PHASE 0 — the gate at
|
||
|
|
1f40014256 |
Power-derive the LODO threshold: hits restored through the gate, rbi/runs
routed as date-driven PHASE 0 — threshold derived BLIND, before any stat was re-read. A reversal is informative only if that date's Brier delta is distinguishable from zero at its row count. Per-row Brier difference d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single stat's verdict could shape the threshold deciding it: pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175 n* = (0.09816/0.01175)^2 = 69.8 -> 70 The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect this removes, and why the previous verdict moved with the number. Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it equals the constant, so it cannot drift from its own justification. The derivation script prints no stat verdict, no date and no reversal. PHASE 1 — LODO at n*, applied cold: hits 5 informative drops, 0 reversals PASS total_bases 4 informative drops, 0 reversals PASS rbi reverses 2026-08-01 (n=99) FAIL runs reverses 08-01 (n=86), 08-05 (244) FAIL hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across 123-272 row dates, favourite sign holding on every testable drop. THIS IS THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the withdrawal at |
||
|
|
6ae11f1193 |
LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration layer without challenging the binding. That floor is a cluster-robust interval bar for a CAUSAL claim. Calibration makes no causal claim, has a bounded failure mode (it can only over- or under-shrink) and consumes no Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and leave-one-date-out tests that directly -- a STRICTER bar, since a cluster count cannot detect a single day carrying the effect. The >=40 floor is retained, correctly scoped as the PROMOTION bar. PHASE 1 — both guards codified, 11 tests, green before Phase 2. Demonstrated on live data: raw population violated=true, mean_p 0.4962, both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The null guard's test demonstrates the trap explicitly, since (null-1)**2 is 1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate. PHASE 2 — LODO: hits n=1140 dates=17 2 reversals (07-22 n=20, 07-26 n=25) FAIL total_bases n=1050 dates=7 0 reversals, 0 sign flips PASS rbi n= 630 dates=5 1 reversal (08-01 n=99) FAIL runs n= 597 dates=5 2 reversals (08-01 n=86, 08-05 n=244) FAIL Threshold sensitivity reported because the verdict moves: total_bases passes at every held-size threshold, runs fails at every one, and hits fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20 before seeing which stats passed and did not move it afterwards to preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a standard error several times the effect, so the instrument is underpowered per-drop -- an argument for pre-registering a higher threshold, which is a Roundtable call, not one to make while holding the results. PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and runs REFUSE. HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService hardcoded it since S91; it fails LODO, so it is out. A stat that cannot survive dropping one day was never calibrated, it was fitted to that day. The consequence is real -- hits props become unstackable for chain.chainAcross -- and it errs toward withdrawing a claim rather than preserving one on a fragile verdict. Deployment is now driven by a frozen, tested CALIBRATION_DEPLOYED set, not a hardcoded stat name. PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither waivable. reverify auto-demotes on the first breach (CI stops excluding zero, or the favourite bias flips sign) and logs the breaking date. Promotion needs the original >=40 bar. A provisional deploy that cannot be taken away is just a deploy. PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The two-bar rule still bites: calibrated YES, proven NO, so they stay a base-rate read, now honestly numbered. Every archetype still collapses to one band -- calibrated p_win separates within archetype no better than raw. PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI, and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did not survive); the refused set is a map of missing inputs; a low-parameter calibrator is queued unbuilt. p_win never mutated; calibration rides as p_win_calibrated with calibration_status provisional. No Bonferroni slot consumed. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
f976df47b8 |
Settle model_snapshots + four-stat calibration: works, deploys nowhere
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.
PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.
PHASE 1, two integrity findings both caught by the gate:
1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').
2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.
THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.
THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.
PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):
hits n=1140 bias +0.0868 brier 0.2626 -> 0.2511 d -0.0115 CI [-0.0139,-0.0097]
TB n=1050 bias +0.0834 brier 0.2490 -> 0.2438 d -0.0052 CI [-0.0061,-0.0045]
rbi n= 630 bias +0.0164 brier 0.2011 -> 0.1965 d -0.0046 CI [-0.0092,-0.0010]
runs n= 597 bias +0.0410 no map fittable (173 fit rows < 200)
ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.
A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.
PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.
PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.
PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.
p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
23d1b13176 |
runs + RBI: mostly-base-rate confirmed, and one level deeper than expected
Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest grade is the POOLED base rate. PREMISE NOTE: the order's closing line says the batter board is per-archetype-graded after this. Nothing has been rescaled for hits or total_bases either -- no archetype slot has ever reached sample and gradeBands remains built, gated and unwired. This is the fourth stat measured, not the completion of three. AUDIT: RBI 935 clean / 43 games; RUNS 617 clean / 33 games. Zero quarantined. No archetype slot reaches 500 -- and the signature archetypes the order names are the two SMALLEST slots on the board, RBI->DRIVER at n=24 and runs->CATALYST at n=9. RUNS is refused structurally before any factor is tested: 33 game clusters against a 40 floor. INPUTS RECONSTRUCTED rather than declared missing. lineup_context only covers 08-04 onward while settled rows start 07-31, so 187/617 runs rows joined. But the play-by-play cache runs from 05-01 and the batting order IS the order batters first appear -- slot, power-behind and reach-base all rebuilt point-in-time, coverage 187 -> 574. RBI, all THEATER: risp_opportunity +0.0047, extra_base_skill +0.0010, risp x extra_base +0.0056. RUNS, all refused on clusters and all pointing the wrong way: +0.0043 / +0.0008 / +0.0054. THE COMPOUND IS THE WORST VERSION IN BOTH STATS. The causally-correct compound was the most promising factor on the sheet and is the most harmful in each. Two multipliers that individually carry nothing do not cancel -- they compound each other's noise. Distinct from the collapsed-sequence lesson: there the product of two REAL effects was too small to use; here the product of two NULL effects is worse than either. THE ARCHETYPE DOES NOT RESCUE IT, and this is where the session nearly went wrong. The base rates look strongly differentiated (RBI DRIVER 0.609 vs BOMBER 0.413; runs GHOST 0.716 vs BOMBER 0.460). Gated directly against the pooled base rate: RBI +0.0010 CI [-0.0034,+0.0050] THEATER; runs -0.0028 CI [-0.0147,+0.0108] candidate at k=33. DRIVER's 0.609 is n=23 -- small-slot noise wearing a decimal point. Read off the table instead of gated, this would have shipped as "archetype differentiation is real and large". It is not. THE CROSS-STAT PATTERN THAT IS REAL -- the counter over-predicts every batter counting stat measured: total_bases p_win 0.5698 vs actual 0.5074 bias +0.0624 rbi p_win 0.4860 vs actual 0.4313 bias +0.0547 runs p_win 0.5949 vs actual 0.5749 bias +0.0200 Across four stats and three sessions, calibration is the systematic defect and factor scarcity is not. TB's held-out isotonic fix (-0.0039) still outperforms every factor tried on any stat, all null or theatre. NO RESCALE. Nothing proved, nothing certified calibrated, no slot at sample, and for RBI the archetype split is itself theatre -- so the honest band is the pooled base rate, which gradeBands returns by construction. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
6a327d9114 |
total_bases: every power factor is THEATER, and a units bug nearly hid it
PREMISE CORRECTION: the per-archetype rescale is not "proven and live on hits". gradeBands was built, gated and explicitly NOT wired two orders ago -- no hits archetype slot reached sample, every band came back base-rate, and only defense_by_direction proved pooled. This applies an unvalidated-at-archetype-level method to a second stat. FULL-HISTORY AUDIT: 988 clean settled TB rows (101 quarantined, 948 with p_win), 341 players, and only 9 DISTINCT GAME DATES. No archetype slot reaches 500 -- BOMBER 340, GHOST 147, BRUSH 55. Confirmed short on full history, not a windowed artifact. The 9-date figure matters more than the row count: ~49 games means any game- or venue-borne factor has almost no replication here. THE BASELINE HAD TO CHANGE, to a harder null. TB lines vary (1.5 on 559 rows, 0.5 on 345), so a per-line personal base rate would rest on ~2 rows per player-line and would have to be invented. The null is the counter's own p_win, which already prices the line -- beating the champion, not beating "he's due". THE UNITS BUG, caught, and it had produced the best result in the programme. The first run reported barrel_rate at Brier -0.0095, the largest improvement ever measured here. fromStatcastRow returns barrel_pct as a FRACTION (0.06) while the raw table stores 0-100, so (0.06 - 7.8) * 0.018 clamped EVERY row to the maximum negative shift. That uniform downward push "improved" Brier purely by leaning on the counter's over-prediction and contained no barrel information at all. Same family as the S80 trap, inverted. exit_velo was a second bug -- the column is avg_exit_velo, so it read null on every row and reported n=0. A zero is a wiring bug until proven an honest absence. GATE with units fixed, 138 cumulative tests: barrel_rate n=707 shift 0.0364 brier +0.0036 THEATER exit_velo n=707 shift 0.0229 brier +0.0022 THEATER hard_contact_allowed n=707 shift 0.0260 brier +0.0033 THEATER park_weather_hit_type n=651 36 entities PENDING (k<40) platoon_severity n=481 PENDING (n<500) THE PREDICTED INVERSION WENT THE OTHER WAY. BOMBER x barrel_rate is +0.0114, the single most harmful cell in the table, exactly where the strongest proof was predicted. GHOST +0.0012. All sample-blocked so not a verdict, but recorded so it is not claimed later. AND IT IS NOT DOUBLE-COUNTING -- tested and refuted: corr(barrel, p_win) = -0.061, the counter is not pricing barrel at all. The duller answer is corr(barrel, counter RESIDUAL) = -0.012. Barrel is a real skill that carries no information about what the counter gets wrong at this line. That also closes the S81 lead: hard_hit r=0.153 at n=295 drifted to 0.135 at n=383 and is THEATER at n=707. THE REAL FINDING: TB is miscalibrated, not under-factored. mean p_win 0.5698 vs actual 0.5074, bias +0.0624. Held out on a strict time split (fit < 2026-08-02, eval 651 unseen rows): raw 0.25007, constant de-bias 0.24740 (-0.00267), isotonic 0.24621 (-0.00386). Worth more than any factor tested and the only intervention pointing the right way -- and still refused at the corrected bar on 32 clusters. A CANDIDATE, not a result. It also explains the units bug's fake success exactly: a blanket downward shift is a crude de-bias. NO RESCALE. Nothing proved, nothing certified calibrated, no slot at sample -- every band would be the honest base-rate band gradeBands already returns by construction. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
8ab6557faa |
Collapsed sequence edge: two proven links whose product is too small to use
Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.
FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.
TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.
GATE on the concentrated subset, 114 cumulative tests:
early-exit x WEAK pen n=1931 brier -0.0001 CI [-0.0014,+0.0010] NOT_PROVEN
early-exit x STRONG pen n=2574 brier 0.0000 CI [-0.0011,+0.0010] THEATER
all early-exit later ABs n=6869 brier -0.0001 CI [-0.0007,+0.0005] NOT_PROVEN
pooled all later ABs n=17891 brier 0.0000 CI [-0.0004,+0.0003] THEATER
Not pooled-diluted -- the concentrated subset was gated alone and is no
better.
THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925
hit-rate swing across pen quality 0.0394
MAX JUSTIFIABLE ADJUSTMENT 0.00365
actually applied 0.01930 -> 5.3x over-movement
A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.
A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.
NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.
THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.
Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
b2e4c6c4fb |
Link 2 at the coarse grain: pen QUALITY proves, archetype does not
The refinement was right. Naming the individual reliever failed; the same
question at the grain the chain needs passes, and it transmits more than
anything else measured in this chain.
WHY IT WAS WORTH RE-ASKING: last session's null (the pen is on average no
softer, +0.0010 on 35,760 PAs) does NOT rule this out, and treating it as
though it did would have been the error. An average washing out is fully
consistent with quality VARIATION mattering. It does -- actual arm quality
moves the hit rate monotonically across quartiles, 0.2244 / 0.2293 /
0.2410 / 0.2501, a 2.57pp spread, larger than the whole times-through-
the-order effect.
CLUSTER UNIT CORRECTED, THEN CHECKED RATHER THAN ARGUED. Last session
refused Link 2 partly as team-borne (30 bullpens, the park ceiling). My
first re-check was that 76% of pen-quality variance is within-team -- but
that is a statement about TREATMENT variance, not about where errors
correlate, and stopping there would have been picking the convenient
answer. Measured the actual thing: ICC of prediction error by team =
0.0261, design effect 1.41, SEs inflated ~19%. So the verdict was run
three ways:
unclustered CI [-0.0067,-0.0010] excludes zero
team-clustered (30) CI [-0.0086,-0.0003] excludes zero (below the
40-cluster floor -- indicative, not a pass)
design-effect adjusted CI [-0.0072,-0.0005] excludes zero
QUALITY GRAIN PROVES on the concentrated elevated-early-exit subset:
n=501 team-games, 426 clusters, MAE 0.0294 -> 0.0260, delta -0.0034, CI
[-0.0063,-0.0005] at 110 cumulative tests. Pooled also proves, so it is
not a subset artefact.
ARCHETYPE GRAIN DOES NOT: 0.5669 vs a 0.5309 modal-guess baseline,
corrected interval [-0.1073,+0.0268] spans zero. Two grains tested, one
earned a place -- penQuality.js exposes no archetype and a test asserts
it.
WHAT LINK 3 RECEIVES, which is the number that actually matters -- not
the MAE gain but realized outcome separation, prediction strictly
point-in-time:
predicted BEST pen 167 games 2,044 PAs hit rate 0.2231 +/-0.0180
predicted WORST pen 167 games 1,799 PAs hit rate 0.2501 +/-0.0200
2.70pp separated, intervals non-overlapping, capturing nearly all the
2.57pp available at the quartile grain. Caveat stated not buried: the
tercile cut is chosen in-sample; the prediction driving it is not.
BUILT: penQuality.js + 9 tests. Abstains below 5 prior club games and 40
arm appearances -- a league-average stand-in would assert "this is an
ordinary bullpen", which is a claim, and usually the wrong one for exactly
the clubs whose pens just turned over.
Link 3 is unblocked on a proven Link 2 at the quality grain only. Not run
here; this order scopes to building and gating Link 2.
Parallel track logged unchanged: TB n=948 pooled, BOMBER x TB 340, short
by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
e4dae0e6b0 |
Reliever chain: Link 1 proves, Link 2 does not, and the premise inverts
The causal insight is right -- the game is a sequence and the matchup does shift mid-game. The direction is backwards, measured on 93,663 plate appearances from 1,238 games pulled free from statsapi. LINK 1 PROVES. Starter batters-faced, point-in-time from his own prior starts only, clustered on the pitcher: MAE 3.2226 -> 2.7990, delta -0.4236, CI [-0.6006,-0.2731] at 0.9995 corrected for 107 tests, 1,706 starts across 204 pitchers. It finds the tail the chain needed -- early exits are a 23.2% base rate, model-flagged starts are 34.0% early, lift +10.8pp. Scope correction inside Link 1: the order specifies fatigue x GAME SCRIPT, but game script is not available at grade time -- whether he gets hit tonight is the thing being projected, not an input to it. Only the workload half is measured; the in-game half is recorded as a live feature, out of scope, rather than quietly folded in. LINK 2 DOES NOT PROVE, twice over. Model accuracy 17.2% vs an 8.6% baseline -- doubling it sounds good and is not, since naming a specific arm is wrong five times in six. And structurally the entity is the BULLPEN: 39,629 post-starter plate appearances across 30 clubs is 30 readings, below the 40-cluster floor, the same permanent ceiling as park geometry and team defence. LINK 3 NOT RUN, per the order's own rule. THE PREMISE IS REFUTED, and this chains on nothing so it was safe to measure: vs STARTER n=48,492 hit rate 0.2444 +/-0.0038 vs BULLPEN n=35,760 hit rate 0.2373 +/-0.0044 The pen is 0.7pp HARDER. The specific effect the chain exists to exploit -- early exit making later at-bats softer -- is +0.0010 on 35,760 PAs. A well-powered null, not a sample problem. What IS real is times through the order: TTO1 0.2351 -> TTO2 0.2515 -> TTO3 0.2518. A starter does decay as the lineup sees him again, but that advantage is SURRENDERED when he leaves, not extended -- the pen is harder than his second and third time through. A modern bullpen is a queue of fresh specialists throwing one inning each; there is no tiring arm to punish. So the insight survives inverted, and Link 1 stays valuable for the opposite reason it was built: a likely early hook predicts the hitter LOSES his third-time-through look (0.2518 -> 0.2373 on that PA). The mispricing is on hitters who get an EXTRA look at a starter going deep. BUILT: predictionGate.js + tests -- the two-part gate for a continuous prediction. factorGate binarises outcomes for Brier, which would destroy a target like batters faced. Same discipline, same THEATER verdict, real scale. PRE-REGISTERED NOT RUN: Link 2' using a PA-weighted bullpen AGGREGATE rather than a named arm. Recorded rather than substituted in -- running Link 3 on a swapped-in Link 2 is the assumed-link failure the order forbids. Given the premise result its expected value is now low. PARALLEL TRACK logged: total_bases n=948 pooled, BOMBER x TB 340, short by 160. Sample-readiness only, not a verdict. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
3081c92e00 |
Per-archetype grade bands: built, gated, and the rescale blocked twice
The premise does not hold. proven-status.js run fresh: PROVEN_SET is EMPTY, no archetype x stat reaches the gate. pitcher_contact_profile has a CI upper bound of exactly 0.0000 and platoon_severity is held on 4.5%-contaminated splits, so the proven set is one factor, pooled, not three archetype-conditioned ones. The specific pattern the order names -- defense strong for GHOST/BRUSH, null for BOMBER -- is the one I measured running the OTHER WAY yesterday, both noise-dominated. But the second blocker is new and matters more, because it would stop the rescale even if the factors had proved: the grade does not separate within any archetype. Every archetype collapses to ONE band at the corrected bar, because bands merge when their intervals overlap and publishing two letters we cannot tell apart is a distinction we have not measured. Uncorrected, so the ranking is visible rather than hidden by the bar, this INVERTS the order's design. The order gives contact types the factor-rich treatment and power types honest base-rate, reasoning that single-game hits are variance for a power profile. Measured: BOMBER n=466 corr(p_win,outcome) +0.207 quintiles 0.75 0.62 0.60 0.48 0.48 GHOST n=192 corr(p_win,outcome) -0.007 quintiles 0.47 0.63 0.74 0.58 0.45 BOMBER is the one archetype the model ranks, and it splits into a real A 0.660 / B 0.481 at 95%. GHOST is flat, and non-monotone -- its most confident reads hit 47% while its middle reads hit 74%. Shipping as specified would have given the factor-rich treatment to the archetype the model reads worst and left base-rate on the one it reads best. That is mechanically sensible in hindsight: a power hitter's hit tracks whether he can damage the arm, a contact hitter's depends on balls finding holes. BOMBER's split does not survive the cumulative correction at 106 tests. Exposing it by loosening the correction is the curve-to-make-A's the order forbids, so it stays one band. BUILT: gradeBands.js -- lift against the archetype's OWN base rate (the same 62% is lift for a 45% profile and a deficit for a 68% one), indistinguishable neighbours merged, thin bands PROVISIONAL not dropped, Wilson intervals widened by the cumulative correction. The two-bar rule is structural: proven-alone, calibrated-alone and neither all return base_rate with the reason stated, so with nothing proven no factor-informed band can be produced at all. reasoning() is built and tested but NOT wired to the card -- there is no per-archetype band being served, so attaching the copy now would ship product language for a rescale that does not exist. NOT BUILT: the specified power-type reason "the matchup edge is in total_bases". total_bases is recorded INCONCLUSIVE (+0.0038, CI [-0.068,+0.075]). Wiring it would assert an edge measured as indistinguishable from zero -- the exact fabricated-reason failure this module exists to prevent. BOMBER x hits is 29 rows short of the gate and is the archetype the model actually reads. That is the first slot to test, not GHOST. Counter and frozen clusters byte-identical. No letter was moved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
6b17f79367 |
Per-archetype re-audit: no slot reaches 500, and the replication unit
decided everything The premise does not hold. prove-hit-factors.js has no date filter anywhere in it and pages the full table -- there was never a window to widen. Full clean history is 1,266 rows, not 2,715. platoon was not "proved" last session, it was explicitly held on 4.5%-median-contaminated season-to-date splits, and pitcher_contact_profile was demoted. The proven set going in was one factor, not three. STEP 1: no archetype slot reaches n>=500 on full history. Best is BOMBER at 408, and BOMBER is the most common archetype on the board. GHOST 173, BRUSH 64, DRIVER 43, CATALYST 16. These are confirmed genuinely short, not artifacts. STEP 2 is where the real finding is. park_hits initially PROVED at 619 rows across 45 games -- but those games only ever visited 14 distinct park values. A park effect is replicated across parks, and unmodelled park heterogeneity is confounded with the thing being estimated. Each factor is now clustered on the coarser of the game and the entity its treatment rides on. That flipped two verdicts and confirms Kev's causal-correctness thesis from a new direction: defense_by_direction has 442 hitter-team units of replication where crude team defense has 26. The correct atom is not just more accurate, it is the only one measurable at all. park_hits (14) and defense (26) can never be validated however long the ledger runs -- the same ceiling as park dimensions, reached independently. Also fixed a bar I got wrong last session: I transplanted the 500-row floor onto clusters, which refused a factor with 1,059 rows over 85 games while answering neither question. Two floors now -- rows>=500 for a stable estimate, clusters>=40 for a trustworthy interval. Not a lowered bar: park_hits and defense are still refused. PROVEN: defense_by_direction only, pooled, [-0.0054,-0.0012] at 99 tests. It stays POOLED-ONLY -- no per-archetype reasoning wired, nothing grandfathered. The card must not say "GHOST: defence matchup strong" because we have not earned that sentence. The predicted fingerprint did not appear either: BOMBER -0.0036 vs GHOST -0.0024, the opposite direction, both noise-dominated. Recorded so it is not claimed later. RESCALE: NOT READY. One proven factor worth -0.0031 Brier. Rescaling on that is relabelling. Counter and frozen clusters byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
7b85934dc3 |
Under-querying vs out of data: the answer depends on the unit
The platoon test's n=452 described how much of the JOIN survived, not how much data exists. There are 1,266 clean settled hits rows and zero quarantined ones. platoon_splits had been ingested from tonight's lineups only (315 players), so any hitter who settled a prop without appearing in an ingest-day lineup was silently absent from every test. Backfilled all 380 hitters (81 fetched, 0 unresolved). Re-ran on 1,059 rows, up from 452. THE DEMOTION IS THE HEADLINE. pitcher_contact_profile, the strongest proven factor in the programme (-0.0064, CI [-0.0113,-0.0014]), roughly halved to -0.0034 on more than double the sample and its corrected interval now spans zero. The Bonferroni denominator also rose to 55, which widens every interval -- but a denominator cannot move a point estimate, and that halved on its own. platoon and platoon_severity now clear the bar and are NOT promoted. Upper bound -0.0001, on season-to-date splits that contain the games they predict: measured contamination is 4.5% median, 12.4% at p90, 137% worst. I had assumed ~1%. They stay CANDIDATE pending point-in-time splits. GAME-LEVEL IS A DIFFERENT PROBLEM. game_context held zero weather rows ever -- not because the fetcher was wrong (it correctly targets Open-Meteo's archive) but because ledger_entries keys a game as mlb:2026-08-03:Away@Home and game_context keys it as mlb:823437. Every lookup missed and NULL columns read as honest absence. Third occurrence of that class. Fixed the join: 96/101 settled games now carry actual archived weather, park dimensions backfilled 15 -> 30 venues. But 928 total_bases rows sit on 47 games at 17.6 rows per game. Park and weather assign one value per game, so resampling rows would have manufactured a pass. factorGate now resamples clusters when rows carry one and judges sample against effective_n; unclustered rows keep the original path byte-for-byte. Verdict: 47 clusters < 500, and the point estimate is +0.0011 -- worse, not merely unproven. Weather needs ~57 more days. Park dimensions need never: there are 30 ballparks in MLB, so a venue-constant factor can never reach 500 independent units. That bar was built for player-level factors and does not transfer. Wind is refused. We have speed and bearing for all 96 games; we lack park orientation, and 220 degrees is blowing out at one park and in at another. Using speed alone would assert an effect while discarding the sign that decides what it is. Counter and frozen clusters untouched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
6452926732 |
Retain raw weather, and record platoon severity 48 rows short
Two fixes in the weather path, and the second was hiding behind the first. The scalar weather_mod cannot express a hit-TYPE conversion at all -- wind out and warm turning fly balls into extra bases, and cold heavy air turning them into outs, collapse to the same number once multiplied -- so the raw temperature, wind speed and wind direction are now retained alongside it. And the old guard only kept the environment when the multiplier was not 1, which silently discarded the forecast for every ordinary night. That is the majority of games, and precisely the rows a hit-type model would need in order to learn what ordinary looks like. Platoon severity is built and measured at n=452, which is 48 rows short of the gate: CANDIDATE_PENDING, neither proven nor theatre. It moves less than flat platoon (0.021 against 0.026), consistent with the pattern, and its Brier point estimate is favourable but the corrected interval still spans zero. Worth naming: the refusal costs sample, and that is the design working. Flat platoon scores 741 rows because it will happily apply a boost to anyone; severity scores 452 because the other 289 are hitters whose split we cannot actually read at 60 plate appearances on the short side. Buying those rows back by shrinking instead of refusing would have produced a number indistinguishable from a measured league-average split, which is a different claim from the one the data supports. Park dimensions are ingested and verified in production across fifteen venues, joined by the venue the game is actually at rather than inferred from the home team -- neutral-site and international games break that assumption without surfacing an error. The park-and-weather-to-hit-type atom is NOT built. Its inputs landed this session and carry a single as_of date, so testing it on total_bases would be scoring games with inputs that postdate them. Building it now would produce something plausible rather than something proven. Proven factors for hits remain pitcher_contact_profile and defense_by_direction. 4,307 tests green (344 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
20c45cbcd1 |
The causally-correct defence atom proves where the crude one did not
Kev's insight holds, and the data says so cleanly. defense_by_direction PROVES on hits -- n=528, Brier -0.0034, interval [-0.0059, -0.0009] at the 99.9% level the cumulative correction now demands -- while team-average defence remains not proven, its interval still spanning zero. Same signal, same rows, different unit. The detail worth keeping is that the causally-correct atom moves the number LESS THAN HALF as much as the crude one, 0.013 against 0.030, and is the one that is reliably right. The team average was moving more and knowing less. Big movement is not evidence of a good factor; it is frequently the tell. Both halves turned out to be free, as the order expected. Savant's batted-ball leaderboard carries pull/straight/oppo crossed with ground/air for 609 hitters -- the statcast leaderboard we already pull does not, it has nineteen columns and no direction at all -- and the OAA feed already carries each fielder's position, so per-position defence is a regrouping of last week's ingest rather than a new source. Verified in production: 609 spray profiles, 31 teams. Handedness is what joins them and getting it backwards would have been invisible. Pull for a right-handed hitter is the left side; for a left-handed hitter it is the right side. A model that ignored `bats` would send half the league's grounders to the wrong infielders and still look like it was reading defence, and nothing downstream would have caught it. Switch hitters bat opposite the pitcher, which this does not resolve, so they are unreadable rather than guessed. Unmeasured zones are renormalised away rather than contributing a zero, since a zero asserts an exactly-average fielder standing there, and coverage states honestly what share of a hitter's contact we could actually read. ATOM 2 is input-blocked rather than sample-blocked, and the distinction matters because waiting will not fix it. The weather free-source check passes -- Open-Meteo is already wired and exposes temperature, wind speed, wind direction and precipitation -- but those raw fields are collapsed into a single scalar modifier and wx_forecast is empty on all 1,119 settled rows. Park DIMENSIONS are not ingested at all; parkFactors holds coefficients, not wall heights or fence distances. A park-and-weather-to-hit-type conversion needs both, so it is scoped rather than half-built: retaining the raw weather fields is the cheap half, dimensions are the missing one. Proven factors for hits are now pitcher_contact_profile and defense_by_direction, both pooled; every per-archetype slot remains sample-blocked. 4,297 tests green (342 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
a9ee55550b |
Build the two-part factor gate: one factor proves, and zero are theatre
The question was whether the hit grade reads tonight's game or just says he is due. Answering it needed a gate that correlation cannot provide, because correlation cannot separate the two ways a factor looks alive: it reads the game, or it moves the number and reads nothing. The second is what a product ships by accident -- arch-v1 moved 76% of rows by 2.5 points, changed resolution by 0.0000, and was live for months, and no user could have told. So a factor must now clear both conditions: move the prediction off the player's own leave-one-out base rate, AND improve out-of-sample Brier. Brier rather than correlation, because correlation asks whether the ordering improved and this asks whether the NUMBER got closer to what happened -- and for a graded probability the number is the product. The correction applies to the interval itself, which turned out to matter more than expected. A plain 95% CI is the right bar for one test; at fifty cumulative tests roughly two or three intervals exclude zero by chance alone. Widening to 1 - 0.05/tests, currently 99.9%, flipped both defence and platoon out of "proves". A 95% interval would have shipped two unproven factors into the grade, with reasoning text explaining them to users. That forced a distinction I had initially collapsed. Defence and platoon have FAVOURABLE point estimates whose corrected intervals merely span zero, and calling that THEATER would repeat the error this codebase keeps correcting: insufficient evidence is not evidence of absence. THEATER is now reserved for its one real meaning -- moves the number, reads nothing -- and NOT_PROVEN_AT_CORRECTED_BAR names a real candidate held to a bar that rises with every hypothesis the programme tests. Result on 741 settled hits rows: pitcher_contact_profile PROVES, improving Brier by 0.0066 with a 99.9% interval of [-0.0114, -0.0016]. Defence (-0.0043) and platoon (-0.0039) are not proven at the corrected bar. Park is sample-blocked at n=405. Zero factors are theatre, which is the genuinely good news: nothing decorative is being wired. Per-archetype every slot is sample-blocked (BOMBER 252-294, GHOST 67-125). Two spec gaps worth recording. The approach identities the order names -- SPRAY, DAMAGE-DEALER, COUNT-WORKER -- do not exist in the registry; the MLB batter archetypes are BOMBER, GHOST, TORCH, BRUSH, DRIVER, FLEX, ALPHA, HYBRID and CATALYST. And parkFactors maps hits to run_base, so there is no hits-specific park factor at all: a park that turns outs into hits without producing runs is invisible to the input we have. The grade rescale is NOT run. It was explicitly gated on the factor proving, and one pooled factor worth 0.0066 of Brier is not a factor-informed distribution -- rescaling on it would dress a base-rate model as a matchup model, which is the exact thing this gate was built to prevent. 4,286 tests green (340 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
4d1803f6d7 |
Calibrate hits point-in-time: partial pass, and an honest ceiling of 0.667
Fitted the isotonic map on game_date < 2026-08-02 (n=589) and evaluated it on everything from that date forward (n=383). The map never saw the evaluation rows, which is the only thing that makes the result mean anything -- fitting and evaluating on the same rows always looks perfectly calibrated, because the map is reciting the answers it was built from. It works, on most of the distribution. Held-out after correction: 0.477 comes back 0.506, 0.587 comes back 0.580, 0.667 comes back 0.603 -- against raw errors of +0.191, +0.279 and +0.246 in the same bins. Ordering survived, and that was verified pairwise rather than assumed, because a broken map would silently destroy the one thing this model does well. Two findings matter more than the pass. First, the honest ceiling is 0.667. Once the numbers are truthful this model has no 80%-plus hit reads at all -- the top of its range was miscalibration, not confidence. A four-leg ticket at the ceiling is 0.198, where the raw numbers implied 0.686. The high-floor parlay is a two-thirds-per-leg proposition, and that is the number to say out loud. Second, calibration is certified BY BAND rather than by a blanket flag. Held-out error was -0.029 and +0.007 through the middle but -0.167 at the bottom and +0.063 at the top: the model is trustworthy over most of its mass and untrustworthy at both edges. A single true/false would either throw away the 72% that works or ship the edges that do not. Only a probability inside a certified band is marked stackable, and that flag is what chainAcross requires before it will compound anything. The certified band is 0.40 to 0.60, n=276. A methodological catch on the way: my first pass condition demanded honest bins at 0.70 and above -- but honest calibration REMOVES those bins, since the ceiling drops to 0.667. The gate would have failed the repair for succeeding. It now tests the highest remaining band instead of a fixed threshold. Wired forward with the same discipline: calibrationService fits strictly before today, splits by time rather than at random, and returns null on thin history so that "no calibrator" means nothing is stackable rather than "trust the raw numbers". p_win is never mutated -- the calibrated value rides beside it as p_win_calibrated, because a calibration map is a correction to a forecast, not a different forecast, and the counter stays byte-identical. 4,275 tests green (339 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
ff037e40c2 |
Re-adjudicate: nothing to demote, and close the hole that would have mattered
There is nothing to re-adjudicate. The proven set is empty and always has
been -- verified three ways: proven-status reports EMPTY, validatedSkills()
returns {} for every archetype, and zero conditioning entries have ever
reached PROVEN. The one PROVEN feature is recent_frequency_prior, which is the
incumbent counter itself, proven by the S78 ablation as ~100% of the
champion's resolution. It is the baseline every challenger is measured
against, not a conditioning interaction, and demoting it would leave the model
with nothing to grade from.
A correction to the premise: the cumulative gate did NOT catch a false
positive last session. It caught nothing, because there was nothing in the
proven set to catch. What it did was tighten alpha from 0.0026 to 0.0013
within one session, which demonstrated the mechanism working rather than a
demotion. So steps 3 and 4 -- demote, recalibrate -- are vacuous here, and
readjudicateAll says so plainly rather than glossing a no-op.
But the worry behind the order was well founded, and the audit found the real
exposure: promote() did not require the cumulative denominator. It checked n,
lift and CI, and nothing stopped a future session from testing eight
hypotheses, correcting by eight, and promoting on a p-value that would not
survive the programme's real denominator. That is precisely the hole that
makes a retroactive re-adjudication pass necessary later, so it is closed at
promotion time instead. isSufficient now refuses evidence carrying no
correction, evidence corrected against fewer tests than the cumulative count,
and any p-value that does not clear 0.05 over its own test count. The same
rule guards a PROVEN conditioning entry.
The second audit found two of four analysis scripts still correcting
per-session; pitcher-prove-k and tb-solo-and-interactions now use the
cumulative ledger, so the correction is native on every path.
reAblation.js is the standing second line: pure and injectable, so the
decision rule cannot drift from the gate's, and every verdict records both
p-values and both test counts so a demotion is re-derivable by anyone. A
feature promoted at alpha 0.05/20 can demote on the same p-value once the bar
is 0.05/60 -- correct, because the bar rose only after the programme had more
chances to get lucky. No fresh measurement is PENDING_RETEST and never a
demotion: absence of a re-test is not evidence, and demoting on it would
punish whichever stat happens to be off-season.
Net effect on the proven set is zero. No demotions, no recalibrations, and no
public ledger event -- announcing "recalibrated after re-adjudication" when
nothing changed would itself be a false signal of rigour.
4,238 tests green (337 suites); web build exit 0; counter byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
ece2b9f5f9 |
Ingest defence, and make Bonferroni cumulative across the programme
Two things shipped that stand regardless of sample. DEFENCE. Statcast Outs Above Average is free on the host we already pull six feeds from, so there was nothing to decide. 514 fielders, aggregated to team level -- the unit a batter's prop actually needs, the defence behind the pitcher he faces -- and persisted as 31 team rows. Verified in production. Cubs +56 best, Mariners -29 worst. Unknown is not zero, and it bites unusually hard here: an OAA of 0 is a REAL reading meaning exactly average, so coercing absence to 0 would assert that every unmeasured fielder is league-average, which is the commonest defensive profile there is. team_defense also carries as_of_date in its primary key from the first row -- statcast_aggregates was built upsert-in-place and that silently made every backtest leak the games it predicted, so point-in-time is available here before it is needed rather than after a wrong answer. A bug worth recording as a class: BASE already ends in /leaderboard, so the new feed built a doubled path and 404'd. Because a failing feed degrades to an empty index by design -- correct, so one broken source cannot fail the whole pull -- it surfaced as "fielding_oaa: 0 rows", which reads exactly like "Statcast has no fielding data". Graceful degradation makes a wiring bug look like an honest absence. CUMULATIVE CORRECTION. Bonferroni had been applied per session throughout: a run testing eight features corrected by eight. Across a programme's lifetime that is wrong in the dangerous direction, because every order gets a fresh generous alpha and the false-positive rate compounds quietly. Correcting by 8 when sixty have been tried is how a noise result eventually gets recorded as PROVEN with a p-value to point at. The denominator is now distinct hypotheses ever tested, persisted, and it moved 19 -> 38 within this session alone, alpha 0.0026 -> 0.0013. Re-tests deliberately do not inflate it: re-asking the same question on more data is not a new shot on goal, and counting it would punish the discipline of waiting for sample. THE MEASUREMENT. The differential the theory predicted is present: defence correlates with the counter's residual at +0.130 for GHOST, the contact and speed archetype, and -0.018 for BOMBER, the power archetype. A GHOST's hits depend on whether anyone can range to the ball; a BOMBER's barrels clear the defence entirely. So a flat BOMBER result is the theory working rather than the test failing. It is not a result. GHOST is n=104 against a 500 bar, with p=0.188 against a corrected alpha of 0.0013 -- three orders of magnitude short. Both are recorded as CANDIDATE with their measured lift, tagged contact-skill, so the re-run at full sample compares against a recorded baseline. Nothing proved, so nothing was recalibrated and nothing shipped. 4,228 tests green (336 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
ac1361486e |
Build the conditioning registry, and a probe so "proven" stops drifting
The order opens with "two proven clusters live". They are not proven -- the
proven set is empty -- and this is the fourth consecutive order to start from
a stronger claim than the measurements support. Correcting that in prose four
times has not worked, so this session adds scripts/proven-status.js, which
recomputes the answer from the ledger: hits LOSES (-0.096, CI excluding zero),
total_bases INCONCLUSIVE (+0.004), strikeouts INCONCLUSIVE (+0.259 at n=57).
It deliberately reports sample readiness separately from recorded verdicts, so
"n>=500" can never again be read as "passed".
A counting error worth recording. The first read of the top-volume archetype
said BOMBER x hits was 641 rows -- gate-ready. It is 287. model_snapshots
holds one row per prop PER SNAPSHOT CYCLE, so joining it to ledger_entries
counts each ledger row once per cycle it appeared in. Deduping on the ledger
row id gives the true figure, and my own status script had the same bug until
it was fixed. That is the difference between running the gate and being short
by 213.
So no archetype x stat combination reaches the gate. BOMBER x hits at 287 is
the closest; pitcher archetypes are untestable at 58 settled strikeout rows
across all of them, so the pitcher half of this order could not be run.
The registry is built: recordConditioning keys archetype x underlying-skill x
interaction x status with measured lift, and the skill tag is MANDATORY and
enforced -- untagged entries are refused, and PROVEN without sufficient
evidence is refused. validatedSkills() returns the coherent profile as it
stands, which is {} for every archetype, by design.
BOMBER x hits conditioning was tested across the order's categories and every
result is underpowered: arsenal (barrel x breaking share) incremental +0.043,
batted-ball (launch x pitcher GB) +0.001, contact quality -0.020 and -0.015,
K x K -0.063. Within BOMBER the counter still leads on hits, 0.218 to 0.160,
consistent with the closed pooled negative.
One bug fixed mid-run: fromStatcastRow maps percentage and raw fields only and
does not carry pitch_mix, so the arsenal category first reported n=0 for every
row -- it was measuring nothing rather than failing. Without catching it,
"arsenal doesn't matter" would have been recorded from a column that was never
populated.
On defense: I looked for a derivable proxy before calling it unsourceable, and
there isn't one. We ingest no fielding data at all, and opposing pitchers'
hits-allowed conflates pitching with defense, so it would validate the wrong
skill. It needs Savant's fielding endpoint -- free, same host as the five
feeds already ingested -- and it is not sourced here, because sourcing it to
test at n=282 would answer nothing.
Nothing proved, so nothing was recalibrated and nothing shipped.
4,221 tests green (335 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
9538e11198 |
Derive the lineup K-rate free, and fingerprint the cap fix
Two premise corrections first. Pitcher stuff features have NOT proven solo through the gate -- every one was refused on sample (n=57 against 500). Four exceed the effect-size bar (arm angle -0.250, whiff +0.213, k rate +0.206, chase +0.195), which is why they are worth pursuing, but clearing one of three thresholds is not passing. And the carrier was not blocked only on the lineup input: that input was built and measured last session at 94.7% coverage. What blocks it is n, and n was being throttled by the grading cap. RUNG 1 IS DERIVED AND COSTS NOTHING. Opposing-team K-rate comes from joining the opposing roster to the batter k_pct values already in statcast_aggregates -- no new feed. The improvement this session is that it is PA-WEIGHTED: an unweighted roster mean counts a 12-PA callup the same as an everyday starter, which is not the lineup a pitcher faces. That change alone reversed the term's sign. Unweighted, the lineup term HURT the model (0.1738 -> 0.1285). PA-weighted, it HELPS (0.1738 -> 0.1953). Same hypothesis, same data -- the derivation was the problem, not the signal, which is the entire argument for deriving the best honest version before sourcing anything. Head-to-head is now +0.2592 with a CI of [-0.0167, +0.5645], very nearly excluding zero, at n=57. Within archetype, the two strata come out with OPPOSITE signs -- FLAME incremental -0.152, non-FLAME +0.145 -- and the pooled value (+0.077) sits between them, which is the shape a conditional effect makes and is invisible when pooled. That is what stratifying was for. But n is 20 and 24, the standard error on a correlation there is about 0.22, and the direction contradicts the theory that predicted a stronger effect for finesse arms. It is recorded as a structure to re-test, not as a finding. Rungs 2 and 3 are NOT triggered. A rung fails only once it has been fairly tested, and Rung 1 is n-blocked rather than failed. Sourcing confirmed lineups now would be paying for precision on top of a proxy we have not yet measured. THE RESULT THAT DECIDES THE TIMELINE: yesterday's cap raise is fingerprinted in production at 907 grades per snapshot, up from 334, with strikeouts going 6 to 17. That puts n>=500 for pitcher Ks about a week out instead of three months. Operational note: the manual internal snapshot endpoint now 524s at the Cloudflare edge because grading the full board exceeds 100s -- the run still completes server-side (this very snapshot was written by a 524'd request) and the cron is in-process, so a 524 there is not a failure. Nothing proven, nothing calibrated, nothing shipped. The counter remains anti-predictive on strikeouts at -0.064 and the skill model leads it by 0.26. 4,221 tests green (335 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
843c8c6d4b |
Build the pitcher engine, and find the cap was eating the whole board
Strikeouts are NOT proven -- n=57 against a bar of 500. But the finding that matters is not a correlation. THE CAP. Measured on the live slate via the refusal diagnostic: 1,244 unique gradeable props exist, the 500 cap graded about 334, and because dedupeProps takes first-row-wins in FEED ORDER, what survives is decided by feed position rather than value. Pitchers are 2.6% of a batter-dominated feed, so we were grading SIX strikeout props a slate against 32 available -- putting n>=500 three months away for every pitcher stat. Pitcher props were never being refused (graded 5, refused 0, suppressed 0); it was truncation. Raised 500 -> 1500 on measured cost: 721ms per prop at concurrency 5 is about 179 seconds for the full board, against a cron that runs five times a day and a fire-and-forget caller that never holds an HTTP response. statsapi is free and unlimited. Concurrency stays at 5 -- one variable at a time. This unblocks every n-blocked stat in the programme, not just pitchers. THE ENGINE. pitcherEngine.js is its own engine, not the batter engine pointed at pitchers: the batter model asks whether contact becomes a hit and reads contact quality, the pitcher model asks whether the plate appearance ends without contact at all and reads stuff. Archetypes are FLAME (whiff-led), SCALPEL (chase-led), SINKER (pitches to contact) and DEFAULT, and a test asserts the weight keys are not the batter engine's. The projection is K% by log5 against THIS lineup, times batters faced, through a binomial. An unclassifiable arm gets the balanced map, never a guessed archetype. THE MEASUREMENT, at n=57 and contaminated. Four solo features clear the 0.15 effect bar and fail only on sample: arm angle at -0.250 -- the largest correlation measured anywhere in this programme -- then whiff +0.213, k rate +0.206, chase +0.195. The batter cluster's best was 0.135. Head to head, pitch-v1 resolves 0.1285 against the counter's -0.0639, delta +0.192 with a CI spanning zero. That negative is the interesting number. The counter is ANTI-PREDICTIVE on strikeouts: counting a pitcher's recent Ks is worse than useless, because his recent totals track which lineups he drew and how long he was left in rather than his skill. It is the one stat where the incumbent has no defensible edge. A bug caught on the way. resolveTeam wants an abbreviation and the game log supplies full team names, so the roster join silently resolved nothing and the first run reported 0% lineup coverage -- the theorized stuff x lineup carrier was never being tested, not failing. Fixed; coverage is now 94.7%. The carrier still shows no incremental signal over whiff alone, and adding the lineup term lowered head-to-head resolution, which is recorded rather than dropped. Calibration was not reached: nothing passed the first bar. The batter model and the counter are byte-identical, verified by diff. 4,221 tests green (335 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
c0621e7aa2 |
Measure the batter cluster: the proven set is empty, and hits is closed
PREMISE CORRECTION FIRST, because it defines the bar. total_bases has not
passed BAR 1. Its head-to-head is inconclusive at parity -- delta +0.004 to
+0.007 with a CI spanning zero -- and it is contaminated, and no feature of
its passed the gate. It was described last session as the first challenger
that did not LOSE, which is not the same as proven. If it is installed as the
frozen proven reference and every other stat is held to "the identical bar
total_bases cleared", the bar becomes "be inconclusive at parity" and the
whole cluster passes on a null result. The proven set is EMPTY.
HITS IS NOW A FINAL ANSWER. At n=803 it clears the gate's sample requirement,
so its features were properly TESTED rather than refused: every one fails on
effect size (max marginal |r| 0.053 against a 0.15 bar), every interaction's
incremental contribution collapses to about zero, and the model loses
head-to-head by 0.096 with a CI excluding zero. That is a well-powered
negative and hits should be closed rather than retried.
The rest are n-blocked: total_bases 383, rbi 391, home_runs 228, runs 188,
against a bar of 500. Two leads are worth carrying. home_runs barrel rate has
a marginal r of -0.135, and the sign matters -- higher barrel rate goes with
the counter OVER-predicting, which would be a correction rather than a new
predictor. And runs batterK x pitcherK has the largest incremental in the
cluster at +0.132, with a clean mechanism: strikeouts destroy plate
appearances, and a PA that never happens cannot score.
RBI deserves a caveat rather than a verdict. It is power times OPPORTUNITY,
and we ingest no baserunner state at all, so half its mechanism is missing. A
weak RBI result is evidence that we are modelling half the stat.
total_bases was held frozen: git diff on skillProjection against the prior
commit is empty. The counter is untouched.
Also fixed and verified in production: the point-in-time retention shipped
after yesterday's refresh had already run, so statcast_history was empty, and
its first run then failed on a hand-enumerated schema that had already drifted
from its source ("could not find the 'swing_pct' column"). The refresh itself
still succeeded and wrote all 1,387 aggregate rows, which confirmed the
best-effort guard in prod. The table now mirrors the source via LIKE and the
writer passes rows through whole. Verified live: 1,387 rows retained at as_of
2026-08-03. A usable point-in-time window starts 2026-08-04.
Stage B has nothing to calibrate. Everything now waits on a point-in-time
window and on sample -- both waiting problems, not building problems.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
|
||
|
|
4aab18096f |
Prove both on total bases -- and find that my own fix destroyed the backtest
Nothing passed. Nothing promoted. Counter byte-identical. THE BLOCKER, which is the real finding. statcast_aggregates is upserted in place and holds exactly one as-of date. Yesterday's skill backtest was honest only by accident: the nightly refresh was unreachable code, so the profiles sat frozen at 2026-07-21 -- before the settled window. Repairing that cron was right for production and it refreshed them to today, destroying every prior version. Scoring a 2026-07-25 game now uses a season aggregate that contains that game. Point-in-time validation is structurally impossible from that table, so every number in this run is contaminated and directional, and none of it is a gate verdict. Fixed forward: statcast_history retains a dated snapshot on every refresh, so point-in-time becomes "as_of_date < game_date, most recent". Retention is best-effort and cannot fail the refresh; both properties are unit-tested. It has one day of data, which is not yet a window. SOLO BASELINE, n=383, Bonferroni across 12 tests (alpha 0.00417): nothing passes. hard_hit_pct is closest at marginal r 0.135 with p 0.0080, failing both the 0.15 effect bar and the corrected alpha. And it drifted DOWN from 0.153 at n=295 -- an estimate regressing as noise averages out, not an effect firming up. I called that number encouraging yesterday; on 88 more rows it is fading, and it should not keep being quoted at its best value. INTERACTIONS, each scored by partial correlation against the counter residual controlling for both of its own components: none pass. Only barrel x power archetype has an incremental exceeding its parts (-0.101 against 0.019) at n=260 -- the shape Discipline 2 predicts, but a lead, not a finding. A methodological catch worth keeping. The archetype conditioner was first built as barrel_pct over league barrel -- a monotone transform of one of its own components -- so the "interaction" was barrel squared, measuring nonlinearity in barrel rate rather than any archetype effect, and it produced this run's only positive result. A Gauss-Jordan pivot test does not catch that, because the two columns differ by a scale factor. Fixed with a scale-free collinearity check plus real archetype labels joined from model_snapshots. Without it this document would have reported a fabricated interaction as the session's finding. COMBINED vs COUNTER on total bases: 0.2718 against 0.2647, delta +0.0071, CI [-0.065, +0.079] -- inconclusive, and the first time a challenger has not lost. The same engine on hits was -0.116 with a CI excluding zero. That contrast is the whole argument for total bases, and it is what the physics said: contact quality governs extra bases, not whether a grounder finds a hole. Also built: the compound TB projection. skillProjection no longer refuses total bases -- a deterministic bases-per-hit multiplier had made P(TB>=2) exactly P(hits>=1), a relabelled hits curve. It is now a convolution over per-PA base outcomes with hit-type shares shifted by skill. Non-degeneracy is locked by test. 4,204 tests green (334 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
c7cc8f5e52 |
Build the gate, run it, and find we were proving things on the wrong stat
PREMISE CORRECTION FIRST. statModel.js and correlateValidator.js do not exist in this repository. The validation spec's only prior form is src/services/python/blueprints/unconventional.py -- a Flask blueprint in the Python service that is offline in production, scoring NBA factors against a warehouse that was never populated -- and tests/unit/supplementSystems.test.js requires only fs and path while defining its own validateFactor inline at line 368. Those tests assert a re-implementation of the thresholds, not an implementation, which is exactly why they passed for months while nothing was connected. The diagnosis behind the order is right -- every challenger was measured without a gate -- but the cause is that there was no gate on the Node side to import. So it is built, to the exact spec. correlateValidator: n>=500, |r|>=0.15, p<0.05, Bonferroni across the sweep. The p-value is exact rather than approximated (t-transform through a regularized incomplete beta) and is verified in the suite against known values, because scipy is not available here. Pairs with an unknown side are dropped, never zero-filled -- a zero-fill inside a correlation does not add noise, it invents a point at the origin. THE RUN, hits, n=570, Bonferroni-8: every skill feature fails, and not narrowly. The strongest marginal correlation against the counter's residual is 0.062 against a 0.15 bar. That is an effect-size failure at a sample that would have found a real effect comfortably -- a clean, well-powered negative. The head-to-head agrees: value engine 0.0499 against the counter's 0.166, delta -0.116 with CI [-0.189, -0.043]. Not promoted. THE RUN, total bases, n=295: cannot be tested, and that is the finding. hard_hit_pct shows a marginal r of 0.153 -- above the threshold -- and exit velo 0.124, refused solely because n is 205 short of 500. It is the most encouraging number this work has produced, and it is what the physics predicts: contact quality governs extra bases, not whether a grounder finds a hole. We have been testing skill inputs on the one stat where they should not matter much. Two things the run forced. Feature verdicts are now PER STAT, because marking these DEAD sport-wide on hits evidence would have killed, for total bases, the features that look most alive there -- per-sport doctrine one level deeper. And the gate now reports r and p even when underpowered, because "not enough data yet" and "nothing here" demand opposite decisions and a bare refusal was hiding the best signal on the board. Next: build the compound TB projection (skillProjection still refuses total bases by design, since a deterministic bases-per-hit made P(TB>=2) identical to P(hits>=1)), accrue to n>=500, re-run this gate. Leave hits alone. 4,200 tests green (334 suites); web build exit 0; counter byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
258d8a6655 |
The skill engine: built, gated by construction, and Stage A honestly lost
Built src/services/model/ -- the forward, archetype-selected, skill-based projection, as a challenger. The champion is untouched. featureRegistry makes "earn its place or it's out" structural rather than aspirational: CANDIDATE / PROVEN / DEAD per feature per sport, liveFeatures() returns PROVEN only, promotion requires n>=200 with positive lift and a CI excluding zero, and there is deliberately no override argument. It ships with exactly ONE proven feature -- the incumbent counter, because it is the only one with a measurement. A test asserts that with only PROVEN features allowed the projection returns null, so an unproven model cannot reach a user by accident. The three champion adjustment layers are registered DEAD with their reasons so they cannot be silently rebuilt. skillProjection is a PA outcome tree: K and BB combined by log5 odds-ratio against league (both identities unit-tested), then archetype-weighted contact quality against contact allowed, then Binomial(PA, p_hit) mixed over a PA distribution. Archetype is a FEATURE SELECTOR, not a nudge -- BOMBER reads barrels at 0.50 and ground-ball speed at 0.00, GHOST inverts it -- and a test locks that the same hitter read two ways moves more than 0.15. STAGE A: IT LOSES. Out-of-sample on 570 settled hits props with 91.9% opposing-pitcher coverage, resolution 0.0499 against the champion's 0.166, delta -0.116 with CI [-0.189, -0.043]. It is not selective either: its eight most confident picks hit 50%, a lift of -0.065. Not promoted. The gate did its job on its first real test, which is the point of having built it that way. Two false starts, both recorded because they nearly produced a wrong verdict: statcast_aggregates stores PERCENTAGES, so raw rows made bip = 1-29.6-17.1 and refused 568 of 576 -- the honest-absent guards made a units bug loud instead of silent, and the conversion now lives at one chokepoint. And the first run resolved an opposing pitcher for 1 of 570 rows, because ledger team/opponent are NULL, so it would have reported "skill-v1 loses" while measuring a batter-only model with no matchup in it at all. The verdict above is from the corrected run. The loss is real but partial: park was passed as 1.0, handedness and opportunity_drift never fired, PA is season-PA over a constant, and the skill profiles carry no recency at all while the champion has a last-5 term. Also fixed: the Statcast nightly refresh was unreachable code. It sat inside tick() below "if (!HOURS_UTC.includes(h)) return" while testing h === 11, so it had never run once; the aggregates were 13 days stale and both of its alerts were in the same dead branch. It now runs on its own tick, and the test that passed happily throughout -- it only checked the string existed -- is replaced by one that asserts it is not behind the guard. 4,182 tests green (333 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
d8bf7765db |
Decompose the champion: its whole edge is a hit-rate counter
READ-ONLY. src/ and web/ untouched; 4,159 tests still green. WHAT THE CHAMPION IS. probabilityEstimator is five lines of arithmetic: the empirical frequency of (stat > THIS line) over the game log, blended 0.6/0.4 with the last-5 frequency, then +/-0.03 opponent, +/-0.015 home/away, a cv>0.40 pull toward 0.50, and a clamp to [0.10, 0.95]. It reads three features. featureCache retains a dozen more that p_win never touches. THE ABLATION IS EXACT, NOT A REFIT. Every adjustment is closed-form from stored features and the consistency step is linear, so each layer subtracts algebraically out of the stored p_win -- no re-estimation, no re-fetch, no lookahead possible. Per stat, paired bootstrap: removing ALL THREE adjustments changes resolution by NOTHING on every stat hits -0.0059 total_bases -0.0015 rbi +0.0106 runs +0.0130 walks +0.0008 and rbi's home/away is mildly HARMFUL (+0.0053, CI excludes zero). So ~100% of the champion's resolution is base+recency: how often this player has cleared this number lately. Everything else is decoration. A CORRECTION. Pooled, the champion resolves 0.46; per stat it is 0.196 (hits) to 0.499 (rbi). Pooling stats with different base rates inflates correlation, so 0.46 should not be quoted as the champion's resolution. Last session's paired differences remain valid; only the absolute level was inflated. THE BIGGEST LOSS IS NOT A MISSING FEATURE -- IT IS THE CLAMP. 358 of 1,741 settled rows (20.6%) sit on the boundary, so the model emits a constant there and cannot rank a fifth of the book at all. And that constant hides two opposite failures: 0.900 covers home_runs-under truly winning 99.5% (9.5pts under-confident) next to hits-under truly winning 51.9% (38.1pts over- confident). PROB_CEIL=0.95 makes the 99.5% case inexpressible. Global over-prediction is +3.5pts, +7.6 on total_bases. None of this needs new data. ONE REAL MISSING-WEIGHTING LEAD: opportunity_drift, residual corr +0.156 on hits and +0.145 on total_bases -- it REPEATS across independent stats, unlike the weather hits on TB which sit inside the expected false-positive count (70 tests at alpha .05 expects 3-4). And we already compute it: arch-v1's opportunity axis uses it and extracts nothing (delta +0.0001). Wrong implementation, not a missing feature -- opportunity must scale the rate, not nudge the probability. ARCHETYPE IS UNMEASURABLE, NOT REFUTED. Only 2 of 41 labels (BOMBER, GHOST) reach n>=40 settled rows and every mean residual straddles zero. That is "we have not measured it", and it does not license acting in either direction. Why every challenger has failed is now legible: the ladder and hits-v1 REPLACE the frequency question with a fitted distribution; the environment axis adds inputs the champion ignores. Asking the frequency question at the traded line is the thing that works. Flagged, not fixed: model_snapshots.outcome is NULL on all 22,032 rows -- the retention table built for exactly this replay was never settled, so labels had to be joined from ledger_entries. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
3ba3dd28f3 |
Scoreboard every challenger; diagnose the 429 as odds-api, not PropLine
PROMOTE-THE-EARNED. Nothing was promoted, because nothing earned it -- not because the bar was held high. Measured on the same bar that refuted hits-v1: own rows only, direction-aligned, paired bootstrap, promote only on a CI excluding zero. arch-v1 n=1741 delta 0.0000 CI[-0.0050,+0.0054] inconclusive contact-v1 n=1055 delta +0.0008 CI[-0.0052,+0.0069] inconclusive proj-v1.1 n=1664 delta -0.0301 CI[-0.0543,-0.0060] reliably WORSE matchup/tb-v1/hits-v1 n=0 genuinely pending (rows dated 08-02+) arch-v1 is the interesting one: it MOVED 76% of rows by 2.5 points on average and resolution is identical to the champion to four decimals, on the moved rows too. That is active movement carrying no information -- a finding, not a pending verdict. These are true prospective holdouts: arch-v1 and contact-v1 wrote p_win at grade time into their own columns before the game. Nothing recomputed. THE 429, read-only. The premise was that we re-pull the full picture every slot and blow the quota. Measured: PropLine is at 5 calls of 3,000/day -- 0.17%. One snapshot is ONE PropLine call per sport, all markets comma-joined. There is no request-pattern problem, so a change-based pull cannot fix it and no tier upgrade is needed. The 429 is odds-api: 478/500 MONTHLY, blocked at 95%. oddsService falls through silently when PropLine returns empty, and the backup's quota gate throws the error -- so an empty slate is indistinguishable from an outage and the message names the wrong provider. Flagged for its own order. Could NOT verify PropLine movement endpoints: docs are auth-gated and the keys are production-only. Not asserted either way. The movement-as-data argument stands on its own merits and should be justified that way, not as a quota fix it isn't. Book-breadth invariant written down: we never discard books. All are kept and shown (DISPLAY_BOOKS = MODEL + REFERENCE + DFS); DFS pick'em is excluded from PRICING only, because a fixed-payout shaded number is not a market price. Verified this is already what bookRoles.js does. Champion byte-identical; every challenger stays wired. 4,159 tests green (332 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
07626de3de |
hits-v1: built on the right structure, measured honestly, REFUTED
Hits was diagnosed as a family mismatch: 84% of hits rows trade at 0.5, so the stat rides on P(0), and a negative binomial has unbounded support and no notion of opportunity at all. hits-v1 models it as the bounded conversion it is -- N ~ the player's empirical at-bat distribution, hits|N ~ Binomial(N,q), with the multiplier scaling q (conversion) and never N (opportunity). STEP 0 confirmed the inputs before the model existed: 30/30 real ledger players, 100% combined-input coverage. Every read goes through knownRate -- a row with no atBats is dropped, never counted as a 0-at-bat game. It FIRES: 158/159 hits props (99.4%) on the live production snapshot, through the real attachProjection path. Scoping by book IDENTITY rather than price shape kept 94 out-of-promotion-band props on the board, 93 of them modelled -- 59% that a price rule would have deleted. And it LOST. Point-in-time replay (game log truncated strictly before each row's game_date, real grade-time multiplier), hits-only, direction-aligned, n=242: resolution champion 0.195 / ladder 0.048 / hits-v1 0.026. Paired bootstrap on the same rows: hits-v1 - ladder = -0.022, CI95 excluding zero. Not promoted. The value is in what it eliminates. The family was wrong AND the mean was not the constraint -- hits-v1 moved the line-0.5 mean 0.554 -> 0.581 toward a 0.598 base rate while resolution fell. What is left is per-prop discrimination: the ladder's inputs, not its distribution. The pre-registered fallback is recorded as WRONG rather than deleted. It said hits might be genuinely low-resolution for anyone; the champion scores 0.276 on the identical 189 rows, so there is real signal and the ceiling claim was the comfortable reading, not the honest one. Its own control refuted it, and that control was already in hand when the branch was written. hits-v1 stays wired as a challenger writing its own ledger columns so the forward accrual can confirm the backtest. Champion, ladder, ranking, calibration, reference ruler and the four accruing verdicts are byte-identical -- the diff has zero deleted lines. Tests 4,156 green (332 suites); web build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
e29ab6fd6a |
Takeable enforcement: verified on real rows, 1,006 tagged, re-stamp call ready
PART 1 verified by inducing the REAL rowsFromSnapshot over REAL lock_lines
rows from prod. Three cases, 0 non-takeable anchors:
Narvaez (dabble/kalshi/prizepicks/smarkets, NO takeable book)
-> book=null, price=null, takeable=null [honest absent]
Schwarber(bovada/dabble/novig/PINNACLE before draftkings)
-> draftkings +102 [pinnacle SKIPPED, proving TAKEABLE not MODEL]
Ohtani (dabble/onexbet before draftkings) -> draftkings -266
Narvaez is the case that matters: pre-fix he was stamped dabble +104
takeable=true; he is now honestly absent.
A HARNESS BUG RECORDED: my first verification pulled live /api/odds/mlb,
which returned {"error":"Odds data temporarily unavailable"}. The script
read that as 0 props and printed "all from takeable books? true" -- a
VACUOUSLY TRUE pass. I caught it only because I also printed the book list
and it was empty. Same family as the silent-false traps: a probe that finds
nothing looks identical to a probe that finds nothing wrong.
PART 2: 1,006 rows tagged via the purpose-built quarantine_reason at ROW
level with three sub-cases (recoverable_same_line 936, no_takeable_quote
49, takeable_line_differs 21). getModelAggregate ALREADY excluded
quarantined rows, so the public record and the n>=20 gate were clean
automatically; all five committed holdout scripts now carry the exclusion
explicitly.
PART 3 -- the re-stamp call is now fact-based. The takeable LOCK-TIME price
is recoverable for 936/1,006 (93.0%) from lock_lines, the correct
instrument. Only 431 appear in closing_captures, which is the wrong timing
for a lock price anyway.
LINE CONTAMINATION ANSWERED (previously unverified): the stored line
MATCHES a takeable book's line on 936 (93.0%), DIFFERS on 21 (2.1%), and is
unverifiable on 49 (4.9%) where no takeable book quoted the prop at all.
That makes it cleanly row-level: re-stamp the 936 as an honest JOIN and
recover 886 pending rows for the holdouts, or leave all 1,006 excluded.
Either way the 21 + 49 stay out -- re-stamping those would invent a lock
price, or a line, we never captured. Nothing re-stamped; Kev's call.
Gates: 4,111 tests / 330 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
|
||
|
|
eabf3b5bcf |
tb-v1: model total_bases as a compound outcome (challenger)
Current ladder (proj_p_over_line) and champion p_win are BYTE-IDENTICAL. tb-v1 writes alongside them, on total_bases props only. STEP 0 -- components confirmed on real data, not assumed. statsapi has no singles field, but hits - doubles - triples - homeRuns reproduces stored totalBases EXACTLY on a real 10-game log. So the decomposition is exact, not an approximation. THE MODEL. Each component gets its own per-game Poisson rate; TB is their weighted sum, and the PMF is built by exact convolution rather than simulated (TB support is small). It inherits the SAME combined multiplier proj-v1.1 computes, so the two models differ only in STRUCTURE. Why this is the fix: with identical mean TB of 1.0, a pure-HR hitter and a pure-singles hitter get P(TB>=4) of 0.221 vs 0.019 -- a 12x difference an NB on TB alone cannot express, because it treats one home run as four events. A test asserts that separation, and asserts P(TB>=4) for a pure-HR hitter equals P(at least one HR) exactly. INDEPENDENCE IS AN APPROXIMATION AND IS LABELLED AS ONE: a plate appearance that becomes a double cannot also become a single, so the components are weakly negatively correlated and independent Poissons slightly overstate the tail. Closer to the truth than what it replaces; not a solved problem. HONEST-ABSENT throughout: fewer than 3 usable games, or no derivable component, returns null and the prop keeps the current ladder value. An inconsistent row (hits < extra-base hits) is SKIPPED rather than clamped to zero -- clamping would invent a plausible line out of a broken one. I HIT THE Number(null)===0 TRAP IN MY OWN CODE and a test caught it: a null rate passed a naive finite check and was treated as a measured zero, which is the difference between "this player never triples" and "we do not know his triple rate". Both tbPmf and tbMean now reject null/''/boolean strictly. Holdout committed: TB ROWS ONLY (49 of 437 settled -- averaging into other stats would hide the effect) and DIRECTION-ALIGNED, since the unaligned comparison is the artifact that accounted for 41% of the ladder's apparent loss. If tb-v1 does NOT improve, the family-mismatch hypothesis is wrong and the mean/similarity branch reopens -- recorded in the query header. Migration applied: proj_tb_p_over + proj_tb_meta, NULL-meaningful. Gates: 4,104 tests / 329 suites green; next build exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |