Compare commits

..

273 Commits

Author SHA1 Message Date
builtbykev 387ae4d54e Handoff doc: ground-truth build state at 71d3b7b
Written from the repo, not summary -- every claim grep- or run-verified,
UNKNOWN where the repo cannot say.

State: HEAD and gitea main in sync at 71d3b7b, tree clean, 4,539 tests
passing, web build exit 0. Push remote is gitea, never origin.

DONE this arc, existence-verified: E1 movement strip, F9-F11 offseason hub
shell, E10 Report template, E12 /report archive, content engine + studio
API + preview, Wave-D1 motion primitives, honest served grade, archetype
doctrine. Also verified done and previously mis-boarded: D1 glyphs (39 of
83 -- every one that maps to a real archetype), A1 card token, B1 boundary
channel, book comparison, THE WIRE.

GATED with each gate named and each absence verified by 0-file grep: F5
article media and E16/F8 on the card-system reconciliation, in-season hub
IA on the content formula, E9/E15 on model accrual, E2/E6 on licensing,
E3 crown on measurement, E13 on another order. NO UNGATED WAVE-2 TARGETS
REMAIN -- the next move is a decision, not a build.

The two open decisions are stated with what each unblocks. Card-system
reconciliation is the cheapest: it frees F5 and E16/F8 together, and it
exists because I built a 1080x1350 renderer without checking whether a
designed card system existed. It did.

Accrual clock recorded at 0 eligible dates on all four items, with the
first trigger (10 calibration dates) and the attempt-floor-not-trust-floor
caveat carried forward.

Standing doctrine carried: Truth Law including designer samples and the
inverse Number(null) breach (zero is a real fact); the 83-glyph taxonomy
with 44 dormant slots; the repaired champion and the standing flag that
every prior factor verdict was measured against the broken baseline;
calibration withdrawn; unissuable A with a B+ 0.663 ceiling; both CI
guards; and compose-don't-fork, which this arc failed three times and
caught twice.

Both parallel workstreams written as seed prompts.

specs/HANDOFF-2026-08-08.md

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-08 00:10:34 -04:00
builtbykev 71d3b7b786 E10 Report issue template + E12 /report archive, to spec
PHASE 0 — the spec, read not recalled. E10: "Hybrid: dark billboard header
that survives every client, light paper body Gmail can't wreck. 600px,
stacked, no webfont dependence." Content law: "One email per slate day.
Top read, what changed, the record. Nothing else." E12: "EVERY ISSUE SHOWS
ITS OWN DAY RECORD -- THE ARCHIVE IS A LEDGER TOO."

COMPOSED, NOT FORKED. The audit had E10 as PARTIAL, not absent:
newsletterService already builds the daily report's CONTENT and lints its
voice. What was missing is the designed hybrid SHELL, so reportTemplate.js
is a template over that builder rather than a second report -- the same
call made for the movement strip, and for the same reason.

PHASE 1 — the hybrid shell is an ENGINEERING constraint, not a look, and
the tests say so: Gmail strips style blocks, Outlook ignores flexbox, and
a dark body renders as a black rectangle in several clients. Hence tables,
inline styles, 600px fixed, system fonts, no image required to read, and
the green SHIFTS from #00D4A0 to #00A57D on paper because the dark-mode
green is unreadable there.

FACT-CONTRACTED: a section whose data is absent is OMITTED and NAMED in
`omitted`, never filled. There is no code path producing a placeholder
figure. The honesty block carries the real numbers -- graded count,
cleared-ceiling count, the realized rate against baseline, and that we do
not issue A grades.

E1'S LAW TRAVELS EVEN THOUGH ITS RENDERING CANNOT. An SVG strip is not
reliable in email, so movementText carries the RULE: green only when the
move favours the read, and a flat market says FLAT · [N]D rather than
showing nothing.

NO DESIGNER SAMPLE DATA. Nabers 1,120.5, No 128, DAY RECORD 9-4 are a spec
for what a live issue renders; pasting them in would be fabrication
carrying a designer's authority and would look entirely correct. Tested.

PHASE 2 — /report is now the real archive, REPLACING the S41 redirect to
/blog. That redirect existed because the surface did not; E12 built it, so
the placeholder is correctly gone and the S41 test is updated rather than
worked around. Every row carries its own day record, and an unknown record
says UNSETTLED -- never a dash that reads as zero. Empty archive is an
honest state.

Backend: public read-only /api/report over Redis issues, plus the Next
proxy. Both surfaces registered under the reachability guard.

A test bug I made twice now: my check for forbidden sample values matched
the template's own doc block, which NAMES those values as things never to
paste. Documentation worth keeping, so both suites strip comments before
matching -- a guard that reads its own warning is not reading the code.

WAVE-2 STATUS: E1, F9-F11, E10, E12 done. Still gated -- F5 article media
and E16/F8 on the card-system reconciliation; the in-season hub IA on the
social chat's formula; E9/E15 on model; E2/E6 on licensing.

Read-only throughout; serving fingerprint unchanged including
newsletterService; accrual clock unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 22:17:59 -04:00
builtbykev 49e76068da Doctrine + E1 movement strip + F9-F11 offseason hub shell, to spec
PHASE 0 — specs/ARCHETYPE-TAXONOMY-DOCTRINE.md records the ruling as
shared law: 83 designed glyphs are the full four-sport taxonomy; a glyph
renders ONLY where its archetype is modeled and proven. 39 of 83 map to a
real archetype and are wired; the 44 unmapped are DORMANT SLOTS for
WNBA/NBA/Soccer, not a wiring gap. Wiring them would mean inventing 44
archetypes to consume artwork -- decoration presented as classification,
which is forbidden. DUAL THREAT and PAINT BOSS are modeled archetypes with
no mark: the mirror gap, flagged to the design side. When a sport's
archetypes ship, activation is a MANIFEST lookup, not new art.

PHASE 1 — E1 movement strip. The spec's own line is "the movement strip is
defined once here and reused everywhere a line has a past", so it is a
primitive, not a fourth chart.

RECONCILED RATHER THAN FORKED: lib/gradeShift.js ALREADY implements E1's
colour law -- toward/against/flat, including the direction flip that makes
an UNDER's favourable move the opposite sign of an OVER's. MovementStrip
CONSUMES buildGradeTimeline instead of reimplementing it, and a test
asserts it never redefines isUnder. GradeShift stays the grade-history
view; this is the reusable strip. That is the card-fork lesson applied
before it could happen again.

Spec laws honoured: STEPS NOT CURVES (H then V, no smoothing -- a curve
invents prices that never traded, and a test rejects any C/S/Q/T command);
green only when the move FAVOURS the read; FLAT renders as a hairline plus
FLAT · [N]D because a flat market is a finding; and too little history
says NO MOVEMENT HISTORY rather than rendering blank.

PHASE 2 — F9-F11 offseason hub shell, built from Vyndr Offseason.dc.html.
The spec's load-bearing words are used verbatim: "OUTLOOKS REPRICE ON NEWS
· NOT GAME ODDS" (an offseason number is not a game line), the QUIET WIRE
empty state ("No outlook-moving news since X. We don't manufacture
movement."), WHAT CHANGED TODAY as the hero with the countdown ambient and
top-right, the tag-colour-is-meaning row anatomy, the open -> NOW -> FAIR
triplet with the movement strip embedded, and the OUTLOOK ONLY block where
every row carries NOT GRADED.

THE DESIGN FILE'S SAMPLE DATA IS NOT IN THE COMPONENT. Wembanyama +420 ->
+330, Nabers cleared 11:42 AM, the Summer League names -- all of it is a
SPEC for what a live feed renders, and copying it in would be fabrication
carrying a designer's authority. A test asserts none of those strings
appear.

The IN-SEASON information architecture is NOT invented here. The spec
covers an offseason hub; nothing specifies how content, articles, wire and
the live slate share year-round navigation. That remains the open design
gap, and the route notes it.

Two test bugs caught and fixed: my first assertions matched my own doc
comments -- the ordering check found "WHAT CHANGED TODAY" in the header
block and the no-curves check caught the word "curve" in the sentence
explaining why curves are wrong. A guard that reads its own explanation is
not reading the render; both now strip comments first.

PHASE 3 — both surfaces registered under the reachability guard. Read-only
throughout, serving fingerprint unchanged, accrual clock unchanged at 0
eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 21:24:36 -04:00
builtbykev 60469422af Wave D1 primitives — and the audit says E1/E9 were not Wave 1
PHASE 0 — the order proposed E1 + E9 as Wave-1 and told me to follow the
audit if it disagreed. It disagrees: E9 calibration curve is WAVE D3,
gated on MODEL work ("resolve n>=20 vs N30, accrue buckets"), and E1
movement strip is WAVE D6, a large surface build. E9's gate is live right
now -- calibration is WITHDRAWN at 0 eligible dates, so the curve could
only render its empty state today. Building it would ship a component
whose entire purpose is unavailable.

PHASE 1 — three of the five real Wave D1 items were ALREADY DONE, and the
2026-07-31 audit has aged:

  D1 glyph library  audit: 38/83 wired (46%)
                    now:   COMPLETE for everything wireable -- 39 of 83
                           designed glyphs map to a real archetype, all 39
                           are wired, colours match the registry exactly
                           (0 disagreements).
  A1 card token     audit: BUILT-BUT-DRIFTED, "in only 1 file"
                    now:   BUILT-TO-SPEC -- it IS the --bg-1 token,
                           consumed by 32 files. The audit counted literal
                           hex, which is what a correctly tokenised value
                           looks like.
  B1 boundary blue  audit: PARTIAL, hex in 2 files
                    now:   BUILT -- --priced-out/#8fb2de is a token with a
                           documented colour law, 4 consumers.

The 44 unwired glyphs are NOT a wiring gap: they have no backend
archetype, so wiring them means inventing 44 archetypes to consume
artwork -- the fabrication this programme refuses. That is the 41-vs-74
scope question and it is Kev's call. Separately, 2 registry archetypes
have NO designed glyph (DUAL THREAT, PAINT BOSS) -- a design gap.

PHASE 2 — what was genuinely absent is now built. web/src/lib/motion.js:
nudge() capped at 180ms so it reads as acknowledgement rather than
latency; bootStagger capped at 240ms because uncapped, row 40 waits 1.1s
and the stagger BECOMES the latency it exists to disguise; rowHover
returns handlers not CSS so touch cannot stick a hover state; and
revealOnIntersect returns an unobserve in every path and reveals
IMMEDIATELY when there is no IntersectionObserver or motion is reduced --
content is never hidden behind a capability check.

Reduced motion is honoured, not softened. The sharpest of the 10 tests:
bootStagger under reduced motion returns opacity 1, not merely delay 0 --
if the CSS animation supplies the opacity, skipping it leaves the row
invisible forever.

PHASE 3 — the reachability guard gains a PRIMITIVES section: a module
built to be embedded must declare its exports AND name its intended
consumers, because a primitive imported by nothing is the same
built-but-unread class as an unmounted component.

WAVE-2 UNGATED: F9-F11 offseason hub, F5 article media, E10/E12 Report,
E1 movement strip. GATED: E9 + E15 on model, E16/F8 on the resolution tail
and the card-system reconciliation, E2/E6 on licensing, E13 on another
order.

Read-only throughout; serving fingerprint unchanged; accrual clock
unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 20:31:55 -04:00
builtbykev 7a318ccce0 Media/hub design inventory: the specs already exist, and so does a gap audit
READ-ONLY. Nothing designed, built or mounted.

THE HEADLINE: this inventory was largely already done. specs/
design-vs-build-gap-audit.md is a 61-item design-vs-build audit from
2026-07-31, every claim grep-verified, with a wave-ordered build plan. The
useful work is reconciling it, not redoing it.

WHAT ALREADY HAS A FULL DESIGN SPEC (and is NOT built):
- Offseason hub -- Vyndr Offseason.dc.html, 119KB: hub home desktop+390, an
  NBA Summer-League variant with an OUTLOOK ONLY / NOT GRADED honesty
  block, season-long board, season-read reveal with WHAT WOULD CHANGE THIS
  READ, news/outlook feed row anatomy, quiet-wire empty state.
- Article media (S3) -- hero template plus four hero graphic archetypes,
  all GENERATED data visuals never stock, inline figures with a one-pull-
  quote max and a caption law that every figure names its data, article
  card, OG 1200x630 with the master already rasterised.
- Movement strip (E1) -- steps not curves, green only when the move favours
  the read, FLAT as hairline. THIS CORRECTS MY 2026-08-07 BOARD, which
  listed it UNKNOWN and possibly satisfied by GradeShift. It is a specified
  primitive that does not exist.

THE WIRE is BUILT (vyndr/Ticker on real ticker exhaust) and is a DIFFERENT
thing from NewsWire.tsx, which is the offseason news/outlook feed.

THE ONE REAL DESIGN GAP: the hub's information architecture across seasons.
The Offseason file specifies an OFFSEASON hub; nothing specifies what that
surface is IN-season, or how content, articles, wire and the live slate
share one navigation. Every component exists on paper; their composition
into a year-round media surface does not.

A FINDING AGAINST MY OWN RECENT WORK: I built a card renderer at 1080x1350
without checking whether a designed card system existed. It does -- five
master sizes with defined layouts and rasterised PNGs in exports/. The
content engine's cards are a parallel invention. Not wrong, but they should
conform to E16/F8 rather than diverge, and that is a design decision.

CORRECTION: the order lists content API, preview page, book-comparison
mount and guard widening as still pending. All four shipped at c575a70,
and book comparison was never unmounted -- my earlier board grepped only
web/src/app and missed component-level mounting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 19:18:36 -04:00
builtbykev c575a708c7 Content studio API + preview page; widen the reachability guard; correct
two inventory errors

INVENTORY CORRECTION, and it was mine. Phase 2's two "orphans" are NOT
orphans -- my board grepped only web/src/app and missed component-level
mounting. The transitive check says both are already mounted:

  BookComparisonPanel -> GradeResultCard -> app/scan/page.tsx
  NewsWire            -> ExploreHub      -> app/explore/page.tsx

So book comparison is DONE (wired to /api/books, rendering on the grade
card) and THE WIRE is DONE-BY-DESIGN, mounted in ExploreHub. Its header
names an "Offseason Hub" as its home, and that hub genuinely does not
exist -- but that is board item #8, not a mounting bug, and inventing a
surface to satisfy a comment would be the wrong fix.

The lesson is the same one this session keeps teaching: I checked one
directory and reported a conclusion the check could not support.

ALSO CAUGHT: I overwrote src/routes/content.js, which was the Session-29
content-templates route, by picking a filename without looking. Restored
from git with no work lost; the new surface lives at
/api/content-studio and both now coexist.

PHASE 0/1 — /api/content-studio serves finished posts (copy, branded card,
card_svg, the fact_contract each was REQUIRED to have, and the facts that
actually backed it) plus a POST for editorial status in Redis. Private via
internal key; the Next proxy holds the key server-side so the browser
never does. /studio renders it as a thin client -- copy and card side by
side with the fact contract visible, because reviewing copy by reading it
is exactly how a wrong number ships. Never-blank: a night with nothing
generated says so.

API-FIRST is the point: the endpoint an autonomous poster will call is the
one the page already renders, so the agent handoff is a pointer change,
not a rebuild. Contract documented at docs/CONTENT-STUDIO-API.md.

EXPRESS 5 BROKE 23 SUITES at first: `router.get('/:date?')` throws at
mount time in Express 5, taking down everything that imports app.js. Two
explicit routes instead.

PHASE 3 — the reachability guard is widened from grade-fields-only to a
general built-but-unread check. Book comparison, THE WIRE and the content
studio are now registered surfaces; a page counts as its own entry point
(Next mounts it by convention) while everything else must trace to one.
22 checks green; a registered-but-unimported surface still goes red.

FULLY ISOLATED: read-only on model/slate/ledger, serving fingerprint
verified unchanged, accrual clock unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 18:51:52 -04:00
builtbykev 74aa75945e Content engine: posts that structurally cannot lie
PHASE 0 — contentEngine makes Truth Law structural, not careful. Copy is
token-substituted and an unbacked {token} REFUSES to render -- there is no
code path that produces a plausible default. The fact contract is asserted
before any string is built. Card and copy render from ONE fact object, so
a caption and a card cannot disagree. No live model writes factual claims:
the voice is in the template, the facts are pulled, and the voice-polish
port is deliberately unwired, because an LLM that can rewrite a sentence
can rewrite a number.

18 tests carry the proof. The one that matters most: ZERO IS PRESENT.
"0 cleared B+" is our most honest possible post, and treating 0 as missing
would be the Number(null)===0 breach wearing its opposite coat -- it would
silently delete exactly the post the brand is built on.

PHASE 1 — three templates, generating real posts from tonight's data:
hot hitters off the repaired full-season log, the honesty flex off the
real servedGrade distribution (2,140 graded / 70 cleared B+ / 42% not
separable / A unissuable), and streaks verified from settled outcomes only.

THE ENGINE CAUGHT A BUG IN ITSELF, and it is the sharpest lesson here. The
first run published "No hitter is meaningfully hot tonight -- we could
dress up a middling week as a streak. We don't." That was FALSE: the
box-score cache spans only the settled window, every player had under 20
games, and the pool was empty. A broken pull was publishing as considered
editorial judgement -- the fourth appearance of this class tonight and the
first where our OWN HONESTY COPY was the disguise.

Fixed structurally rather than by patching the number: an absent() variant
may now DECLINE to speak, and the template separates "no candidates at
all" (SKIP with a reason) from "candidates judged, none hot" (honest
absence). Both locked by test. Source corrected to mlbStatsAdapter.fullLog,
the same log the repaired champion reads.

PHASE 2 — cardRenderer emits SVG rather than canvas: it is text, so it
diffs in review and its numbers are greppable, which matters when the
whole claim is that the numbers are real. VYND white + R green, slashed-Y,
scanlines, mono. The card never formats its own facts -- every string
arrives pre-rendered and gate-checked.

PHASE 3 — scripts/generate-content.js writes copy + card per template to
.content-out/<date>/. Template N+1 is a registry entry: requires, pull,
copy, card, absent. Queued as stubs, not built: hot takes, daily reads,
"grades we DIDN'T give", cross-sport streak variants (the streak template
is already sport-agnostic -- settled outcomes and a noun).

FULLY ISOLATED: read-only on every source, zero writes to serving, model
or ledger tables. Serving fingerprint verified unchanged. The accrual clock
is untouched at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 16:34:57 -04:00
builtbykev 08791520fc Board: mechanical inventory of every outstanding item, read from the repo
Inventory only, no build. Every line cites a file or a query; unverifiable
items are marked UNKNOWN rather than asserted.

Corrections to the assumed board:
- Scanner is DONE on the blue-boundary format (scan/page.tsx:715-717), not
  still amber.
- Screen conversion is DONE except ONE file -- RouteStub survives only in
  app/notifications/page.tsx across 47 route dirs.
- Terminal is deliberately retired (S57 redirect), not debt.
- ev_pct DOES render, in 5 frontend files. That flag was stale.
- edge_pct is NOT a scale bug: (model-line)/line is correct arithmetic that
  explodes on small lines (a 5.5 projection on a 0.5 line is a legitimate
  1000%). Live top values 900/860/700 on 68,364 rows. It is a display
  defect, not a broken computation, and it reaches a user only via
  alt_lines typing -- the grade card computes edge independently.
- All FOUR sports have real archetype registries (nba 15, mlb 15, soccer 6,
  wnba 5) and ACTIVE_SPORTS runs all four. Only MLB BATTERS is MODELED;
  everything else is scaffolding with a sound base rate and no proven
  factors. MLB pitchers: pitcherEngine.js exists at 261 lines and is read
  by ZERO serving code.
- Book comparison is built and mounted NOWHERE -- the same built-but-unread
  class the reachability guard exists for, and NOT covered by it, since
  that contract is grade fields only.

UNKNOWN and marked as such: opp_rank_stat prod population, internal-key
rotation, the ~9% void/DNP rate, and whether MovementStrip is a distinct
spec item or satisfied by GradeShift/MarketBreadth.

The ordered board puts five PARTIAL items first and flags accrual
sensitivity: seven items are isolated design/guard work buildable today
with zero clock impact; five touch the serving path and would muddy the
accruing repaired-champion dates. edge_pct is the sharpest tension -- a
real honesty defect a user can see, whose fix touches serving mid-accrual.

Clock today: 0 eligible dates, all four re-audit items WAITING.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 15:13:02 -04:00
builtbykev 55b210cb95 Fix the dormant basketball window-bug before it ships; guard the class
PHASE 0 — audit. espnStatsAdapter's slice(0,20) was already fixed at
494c83c, so the ESPN branch of getStatRows inherits a full log and is
CORRECT. MLB's l20_avg is a real seasonTotal/games aggregate, verified,
also CORRECT. TWO basketball defects remained:

  DEFECTIVE  nbaGameLogFeatures   m20 = avg(vals.slice(0, 20))
             -- l20_avg is the season reference projectionFor reads, and
             slicing to 20 made it a twenty-game average wearing a season
             label. The MLB l20 bug, unfixed for basketball.
  DEFECTIVE  getStatRows python   getGameLogs(playerName, sp, 20)

PHASE 1 — both fixed via a named SEASON_LOG_DEPTH = 100, past an 82-game
season so a request can never truncate one. API COST: ZERO. The count is a
request parameter, so asking for a season is the same single call. No
extra request, no extra quota.

MEASUREMENT DEFERRED, EXPLICITLY: basketball is offline, there are no
settled basketball rows, and resolution before/after CANNOT be measured
now. This is a code fix, not a measured improvement -- exactly like the
deferred MLB sibling paths.

PHASE 2 — the window guard asserts the property on source: no fixed N may
stand in for a season. It immediately caught TWO MORE instances I had
missed in Phase 1 -- a hardcoded 20 at featureCache:377 and
gameLogService's own `count = 20` DEFAULT, which would have handed a
twenty-game window to any caller that omitted the argument. That is a
seventh path, found by the guard rather than by me.

Retro-confirmed: run unchanged against 981a05c it goes 3 failed / 6
passed, flagging the basketball slice, the game-log request and the
season-reference check.

The class in full, now six paths across two sports, every one of which
looked like ordinary code. `slice(0, 20)` is unremarkable; what made it a
defect was the QUESTION it answered -- "what is this player's season
rate?" -- and no test could see that mismatch because the value produced
was always a plausible number.

PHASE 3 — THIS CLOSES THE NON-ACCRUAL ARC. Everything buildable without
settled rows is built: push unblocked (it was the wrong remote, not the
firewall), grade surface honest and rendering, reachability guarded,
champion repaired across every path including dormant basketball, and the
window class guarded so it cannot return.

The program is now correctly IDLE on modeling. Today's count: 0 eligible
dates, all four re-audit items WAITING. FIRST TRIGGER: 10 eligible
calibration dates on the repaired champion, at which point the resumption
order is calibration re-fit, hits factor lift, prior verdict re-audit,
rbi lineup-slot gate.

No basketball measurement claimed. No NBA chain/archetype build. MLB
serving verified byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 15:00:24 -04:00
builtbykev 981a05cbd6 Render-reachability guard: make built-but-unread a CI failure
Three consecutive orders shipped a backend-correct field that never
reached a screen, all on a green suite: gradeBands (required by no
serving code), served_grade (dropped at the adapter boundary),
GradeScaleLegend (imported by nothing). Each was caught by luck on a later
re-check, and in two of the three I had already reported the wiring done.

WHY GREEN TESTS COULD NOT SEE IT: backend tests stop at the API payload.
They prove a field is PRODUCED and say nothing about whether it is
CONSUMED. Invisible by construction, not an oversight in any one test.

THE TRAP, NAMED: the difficulty pools in the backend, so by the time a
field exists on the payload it feels finished. What remains is a
three-line adapter change nobody considers worth verifying, so it gets
claimed rather than traced. The last inch is the one with no friction,
which is exactly why it gets skipped. "I added the field" and "a user can
see it" are different claims and only the first is fun.

THE GUARD traces each promised field the whole way: payload -> adapter
consumes -> component renders -> component is MOUNTED. Mounted is
transitive to a Next entry point (page/layout/template), the only thing
that puts a pixel on screen, depth-limited so an import cycle cannot hang
the suite.

Container rows are exempted EXPLICITLY, not silently: served_grade carries
container:true plus a rendersVia list, and a separate assertion checks
every named part actually renders. The exemption is auditable and cannot
hide an unrendered field.

The guard tests itself -- an orphan component must report unmounted, and
the contract must be non-empty, since an empty contract passing vacuously
is how this would most plausibly rot.

RETRO-PROOF: run unchanged against 3591c76, before the wiring, it goes
11 failed / 8 passed and names the exact bugs -- "the ceiling stance /
grade scale legend - its component is MOUNTED, not merely written", "the
served grade object - the adapter consumes it", "whether the band
separates from the baseline - a component actually renders it". Green on
the current tree.

HONEST SCOPE LIMIT: gradeBands is NOT in the contract and would not be
caught. It is a backend module, not a promised user-facing field, and it
is correctly unwired -- every band collapses to base-rate at current
resolution. Out of scope by design, not oversight.

Now in the standing suite, so the three-gate floor is tests green
(including reachability) + build exit 0 + fingerprint. No serving or model
change. p_win never mutated. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:44:17 -04:00
builtbykev 09186ea609 Wire the composed surface: the honest fields now reach a user
PHASE 0 re-check caught the same failure a THIRD time, mine again. Last
turn I created GradeScaleLegend.tsx and never mounted it, and I reported
that the separation flag "renders per band" -- it did not. grep: zero
frontend references to separates_from_base_rate, served_grade or
factor_adjustment. The honest grade was reaching the API payload and dying
at the adapter boundary.

That is three occurrences in three orders of the same shape: built,
correct, unread. gradeBands, then served_grade, then the legend.

PHASE 2 — the composed surface is wired end to end:
  analyzeViaEngine1 -> served_grade + factor_adjustment on the payload
  scan/page.tsx     -> ScanResponse types them and forwards them
  gradeAdapter      -> gradeMeaning, separatesFromBaseRate,
                       bandRealizedRate, factorsApplied, refusalReason
  GradeResultCard   -> renders "WHAT THIS GRADE MEANS"

What a user now sees that they could not before: what the band has
actually realized, an amber note when the read CANNOT be separated from
the baseline, and -- only where a factor actually fired, with its proven
sign -- what moved the read, in plain language rather than feature names
("where he hits it vs who is standing there").

No narrative on props where nothing fired: factorsApplied is empty and the
block self-hides. Only the three PROVEN hits factors have labels, so an
unproven factor cannot acquire prose by being added to the map.

PHASE 1 — GradeScaleLegend is now MOUNTED on the grade card (compact). The
ceiling is a stated position where the grade is, not a page a user would
have to find.

PHASE 3 hand-verified through the real adapter across twelve states: B+
with 3/0 factors, B with 1, C+/C/C- all flagged not-separable, D, F, a
switch-hitter case where 2 of 3 factors fire, two refusals and a
no-forecast. never-blank PASS, no-manufactured-A PASS, no-narrative-when-
nothing-fired PASS, separation-flag-reaches-card PASS.

No A-threshold loosening. No calibrated number leaks. p_win never mutated.
engine_grade still read by zero serving code. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:36:03 -04:00
builtbykev 3591c7626e Total grade cutover + the ceiling stated as a position
PHASE 0 caught my own repeat of the failure I diagnosed one order ago.
91927a4 attached `served_grade` BESIDE the old letter and left `grade`
alone -- so the honest grade reached nobody, exactly as gradeBands had
been built-correct-and-unread. grep showed served_grade appearing in one
file (where I set it) and all 14+ consumers -- scan route, dashboard,
parlay, newsletter, desk, content templates, retention -- still reading
`.grade`, i.e. still the dishonest letter.

CUTOVER IS NOW TOTAL: legacy.grade IS the honest letter. Overwriting the
one field every consumer already reads cuts every surface over at once
instead of editing fourteen call sites and missing one. engine1's index is
preserved as `engine_grade` and verified read by ZERO serving code.

Confidence follows the letter: it came from a grade-band midpoint of the
OLD letter, so leaving it would have paired a served B+ with a C's
confidence. Both now derive from p_win, kept on the existing 0-100 scale.

MEASURED BLAST RADIUS before shipping: 303 of 47,991 non-refused
snapshots (0.6%) have a grade but no p_win, and now render NO READ instead
of a letter. That is correct -- their old letter came from the retired
index carrying 0.48% resolution, i.e. noise -- and NO READ is a rendered
state with a reason, so never-blank holds.

PHASE 1 — the ceiling is now a STATED POSITION, not a confusing absence.
servedGrade.SCALE_LEGEND plus web GradeScaleLegend.tsx say it plainly: we
do not issue A grades, no band has hit at a rate that would justify one,
our honest ceiling is a strong B+ (~66% realized vs ~60% baseline), and if
the model earns an A the legend changes and we say why. The
separates_from_base_rate flag renders per band -- C+/C/C- are labelled
"we cannot separate this from the baseline", which is most of any slate.

PHASE 3 hand-verified across every state: B+ with 3 factors (basis
forecast_plus_matchup_factors), B+ with none (forecast_only), C flagged
not-separable, F, and three refusal states rendering NO READ with reasons.
never-blank PASS, no-manufactured-A PASS.

Test fallout was real and is documented rather than papered over: engine
BEHAVIOUR assertions moved to engine_grade, suppression assertions stayed
on grade (a suppressed prop has no letter either way), and the confidence
78 -> 95 change is the grade-band midpoint being replaced by p_win.

No A-threshold loosening. No calibrated number leaks (deployed set empty).
p_win never mutated. Ten frozen modules verified unchanged including
engine1 and probabilityEstimator. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:07:02 -04:00
builtbykev 91927a4a8a Serve an honest grade: the letter was carrying 1/6 the information of the
number beside it

PHASE 0 corrects the order's premise. A grade letter has been served all
along -- engine1.gradeProp builds it from an additive factor index,
computed INDEPENDENTLY of p_win. gradeBands is orphaned for a different
reason than assumed: it defines what a letter MEANS from realized
outcomes, and every band collapses to base-rate at current resolution.

The measurement that changed this order, on 3,417 settled props:

  grade  n      realized  mean p_win
  A         8    0.500      0.647     <- the TOP grade did WORST
  B       985    0.640      0.700
  C     1,695    0.602      0.676
  D       303    0.558      0.604
  F       426    0.535      0.588

  letter resolution 0.00116 (0.48% of variance)
  p_win  resolution 0.00715 (2.98%)
  -> the letter carried 0.16x the information of the number beside it

Concretely, from the hand-verify: Christian Encarnacion's 0.95 over
graded C and his 0.05 under ALSO graded C -- same hitter, opposite
forecasts, same letter. The gap was never that grades don't ship; it is
that the weaker of two available signals shipped as the headline.

PHASE 1 — model/servedGrade.js derives the letter from p_win with bands
anchored on MEASURED realized rates (B+ 0.663 / B 0.646 / C+ 0.615 /
C 0.589 / C- 0.548 / D 0.512 / F 0.447, base 0.6005).

NO MANUFACTURED A, structurally: A+/A/A- are UNISSUABLE, not rare. The
realized rate plateaus at 0.65-0.68 above p_win 0.70, so no band has
earned a top letter; a test sweeps every p_win 0..1 and asserts none
produces one. Even 0.99 tops out at B+ with its realized 0.663 attached.
Raising that ceiling later is a deliberate, visible act.

Bands that cannot separate SAY so -- C+/C/C- carry
separates_from_base_rate false and copy naming it, which is the honest
description of a forecast explaining 3% of variance. Every grade states
its basis (forecast_only vs forecast_plus_matchup_factors, naming which
factors fired) and calibrated:false. engine1.grade is preserved as
engine_grade so nothing downstream breaks.

PHASE 2 — refusals render real states: insufficient_data -> "not enough
history to call this one"; juiced_no_edge -> "the book has priced the vig
past any edge on this side". 1,870 refused snapshots carry exactly those
two reasons and both now surface.

PHASE 3 — hand-verified on 12 real served props. Freeman/Rice/Encarnacion
0.95 overs now B+ (was B, C, B); the 0.05 unders now F (was C). Refused
doubles render NO READ with their reason. never-blank PASS,
no-manufactured-A PASS.

Serving change; nine frozen model modules unchanged including engine1;
p_win never mutated; no calibrated number leaks (deployed set empty); no
Bonferroni slot.

STILL TRUE: the forecast explains ~3% of outcome variance. This order did
not make the model better. It made the letter stop overstating it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 04:53:02 -04:00
builtbykev 9159b7e1b9 Diagnose the push block: it was the wrong remote, not the firewall
The GATE-0 theory is REJECTED for VYNDR, tested rather than assumed. Both
remotes are ALREADY HTTPS -- there is no git@ SSH URL in this repo to be
blocked -- and both hosts answer on 443: git.builtbykev.com HTTP 200 in
0.64s, github.com HTTP 200 in 0.17s, Gitea's git endpoint 200, GitHub's
401 (auth required, reachable). No connectivity failure of any kind.
COLYRA's HTTPS-remote fix was right for COLYRA; VYNDR was already in the
state that fix produces, so applying it would have minted a new token to
solve a problem that did not exist -- and the pre-existing credential
would have made the new token look like the cure.

The real cause is in the error text, which named it exactly and was
misread all session: "could not read Username for 'https://github.com'".
Every attempt used `git push origin main`, and origin is GitHub
(kev3109/betonblk) with no stored credential. The WORKING remote is
`gitea` (git.builtbykev.com/builtbykev/vyndr.git), and a valid credential
for it sat in ~/.git-credentials the entire time. The habit of typing
`origin` is what kept twenty commits local. Nothing was blocked.

FIX: git push gitea main. Pushed 6452926..ecf78b9, 21 commits, verified by
git ls-remote matching local HEAD exactly.

No firewall rule touched, no token created, no code or model change --
infra only. The bundle and patch series stay as belt-and-braces; they are
no longer the only copy.

docs/GIT-PUSH-DIAGNOSIS.md records it so no future session re-reaches for
an infrastructure theory when the error message names a credential and a
specific host.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 04:24:21 -04:00
builtbykev ecf78b911c Accrual watch + pre-registered resumption; the program is idle on modeling
PHASE 1 — the existence risk is resolved. Nineteen commits were local-only
with no push credentials. Three artifacts now exist off the working tree:

  ~/vyndr-full-history-2026-08-07.bundle   6.8M  SELF-CONTAINED, clone it
  ~/vyndr-session-2026-08-07.bundle        244K  needs existing history
  ~/vyndr-session-patches/  (20 patches)   1.5M

Use the full-history bundle -- `git bundle verify` shows the session bundle
requires ref 6452926, so it only applies onto a repo that already has this
history. The full one clones standalone. Copying one off the machine is
the top non-accrual action item.

PHASE 2 — the watch, today's starting line: 0 eligible rows, 0 eligible
dates, everything WAITING. Zero is correct, since the repair ships in this
session's commits and no settled row can carry the marker yet.

THRESHOLDS ARE ATTEMPT FLOORS, NOT TRUST FLOORS, and this is written into
the ledger so a future session cannot misread it. Reaching 10 dates means
the calibration re-fit CAN BE MEASURED, not that it is trustworthy -- we
lived that distinction tonight with 2-4 cluster CIs, a LODO gate at
1.4-9.3% power, and a >=40 bar that was right for promotion and wrong for
deploy. A 10-date map is thin, deploys PROVISIONAL with auto-demotion, and
its interval will still be wide.

Real-time estimates stated so the wait reads as designed: ~2 weeks for
calibration and hits lift, 3+ weeks for the verdict re-audit and the rbi
gate, since rbi accrues slowest and a date only counts once its props
settle.

PHASE 3 — resumption order fixed, triggered purely by date thresholds:
calibration re-fit @10, hits composed-lift re-measure @10 (the 1.39%
figure is VOID, direction unknown), prior verdict re-audit @14 (not
pre-priced), rbi lineup-slot two-part gate @14. FIRST TRIGGER: when
reAuditEligibility reports 10 eligible calibration dates.

Until then the program is honestly IDLE on modeling. That is the correct
state, not a gap to fill -- the alternative is measuring on reconstructions
of a retired forecast, which this program has now refused by name three
times.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:50:53 -04:00
builtbykev 2391574f00 Live-surface integrity on the repaired champion; fix the label my own
repair falsified

PHASE 0 — four checks PASS, one defect found and fixed.

PASS  gradeBands is required by NO serving code -- built across several
      orders, never wired. No stale band derived from the retired
      ten-game champion can reach a user, because none reaches a user
      at all.
PASS  CALIBRATION_DEPLOYED is [] and the calibrate loop iterates it, so
      calibrate() is never called and p_win_calibrated is never set. The
      only assignment site sits inside that empty loop. No withdrawn map
      can leak.
PASS  projectionFor reads l20_avg, which mlbGameLogFeatures now builds
      from fullLog -- so refusals are computed on the repaired
      full-window reference, not the retired ten-game one.
PASS  factors still fire with correct sign across the repaired base
      range (0.35/0.50/0.65/0.80): defense lowers, pitcher-contact
      raises, platoon raises at every point. Mechanical firing check
      only -- NOT a lift re-measurement, which waits for accrual.

DEFECT FIXED — my own repair falsified a user-facing sentence. The grade
card rendered "Last 20 games average: X" from l20_avg, and l20_avg is now
a FULL SEASON average. The number changed and the label did not, so the
surface was stating something the data no longer supported. Copy now reads
"Season average"; trapDetection's L20 explanations likewise. The field
name is kept -- it is read in many places -- but no rendered sentence
claims a window that isn't there.

That is the same class as everything else tonight, one layer out: a
correct-looking string describing data that moved underneath it.

Serving change; frozen model modules unchanged; p_win never mutated; no
Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:49:54 -04:00
builtbykev 494c83cf76 Hunt the window-bug class: three more paths, and the forward re-audit rule
in code

PHASE 0 — getStatRows is the single base-rate path, so every branch is
audited, plus the feature builders since l20_avg is the season reference
projectionFor reads:

  getStatRows MLB -> estimator base    fullLog            CORRECT (929fd81)
  mlbGameLogFeatures l5/l10/l20        last10 = 10        DEFECTIVE
  espnStatsAdapter.parseGameLog        slice(0,20)        DEFECTIVE
  getStatRows NBA/WNBA ESPN branch     inherits 20-cap    DEFECTIVE via source
  getStatRows NBA/WNBA python branch   getGameLogs(...,20) dormant (offline)
  pitcherEngine / skillProjection      statcast profiles  N/A
  pitcher props via getStatRows MLB    fullLog            CORRECT
  settleSource                         full log (S64)     CORRECT

THE PITCHER ANSWER IS GOOD NEWS: pitcher props run through the same
getStatRows MLB branch, so 929fd81 repaired them too. There is no separate
defective pitcher base-rate path.

THE ONE HIDING IN PLAIN SIGHT: mlbGameLogFeatures carries the comment
"l20 = all available (the season per-game reference projectionFor needs)"
while building from last10 -- so l20_avg was a TEN-GAME AVERAGE WEARING A
SEASON LABEL, feeding both the consistency pull inside the estimator and
projectionFor, which decides refusals. It survived the previous repair
because that fix touched only getStatRows.

PHASE 1 — mlbGameLogFeatures now reads fullLog; espnStatsAdapter drops its
slice(0,20) cap. ZERO new API calls on both: each widens data already
fetched and then discarded, the same shape as the original repair. The
python branch is left alone -- the service is offline in prod and fixing it
would be speculative.

Their before/after resolution is NOT measured, deliberately: the only way
to measure today is to reconstruct the repaired forecast over old rows,
which is the reconstruction-vs-served trap this order refuses. Code fix
now, measurement at accrual.

PHASE 2 — MODEL_VERSION bumped to engine1@2026-08-07-fullwindow, so every
forward snapshot is self-identifying (retentionService already stamps it;
no new plumbing). model/reAuditEligibility.js encodes the rule: isEligible
accepts only the repaired marker, assess counts eligible DATES not rows,
and ACCRUAL is frozen at calibration 10 / hits-lift 10 / verdict-reaudit
14 / rbi-gate 14. A test locks the invisible case -- a MIXED table of 330
rows with 30 repaired returns eligible_dates 3, not 330 rows of false
confidence. Once both generations share a table a naive count would fit a
map on a blend of two forecasters.

PHASE 3 — the board, each consequence labelled: calibration WITHDRAWN
(refits at 10 dates, never on reconstructions); factor verdicts SUSPECT
(all measured against a champion worse than a frequency table, direction
UNKNOWN, not pre-priced, 14 dates); hits factor lift UN-REMEASURABLE (10
dates, factors still wired and transmitting); rbi lineup-slot RE-QUEUED
(14 dates). Pre-registered order: calibration, hits lift, verdict
re-audit, rbi gate.

Then STOP and accrue. Nothing further can be honestly measured until the
board fills with rows the repaired champion produced.

Serving-path changes by design for the MLB feature path and NBA/WNBA logs;
eleven frozen model modules verified unchanged. p_win never mutated. No
Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:40:20 -04:00
builtbykev 929fd81940 Repair the champion: it was reading ten games, not a season
PHASE 0 — the defect is real past the peek. Against a FAIR point-in-time
baseline (each player's rate over games strictly before that date, >=10
prior games, box scores back to 05-01), the served champion LOSES on all
four stats, three of four CIs excluding zero:

  hits  0.00251 vs 0.00774  CI [-0.0074,-0.0011]
  TB    0.00393 vs 0.00619  CI [-0.0055,-0.0003]
  rbi   0.02481 vs 0.03133  CI [-0.0153,-0.0005]
  runs  0.00181 vs 0.00683  CI [-0.0114,+0.0008]

PHASE 1 — the cause is the WINDOW, not the weights. estimateProbability
builds its base rate as the frequency over every row it is handed, and
featureCache.getStatRows handed it res.last10. So the "season rate" was a
TEN-GAME rate, and 0.4 of the forecast was the last five OF THOSE TEN. The
0.40 recency weight costs resolution on all four stats (-0.00086,
-0.00107, -0.00562, -0.00365). Nudges are mixed and small -- harmful on
hits and rbi, marginally helpful on TB and runs -- so they are left alone.

PHASE 2 — two lines, no new data, no extra API call, because fullLog was
already fetched by the same adapter call that produced last10:
getStatRows now reads fullLog, and RECENCY_WEIGHT goes 0.40 -> 0.20.

  hits  0.00251 -> 0.00817  (tripled; now above the fair baseline)
  TB    0.00393 -> 0.00734  (above baseline; vs old CI [0.0020,0.0067])
  rbi   0.02481 -> 0.02727  (still below baseline, CI includes zero)
  runs  0.00181 -> 0.00436  (still below baseline, CI includes zero)

Gate stated exactly: hits and TB now exceed the fair baseline on the point
estimate; rbi and runs remain below but EVERY CI now includes zero, so no
stat reliably loses to a frequency table. That is a tie on rbi/runs, not a
win, and it is reported as one. Only TB's improvement over the old
champion is CI-confirmed; the rest are directional.

STALE-FIT GATE: CALIBRATION_DEPLOYED is now EMPTY. The low-param maps were
fitted on the retired forecast and fromLedger cannot rescue them -- settled
ledger rows still carry OLD p_win, so refitting today would refit the
retired forecast. Nothing is served calibrated until dates settle under
the repaired champion, and the favourite-longshot bias must be re-measured
rather than assumed to survive. The shadow duel is void.

PHASE 3 — the hits factor lift is NOT re-measured, and cannot be yet: it
needs settled rows produced BY the repaired champion, which ships in this
commit. Replaying would score the factors against a reconstruction rather
than the served forecast. Deferred, explicitly. The factors remain wired
and transmitting; only their lift is unquantified on the new baseline.

PHASE 4 — standing flag, and it is large: EVERY factor verdict in this
programme, every null and every THEATER, was measured against a champion
worse than a frequency table. Signal added to noise reads as noise. Prior
verdicts may deserve re-audit. Logged, not re-run.

Re-queued not built: rbi lineup-slot / RISP opportunity through the
two-part gate, now landing on a repaired champion.

Serving-path change by design; the byte-identical invariant inverted and
all four stats move. Nine frozen model modules verified unchanged. No
Bonferroni slot -- resolution accounting on the champion's own knobs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:28:33 -04:00
builtbykev 65ca6493db Decompose the rbi anomaly: it is lineup ROLE, and the counter is
out-resolved by a frequency table on three of four stats

PHASE 0 — the 14.51% is REAL. Re-derived with a paged pull asserted
against an exact count (rbi 7,930 == 7,930; hits 11,690; TB 12,086; runs
6,440), since this harness produced a false null three times tonight. rbi
resolution 0.03268 reproduces, deciles are monotone through the middle,
and 20 raw rows are in the artifact for hand audit.

CAVEAT GOVERNING EVERYTHING BELOW: the naive forecasts are leave-one-out
ON THE EVALUATION WINDOW, so they see the rows they are scored on while
the model is strictly point-in-time. They are upper bounds on available
resolution, not fair competitors, and every comparison is read that way.

PHASE 1 — the split:

  stat  MODEL    (a)player-base  (b)lineup-slot  (c)within-stratum
  rbi   0.03268     0.01167         0.03608          0.01908
  hits  0.00252     0.00446         0.00100          0.00473
  TB    0.00442     0.01331         0.03448          0.00607
  runs  0.00130     0.00262         0.01156          0.01170

FINDING 1 — rbi's resolution is LINEUP ROLE almost exactly. Batting-order
slot alone resolves 0.03608 against the model's 0.03268. A single integer
accounts for the whole anomaly and slightly more. That is opportunity, not
skill -- the cleanup hitter bats with runners on. 36% is matched by player
identity alone. Within similar-base-rate strata the model still resolves
0.01908, 58% of its total and higher than any other stat's ENTIRE model
resolution, so genuine within-role discrimination exists on top.

FINDING 2 — on three of four stats the model is beaten by "he's a .270
hitter". Base-rate-only out-resolves the model 1.8x on hits, 3.0x on TB,
2.0x on runs. Even allowing for the window-peeking advantage, a 1.8-3.0x
gap is not explained by that alone: the served counter appears to DESTROY
discrimination relative to the player's own rate. rbi is the one stat
where the model beats the naive baseline.

FINDING 3 — lineup slot out-resolves the MODEL on three stats: TB 7.8x,
runs 8.9x, rbi 1.1x. Hits is the only stat where batting order carries
less, which is mechanically right -- a hit is a hit wherever you bat, but
runs, RBI and total bases all scale with opportunity.

PHASE 2 — all three worlds are partly true, in measured proportions.
World A ~90% true (slot covers rbi's entire resolution). World B ~36% true
for rbi, but the WHOLE story for hits/TB/runs where base rate alone wins.
World C true with a low ceiling: hits' total available spread resolution
is 0.00446, i.e. 1.8% of variance from a forecast that has seen the
answers.

PHASE 3 — the next arc is NOT "strengthen hits factors". Hits has the
lowest available resolution on the board and last order's wiring already
took it to 1.39% of a ~1.8% ceiling. Named first factor order for next
session: LINEUP SLOT / RISP OPPORTUNITY on rbi through the two-part gate --
input already ingested and prod-verified (S89), resolution measured not
hypothesised, causally-correct unit is plate appearances with runners on.
Measured availability is not a pass; it still faces the gate.

And higher-value than either: the counter being out-resolved by a
frequency table on three of four stats is a defect in the CHAMPION, not a
factor problem, and it costs nothing to test -- the recency blend and the
+/-0.03 / +/-0.015 nudges are three lines in probabilityEstimator.

The hits transmission win from 43f65d3 stands: the conduit is real and
permanent. This order changes only which stat has the most worth flowing
through it.

Diagnostic only -- no factor wired, no serving path changed, p_win
untouched, all frozen modules byte-identical. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:09:44 -04:00
builtbykev 43f65d30cb Wire the three proven hits factors pre-grade: transmission proven, gain
inconclusive

THE BUG THIS NEARLY SHIPPED AS A FINDING. The first audit reported 0
factors fired on all 1,140 rows. Not a result -- my paging helper ordered
by `id`, and batter_spray, team_defense, platoon_splits and
statcast_aggregates have composite primary keys with NO id column. The
query errored, the loop broke on error, and four fully-populated tables
read as empty. hitsFactorContext.js -- the PRODUCTION loader -- had the
identical defect, so live wiring would have loaded nothing and served
unadjusted while logging success. Third occurrence of this class in one
session. Both loaders now order by a real column and THROW rather than
degrade. The Phase 2 gate is what caught it: no resolution number was
quoted until transmission was proved.

PHASE 1 — pipeline is now base -> FACTORS -> CALIBRATE -> GRADE. Context
built in snapshotService BEFORE gradeAndCacheSlate (was line 640+, grade
at 454), threaded per prop, applied to p_over before p_win is set with
p_win_prefactor and a full trace retained. Hits only. Coverage 859/1140
rows (75%): 474 with all three factors, 256 two, 129 one, 281 none.

PHASE 2 — TRANSMISSION PROVEN, 12/12 sign-correct, 4/4 per factor, each
applied IN ISOLATION. My first table compared each factor's expected sign
against the COMPOSITE change and showed 3 false failures -- with three
factors firing the net can oppose any single member; that was a flaw in
the test, not the wiring. Two under-side rows confirm the flip is handled:
a factor raising p(over) correctly lowers p_win. Switch hitters (Bailey,
Bell, Rocchio) took no spray adjustment while their other factors fired
normally -- the refusal is selective, not a blanket skip.

PHASE 3/4 — both maps refit on the factor-adjusted forecast; the
shadow-duel baseline is VOID and restarts, since it accumulated against a
different forecast. Point-in-time, 765 held-out rows:

  reliability 0.00795 -> 0.00828
  RESOLUTION  0.00229 -> 0.00345   (variance explained 0.93% -> 1.39%)
  Brier       0.25398 -> 0.25305   delta -0.00093  CI [-0.00225,+0.00002]

Resolution rose 51% relative. The CI TOUCHES ZERO on 4 eval dates, so the
composition does NOT earn a proven keep -- three isolated passes did not
grant a composed pass. INCONCLUSIVE, reported as such. The gain is far
below the sum of the isolated effects, which is expected: all three run
through the same pitcher-batter confrontation and share signal.

PHASE 5 — 1.39% of variance is still far below what band separation
needs. The pivot was correct and incomplete: the plumbing defect was real
and is fixed, three proven factors reach the served number for the first
time, and transmission alone did not buy grade separation. Next arc is
factor STRENGTH and BREADTH, not more plumbing.

PHASE 6 — rbi anomaly logged, not chased: 14.51% variance explained vs
hits 1.03%, on the stat we do not serve corrected and which has no proven
factors. Either the biggest lever on the board or a mirage; it deserves
its own order.

The byte-identical invariant INVERTED for hits by design. All 13 frozen
non-hits modules verified unchanged, probabilityEstimator included -- the
factors ride outside it. No new Bonferroni slot; the composed OOS claim is
reported with its CI and not claimed as a pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 02:53:48 -04:00
builtbykev e872eff4ce Instrument the calibration duel forward; diagnose the resolution ceiling
— the proven factors were never wired in

PHASE 0 — two truths recorded. The swap is a BET, not an OOS win:
isotonic beat low-param on identical held-out rows (hits +0.0028, rbi
+0.0042, TB tied) and we serve low-param anyway on an untestable prior
about shared daily structure. At 19 dates nothing here can test it. And
the MIN_SLOPE catch is preserved as standing rationale: a near-zero or
negative slope collapses toward base-rate-for-everything, which LOWERS
Brier while destroying all resolution -- a metric win that guts the
product.

PHASE 1 — the duel is now falsifiable. Both corrections computed on every
hits/TB prop; p_win_lowparam served, p_win_isotonic_shadow logged in its
own try so it can never break serving. calibrationDuel.adjudicate encodes
the rule IN CODE before any forward date exists: >=10 forward dates and
isotonic winning with a date-block CI excluding zero => REFUTED, revert;
otherwise UPHELD; under 10 dates PENDING regardless of the numbers. A
date counts as forward only if NEITHER map was fitted on it -- otherwise
we would be scoring which map memorised better. Nothing swaps now.

PHASE 2 — the ceiling, quantified via Murphy decomposition:

  stat   reliability  RESOLUTION  uncertainty  variance explained
  hits      0.01353     0.00252      0.24532        1.03%
  TB        0.01419     0.00442      0.24329        1.82%
  rbi       0.00654     0.03268      0.22531       14.51%
  runs      0.00788     0.00130      0.23182        0.56%

Calibration did exactly what theory says and nothing more: hits
reliability 0.01353 -> 0.00233 (-0.0112, 83% of the error removed) while
resolution moved -0.0002. Unexpected: rbi has 13x the resolution of hits
and is the one stat we do NOT serve corrected -- it needs calibration
least and discriminates most.

PHASE 2 DIAGNOSIS — NOT-TRANSMITTED, and not weak, ABSENT. Traced in code:
sprayDefense.js and platoonSeverity.js are required by NOTHING in src/,
only by analysis scripts and their own tests. The served p_win
(intelligence/probabilityEstimator.js:54) reads exactly four inputs --
game-log frequency, opp_rank_stat +/-0.03, home_away +/-0.015, and a cv
pull -- with zero occurrences of spray, platoon, hard-hit or
contact-profile. And snapshotService grades at line 454 while computing
challenger/context at 640+, so everything proven is computed DOWNSTREAM of
the grade it would inform. The three proven hits factors have never once
moved a served number.

That reframes the recent nulls: "calibrated p_win does not separate within
archetype" was never a statement about factors. The factors were not in
the forecast.

PHASE 3 — bands rebuilt on SERVED values (hits/TB low-param, rbi/runs
raw): 28 archetype slots across four stats, ZERO show lift. No longer an
open shrug -- it is the arithmetic of resolution 0.0013-0.0327 against
uncertainty ~0.23. A forecast explaining 1% of variance cannot produce
separating bands, and no correction to its numbers will change that.

HEADLINE: calibration is complete, delivered honest numbers on two stats
and zero grade separation, because the counter has no resolution -- and
the proven factors are not wired into the forecast at all. The second is
the reason for the first, and it is plumbing rather than a modelling wall.
Per-archetype grades need proven factors that actually reach p_win. Last
calibration order.

Serving unchanged from 74cf1ce. p_win never mutated. No Bonferroni slot.
Counter and frozen clusters verified file-by-file (15 modules).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 02:25:45 -04:00
builtbykev 74cf1ce974 Robust bias established; low-parameter correction replaces isotonic
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.

PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.

PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:

  hits         a=0.406 shrink 0.565  0.2626 -> 0.2540  CI [-0.0112,-0.0069]  DEPLOY
  total_bases  a=0.472 shrink 0.333  0.2490 -> 0.2429  CI [-0.0062,-0.0059]  DEPLOY
  rbi          a=0.775 shrink 0.231  0.2011 -> 0.2007  CI [-0.0007, 0]       REFUSE
  runs         a=-0.032                                                      REFUSE

A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.

STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.

PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.

PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.

p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 23:20:05 -04:00
builtbykev ced40421ed Audit the LODO instrument: it cannot evaluate any stat, and both prior
FAILs were false

PHASE 0 — the gate at 1f40014 was mine and was an incoherent pair. A 1-SE
informativeness bar with a ZERO-reversal rule: at exactly 1 SE a stable
stat's drop reverses with prob Phi(-1)=0.1587, so on four informative
drops P(>=1 reversal | perfectly stable) = 1 - 0.8413^4 = 0.50. It failed
stable stats half the time by construction. And the pooled n*=70
mis-credited EVERY stat -- too low for hits (own 77) and runs (81), too
high for total_bases (60) and rbi (54).

PHASE 1, blind. Per-stat (g, sigma_row): hits -0.01288/0.11251, TB
-0.01380/0.10680, rbi -0.00884/0.06459, runs -0.00902/0.08080. All four
clear z=1.96 at full n, so none is NO-EFFECT. Committed k=1 with per-stat
n* and a binomial cutoff holding FP at 0.004-0.031.

THE FINDING THAT DOMINATES: the test has no power. Against a strong
instability (date-to-date SD equal to the effect) it detects a failure
1.4%-9.3% of the time, and across every k from 1.0 to 2.0 the best any
stat reaches is 0.337. A gate that cannot fail cannot pass, so
LODO_POWER_FLOOR=0.50 makes UNTESTABLE structural -- "could not test" can
never read as "passed".

PHASE 2/3 cold, at each stat's OWN n*:

  hits  5 informative, 0 reversals, cutoff 2, power 0.093  UNTESTABLE
  TB    5 informative, 0 reversals, cutoff 2, power 0.093  UNTESTABLE
  rbi   4 informative, 1 reversal,  cutoff 2, power 0.045  UNTESTABLE
  runs  3 informative, 2 reversals, cutoff 2, power 0.014  UNTESTABLE

Setting the power floor aside entirely, NOT ONE STAT EXCEEDS ITS CUTOFF.

PHASE 4 — rbi's FAIL was false, as the order suspected. So was RUNS' --
which the order did not anticipate, having classified it DATE-DRIVEN on a
244-row reversal; two reversals in three drops does not clear a cutoff of
2. TB's PASS was vacuous: the test could not have failed it. hits' own n*
is LARGER than the pooled one (77 vs 70), and it remains untestable.

PHASE 5 — deploy basis is now the date-clustered CI alone:

  hits  CI [-0.0139,-0.0097], 4 date clusters   relabelled ci_only
  TB    CI [-0.0061,-0.0045], 2 date clusters   RELABELLED, kept
  rbi   CI [-0.0092,-0.0010], 2 date clusters   NEWLY DEPLOYED
  runs  no fittable map at its split            REFUSE, no CI either

Every deployed stat carries calibration_basis ci_only_lodo_untestable and
auto-demotion is the SOLE stability guard, not a backstop to a passed
test. Stated plainly: those intervals rest on 2-4 date clusters, which is
thin, and it is now the only support. rbi gains chainAcross stackability;
its bands rebuilt on p_win_calibrated (425 rows) are every-archetype
base_rate. runs is queued for the low-param calibrator for the ordinary
reason -- no fittable map -- not on the date-driven finding, which was an
artefact.

PHASE 6 — the deploy set was set by a coin-flip-power ruler; it is now set
by a per-stat power-coherent pre-committed test whose first act was to
report that it cannot evaluate anything. The audit was permitted to wound
the live deploy and did: total_bases lost its LODO claim. Standing
question unchanged -- 18 archetype slots across three deployed stats, every
one a single band indistinguishable from base rate.

Blind ordering held. p_win never mutated. No Bonferroni slot. Counter and
frozen clusters verified file-by-file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 23:01:57 -04:00
builtbykev 1f40014256 Power-derive the LODO threshold: hits restored through the gate, rbi/runs
routed as date-driven

PHASE 0 — threshold derived BLIND, before any stat was re-read. A
reversal is informative only if that date's Brier delta is
distinguishable from zero at its row count. Per-row Brier difference
d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and
n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single
stat's verdict could shape the threshold deciding it:

  pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175
  n* = (0.09816/0.01175)^2 = 69.8 -> 70

The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect
this removes, and why the previous verdict moved with the number.
Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with
LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it
equals the constant, so it cannot drift from its own justification. The
derivation script prints no stat verdict, no date and no reversal.

PHASE 1 — LODO at n*, applied cold:

  hits         5 informative drops, 0 reversals   PASS
  total_bases  4 informative drops, 0 reversals   PASS
  rbi          reverses 2026-08-01 (n=99)         FAIL
  runs         reverses 08-01 (n=86), 08-05 (244) FAIL

hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across
123-272 row dates, favourite sign holding on every testable drop. THIS IS
THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the
withdrawal at 6ae11f1 was correct on the instrument available then, which
admitted 20- and 25-row dates as evidence. Nothing about hits changed;
the threshold stopped being chosen.

PHASE 2 — both failures are DATE-DRIVEN, not underpowered. Every
reversal sits above n*=70 (99, 86, 244), so no threshold and no further
accrual rescues either: isotonic is fitting day-structure. Routed to the
low-parameter calibrator queue (Platt/beta), not built here.

PHASE 3 — CALIBRATION_DEPLOYED is now ['hits','total_bases'], frozen and
tested, both PROVISIONAL with auto-demotion armed and the >=40
date-cluster promotion bar unchanged. hits stackability for
chain.chainAcross is RESTORED, and the record shows it returned through
the powered gate rather than by fiat. hits bands rebuilt on
p_win_calibrated (765 eval rows): every archetype still one band, still
base_rate -- calibrated YES, proven-per-archetype NO.

PHASE 4 logged: the deploy set is now set by a power-derived,
pre-committed, tested constant rather than an operator-chosen number. At
6ae11f1 that rule moved the live path AGAINST the operator; it has now
moved it back on the same evidence because the instrument changed. Both
directions are the rule working. And calibrated p_win separates within
archetype no better than raw across 13 archetype slots on two deployed
stats -- per-archetype separation will come from proven factors or not at
all.

p_win never mutated; no Bonferroni slot consumed; counter and frozen
clusters verified byte-identical file by file.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 20:15:32 -04:00
builtbykev 6ae11f1193 LODO-gated provisional calibration: total_bases deploys, hits withdrawn
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration
layer without challenging the binding. That floor is a cluster-robust
interval bar for a CAUSAL claim. Calibration makes no causal claim, has a
bounded failure mode (it can only over- or under-shrink) and consumes no
Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and
leave-one-date-out tests that directly -- a STRICTER bar, since a cluster
count cannot detect a single day carrying the effect. The >=40 floor is
retained, correctly scoped as the PROMOTION bar.

PHASE 1 — both guards codified, 11 tests, green before Phase 2.
Demonstrated on live data: raw population violated=true, mean_p 0.4962,
both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The
null guard's test demonstrates the trap explicitly, since (null-1)**2 is
1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate.

PHASE 2 — LODO:

  hits         n=1140 dates=17  2 reversals (07-22 n=20, 07-26 n=25)  FAIL
  total_bases  n=1050 dates=7   0 reversals, 0 sign flips             PASS
  rbi          n= 630 dates=5   1 reversal  (08-01 n=99)              FAIL
  runs         n= 597 dates=5   2 reversals (08-01 n=86, 08-05 n=244) FAIL

Threshold sensitivity reported because the verdict moves: total_bases
passes at every held-size threshold, runs fails at every one, and hits
fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20
before seeing which stats passed and did not move it afterwards to
preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a
standard error several times the effect, so the instrument is
underpowered per-drop -- an argument for pre-registering a higher
threshold, which is a Roundtable call, not one to make while holding the
results.

PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and
runs REFUSE.

HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService
hardcoded it since S91; it fails LODO, so it is out. A stat that cannot
survive dropping one day was never calibrated, it was fitted to that day.
The consequence is real -- hits props become unstackable for
chain.chainAcross -- and it errs toward withdrawing a claim rather than
preserving one on a fragile verdict. Deployment is now driven by a frozen,
tested CALIBRATION_DEPLOYED set, not a hardcoded stat name.

PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither
waivable. reverify auto-demotes on the first breach (CI stops excluding
zero, or the favourite bias flips sign) and logs the breaking date.
Promotion needs the original >=40 bar. A provisional deploy that cannot be
taken away is just a deploy.

PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The
two-bar rule still bites: calibrated YES, proven NO, so they stay a
base-rate read, now honestly numbered. Every archetype still collapses to
one band -- calibrated p_win separates within archetype no better than raw.

PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI,
and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did
not survive); the refused set is a map of missing inputs; a low-parameter
calibrator is queued unbuilt.

p_win never mutated; calibration rides as p_win_calibrated with
calibration_status provisional. No Bonferroni slot consumed. Counter and
frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 18:31:19 -04:00
builtbykev f976df47b8 Settle model_snapshots + four-stat calibration: works, deploys nowhere
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.

PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.

PHASE 1, two integrity findings both caught by the gate:

1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').

2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.

THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.

THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.

PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):

  hits  n=1140  bias +0.0868  brier 0.2626 -> 0.2511  d -0.0115  CI [-0.0139,-0.0097]
  TB    n=1050  bias +0.0834  brier 0.2490 -> 0.2438  d -0.0052  CI [-0.0061,-0.0045]
  rbi   n= 630  bias +0.0164  brier 0.2011 -> 0.1965  d -0.0046  CI [-0.0092,-0.0010]
  runs  n= 597  bias +0.0410  no map fittable (173 fit rows < 200)

ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.

A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.

PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.

PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.

PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.

p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 15:30:30 -04:00
builtbykev 23d1b13176 runs + RBI: mostly-base-rate confirmed, and one level deeper than expected
Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest
grade is the POOLED base rate.

PREMISE NOTE: the order's closing line says the batter board is
per-archetype-graded after this. Nothing has been rescaled for hits or
total_bases either -- no archetype slot has ever reached sample and
gradeBands remains built, gated and unwired. This is the fourth stat
measured, not the completion of three.

AUDIT: RBI 935 clean / 43 games; RUNS 617 clean / 33 games. Zero
quarantined. No archetype slot reaches 500 -- and the signature
archetypes the order names are the two SMALLEST slots on the board,
RBI->DRIVER at n=24 and runs->CATALYST at n=9. RUNS is refused
structurally before any factor is tested: 33 game clusters against a 40
floor.

INPUTS RECONSTRUCTED rather than declared missing. lineup_context only
covers 08-04 onward while settled rows start 07-31, so 187/617 runs rows
joined. But the play-by-play cache runs from 05-01 and the batting order
IS the order batters first appear -- slot, power-behind and reach-base
all rebuilt point-in-time, coverage 187 -> 574.

RBI, all THEATER: risp_opportunity +0.0047, extra_base_skill +0.0010,
risp x extra_base +0.0056. RUNS, all refused on clusters and all pointing
the wrong way: +0.0043 / +0.0008 / +0.0054.

THE COMPOUND IS THE WORST VERSION IN BOTH STATS. The causally-correct
compound was the most promising factor on the sheet and is the most
harmful in each. Two multipliers that individually carry nothing do not
cancel -- they compound each other's noise. Distinct from the
collapsed-sequence lesson: there the product of two REAL effects was too
small to use; here the product of two NULL effects is worse than either.

THE ARCHETYPE DOES NOT RESCUE IT, and this is where the session nearly
went wrong. The base rates look strongly differentiated (RBI DRIVER 0.609
vs BOMBER 0.413; runs GHOST 0.716 vs BOMBER 0.460). Gated directly
against the pooled base rate: RBI +0.0010 CI [-0.0034,+0.0050] THEATER;
runs -0.0028 CI [-0.0147,+0.0108] candidate at k=33. DRIVER's 0.609 is
n=23 -- small-slot noise wearing a decimal point. Read off the table
instead of gated, this would have shipped as "archetype differentiation
is real and large". It is not.

THE CROSS-STAT PATTERN THAT IS REAL -- the counter over-predicts every
batter counting stat measured:

  total_bases  p_win 0.5698 vs actual 0.5074  bias +0.0624
  rbi          p_win 0.4860 vs actual 0.4313  bias +0.0547
  runs         p_win 0.5949 vs actual 0.5749  bias +0.0200

Across four stats and three sessions, calibration is the systematic
defect and factor scarcity is not. TB's held-out isotonic fix (-0.0039)
still outperforms every factor tried on any stat, all null or theatre.

NO RESCALE. Nothing proved, nothing certified calibrated, no slot at
sample, and for RBI the archetype split is itself theatre -- so the
honest band is the pooled base rate, which gradeBands returns by
construction.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 14:42:33 -04:00
builtbykev 6a327d9114 total_bases: every power factor is THEATER, and a units bug nearly hid it
PREMISE CORRECTION: the per-archetype rescale is not "proven and live on
hits". gradeBands was built, gated and explicitly NOT wired two orders
ago -- no hits archetype slot reached sample, every band came back
base-rate, and only defense_by_direction proved pooled. This applies an
unvalidated-at-archetype-level method to a second stat.

FULL-HISTORY AUDIT: 988 clean settled TB rows (101 quarantined, 948 with
p_win), 341 players, and only 9 DISTINCT GAME DATES. No archetype slot
reaches 500 -- BOMBER 340, GHOST 147, BRUSH 55. Confirmed short on full
history, not a windowed artifact. The 9-date figure matters more than the
row count: ~49 games means any game- or venue-borne factor has almost no
replication here.

THE BASELINE HAD TO CHANGE, to a harder null. TB lines vary (1.5 on 559
rows, 0.5 on 345), so a per-line personal base rate would rest on ~2 rows
per player-line and would have to be invented. The null is the counter's
own p_win, which already prices the line -- beating the champion, not
beating "he's due".

THE UNITS BUG, caught, and it had produced the best result in the
programme. The first run reported barrel_rate at Brier -0.0095, the
largest improvement ever measured here. fromStatcastRow returns
barrel_pct as a FRACTION (0.06) while the raw table stores 0-100, so
(0.06 - 7.8) * 0.018 clamped EVERY row to the maximum negative shift.
That uniform downward push "improved" Brier purely by leaning on the
counter's over-prediction and contained no barrel information at all.
Same family as the S80 trap, inverted. exit_velo was a second bug -- the
column is avg_exit_velo, so it read null on every row and reported n=0. A
zero is a wiring bug until proven an honest absence.

GATE with units fixed, 138 cumulative tests:

  barrel_rate           n=707  shift 0.0364  brier +0.0036  THEATER
  exit_velo             n=707  shift 0.0229  brier +0.0022  THEATER
  hard_contact_allowed  n=707  shift 0.0260  brier +0.0033  THEATER
  park_weather_hit_type n=651  36 entities   PENDING (k<40)
  platoon_severity      n=481  PENDING (n<500)

THE PREDICTED INVERSION WENT THE OTHER WAY. BOMBER x barrel_rate is
+0.0114, the single most harmful cell in the table, exactly where the
strongest proof was predicted. GHOST +0.0012. All sample-blocked so not a
verdict, but recorded so it is not claimed later.

AND IT IS NOT DOUBLE-COUNTING -- tested and refuted: corr(barrel, p_win)
= -0.061, the counter is not pricing barrel at all. The duller answer is
corr(barrel, counter RESIDUAL) = -0.012. Barrel is a real skill that
carries no information about what the counter gets wrong at this line.
That also closes the S81 lead: hard_hit r=0.153 at n=295 drifted to 0.135
at n=383 and is THEATER at n=707.

THE REAL FINDING: TB is miscalibrated, not under-factored. mean p_win
0.5698 vs actual 0.5074, bias +0.0624. Held out on a strict time split
(fit < 2026-08-02, eval 651 unseen rows): raw 0.25007, constant de-bias
0.24740 (-0.00267), isotonic 0.24621 (-0.00386). Worth more than any
factor tested and the only intervention pointing the right way -- and
still refused at the corrected bar on 32 clusters. A CANDIDATE, not a
result. It also explains the units bug's fake success exactly: a blanket
downward shift is a crude de-bias.

NO RESCALE. Nothing proved, nothing certified calibrated, no slot at
sample -- every band would be the honest base-rate band gradeBands
already returns by construction.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 03:13:35 -04:00
builtbykev 8ab6557faa Collapsed sequence edge: two proven links whose product is too small to use
Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.

FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.

TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.

GATE on the concentrated subset, 114 cumulative tests:

  early-exit x WEAK pen    n=1931  brier -0.0001  CI [-0.0014,+0.0010]  NOT_PROVEN
  early-exit x STRONG pen  n=2574  brier  0.0000  CI [-0.0011,+0.0010]  THEATER
  all early-exit later ABs n=6869  brier -0.0001  CI [-0.0007,+0.0005]  NOT_PROVEN
  pooled all later ABs    n=17891  brier  0.0000  CI [-0.0004,+0.0003]  THEATER

Not pooled-diluted -- the concentrated subset was gated alone and is no
better.

THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:

  P(faces pen | early-exit flagged)  0.8075
  P(faces pen | starter goes deep)   0.7149
    exposure the flag actually buys  0.0925
  hit-rate swing across pen quality  0.0394
  MAX JUSTIFIABLE ADJUSTMENT         0.00365
  actually applied                   0.01930   -> 5.3x over-movement

A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.

A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.

NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.

THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.

Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 02:43:23 -04:00
builtbykev b2e4c6c4fb Link 2 at the coarse grain: pen QUALITY proves, archetype does not
The refinement was right. Naming the individual reliever failed; the same
question at the grain the chain needs passes, and it transmits more than
anything else measured in this chain.

WHY IT WAS WORTH RE-ASKING: last session's null (the pen is on average no
softer, +0.0010 on 35,760 PAs) does NOT rule this out, and treating it as
though it did would have been the error. An average washing out is fully
consistent with quality VARIATION mattering. It does -- actual arm quality
moves the hit rate monotonically across quartiles, 0.2244 / 0.2293 /
0.2410 / 0.2501, a 2.57pp spread, larger than the whole times-through-
the-order effect.

CLUSTER UNIT CORRECTED, THEN CHECKED RATHER THAN ARGUED. Last session
refused Link 2 partly as team-borne (30 bullpens, the park ceiling). My
first re-check was that 76% of pen-quality variance is within-team -- but
that is a statement about TREATMENT variance, not about where errors
correlate, and stopping there would have been picking the convenient
answer. Measured the actual thing: ICC of prediction error by team =
0.0261, design effect 1.41, SEs inflated ~19%. So the verdict was run
three ways:

  unclustered            CI [-0.0067,-0.0010]  excludes zero
  team-clustered (30)    CI [-0.0086,-0.0003]  excludes zero (below the
                         40-cluster floor -- indicative, not a pass)
  design-effect adjusted CI [-0.0072,-0.0005]  excludes zero

QUALITY GRAIN PROVES on the concentrated elevated-early-exit subset:
n=501 team-games, 426 clusters, MAE 0.0294 -> 0.0260, delta -0.0034, CI
[-0.0063,-0.0005] at 110 cumulative tests. Pooled also proves, so it is
not a subset artefact.

ARCHETYPE GRAIN DOES NOT: 0.5669 vs a 0.5309 modal-guess baseline,
corrected interval [-0.1073,+0.0268] spans zero. Two grains tested, one
earned a place -- penQuality.js exposes no archetype and a test asserts
it.

WHAT LINK 3 RECEIVES, which is the number that actually matters -- not
the MAE gain but realized outcome separation, prediction strictly
point-in-time:

  predicted BEST pen   167 games  2,044 PAs  hit rate 0.2231 +/-0.0180
  predicted WORST pen  167 games  1,799 PAs  hit rate 0.2501 +/-0.0200

2.70pp separated, intervals non-overlapping, capturing nearly all the
2.57pp available at the quartile grain. Caveat stated not buried: the
tercile cut is chosen in-sample; the prediction driving it is not.

BUILT: penQuality.js + 9 tests. Abstains below 5 prior club games and 40
arm appearances -- a league-average stand-in would assert "this is an
ordinary bullpen", which is a claim, and usually the wrong one for exactly
the clubs whose pens just turned over.

Link 3 is unblocked on a proven Link 2 at the quality grain only. Not run
here; this order scopes to building and gating Link 2.

Parallel track logged unchanged: TB n=948 pooled, BOMBER x TB 340, short
by 160.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 02:17:00 -04:00
builtbykev e4dae0e6b0 Reliever chain: Link 1 proves, Link 2 does not, and the premise inverts
The causal insight is right -- the game is a sequence and the matchup does
shift mid-game. The direction is backwards, measured on 93,663 plate
appearances from 1,238 games pulled free from statsapi.

LINK 1 PROVES. Starter batters-faced, point-in-time from his own prior
starts only, clustered on the pitcher: MAE 3.2226 -> 2.7990, delta
-0.4236, CI [-0.6006,-0.2731] at 0.9995 corrected for 107 tests, 1,706
starts across 204 pitchers. It finds the tail the chain needed -- early
exits are a 23.2% base rate, model-flagged starts are 34.0% early, lift
+10.8pp.

Scope correction inside Link 1: the order specifies fatigue x GAME
SCRIPT, but game script is not available at grade time -- whether he gets
hit tonight is the thing being projected, not an input to it. Only the
workload half is measured; the in-game half is recorded as a live feature,
out of scope, rather than quietly folded in.

LINK 2 DOES NOT PROVE, twice over. Model accuracy 17.2% vs an 8.6%
baseline -- doubling it sounds good and is not, since naming a specific
arm is wrong five times in six. And structurally the entity is the
BULLPEN: 39,629 post-starter plate appearances across 30 clubs is 30
readings, below the 40-cluster floor, the same permanent ceiling as park
geometry and team defence. LINK 3 NOT RUN, per the order's own rule.

THE PREMISE IS REFUTED, and this chains on nothing so it was safe to
measure:

  vs STARTER  n=48,492  hit rate 0.2444 +/-0.0038
  vs BULLPEN  n=35,760  hit rate 0.2373 +/-0.0044

The pen is 0.7pp HARDER. The specific effect the chain exists to exploit
-- early exit making later at-bats softer -- is +0.0010 on 35,760 PAs. A
well-powered null, not a sample problem.

What IS real is times through the order: TTO1 0.2351 -> TTO2 0.2515 ->
TTO3 0.2518. A starter does decay as the lineup sees him again, but that
advantage is SURRENDERED when he leaves, not extended -- the pen is
harder than his second and third time through. A modern bullpen is a
queue of fresh specialists throwing one inning each; there is no tiring
arm to punish.

So the insight survives inverted, and Link 1 stays valuable for the
opposite reason it was built: a likely early hook predicts the hitter
LOSES his third-time-through look (0.2518 -> 0.2373 on that PA). The
mispricing is on hitters who get an EXTRA look at a starter going deep.

BUILT: predictionGate.js + tests -- the two-part gate for a continuous
prediction. factorGate binarises outcomes for Brier, which would destroy
a target like batters faced. Same discipline, same THEATER verdict, real
scale.

PRE-REGISTERED NOT RUN: Link 2' using a PA-weighted bullpen AGGREGATE
rather than a named arm. Recorded rather than substituted in -- running
Link 3 on a swapped-in Link 2 is the assumed-link failure the order
forbids. Given the premise result its expected value is now low.

PARALLEL TRACK logged: total_bases n=948 pooled, BOMBER x TB 340, short
by 160. Sample-readiness only, not a verdict.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 01:58:38 -04:00
builtbykev 3081c92e00 Per-archetype grade bands: built, gated, and the rescale blocked twice
The premise does not hold. proven-status.js run fresh: PROVEN_SET is
EMPTY, no archetype x stat reaches the gate. pitcher_contact_profile has
a CI upper bound of exactly 0.0000 and platoon_severity is held on
4.5%-contaminated splits, so the proven set is one factor, pooled, not
three archetype-conditioned ones. The specific pattern the order names --
defense strong for GHOST/BRUSH, null for BOMBER -- is the one I measured
running the OTHER WAY yesterday, both noise-dominated.

But the second blocker is new and matters more, because it would stop the
rescale even if the factors had proved: the grade does not separate
within any archetype. Every archetype collapses to ONE band at the
corrected bar, because bands merge when their intervals overlap and
publishing two letters we cannot tell apart is a distinction we have not
measured.

Uncorrected, so the ranking is visible rather than hidden by the bar,
this INVERTS the order's design. The order gives contact types the
factor-rich treatment and power types honest base-rate, reasoning that
single-game hits are variance for a power profile. Measured:

  BOMBER n=466  corr(p_win,outcome) +0.207  quintiles 0.75 0.62 0.60 0.48 0.48
  GHOST  n=192  corr(p_win,outcome) -0.007  quintiles 0.47 0.63 0.74 0.58 0.45

BOMBER is the one archetype the model ranks, and it splits into a real
A 0.660 / B 0.481 at 95%. GHOST is flat, and non-monotone -- its most
confident reads hit 47% while its middle reads hit 74%. Shipping as
specified would have given the factor-rich treatment to the archetype the
model reads worst and left base-rate on the one it reads best. That is
mechanically sensible in hindsight: a power hitter's hit tracks whether
he can damage the arm, a contact hitter's depends on balls finding holes.

BOMBER's split does not survive the cumulative correction at 106 tests.
Exposing it by loosening the correction is the curve-to-make-A's the
order forbids, so it stays one band.

BUILT: gradeBands.js -- lift against the archetype's OWN base rate (the
same 62% is lift for a 45% profile and a deficit for a 68% one),
indistinguishable neighbours merged, thin bands PROVISIONAL not dropped,
Wilson intervals widened by the cumulative correction. The two-bar rule
is structural: proven-alone, calibrated-alone and neither all return
base_rate with the reason stated, so with nothing proven no
factor-informed band can be produced at all.

reasoning() is built and tested but NOT wired to the card -- there is no
per-archetype band being served, so attaching the copy now would ship
product language for a rescale that does not exist.

NOT BUILT: the specified power-type reason "the matchup edge is in
total_bases". total_bases is recorded INCONCLUSIVE (+0.0038, CI
[-0.068,+0.075]). Wiring it would assert an edge measured as
indistinguishable from zero -- the exact fabricated-reason failure this
module exists to prevent.

BOMBER x hits is 29 rows short of the gate and is the archetype the model
actually reads. That is the first slot to test, not GHOST.

Counter and frozen clusters byte-identical. No letter was moved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-06 01:03:10 -04:00
builtbykev 6b17f79367 Per-archetype re-audit: no slot reaches 500, and the replication unit
decided everything

The premise does not hold. prove-hit-factors.js has no date filter
anywhere in it and pages the full table -- there was never a window to
widen. Full clean history is 1,266 rows, not 2,715. platoon was not
"proved" last session, it was explicitly held on 4.5%-median-contaminated
season-to-date splits, and pitcher_contact_profile was demoted. The
proven set going in was one factor, not three.

STEP 1: no archetype slot reaches n>=500 on full history. Best is BOMBER
at 408, and BOMBER is the most common archetype on the board. GHOST 173,
BRUSH 64, DRIVER 43, CATALYST 16. These are confirmed genuinely short,
not artifacts.

STEP 2 is where the real finding is. park_hits initially PROVED at 619
rows across 45 games -- but those games only ever visited 14 distinct
park values. A park effect is replicated across parks, and unmodelled
park heterogeneity is confounded with the thing being estimated. Each
factor is now clustered on the coarser of the game and the entity its
treatment rides on.

That flipped two verdicts and confirms Kev's causal-correctness thesis
from a new direction: defense_by_direction has 442 hitter-team units of
replication where crude team defense has 26. The correct atom is not just
more accurate, it is the only one measurable at all. park_hits (14) and
defense (26) can never be validated however long the ledger runs -- the
same ceiling as park dimensions, reached independently.

Also fixed a bar I got wrong last session: I transplanted the 500-row
floor onto clusters, which refused a factor with 1,059 rows over 85 games
while answering neither question. Two floors now -- rows>=500 for a stable
estimate, clusters>=40 for a trustworthy interval. Not a lowered bar:
park_hits and defense are still refused.

PROVEN: defense_by_direction only, pooled, [-0.0054,-0.0012] at 99 tests.
It stays POOLED-ONLY -- no per-archetype reasoning wired, nothing
grandfathered. The card must not say "GHOST: defence matchup strong"
because we have not earned that sentence. The predicted fingerprint did
not appear either: BOMBER -0.0036 vs GHOST -0.0024, the opposite
direction, both noise-dominated. Recorded so it is not claimed later.

RESCALE: NOT READY. One proven factor worth -0.0031 Brier. Rescaling on
that is relabelling.

Counter and frozen clusters byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-05 19:58:37 -04:00
builtbykev b818626870 BUILD-STATE: under-querying vs out-of-data session
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-05 19:30:56 -04:00
builtbykev 7b85934dc3 Under-querying vs out of data: the answer depends on the unit
The platoon test's n=452 described how much of the JOIN survived, not how
much data exists. There are 1,266 clean settled hits rows and zero
quarantined ones. platoon_splits had been ingested from tonight's lineups
only (315 players), so any hitter who settled a prop without appearing in
an ingest-day lineup was silently absent from every test.

Backfilled all 380 hitters (81 fetched, 0 unresolved). Re-ran on 1,059
rows, up from 452.

THE DEMOTION IS THE HEADLINE. pitcher_contact_profile, the strongest
proven factor in the programme (-0.0064, CI [-0.0113,-0.0014]), roughly
halved to -0.0034 on more than double the sample and its corrected
interval now spans zero. The Bonferroni denominator also rose to 55,
which widens every interval -- but a denominator cannot move a point
estimate, and that halved on its own.

platoon and platoon_severity now clear the bar and are NOT promoted.
Upper bound -0.0001, on season-to-date splits that contain the games they
predict: measured contamination is 4.5% median, 12.4% at p90, 137% worst.
I had assumed ~1%. They stay CANDIDATE pending point-in-time splits.

GAME-LEVEL IS A DIFFERENT PROBLEM. game_context held zero weather rows
ever -- not because the fetcher was wrong (it correctly targets
Open-Meteo's archive) but because ledger_entries keys a game as
mlb:2026-08-03:Away@Home and game_context keys it as mlb:823437. Every
lookup missed and NULL columns read as honest absence. Third occurrence
of that class.

Fixed the join: 96/101 settled games now carry actual archived weather,
park dimensions backfilled 15 -> 30 venues.

But 928 total_bases rows sit on 47 games at 17.6 rows per game. Park and
weather assign one value per game, so resampling rows would have
manufactured a pass. factorGate now resamples clusters when rows carry
one and judges sample against effective_n; unclustered rows keep the
original path byte-for-byte. Verdict: 47 clusters < 500, and the point
estimate is +0.0011 -- worse, not merely unproven.

Weather needs ~57 more days. Park dimensions need never: there are 30
ballparks in MLB, so a venue-constant factor can never reach 500
independent units. That bar was built for player-level factors and does
not transfer.

Wind is refused. We have speed and bearing for all 96 games; we lack park
orientation, and 220 degrees is blowing out at one park and in at
another. Using speed alone would assert an effect while discarding the
sign that decides what it is.

Counter and frozen clusters untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-05 19:30:17 -04:00
builtbykev 6452926732 Retain raw weather, and record platoon severity 48 rows short
Two fixes in the weather path, and the second was hiding behind the first. The
scalar weather_mod cannot express a hit-TYPE conversion at all -- wind out and
warm turning fly balls into extra bases, and cold heavy air turning them into
outs, collapse to the same number once multiplied -- so the raw temperature,
wind speed and wind direction are now retained alongside it.

And the old guard only kept the environment when the multiplier was not 1,
which silently discarded the forecast for every ordinary night. That is the
majority of games, and precisely the rows a hit-type model would need in order
to learn what ordinary looks like.

Platoon severity is built and measured at n=452, which is 48 rows short of the
gate: CANDIDATE_PENDING, neither proven nor theatre. It moves less than flat
platoon (0.021 against 0.026), consistent with the pattern, and its Brier point
estimate is favourable but the corrected interval still spans zero.

Worth naming: the refusal costs sample, and that is the design working. Flat
platoon scores 741 rows because it will happily apply a boost to anyone;
severity scores 452 because the other 289 are hitters whose split we cannot
actually read at 60 plate appearances on the short side. Buying those rows back
by shrinking instead of refusing would have produced a number indistinguishable
from a measured league-average split, which is a different claim from the one
the data supports.

Park dimensions are ingested and verified in production across fifteen venues,
joined by the venue the game is actually at rather than inferred from the home
team -- neutral-site and international games break that assumption without
surfacing an error.

The park-and-weather-to-hit-type atom is NOT built. Its inputs landed this
session and carry a single as_of date, so testing it on total_bases would be
scoring games with inputs that postdate them. Building it now would produce
something plausible rather than something proven.

Proven factors for hits remain pitcher_contact_profile and
defense_by_direction. 4,307 tests green (344 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 20:40:56 -04:00
builtbykev de0077f6f9 Causally-correct platoon + park-dimensions ingest
Applying the method that worked for defence to the two factors the code flagged
as still crude.

PLATOON. The flat version is 'lefty versus righty, add a boost', and it failed
the two-part gate for the same reason team-average defence did: it is not the
unit the causal story runs through. The advantage is only worth what THIS
hitter's split is actually worth -- measured on a real hitter, .284 against
left-handed pitching versus .221 against right-handed, a 63-point split, where
the flat factor applied the same six percent to him and to a hitter with none.

Most of the work is sample discipline, and the second rule matters more than
the first. Severity shrinks toward the league split weighted by the SMALLER
side's plate appearances, because a 500-against-40 split is a 40-PA read. And
below a floor it REFUSES outright rather than shrinking, because a
heavily-shrunk severity is indistinguishable from a measured league-average one
and those are different claims -- without the refusal the atom would quietly
assert a league-typical split about every September call-up in the league.

Switch hitters turn out to be the easy case misread as the hard one. He bats
opposite by choice so the direction is never in doubt, but the per-side value of
his swing is a different question and one this sample cannot answer, so he is
unreadable rather than credited with an automatic edge.

PARK DIMENSIONS. Free from statsapi's venue endpoint, which carries fence
distances, roof, turf and elevation outright -- Wrigley returns 355 down the
left line, 400 to centre, 353 to right, at 595 feet. parkFactors holds run
COEFFICIENTS, which structurally cannot express a park that turns outs into hits
without scoring, and that is why the crude park factor failed.

The park join is by the venue the game is ACTUALLY at, carried from the schedule
feed, never inferred from the home team -- neutral-site and international games
break that assumption and they break it silently. A venue with no geometry at
all is absent rather than a park with zero dimensions.

Both tables dated in the primary key. Venue geometry changes rarely but it does
change, and by now that is the default rather than a lesson.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 20:30:57 -04:00
builtbykev 20c45cbcd1 The causally-correct defence atom proves where the crude one did not
Kev's insight holds, and the data says so cleanly. defense_by_direction PROVES
on hits -- n=528, Brier -0.0034, interval [-0.0059, -0.0009] at the 99.9% level
the cumulative correction now demands -- while team-average defence remains not
proven, its interval still spanning zero. Same signal, same rows, different
unit.

The detail worth keeping is that the causally-correct atom moves the number
LESS THAN HALF as much as the crude one, 0.013 against 0.030, and is the one
that is reliably right. The team average was moving more and knowing less. Big
movement is not evidence of a good factor; it is frequently the tell.

Both halves turned out to be free, as the order expected. Savant's batted-ball
leaderboard carries pull/straight/oppo crossed with ground/air for 609 hitters
-- the statcast leaderboard we already pull does not, it has nineteen columns
and no direction at all -- and the OAA feed already carries each fielder's
position, so per-position defence is a regrouping of last week's ingest rather
than a new source. Verified in production: 609 spray profiles, 31 teams.

Handedness is what joins them and getting it backwards would have been
invisible. Pull for a right-handed hitter is the left side; for a left-handed
hitter it is the right side. A model that ignored `bats` would send half the
league's grounders to the wrong infielders and still look like it was reading
defence, and nothing downstream would have caught it. Switch hitters bat
opposite the pitcher, which this does not resolve, so they are unreadable
rather than guessed.

Unmeasured zones are renormalised away rather than contributing a zero, since a
zero asserts an exactly-average fielder standing there, and coverage states
honestly what share of a hitter's contact we could actually read.

ATOM 2 is input-blocked rather than sample-blocked, and the distinction matters
because waiting will not fix it. The weather free-source check passes --
Open-Meteo is already wired and exposes temperature, wind speed, wind direction
and precipitation -- but those raw fields are collapsed into a single scalar
modifier and wx_forecast is empty on all 1,119 settled rows. Park DIMENSIONS
are not ingested at all; parkFactors holds coefficients, not wall heights or
fence distances. A park-and-weather-to-hit-type conversion needs both, so it is
scoped rather than half-built: retaining the raw weather fields is the cheap
half, dimensions are the missing one.

Proven factors for hits are now pitcher_contact_profile and
defense_by_direction, both pooled; every per-archetype slot remains
sample-blocked.

4,297 tests green (342 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 20:01:09 -04:00
builtbykev 405180e791 Build the causally-correct defence atom: spray x positional OAA
Team-average defence failed the two-part gate for hits, and the reason was the
unit rather than the signal. A left-handed pull-ground hitter meets the first
baseman and the second baseman and almost nobody else, so a team total averages
in five fielders who will never touch his ball.

Both halves were already free on the host we pull from. Statcast publishes
spray x trajectory per hitter -- pull/straight/oppo crossed with ground/air,
608 hitters -- and the OAA feed already carries each fielder's position, so
per-position defence is a regrouping of data ingested last week rather than a
new source. Zero new sourcing, as the order expected.

Handedness is what joins them and getting it backwards would be invisible: pull
for a right-handed hitter is the left side, pull for a left-handed hitter is the
right side, so a model ignoring bats would send half the league's grounders to
the wrong infielders and still look like it was reading defence. A switch hitter
bats opposite the pitcher, which this does not resolve, so he is unreadable
rather than guessed.

Two properties the crude version could not express, both locked by test: two
teams with the SAME total defence read differently for a pull hitter, and a
ground-ball hitter and an air hitter read the same team in opposite directions.

Unmeasured zones are renormalised away rather than contributing a zero, which
would assert an exactly-average fielder standing there, and  states
honestly what share of a hitter's contact we could actually read. Nothing
readable at all returns null, so the caller falls back to the base rate instead
of to an invented 1.0 that looks measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 19:53:07 -04:00
builtbykev a9ee55550b Build the two-part factor gate: one factor proves, and zero are theatre
The question was whether the hit grade reads tonight's game or just says he is
due. Answering it needed a gate that correlation cannot provide, because
correlation cannot separate the two ways a factor looks alive: it reads the
game, or it moves the number and reads nothing. The second is what a product
ships by accident -- arch-v1 moved 76% of rows by 2.5 points, changed
resolution by 0.0000, and was live for months, and no user could have told.

So a factor must now clear both conditions: move the prediction off the
player's own leave-one-out base rate, AND improve out-of-sample Brier. Brier
rather than correlation, because correlation asks whether the ordering improved
and this asks whether the NUMBER got closer to what happened -- and for a graded
probability the number is the product.

The correction applies to the interval itself, which turned out to matter more
than expected. A plain 95% CI is the right bar for one test; at fifty
cumulative tests roughly two or three intervals exclude zero by chance alone.
Widening to 1 - 0.05/tests, currently 99.9%, flipped both defence and platoon
out of "proves". A 95% interval would have shipped two unproven factors into
the grade, with reasoning text explaining them to users.

That forced a distinction I had initially collapsed. Defence and platoon have
FAVOURABLE point estimates whose corrected intervals merely span zero, and
calling that THEATER would repeat the error this codebase keeps correcting:
insufficient evidence is not evidence of absence. THEATER is now reserved for
its one real meaning -- moves the number, reads nothing -- and
NOT_PROVEN_AT_CORRECTED_BAR names a real candidate held to a bar that rises with
every hypothesis the programme tests.

Result on 741 settled hits rows: pitcher_contact_profile PROVES, improving
Brier by 0.0066 with a 99.9% interval of [-0.0114, -0.0016]. Defence (-0.0043)
and platoon (-0.0039) are not proven at the corrected bar. Park is
sample-blocked at n=405. Zero factors are theatre, which is the genuinely good
news: nothing decorative is being wired. Per-archetype every slot is
sample-blocked (BOMBER 252-294, GHOST 67-125).

Two spec gaps worth recording. The approach identities the order names -- SPRAY,
DAMAGE-DEALER, COUNT-WORKER -- do not exist in the registry; the MLB batter
archetypes are BOMBER, GHOST, TORCH, BRUSH, DRIVER, FLEX, ALPHA, HYBRID and
CATALYST. And parkFactors maps hits to run_base, so there is no hits-specific
park factor at all: a park that turns outs into hits without producing runs is
invisible to the input we have.

The grade rescale is NOT run. It was explicitly gated on the factor proving,
and one pooled factor worth 0.0066 of Brier is not a factor-informed
distribution -- rescaling on it would dress a base-rate model as a matchup
model, which is the exact thing this gate was built to prevent.

4,286 tests green (340 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 19:30:18 -04:00
builtbykev 4d1803f6d7 Calibrate hits point-in-time: partial pass, and an honest ceiling of 0.667
Fitted the isotonic map on game_date < 2026-08-02 (n=589) and evaluated it on
everything from that date forward (n=383). The map never saw the evaluation
rows, which is the only thing that makes the result mean anything -- fitting
and evaluating on the same rows always looks perfectly calibrated, because the
map is reciting the answers it was built from.

It works, on most of the distribution. Held-out after correction: 0.477 comes
back 0.506, 0.587 comes back 0.580, 0.667 comes back 0.603 -- against raw
errors of +0.191, +0.279 and +0.246 in the same bins. Ordering survived, and
that was verified pairwise rather than assumed, because a broken map would
silently destroy the one thing this model does well.

Two findings matter more than the pass.

First, the honest ceiling is 0.667. Once the numbers are truthful this model
has no 80%-plus hit reads at all -- the top of its range was miscalibration,
not confidence. A four-leg ticket at the ceiling is 0.198, where the raw
numbers implied 0.686. The high-floor parlay is a two-thirds-per-leg
proposition, and that is the number to say out loud.

Second, calibration is certified BY BAND rather than by a blanket flag.
Held-out error was -0.029 and +0.007 through the middle but -0.167 at the
bottom and +0.063 at the top: the model is trustworthy over most of its mass
and untrustworthy at both edges. A single true/false would either throw away
the 72% that works or ship the edges that do not. Only a probability inside a
certified band is marked stackable, and that flag is what chainAcross requires
before it will compound anything. The certified band is 0.40 to 0.60, n=276.

A methodological catch on the way: my first pass condition demanded honest bins
at 0.70 and above -- but honest calibration REMOVES those bins, since the
ceiling drops to 0.667. The gate would have failed the repair for succeeding.
It now tests the highest remaining band instead of a fixed threshold.

Wired forward with the same discipline: calibrationService fits strictly before
today, splits by time rather than at random, and returns null on thin history
so that "no calibrator" means nothing is stackable rather than "trust the raw
numbers". p_win is never mutated -- the calibrated value rides beside it as
p_win_calibrated, because a calibration map is a correction to a forecast, not
a different forecast, and the counter stays byte-identical.

4,275 tests green (339 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 17:51:15 -04:00
builtbykev 9c5b968351 chaining-v1: the portable chain, and the gate that blocks the parlay surface
The order's own prerequisite for the hit-parlay surface was to verify the hit
probability is calibrated. It is not, and the failure is exactly the shape that
destroys a parlay.

Measured on 972 settled hits props: the model is monotonically over-confident
at the top and flat above 0.70. Predicted 0.911 comes back 0.630. Predicted
0.844 comes back 0.630. Predicted 0.747 comes back 0.605. There is no
discrimination at all in the range a parlay is built from, and the error runs
in the flattering direction. Four "91%" legs are 0.686 by the model and 0.157
in fact -- a 4.4x overstatement that compounds with every leg added.

Single props survive a calibration error of that size. A parlay multiplies it.
So chainAcross REFUSES to compound atoms not marked calibrated, and refusing is
the feature rather than a limitation: a ticket built on these numbers would be
confidently wrong in the direction the user pays for.

calibration.js provides the reliability table, the gate (tolerance 0.05,
weighted to the high end because that is where tickets live) and an isotonic
fit. Isotonic is the honest repair here because it is monotone: the model's
ordering survives untouched while the numbers move to what actually happened.
The fitted map says 0.65 -> 0.594, 0.85 -> 0.639, 0.91 -> 0.639.

chain.js is the portable core -- base events plus context, through a chain
function, into a PLUGGABLE aggregator: across players for a compound ticket, up
to the team for expected scoring. The sport-specific parts are inputs rather
than code paths, so basketball plugs in as content. The archetype
redistribution hook is there now, dormant in baseball because a nine-run lead
does not change who bats next, and live in basketball where a blowout fades the
star and feeds the bench.

Two judgement calls worth naming. Treating same-game legs as independent errs
in the FLATTERING direction, since they share pitcher, park and weather -- so
correlation shifts the compound toward the weakest leg, bounded, and is labelled
an approximation rather than a joint distribution. And market divergence does
NOT downgrade confidence: it flags a contested script whose props are either the
best or the worst on the board, and which one is unknown until settled.
Internal inconsistency does downgrade it, because per-entity reads failing to
sum to the team read means one of them is wrong and we do not know which.

Not built: the independent game-script projection. It needs proven team-level
atoms and out-of-sample validation against actual margins, and no atom has
passed the gate yet. Building it now would produce something plausible rather
than something proven, which is the failure mode this whole programme exists to
avoid.

4,269 tests green (339 suites); web build exit 0; counter and frozen clusters
byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 17:17:14 -04:00
builtbykev a80a775fa0 Record the lineup-context ingest: prod-verified, 153 lineups / 149 opportunity
Verified in production via the new on-demand endpoint: 153 batting-order rows
across 10 games (orders 1-9), and 149 hitter-opportunity rows with RISP shares
ranging 0.170 to 0.528.

The data passes its own coherence check on arrival: the highest RISP-share
hitters all bat fourth and fifth, which is exactly where the mechanism says the
RBI opportunity lives. Nothing was fitted to produce that -- it is the two
tables joining and agreeing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 16:19:18 -04:00
builtbykev 7f69fef14c Add an on-demand endpoint for the lineup-context ingest
The first prod run wrote zero rows while the parser demonstrably works locally
(144 rows, 10 games with lineups posted), so the zero was wiring rather than
absence -- but diagnosing that required a full snapshot, which now takes about
three minutes and 524s at the edge.

Same reasoning as the statcast refresh endpoint: a job is proven by running it
and reading the result, never by waiting for the slot it rides in. This makes
the ingest verifiable in seconds, so 'zero rows' can be told apart from 'no
lineups posted yet' immediately -- which is the exact confusion the defence
ingest hit when a doubled path 404'd and read as 'Statcast has no fielding
data'.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 16:13:53 -04:00
builtbykev 08276c0880 Ingest lineup + baserunner context: the input RBI and runs always needed
RBI is power TIMES opportunity. The same swing drives in one run or three
depending on who is on base, and a hitter batting with the bases empty cannot
drive anyone in however hard he hits it. Every context-free model of RBI here
has failed, and the failure kept being read as 'skill inputs don't work for
RBI' when the truth was that we were modelling half the stat.

Both halves are free from statsapi.mlb.com, which we already call for game
logs, schedules and probable pitchers. No new provider, no key, no quota.

RUNG 1, batting order: schedule?hydrate=lineups returns homePlayers and
awayPlayers as ORDERED arrays of nine, and the order IS the batting order --
index 0 is the leadoff hitter. That single fact gives CATALYST its identity
and supplies lineup-position context for every context-dependent stat.

RUNG 2 turned out cheap, which the cheapest-first rule did not expect. It
looked like it would need play-by-play reconstruction across a season; statsapi
serves situational splits directly, so 'how often does this hitter bat with
runners to drive in' is ONE call per player rather than one per game. Measured
on a real hitter: 87 plate appearances with runners in scoring position
producing 25 RBI, against 302 with the bases empty producing 17. That ratio is
the opportunity half of the stat and it is the thing no amount of exit velocity
can tell you.

Both tables are dated in the primary key. statcast_aggregates was built
upsert-in-place and that silently made every backtest leak the games it was
predicting; a lineup is worse still, because it is a PRE-GAME fact that changes
by the hour, so an in-place table would overwrite what we knew at grade time
with what turned out to be true.

Absent stays absent throughout: no lineup posted is an empty slate rather than
a guessed order, a short lineup records fewer slots rather than padding to
nine, and a hitter with no splits is null rather than a zero RISP share --
which would assert he never bats with runners on, a strong claim and usually a
false one.

Wired into the snapshot best-effort, so a context failure can never break the
pipeline it rides in. The three pre-registered theories are now marked
input-ready rather than input-blocked: DRIVER's power x runners-on and power x
lineup-position, and CATALYST's speed x on-base x power-behind. They are
sample-blocked from here, and the proofs run under native cumulative
correction as sample accumulates -- ingesting is not proving.

Counter and frozen clusters byte-identical. 4,250 tests green (338 suites);
web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 16:07:13 -04:00
builtbykev 52a3142c5f Split infield defence out, and pre-register the DRIVER/CATALYST/SINKER batch
Every slot in this batch is far below the gate -- DRIVER x hits 39, CATALYST
under 22, SINKER not yet gradeable at all -- so per the order's own sample rule
these are CANDIDATE-pending-accumulation, not tested-and-failed. Testing them
now would produce noise and burn cumulative-correction budget on it.

What IS deliverable is the input SINKER's theory needs, and it turned out to be
free. The OAA feed already carries each fielder's position, so infield-only
defence is derivable from data ingested yesterday: 1B/2B/3B/SS summed
separately from the outfield. Team-total OAA is the wrong unit for a
ground-ball pitcher -- he lives on the infield converting grounders and his
outfielders are close to irrelevant to him, so averaging them in dilutes
exactly the signal. On a real team the split shows a +15 infield inside a +2
team total, which is the dilution made visible. Under three measured fielders
in a unit is absent rather than zero, same rule as everywhere else.

The three theories are now PRE-REGISTERED in the registry with their mechanism
and the skill each would validate, marked CANDIDATE. That is the point of
writing them down before the sample exists: the claim is on the record with a
date and cannot be quietly reshaped into whatever the numbers turn out to
support once they arrive.

Two of them are input-blocked rather than sample-blocked, and the distinction
matters because waiting will not fix them. DRIVER's RBI theory needs
baserunner state and CATALYST's runs theory needs both baserunner state and
batting order; we ingest neither, and player_role_profiles is empty. So RBI
does not unblock on DRIVER -- it unblocks on ingesting lineup context, which
is a sourcing question, not an accumulation one. SINKER is the only one of the
three whose inputs are now ready.

Premise note: no registry re-adjudication demoted anything last session. The
proven set was empty, zero features were demoted, nothing was recalibrated,
and nothing is published -- node scripts/proven-status.js confirms it in one
command.

Counter and frozen clusters byte-identical. 4,238 tests green (337 suites);
web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 15:28:16 -04:00
builtbykev ff037e40c2 Re-adjudicate: nothing to demote, and close the hole that would have mattered
There is nothing to re-adjudicate. The proven set is empty and always has
been -- verified three ways: proven-status reports EMPTY, validatedSkills()
returns {} for every archetype, and zero conditioning entries have ever
reached PROVEN. The one PROVEN feature is recent_frequency_prior, which is the
incumbent counter itself, proven by the S78 ablation as ~100% of the
champion's resolution. It is the baseline every challenger is measured
against, not a conditioning interaction, and demoting it would leave the model
with nothing to grade from.

A correction to the premise: the cumulative gate did NOT catch a false
positive last session. It caught nothing, because there was nothing in the
proven set to catch. What it did was tighten alpha from 0.0026 to 0.0013
within one session, which demonstrated the mechanism working rather than a
demotion. So steps 3 and 4 -- demote, recalibrate -- are vacuous here, and
readjudicateAll says so plainly rather than glossing a no-op.

But the worry behind the order was well founded, and the audit found the real
exposure: promote() did not require the cumulative denominator. It checked n,
lift and CI, and nothing stopped a future session from testing eight
hypotheses, correcting by eight, and promoting on a p-value that would not
survive the programme's real denominator. That is precisely the hole that
makes a retroactive re-adjudication pass necessary later, so it is closed at
promotion time instead. isSufficient now refuses evidence carrying no
correction, evidence corrected against fewer tests than the cumulative count,
and any p-value that does not clear 0.05 over its own test count. The same
rule guards a PROVEN conditioning entry.

The second audit found two of four analysis scripts still correcting
per-session; pitcher-prove-k and tb-solo-and-interactions now use the
cumulative ledger, so the correction is native on every path.

reAblation.js is the standing second line: pure and injectable, so the
decision rule cannot drift from the gate's, and every verdict records both
p-values and both test counts so a demotion is re-derivable by anyone. A
feature promoted at alpha 0.05/20 can demote on the same p-value once the bar
is 0.05/60 -- correct, because the bar rose only after the programme had more
chances to get lucky. No fresh measurement is PENDING_RETEST and never a
demotion: absence of a re-test is not evidence, and demoting on it would
punish whichever stat happens to be off-season.

Net effect on the proven set is zero. No demotions, no recalibrations, and no
public ledger event -- announcing "recalibrated after re-adjudication" when
nothing changed would itself be a false signal of rigour.

4,238 tests green (337 suites); web build exit 0; counter byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-04 15:13:39 -04:00
builtbykev ece2b9f5f9 Ingest defence, and make Bonferroni cumulative across the programme
Two things shipped that stand regardless of sample.

DEFENCE. Statcast Outs Above Average is free on the host we already pull six
feeds from, so there was nothing to decide. 514 fielders, aggregated to team
level -- the unit a batter's prop actually needs, the defence behind the
pitcher he faces -- and persisted as 31 team rows. Verified in production.
Cubs +56 best, Mariners -29 worst.

Unknown is not zero, and it bites unusually hard here: an OAA of 0 is a REAL
reading meaning exactly average, so coercing absence to 0 would assert that
every unmeasured fielder is league-average, which is the commonest defensive
profile there is. team_defense also carries as_of_date in its primary key from
the first row -- statcast_aggregates was built upsert-in-place and that
silently made every backtest leak the games it predicted, so point-in-time is
available here before it is needed rather than after a wrong answer.

A bug worth recording as a class: BASE already ends in /leaderboard, so the
new feed built a doubled path and 404'd. Because a failing feed degrades to an
empty index by design -- correct, so one broken source cannot fail the whole
pull -- it surfaced as "fielding_oaa: 0 rows", which reads exactly like
"Statcast has no fielding data". Graceful degradation makes a wiring bug look
like an honest absence.

CUMULATIVE CORRECTION. Bonferroni had been applied per session throughout: a
run testing eight features corrected by eight. Across a programme's lifetime
that is wrong in the dangerous direction, because every order gets a fresh
generous alpha and the false-positive rate compounds quietly. Correcting by 8
when sixty have been tried is how a noise result eventually gets recorded as
PROVEN with a p-value to point at. The denominator is now distinct hypotheses
ever tested, persisted, and it moved 19 -> 38 within this session alone, alpha
0.0026 -> 0.0013. Re-tests deliberately do not inflate it: re-asking the same
question on more data is not a new shot on goal, and counting it would punish
the discipline of waiting for sample.

THE MEASUREMENT. The differential the theory predicted is present: defence
correlates with the counter's residual at +0.130 for GHOST, the contact and
speed archetype, and -0.018 for BOMBER, the power archetype. A GHOST's hits
depend on whether anyone can range to the ball; a BOMBER's barrels clear the
defence entirely. So a flat BOMBER result is the theory working rather than
the test failing.

It is not a result. GHOST is n=104 against a 500 bar, with p=0.188 against a
corrected alpha of 0.0013 -- three orders of magnitude short. Both are
recorded as CANDIDATE with their measured lift, tagged contact-skill, so the
re-run at full sample compares against a recorded baseline.

Nothing proved, so nothing was recalibrated and nothing shipped.

4,228 tests green (336 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 22:20:30 -04:00
builtbykev 25e36c0257 Fix the doubled /leaderboard path — the fielding feed 404'd silently
BASE already ends in /leaderboard, so the new feed built
.../leaderboard/leaderboard/outs_above_average and 404'd. Because a failing
feed degrades to an EMPTY index by design -- correct behaviour, so one broken
source cannot fail the whole mechanism pull -- it surfaced as 'fielding_oaa: 0
rows' rather than as an error, which reads exactly like 'Statcast has no
fielding data'. Worth noting as a class: graceful degradation makes a wiring
bug look like an honest absence.

Verified: 514 fielders, 31 teams. Cubs +56 OAA, Mariners -29.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 22:07:40 -04:00
builtbykev 010a876b3c Ingest free Statcast fielding (OAA) as team defence, dated from day one
Defence was the one conditioning category with no derivable proxy: nothing we
ingest measures fielding, and a team's pitchers' hits-allowed conflates
pitching with defence and would validate the wrong skill. Statcast publishes
Outs Above Average free on the same host as the six feeds already pulled --
verified live at 513 fielders -- so there was nothing to decide.

Added as a seventh feed, indexed per fielder and aggregated to team level,
which is the unit a batter's prop actually needs: the defence behind the
pitcher he faces. Summed OAA is the team's outs converted above average; the
mean rides along because a team with more measured fielders would otherwise
look better merely for being measured more, and a team with under three
measured fielders is absent rather than thin.

Unknown is not zero, and it bites unusually hard here: an OAA of 0 is a REAL
reading meaning exactly average, so coercing absence to 0 would assert that
every unmeasured fielder is league-average -- the most common defensive
profile there is, and a fabricated fact rather than a neutral default.

team_defense carries as_of_date in its primary key from the first row.
statcast_aggregates was built upsert-in-place with a single as_of date, which
silently made every backtest leak the games it was predicting and cost a
session to find; this makes point-in-time available before it is needed
instead of after a wrong answer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 22:03:26 -04:00
builtbykev ac1361486e Build the conditioning registry, and a probe so "proven" stops drifting
The order opens with "two proven clusters live". They are not proven -- the
proven set is empty -- and this is the fourth consecutive order to start from
a stronger claim than the measurements support. Correcting that in prose four
times has not worked, so this session adds scripts/proven-status.js, which
recomputes the answer from the ledger: hits LOSES (-0.096, CI excluding zero),
total_bases INCONCLUSIVE (+0.004), strikeouts INCONCLUSIVE (+0.259 at n=57).
It deliberately reports sample readiness separately from recorded verdicts, so
"n>=500" can never again be read as "passed".

A counting error worth recording. The first read of the top-volume archetype
said BOMBER x hits was 641 rows -- gate-ready. It is 287. model_snapshots
holds one row per prop PER SNAPSHOT CYCLE, so joining it to ledger_entries
counts each ledger row once per cycle it appeared in. Deduping on the ledger
row id gives the true figure, and my own status script had the same bug until
it was fixed. That is the difference between running the gate and being short
by 213.

So no archetype x stat combination reaches the gate. BOMBER x hits at 287 is
the closest; pitcher archetypes are untestable at 58 settled strikeout rows
across all of them, so the pitcher half of this order could not be run.

The registry is built: recordConditioning keys archetype x underlying-skill x
interaction x status with measured lift, and the skill tag is MANDATORY and
enforced -- untagged entries are refused, and PROVEN without sufficient
evidence is refused. validatedSkills() returns the coherent profile as it
stands, which is {} for every archetype, by design.

BOMBER x hits conditioning was tested across the order's categories and every
result is underpowered: arsenal (barrel x breaking share) incremental +0.043,
batted-ball (launch x pitcher GB) +0.001, contact quality -0.020 and -0.015,
K x K -0.063. Within BOMBER the counter still leads on hits, 0.218 to 0.160,
consistent with the closed pooled negative.

One bug fixed mid-run: fromStatcastRow maps percentage and raw fields only and
does not carry pitch_mix, so the arsenal category first reported n=0 for every
row -- it was measuring nothing rather than failing. Without catching it,
"arsenal doesn't matter" would have been recorded from a column that was never
populated.

On defense: I looked for a derivable proxy before calling it unsourceable, and
there isn't one. We ingest no fielding data at all, and opposing pitchers'
hits-allowed conflates pitching with defense, so it would validate the wrong
skill. It needs Savant's fielding endpoint -- free, same host as the five
feeds already ingested -- and it is not sourced here, because sourcing it to
test at n=282 would answer nothing.

Nothing proved, so nothing was recalibrated and nothing shipped.

4,221 tests green (335 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 21:49:51 -04:00
builtbykev 9538e11198 Derive the lineup K-rate free, and fingerprint the cap fix
Two premise corrections first. Pitcher stuff features have NOT proven solo
through the gate -- every one was refused on sample (n=57 against 500). Four
exceed the effect-size bar (arm angle -0.250, whiff +0.213, k rate +0.206,
chase +0.195), which is why they are worth pursuing, but clearing one of three
thresholds is not passing. And the carrier was not blocked only on the lineup
input: that input was built and measured last session at 94.7% coverage. What
blocks it is n, and n was being throttled by the grading cap.

RUNG 1 IS DERIVED AND COSTS NOTHING. Opposing-team K-rate comes from joining
the opposing roster to the batter k_pct values already in statcast_aggregates
-- no new feed. The improvement this session is that it is PA-WEIGHTED: an
unweighted roster mean counts a 12-PA callup the same as an everyday starter,
which is not the lineup a pitcher faces.

That change alone reversed the term's sign. Unweighted, the lineup term HURT
the model (0.1738 -> 0.1285). PA-weighted, it HELPS (0.1738 -> 0.1953). Same
hypothesis, same data -- the derivation was the problem, not the signal, which
is the entire argument for deriving the best honest version before sourcing
anything. Head-to-head is now +0.2592 with a CI of [-0.0167, +0.5645], very
nearly excluding zero, at n=57.

Within archetype, the two strata come out with OPPOSITE signs -- FLAME
incremental -0.152, non-FLAME +0.145 -- and the pooled value (+0.077) sits
between them, which is the shape a conditional effect makes and is invisible
when pooled. That is what stratifying was for. But n is 20 and 24, the
standard error on a correlation there is about 0.22, and the direction
contradicts the theory that predicted a stronger effect for finesse arms. It
is recorded as a structure to re-test, not as a finding.

Rungs 2 and 3 are NOT triggered. A rung fails only once it has been fairly
tested, and Rung 1 is n-blocked rather than failed. Sourcing confirmed lineups
now would be paying for precision on top of a proxy we have not yet measured.

THE RESULT THAT DECIDES THE TIMELINE: yesterday's cap raise is fingerprinted
in production at 907 grades per snapshot, up from 334, with strikeouts going 6
to 17. That puts n>=500 for pitcher Ks about a week out instead of three
months. Operational note: the manual internal snapshot endpoint now 524s at
the Cloudflare edge because grading the full board exceeds 100s -- the run
still completes server-side (this very snapshot was written by a 524'd
request) and the cron is in-process, so a 524 there is not a failure.

Nothing proven, nothing calibrated, nothing shipped. The counter remains
anti-predictive on strikeouts at -0.064 and the skill model leads it by 0.26.

4,221 tests green (335 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 19:47:01 -04:00
builtbykev 843c8c6d4b Build the pitcher engine, and find the cap was eating the whole board
Strikeouts are NOT proven -- n=57 against a bar of 500. But the finding that
matters is not a correlation.

THE CAP. Measured on the live slate via the refusal diagnostic: 1,244 unique
gradeable props exist, the 500 cap graded about 334, and because dedupeProps
takes first-row-wins in FEED ORDER, what survives is decided by feed position
rather than value. Pitchers are 2.6% of a batter-dominated feed, so we were
grading SIX strikeout props a slate against 32 available -- putting n>=500
three months away for every pitcher stat. Pitcher props were never being
refused (graded 5, refused 0, suppressed 0); it was truncation.

Raised 500 -> 1500 on measured cost: 721ms per prop at concurrency 5 is about
179 seconds for the full board, against a cron that runs five times a day and
a fire-and-forget caller that never holds an HTTP response. statsapi is free
and unlimited. Concurrency stays at 5 -- one variable at a time. This unblocks
every n-blocked stat in the programme, not just pitchers.

THE ENGINE. pitcherEngine.js is its own engine, not the batter engine pointed
at pitchers: the batter model asks whether contact becomes a hit and reads
contact quality, the pitcher model asks whether the plate appearance ends
without contact at all and reads stuff. Archetypes are FLAME (whiff-led),
SCALPEL (chase-led), SINKER (pitches to contact) and DEFAULT, and a test
asserts the weight keys are not the batter engine's. The projection is K% by
log5 against THIS lineup, times batters faced, through a binomial. An
unclassifiable arm gets the balanced map, never a guessed archetype.

THE MEASUREMENT, at n=57 and contaminated. Four solo features clear the 0.15
effect bar and fail only on sample: arm angle at -0.250 -- the largest
correlation measured anywhere in this programme -- then whiff +0.213, k rate
+0.206, chase +0.195. The batter cluster's best was 0.135. Head to head,
pitch-v1 resolves 0.1285 against the counter's -0.0639, delta +0.192 with a CI
spanning zero.

That negative is the interesting number. The counter is ANTI-PREDICTIVE on
strikeouts: counting a pitcher's recent Ks is worse than useless, because his
recent totals track which lineups he drew and how long he was left in rather
than his skill. It is the one stat where the incumbent has no defensible edge.

A bug caught on the way. resolveTeam wants an abbreviation and the game log
supplies full team names, so the roster join silently resolved nothing and the
first run reported 0% lineup coverage -- the theorized stuff x lineup carrier
was never being tested, not failing. Fixed; coverage is now 94.7%. The carrier
still shows no incremental signal over whiff alone, and adding the lineup term
lowered head-to-head resolution, which is recorded rather than dropped.

Calibration was not reached: nothing passed the first bar. The batter model
and the counter are byte-identical, verified by diff.

4,221 tests green (335 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 18:43:32 -04:00
builtbykev c0621e7aa2 Measure the batter cluster: the proven set is empty, and hits is closed
PREMISE CORRECTION FIRST, because it defines the bar. total_bases has not
passed BAR 1. Its head-to-head is inconclusive at parity -- delta +0.004 to
+0.007 with a CI spanning zero -- and it is contaminated, and no feature of
its passed the gate. It was described last session as the first challenger
that did not LOSE, which is not the same as proven. If it is installed as the
frozen proven reference and every other stat is held to "the identical bar
total_bases cleared", the bar becomes "be inconclusive at parity" and the
whole cluster passes on a null result. The proven set is EMPTY.

HITS IS NOW A FINAL ANSWER. At n=803 it clears the gate's sample requirement,
so its features were properly TESTED rather than refused: every one fails on
effect size (max marginal |r| 0.053 against a 0.15 bar), every interaction's
incremental contribution collapses to about zero, and the model loses
head-to-head by 0.096 with a CI excluding zero. That is a well-powered
negative and hits should be closed rather than retried.

The rest are n-blocked: total_bases 383, rbi 391, home_runs 228, runs 188,
against a bar of 500. Two leads are worth carrying. home_runs barrel rate has
a marginal r of -0.135, and the sign matters -- higher barrel rate goes with
the counter OVER-predicting, which would be a correction rather than a new
predictor. And runs batterK x pitcherK has the largest incremental in the
cluster at +0.132, with a clean mechanism: strikeouts destroy plate
appearances, and a PA that never happens cannot score.

RBI deserves a caveat rather than a verdict. It is power times OPPORTUNITY,
and we ingest no baserunner state at all, so half its mechanism is missing. A
weak RBI result is evidence that we are modelling half the stat.

total_bases was held frozen: git diff on skillProjection against the prior
commit is empty. The counter is untouched.

Also fixed and verified in production: the point-in-time retention shipped
after yesterday's refresh had already run, so statcast_history was empty, and
its first run then failed on a hand-enumerated schema that had already drifted
from its source ("could not find the 'swing_pct' column"). The refresh itself
still succeeded and wrote all 1,387 aggregate rows, which confirmed the
best-effort guard in prod. The table now mirrors the source via LIKE and the
writer passes rows through whole. Verified live: 1,387 rows retained at as_of
2026-08-03. A usable point-in-time window starts 2026-08-04.

Stage B has nothing to calibrate. Everything now waits on a point-in-time
window and on sample -- both waiting problems, not building problems.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 16:58:57 -04:00
builtbykev c2d6e8ee7d Mirror statcast_history on its source so retention cannot drift
The first production run of the point-in-time retention failed with
"Could not find the 'swing_pct' column of 'statcast_history'" -- the
hand-enumerated column list had already drifted from the table it was copying.
The refresh itself still succeeded and wrote all 1,387 aggregate rows, which
verified the best-effort guard in prod: a retention failure does not fail the
refresh.

The table is now created with LIKE statcast_aggregates, and the writer passes
the row through whole instead of hand-stripping columns, so there is no drift
surface left. Recreating was safe -- nothing had been retained.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 16:49:15 -04:00
builtbykev 4aab18096f Prove both on total bases -- and find that my own fix destroyed the backtest
Nothing passed. Nothing promoted. Counter byte-identical.

THE BLOCKER, which is the real finding. statcast_aggregates is upserted in
place and holds exactly one as-of date. Yesterday's skill backtest was honest
only by accident: the nightly refresh was unreachable code, so the profiles
sat frozen at 2026-07-21 -- before the settled window. Repairing that cron was
right for production and it refreshed them to today, destroying every prior
version. Scoring a 2026-07-25 game now uses a season aggregate that contains
that game. Point-in-time validation is structurally impossible from that
table, so every number in this run is contaminated and directional, and none
of it is a gate verdict.

Fixed forward: statcast_history retains a dated snapshot on every refresh, so
point-in-time becomes "as_of_date < game_date, most recent". Retention is
best-effort and cannot fail the refresh; both properties are unit-tested. It
has one day of data, which is not yet a window.

SOLO BASELINE, n=383, Bonferroni across 12 tests (alpha 0.00417): nothing
passes. hard_hit_pct is closest at marginal r 0.135 with p 0.0080, failing
both the 0.15 effect bar and the corrected alpha. And it drifted DOWN from
0.153 at n=295 -- an estimate regressing as noise averages out, not an effect
firming up. I called that number encouraging yesterday; on 88 more rows it is
fading, and it should not keep being quoted at its best value.

INTERACTIONS, each scored by partial correlation against the counter residual
controlling for both of its own components: none pass. Only barrel x power
archetype has an incremental exceeding its parts (-0.101 against 0.019) at
n=260 -- the shape Discipline 2 predicts, but a lead, not a finding.

A methodological catch worth keeping. The archetype conditioner was first
built as barrel_pct over league barrel -- a monotone transform of one of its
own components -- so the "interaction" was barrel squared, measuring
nonlinearity in barrel rate rather than any archetype effect, and it produced
this run's only positive result. A Gauss-Jordan pivot test does not catch that,
because the two columns differ by a scale factor. Fixed with a scale-free
collinearity check plus real archetype labels joined from model_snapshots.
Without it this document would have reported a fabricated interaction as the
session's finding.

COMBINED vs COUNTER on total bases: 0.2718 against 0.2647, delta +0.0071, CI
[-0.065, +0.079] -- inconclusive, and the first time a challenger has not
lost. The same engine on hits was -0.116 with a CI excluding zero. That
contrast is the whole argument for total bases, and it is what the physics
said: contact quality governs extra bases, not whether a grounder finds a hole.

Also built: the compound TB projection. skillProjection no longer refuses
total bases -- a deterministic bases-per-hit multiplier had made P(TB>=2)
exactly P(hits>=1), a relabelled hits curve. It is now a convolution over
per-PA base outcomes with hit-type shares shifted by skill. Non-degeneracy is
locked by test.

4,204 tests green (334 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 16:27:11 -04:00
builtbykev c7cc8f5e52 Build the gate, run it, and find we were proving things on the wrong stat
PREMISE CORRECTION FIRST. statModel.js and correlateValidator.js do not exist
in this repository. The validation spec's only prior form is
src/services/python/blueprints/unconventional.py -- a Flask blueprint in the
Python service that is offline in production, scoring NBA factors against a
warehouse that was never populated -- and tests/unit/supplementSystems.test.js
requires only fs and path while defining its own validateFactor inline at line
368. Those tests assert a re-implementation of the thresholds, not an
implementation, which is exactly why they passed for months while nothing was
connected. The diagnosis behind the order is right -- every challenger was
measured without a gate -- but the cause is that there was no gate on the Node
side to import. So it is built, to the exact spec.

correlateValidator: n>=500, |r|>=0.15, p<0.05, Bonferroni across the sweep.
The p-value is exact rather than approximated (t-transform through a
regularized incomplete beta) and is verified in the suite against known
values, because scipy is not available here. Pairs with an unknown side are
dropped, never zero-filled -- a zero-fill inside a correlation does not add
noise, it invents a point at the origin.

THE RUN, hits, n=570, Bonferroni-8: every skill feature fails, and not
narrowly. The strongest marginal correlation against the counter's residual is
0.062 against a 0.15 bar. That is an effect-size failure at a sample that
would have found a real effect comfortably -- a clean, well-powered negative.
The head-to-head agrees: value engine 0.0499 against the counter's 0.166,
delta -0.116 with CI [-0.189, -0.043]. Not promoted.

THE RUN, total bases, n=295: cannot be tested, and that is the finding.
hard_hit_pct shows a marginal r of 0.153 -- above the threshold -- and exit
velo 0.124, refused solely because n is 205 short of 500. It is the most
encouraging number this work has produced, and it is what the physics
predicts: contact quality governs extra bases, not whether a grounder finds a
hole. We have been testing skill inputs on the one stat where they should not
matter much.

Two things the run forced. Feature verdicts are now PER STAT, because marking
these DEAD sport-wide on hits evidence would have killed, for total bases, the
features that look most alive there -- per-sport doctrine one level deeper.
And the gate now reports r and p even when underpowered, because "not enough
data yet" and "nothing here" demand opposite decisions and a bare refusal was
hiding the best signal on the board.

Next: build the compound TB projection (skillProjection still refuses total
bases by design, since a deterministic bases-per-hit made P(TB>=2) identical
to P(hits>=1)), accrue to n>=500, re-run this gate. Leave hits alone.

4,200 tests green (334 suites); web build exit 0; counter byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 02:34:02 -04:00
builtbykev 258d8a6655 The skill engine: built, gated by construction, and Stage A honestly lost
Built src/services/model/ -- the forward, archetype-selected, skill-based
projection, as a challenger. The champion is untouched.

featureRegistry makes "earn its place or it's out" structural rather than
aspirational: CANDIDATE / PROVEN / DEAD per feature per sport, liveFeatures()
returns PROVEN only, promotion requires n>=200 with positive lift and a CI
excluding zero, and there is deliberately no override argument. It ships with
exactly ONE proven feature -- the incumbent counter, because it is the only
one with a measurement. A test asserts that with only PROVEN features allowed
the projection returns null, so an unproven model cannot reach a user by
accident. The three champion adjustment layers are registered DEAD with their
reasons so they cannot be silently rebuilt.

skillProjection is a PA outcome tree: K and BB combined by log5 odds-ratio
against league (both identities unit-tested), then archetype-weighted contact
quality against contact allowed, then Binomial(PA, p_hit) mixed over a PA
distribution. Archetype is a FEATURE SELECTOR, not a nudge -- BOMBER reads
barrels at 0.50 and ground-ball speed at 0.00, GHOST inverts it -- and a test
locks that the same hitter read two ways moves more than 0.15.

STAGE A: IT LOSES. Out-of-sample on 570 settled hits props with 91.9%
opposing-pitcher coverage, resolution 0.0499 against the champion's 0.166,
delta -0.116 with CI [-0.189, -0.043]. It is not selective either: its eight
most confident picks hit 50%, a lift of -0.065. Not promoted. The gate did its
job on its first real test, which is the point of having built it that way.

Two false starts, both recorded because they nearly produced a wrong verdict:
statcast_aggregates stores PERCENTAGES, so raw rows made bip = 1-29.6-17.1 and
refused 568 of 576 -- the honest-absent guards made a units bug loud instead of
silent, and the conversion now lives at one chokepoint. And the first run
resolved an opposing pitcher for 1 of 570 rows, because ledger team/opponent
are NULL, so it would have reported "skill-v1 loses" while measuring a
batter-only model with no matchup in it at all. The verdict above is from the
corrected run.

The loss is real but partial: park was passed as 1.0, handedness and
opportunity_drift never fired, PA is season-PA over a constant, and the skill
profiles carry no recency at all while the champion has a last-5 term.

Also fixed: the Statcast nightly refresh was unreachable code. It sat inside
tick() below "if (!HOURS_UTC.includes(h)) return" while testing h === 11, so
it had never run once; the aggregates were 13 days stale and both of its
alerts were in the same dead branch. It now runs on its own tick, and the test
that passed happily throughout -- it only checked the string existed -- is
replaced by one that asserts it is not behind the guard.

4,182 tests green (333 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 02:20:40 -04:00
builtbykev c551bf0340 Reality assessment: the forward model exists, wired to the wrong side of the pipe
READ-ONLY. src/ and web/ untouched.

Inventoried every forward-model component against the real objective -- a
forward matchup projection, not an edge number. The finding is that all of it
already exists and is already loaded in production, and 100% of it sits
DOWNSTREAM of the grade in challenger columns nothing serves. The served p_win
reads three features and a game log; it has never seen a pitcher.

Inputs are HAVE, not missing: statcast_aggregates carries 1,354 rows (750
pitchers, 604 batters) with exit velo, launch angle, barrel, hard-hit, whiff,
chase, pitch mix, GB/FB, arm angle, and handedness complete on every row. Real
gaps are team defense and catcher/umpire. So Stage A is a plumbing-and-
modelling job, not a data-acquisition job.

Found along the way: the Statcast nightly refresh is unreachable code. tick()
returns for any hour not in HOURS_UTC (14,19,22,1,3) and the refresh block
then tests h === 11, which that guard can never admit. The mechanism data has
been frozen at its 2026-07-21 backfill for 13 days, and the block's own
failure alert sits in the same dead branch -- the identical silently-guarded-
out shape as the settlement outage.

Design shows the counter: every factor label the SIGNAL BREAKDOWN renders is a
restatement of recent frequency (l5_hot_vs_line, l20_over_line, back_to_back,
home_game) plus several structurally-NBA labels (referees, coach pace,
starters out) inside a baseball product. Not one names a pitcher, pitch type,
handedness or park. The card's forward-read slots already exist and go
unfilled -- the surface needs feeding, not redesigning.

On what changes: the prior measurements were outcome-accuracy, not edge, so
the metric was right and the question was narrow. proj-v1.1 and hits-v1 stay
correctly refuted as DISTRIBUTION swaps on thin inputs -- neither tested a
matchup-fed projection. arch-v1 is a market-relative nudge by construction and
is the one component genuinely measured on the wrong axis. AT CEILING is
provisional: measured only against features the champion already reads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-03 00:53:56 -04:00
builtbykev d8bf7765db Decompose the champion: its whole edge is a hit-rate counter
READ-ONLY. src/ and web/ untouched; 4,159 tests still green.

WHAT THE CHAMPION IS. probabilityEstimator is five lines of arithmetic: the
empirical frequency of (stat > THIS line) over the game log, blended 0.6/0.4
with the last-5 frequency, then +/-0.03 opponent, +/-0.015 home/away, a
cv>0.40 pull toward 0.50, and a clamp to [0.10, 0.95]. It reads three
features. featureCache retains a dozen more that p_win never touches.

THE ABLATION IS EXACT, NOT A REFIT. Every adjustment is closed-form from
stored features and the consistency step is linear, so each layer subtracts
algebraically out of the stored p_win -- no re-estimation, no re-fetch, no
lookahead possible. Per stat, paired bootstrap:

  removing ALL THREE adjustments changes resolution by NOTHING on every stat
  hits -0.0059  total_bases -0.0015  rbi +0.0106  runs +0.0130  walks +0.0008

and rbi's home/away is mildly HARMFUL (+0.0053, CI excludes zero). So ~100% of
the champion's resolution is base+recency: how often this player has cleared
this number lately. Everything else is decoration.

A CORRECTION. Pooled, the champion resolves 0.46; per stat it is 0.196 (hits)
to 0.499 (rbi). Pooling stats with different base rates inflates correlation,
so 0.46 should not be quoted as the champion's resolution. Last session's
paired differences remain valid; only the absolute level was inflated.

THE BIGGEST LOSS IS NOT A MISSING FEATURE -- IT IS THE CLAMP. 358 of 1,741
settled rows (20.6%) sit on the boundary, so the model emits a constant there
and cannot rank a fifth of the book at all. And that constant hides two
opposite failures: 0.900 covers home_runs-under truly winning 99.5% (9.5pts
under-confident) next to hits-under truly winning 51.9% (38.1pts over-
confident). PROB_CEIL=0.95 makes the 99.5% case inexpressible. Global
over-prediction is +3.5pts, +7.6 on total_bases. None of this needs new data.

ONE REAL MISSING-WEIGHTING LEAD: opportunity_drift, residual corr +0.156 on
hits and +0.145 on total_bases -- it REPEATS across independent stats, unlike
the weather hits on TB which sit inside the expected false-positive count (70
tests at alpha .05 expects 3-4). And we already compute it: arch-v1's
opportunity axis uses it and extracts nothing (delta +0.0001). Wrong
implementation, not a missing feature -- opportunity must scale the rate, not
nudge the probability.

ARCHETYPE IS UNMEASURABLE, NOT REFUTED. Only 2 of 41 labels (BOMBER, GHOST)
reach n>=40 settled rows and every mean residual straddles zero. That is "we
have not measured it", and it does not license acting in either direction.

Why every challenger has failed is now legible: the ladder and hits-v1 REPLACE
the frequency question with a fitted distribution; the environment axis adds
inputs the champion ignores. Asking the frequency question at the traded line
is the thing that works.

Flagged, not fixed: model_snapshots.outcome is NULL on all 22,032 rows -- the
retention table built for exactly this replay was never settled, so labels had
to be joined from ledger_entries.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-02 23:54:51 -04:00
builtbykev f897c7ec06 hits-v1 fingerprint PASSED: 72/72 written in prod, 45 outside the band modelled anyway
The prod-write fingerprint that was blocked by the odds outage has landed on the
first snapshot after deploy. hits-v1 records exactly as the live-board
verification predicted, and the takeable axis behaves as specified -- scope is
book identity, never price shape.

The verdict is unchanged: hits-v1 is REFUTED and stays unpromoted. This confirms
only that it is recording, so the forward accrual can judge the backtest.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-02 22:08:12 -04:00
builtbykev 3ba3dd28f3 Scoreboard every challenger; diagnose the 429 as odds-api, not PropLine
PROMOTE-THE-EARNED. Nothing was promoted, because nothing earned it -- not
because the bar was held high. Measured on the same bar that refuted hits-v1:
own rows only, direction-aligned, paired bootstrap, promote only on a CI
excluding zero.

  arch-v1        n=1741  delta 0.0000  CI[-0.0050,+0.0054]  inconclusive
  contact-v1     n=1055  delta +0.0008 CI[-0.0052,+0.0069]  inconclusive
  proj-v1.1      n=1664  delta -0.0301 CI[-0.0543,-0.0060]  reliably WORSE
  matchup/tb-v1/hits-v1  n=0  genuinely pending (rows dated 08-02+)

arch-v1 is the interesting one: it MOVED 76% of rows by 2.5 points on average
and resolution is identical to the champion to four decimals, on the moved
rows too. That is active movement carrying no information -- a finding, not a
pending verdict.

These are true prospective holdouts: arch-v1 and contact-v1 wrote p_win at
grade time into their own columns before the game. Nothing recomputed.

THE 429, read-only. The premise was that we re-pull the full picture every
slot and blow the quota. Measured: PropLine is at 5 calls of 3,000/day --
0.17%. One snapshot is ONE PropLine call per sport, all markets comma-joined.
There is no request-pattern problem, so a change-based pull cannot fix it and
no tier upgrade is needed.

The 429 is odds-api: 478/500 MONTHLY, blocked at 95%. oddsService falls
through silently when PropLine returns empty, and the backup's quota gate
throws the error -- so an empty slate is indistinguishable from an outage and
the message names the wrong provider. Flagged for its own order.

Could NOT verify PropLine movement endpoints: docs are auth-gated and the keys
are production-only. Not asserted either way. The movement-as-data argument
stands on its own merits and should be justified that way, not as a quota fix
it isn't.

Book-breadth invariant written down: we never discard books. All are kept and
shown (DISPLAY_BOOKS = MODEL + REFERENCE + DFS); DFS pick'em is excluded from
PRICING only, because a fixed-payout shaded number is not a market price.
Verified this is already what bookRoles.js does.

Champion byte-identical; every challenger stays wired.
4,159 tests green (332 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-02 22:07:43 -04:00
builtbykev b06a84af80 Settlement has been dead since 2026-08-01: a 500-id filter overflowed the URL
The self-learning loop stopped two days ago and reported success the whole
time. 1,444 ledger rows from 2026-08-01 sit unsettled with settle_attempts=0
-- never even attempted -- and every accruing challenger has been starved of
settled sample as a result.

ROOT CAUSE. settleLedger fetched open ids, then REFETCHED the full rows with
.in('id', ids). PostgREST puts filters in the URL, so 500 UUIDs became an
18,499-character request that the fetch layer rejects with "TypeError: fetch
failed". The result was destructured as `const { data: rows } = ...` with NO
error binding, so rows came back null, the loop body never executed, and the
function returned {settled:0, voided:0, unrecoverable:0, pending:0} --
byte-identical to a clean "nothing to settle". Reproduced against prod before
changing anything.

WHY IT HID FOR TWO DAYS. It is volume-triggered. Daily volume ran 20-260 rows
and settled perfectly for weeks; 2026-08-01 was the first day past the 500-row
fetch limit. And the zero-settle ops alarm reads these very return values, so
pending:0 told the watchdog the backlog was empty -- the alarm built to catch
exactly this could not see it.

THE FIX. The refetch existed only to add game_date/settle_attempts/
dclv_computed_at. Selecting them in the first query removes the id list
entirely, so there is no URL to overflow at any volume. A failed fetch now
surfaces its error instead of being reported as an empty backlog.

captureClosing carried the same shape one level down -- .in('id', g.ids) on an
UPDATE, which fails identically once a single line|odds group gets large on a
big slate. Its id filters are now chunked at 100 (~3.7 KB).

Tests: the regression is locked by asserting settlement issues NO id-list
filter at 500 rows, and that a failed fetch is never reported as an empty
backlog -- the two properties that would have caught this. Two existing
suites asserted the old two-query shape and were updated to the real one.

4,159 tests green (332 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-02 21:49:30 -04:00
builtbykev 2394fb04a1 Record the hits-v1 fingerprint as PENDING, and why
The prod-write fingerprint did not land: the odds provider is returning 429
(quota exhausted), so the snapshot refuses with gradeCount 0 and the MLB board
has been frozen since 07:30 UTC. The 14/19/22 UTC cron slots failed the same
way, all before this change deployed -- hits-v1 sits inside the snapshot's
existing try/catch, is purely additive, and had zero grades to attach to.

Firing is already verified against the real production snapshot through the
real attachProjection path (158/159). What is pending is only confirmation
that the deployed process writes the columns, which needs a slate the pipeline
can fetch. The exact fingerprint query is recorded in the spec.

The odds quota exhaustion is a live outage of the whole grading pipeline and
is flagged for its own order, not folded into this one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-02 19:07:58 -04:00
builtbykev 07626de3de hits-v1: built on the right structure, measured honestly, REFUTED
Hits was diagnosed as a family mismatch: 84% of hits rows trade at 0.5, so
the stat rides on P(0), and a negative binomial has unbounded support and no
notion of opportunity at all. hits-v1 models it as the bounded conversion it
is -- N ~ the player's empirical at-bat distribution, hits|N ~ Binomial(N,q),
with the multiplier scaling q (conversion) and never N (opportunity).

STEP 0 confirmed the inputs before the model existed: 30/30 real ledger
players, 100% combined-input coverage. Every read goes through knownRate --
a row with no atBats is dropped, never counted as a 0-at-bat game.

It FIRES: 158/159 hits props (99.4%) on the live production snapshot, through
the real attachProjection path. Scoping by book IDENTITY rather than price
shape kept 94 out-of-promotion-band props on the board, 93 of them modelled --
59% that a price rule would have deleted.

And it LOST. Point-in-time replay (game log truncated strictly before each
row's game_date, real grade-time multiplier), hits-only, direction-aligned,
n=242: resolution champion 0.195 / ladder 0.048 / hits-v1 0.026. Paired
bootstrap on the same rows: hits-v1 - ladder = -0.022, CI95 excluding zero.
Not promoted.

The value is in what it eliminates. The family was wrong AND the mean was not
the constraint -- hits-v1 moved the line-0.5 mean 0.554 -> 0.581 toward a
0.598 base rate while resolution fell. What is left is per-prop
discrimination: the ladder's inputs, not its distribution.

The pre-registered fallback is recorded as WRONG rather than deleted. It said
hits might be genuinely low-resolution for anyone; the champion scores 0.276
on the identical 189 rows, so there is real signal and the ceiling claim was
the comfortable reading, not the honest one. Its own control refuted it, and
that control was already in hand when the branch was written.

hits-v1 stays wired as a challenger writing its own ledger columns so the
forward accrual can confirm the backtest. Champion, ladder, ranking,
calibration, reference ruler and the four accruing verdicts are byte-identical
-- the diff has zero deleted lines.

Tests 4,156 green (332 suites); web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-02 19:04:08 -04:00
builtbykev d103ecf4c3 Disambiguate takeable: THREE questions shared one word, now three names
BYTE-IDENTICAL. The audit found no consumer getting the wrong axis, so this
is a disambiguation, not a bug fix. 4,131 tests / 331 suites green.

STEP 1 AUDIT -- and the order's premise was wrong in a useful way:

  the four accruing challengers   read the flag ZERO times (not four)
  the ranking gate                wants PROMOTION, gets promotion  [correct]
  the ledger column               holds the LEDGER band, consumed as such
  the UI (LiveHeroProp)           TYPES a `takeable` field it never renders

THERE ARE THREE DEFINITIONS, NOT TWO -- and I only found the third by
tracing the ranking gate:

  1. IDENTITY    can it be bet?        book identity (takeability)
  2. LEDGER BAND worth recording?      odds >= -160, UNCAPPED plus
  3. PROMOTION   worth crowning?       -160..+200, i.e. band PLUS a ceiling

(2) and (3) genuinely disagree, and I measured it rather than asserting it:
439 rows -- 28.2% of all takeable=true ledger rows -- carry prices above
+200, up to +1300. A +1300 longshot is a real bet worth RECORDING and not
one worth CROWNING. Both are correct for their own purpose.

THE DANGER WAS NEVER THE LOGIC. It was that three questions shared one
word, so a reader could not tell which answer they held -- and hits, which
must model thin/juiced/one-sided REAL markets, would have been the next
reader to guess wrong.

RESOLUTION: all three now have distinct names in config/takeability.js;
gradeRanking calls isWithinPromotionBand so its intent is self-evident (a
test pins it byte-identical to the old valueEngine call across the whole
price range); the ledger dual-writes within_price_band with `takeable`
kept as a documented DEPRECATED MIRROR so nothing breaks. Column comments
in the database now say what each column actually holds.

I did NOT redefine `takeable` in place. Four readers and a ranking gate
sit on it, and silently changing its meaning under cover of a naming
change is exactly the class of move this session keeps removing.

Gates: 4,131 tests / 331 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 18:19:17 -04:00
builtbykev 8c764c22a4 Structural hardening: unknown-is-not-zero + takeability-is-book-identity
Both guards are ADDITIVE. The full suite (4,111 -> 4,126 tests, 331 suites)
passes unchanged through the migration, which is the evidence that no
currently-correct output moved: served path, champion, reference ruler and
the four accruing challengers are byte-identical.

GUARD 1 -- src/utils/known.js. Number(null)===0 has produced at least SIX
separate defects here, including one in a module written the same week its
author documented the trap. Per-module vigilance has demonstrably failed,
so the rule lives in one place and SEVEN sites now delegate: platoonSplits,
projectionChallenger, challengerProjection, contactChallenger,
statcastAggregateService, consensusRuler, gradeRanking -- plus
compoundTotalBases moved onto knownRate.

Two functions, deliberately: knownNumber (any finite number -- a REAL 0 is
a fact and must survive) and knownRate (non-negative, rejects booleans --
for counts/rates where `true` or -1 is broken, not thin). Collapsing them
is how the next variant gets in. firstKnown() exists because `a || b`
discards a measured 0 and `a ?? b` does not.

MY OWN GUARD HAD THE BUG IT EXISTS TO PREVENT, and its own test caught it:
Number([]) === 0, so an empty array coerced to a measured ZERO. Same trap
wearing a different type. Both helpers now reject objects outright.

GUARD 2 -- src/config/takeability.js. Takeability is BOOK IDENTITY and
never price shape. Baseball prop markets are genuinely thin, juiced and
one-sided, and all three are NORMAL structure: betrivers and hardrockbet
legitimately quote one side only (5 such rows surfaced in yesterday's
re-stamp), and a hits-over at -300 is a real placeable bet. A rule that
inferred un-takeability from price extremity or one-sidedness would throw
those away while still admitting a DFS book at an ordinary -119 -- exactly
backwards, because the -119 is the fake one.

THE DISTINCTION THAT MUST NOT COLLAPSE, now enforced by test:
  isTakeableMarket(book)  -- CAN it be bet?     (identity)
  isWithinPriceBand(odds) -- SHOULD we promote? (policy band, floor -160)
A -300 DraftKings prop is takeable AND out of band; a PrizePicks -119 is in
band AND not takeable. Independent axes.

FLAGGED, NOT SILENTLY CHANGED: the ledger's `takeable` column is the
PRICE-BAND answer, and its name predates this distinction. Four challengers
and the ranking gate read it, so renaming or redefining it is its own
order -- doing it here would have changed correct current behaviour under
cover of a hardening change.

Fixtures are REAL prod rows from the 2026-08-02 re-stamp, not invented.

Gates: 4,126 tests / 331 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 17:25:40 -04:00
builtbykev f67245e1e5 Re-stamp A: 862 rows recovered by honest join (not 936 -- see deviation)
Database only; no application code changed, so the served path, champion
and reference ruler are byte-identical.

RESULT: 862 rows re-stamped from the takeable LOCK-TIME price in
lock_lines, 862/862 now anchored to takeable books, tagged
price_source='archive_restamp', quarantine lifted. 812 pending clean rows
recovered into the accruing verdicts. Holdout verification: 2,792 rows,
862 re-stamped included, 0 re-stamped rows non-takeable, 144 still
excluded, 0 quarantined rows leaked, and 0 NON-TAKEABLE rows remain in the
holdout population since 2026-08-01.

DEVIATION, stated rather than buried: the order authorised 936. That
figure came from a takeable book posting the same LINE. Requiring what a
re-stamp actually needs -- that book's price for the GRADED SIDE at LOCK
TIME -- resolves 862. Of the other 74, 73 have a takeable side-price only
OUTSIDE the lock window and 5 are genuinely one-sided markets.

I did not widen the window to reach 936. A takeable price captured hours
after the grade is a later market moment, not a lock price; substituting it
is precisely the reconstruct-vs-join line this order was fenced against,
and it would have been invisible in the totals -- showing only as a
cleaner-looking 936.

Those 74 were also RE-TAGGED, because their old label had become a lie:
recoverable_same_line -> no_takeable_lock_price_for_side. A future attempt
reading the old tag would have been invited to widen the window and call it
recovery.

takeable was RECOMPUTED from the recovered price rather than carried over
-- the old flag was computed FROM the contaminated price and was wrong on
its own terms. 101 rows had their flag change, which is the direct measure
of how wrong it was.

Provenance travels with the data (price_source), on the same principle as
is_proxy: a value recovered by a later join is not identical in kind to one
captured natively at grade time, even when it is the same number.

EVIDENCE FOR THE NEXT ORDER'S INVARIANT: 5 of the excluded rows are
one-sided TAKEABLE markets, and betrivers/hardrockbet legitimately quote
one side only. A guard that inferred takeability from price shape would
throw away real markets while still admitting a DFS book at -119 --
takeability is book IDENTITY, never price extremity or one-sidedness.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 16:22:06 -04:00
builtbykev e29ab6fd6a Takeable enforcement: verified on real rows, 1,006 tagged, re-stamp call ready
PART 1 verified by inducing the REAL rowsFromSnapshot over REAL lock_lines
rows from prod. Three cases, 0 non-takeable anchors:
  Narvaez  (dabble/kalshi/prizepicks/smarkets, NO takeable book)
           -> book=null, price=null, takeable=null  [honest absent]
  Schwarber(bovada/dabble/novig/PINNACLE before draftkings)
           -> draftkings +102  [pinnacle SKIPPED, proving TAKEABLE not MODEL]
  Ohtani   (dabble/onexbet before draftkings) -> draftkings -266
Narvaez is the case that matters: pre-fix he was stamped dabble +104
takeable=true; he is now honestly absent.

A HARNESS BUG RECORDED: my first verification pulled live /api/odds/mlb,
which returned {"error":"Odds data temporarily unavailable"}. The script
read that as 0 props and printed "all from takeable books? true" -- a
VACUOUSLY TRUE pass. I caught it only because I also printed the book list
and it was empty. Same family as the silent-false traps: a probe that finds
nothing looks identical to a probe that finds nothing wrong.

PART 2: 1,006 rows tagged via the purpose-built quarantine_reason at ROW
level with three sub-cases (recoverable_same_line 936, no_takeable_quote
49, takeable_line_differs 21). getModelAggregate ALREADY excluded
quarantined rows, so the public record and the n>=20 gate were clean
automatically; all five committed holdout scripts now carry the exclusion
explicitly.

PART 3 -- the re-stamp call is now fact-based. The takeable LOCK-TIME price
is recoverable for 936/1,006 (93.0%) from lock_lines, the correct
instrument. Only 431 appear in closing_captures, which is the wrong timing
for a lock price anyway.

LINE CONTAMINATION ANSWERED (previously unverified): the stored line
MATCHES a takeable book's line on 936 (93.0%), DIFFERS on 21 (2.1%), and is
unverifiable on 49 (4.9%) where no takeable book quoted the prop at all.

That makes it cleanly row-level: re-stamp the 936 as an honest JOIN and
recover 886 pending rows for the holdouts, or leave all 1,006 excluded.
Either way the 21 + 49 stay out -- re-stamping those would invent a lock
price, or a line, we never captured. Nothing re-stamped; Kev's call.

Gates: 4,111 tests / 330 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 14:24:28 -04:00
builtbykev 5de464330c URGENT: anchor the ledger price/book/takeable to TAKEABLE books
Ships before tonight's settle. Served path, champion, ranking and the
reference ruler are untouched.

TWO leaks, not one. The audit found ledgerService.indexProps; tracing the
lock price found that snapshotService.indexOdds has the SAME defect -- it
also indexed the full props list, so gradedAt.odds (the price a grade is
locked at) could itself be a DFS or exchange price. Fixing only the ledger
would have left the contamination flowing in through the lock.

Both now gate on TAKEABLE_BOOKS -- deliberately NOT MODEL_BOOKS. pinnacle
is model-eligible and correctly not takeable, so a MODEL gate would
re-break this the moment pinnacle's feed recovers. A test asserts pinnacle
cannot anchor a price.

TWO INDEXES, TWO ROLES, because the row needs two different things from a
prop and they have different correctness rules:
  PRICE / BOOK / TAKEABLE -- takeable books only.
  GAME FACTS (game_time, game_date, team/opponent) -- book-INDEPENDENT.
    First pitch is first pitch whichever book listed it, so these still
    come from any book. Gating them too would drop otherwise-valid rows
    for no gain.
Collapsing those roles into one index is precisely the bug.

No takeable quote leaves the key ABSENT and the price null. An honest
missing price beats a price from a book you cannot bet -- and it keeps the
takeable flag from being computed off a DFS number, which is what made it
wrong on its own terms rather than merely mislabelled.

Gates: 4,111 tests / 330 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 14:17:51 -04:00
builtbykev 08e5c908e6 Takeable audit: the ledger is contaminated, and I caused it
READ-ONLY. Nothing enforced or fixed; the five challengers untouched.

VERDICT: gaps exist, and one is LIVE CONTAMINATION of the ledger -- the
exact table every accruing holdout resolves against. book, locked_odds and
the takeable flag ITSELF are being stamped from books you cannot bet: DFS
dabble (707 rows, 24% of all rows), offshore bovada (214), onexbet (42),
exchange kalshi (7, mean |odds| 1120).

0% before 2026-08-01. 47.9% on 08-01. 42.5% on 08-02. It began the day I
widened the books for display.

LEAK LOCATED, not inferred: recordPipelineGrades indexes byKey over the
FULL display-widened props list, then prefers that prop -- book:
(prop && prop.book) || g.book, and locked_odds/takeable both fall back to
oddsForSide(prop). The grade is computed on a MODEL book and the ledger row
is then re-stamped from whatever book indexed first. The takeable flag is
therefore not merely mislabelled: it is computed FROM the contaminated
price, so it is wrong on its own terms.

The served grade path is clean TODAY (428 grades, 100% MODEL books), so
dedupeProps' gate works. But MODEL_BOOKS is NOT a subset of TAKEABLE_BOOKS
-- pinnacle is model-eligible and correctly not takeable -- so the
projection may anchor to a reference line by design. Harmless while
pinnacle returns nothing; live again when it recovers.

BLAST RADIUS bounded but growing: 47 contaminated rows have already
settled (21% of settled rows since 08-01) and ~700 are still pending and
will settle into the holdouts. The damage is mostly ahead of us, which is
what makes this urgent rather than historical.

NOT VERIFIED and not claimed either way: whether the stored `line` is also
contaminated. It traces to the graded prop, but I did not check it
end-to-end; the enforcement order should.

The prediction-vs-reference distinction HOLDS and must not be collapsed:
the prediction target must be takeable, while fair_prob / consensus / edge
stay reference. The bug is not the three-way split -- it is that one write
path ignores it.

Stack sequenced in the plan: (a) takeable enforcement, (b) structural
Number(null)===0 guard (hits will re-trigger it -- its 0.5 lines make P(0)
the whole game), (c) hits. Carry-forward: tb-v1 verdict, the third
pre-registered branch, and the 100s Cloudflare timeout vs a ~115s snapshot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 14:09:14 -04:00
builtbykev aa1228ec42 tb-v1 report + plan: diagnosis on trial, branch pre-registered
Firing verified on a real prod snapshot: 10/10 total_bases props carry
proj_tb_p_over. The snapshot HTTP call returned 524 (Cloudflare's 100s
origin timeout vs a ~115s snapshot) but the work completed server-side --
confirmed from the ledger rather than assumed.

Face validity is good and diagnostic: means agree almost exactly with the
ladder (1.813 vs 1.833), so this is a SHAPE-ONLY intervention, which is
what was intended. Component rates are plausible, and Carroll's triples
rate (0.112, far above his peers) is a clean check -- he is a speed player
and the model sees it.

AN OBSERVATION I AM NOT RESOLVING BY EYE: tb-v1 reads systematically LOWER
than the ladder (0.424 vs 0.540 at the same mean). That is the expected
DIRECTION, since the NB overstates P(>=2) by treating a home run as four
accumulating events -- but whether 0.424 is right or an overcorrection is
not knowable from face validity. A ~1.8-TB hitter clearing 1.5 empirically
sits nearer 45-50%, between the two. I am not claiming tb-v1 is better; the
holdout decides.

BRANCH PRE-REGISTERED, before the result, so the verdict cannot be
reinterpreted afterward: improves -> family-mismatch HOLDS, similarity
stays off the critical path, hits is next; does not improve -> hypothesis
WRONG and the mean-weakness/similarity branch REOPENS.

Also recorded: I hit Number(null)===0 in my own new module -- a null
component rate treated as a measured zero, the difference between "never
triples" and "we don't know his triple rate". A test caught it. Sixth
appearance of this trap in this codebase, and it caught the person writing
the warnings about it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 03:34:14 -04:00
builtbykev eabf3b5bcf tb-v1: model total_bases as a compound outcome (challenger)
Current ladder (proj_p_over_line) and champion p_win are BYTE-IDENTICAL.
tb-v1 writes alongside them, on total_bases props only.

STEP 0 -- components confirmed on real data, not assumed. statsapi has no
singles field, but hits - doubles - triples - homeRuns reproduces stored
totalBases EXACTLY on a real 10-game log. So the decomposition is exact,
not an approximation.

THE MODEL. Each component gets its own per-game Poisson rate; TB is their
weighted sum, and the PMF is built by exact convolution rather than
simulated (TB support is small). It inherits the SAME combined multiplier
proj-v1.1 computes, so the two models differ only in STRUCTURE.

Why this is the fix: with identical mean TB of 1.0, a pure-HR hitter and a
pure-singles hitter get P(TB>=4) of 0.221 vs 0.019 -- a 12x difference an NB
on TB alone cannot express, because it treats one home run as four events.
A test asserts that separation, and asserts P(TB>=4) for a pure-HR hitter
equals P(at least one HR) exactly.

INDEPENDENCE IS AN APPROXIMATION AND IS LABELLED AS ONE: a plate appearance
that becomes a double cannot also become a single, so the components are
weakly negatively correlated and independent Poissons slightly overstate
the tail. Closer to the truth than what it replaces; not a solved problem.

HONEST-ABSENT throughout: fewer than 3 usable games, or no derivable
component, returns null and the prop keeps the current ladder value. An
inconsistent row (hits < extra-base hits) is SKIPPED rather than clamped to
zero -- clamping would invent a plausible line out of a broken one.

I HIT THE Number(null)===0 TRAP IN MY OWN CODE and a test caught it: a null
rate passed a naive finite check and was treated as a measured zero, which
is the difference between "this player never triples" and "we do not know
his triple rate". Both tbPmf and tbMean now reject null/''/boolean strictly.

Holdout committed: TB ROWS ONLY (49 of 437 settled -- averaging into other
stats would hide the effect) and DIRECTION-ALIGNED, since the unaligned
comparison is the artifact that accounted for 41% of the ladder's apparent
loss. If tb-v1 does NOT improve, the family-mismatch hypothesis is wrong
and the mean/similarity branch reopens -- recorded in the query header.

Migration applied: proj_tb_p_over + proj_tb_meta, NULL-meaningful.

Gates: 4,104 tests / 329 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 03:29:23 -04:00
builtbykev 48706210fe Diagnose proj-v1.1: concentrated mean failure, NOT a similarity problem
READ-ONLY. Nothing built or fixed; the four challengers untouched.

41% OF THE REPORTED GAP WAS A MEASUREMENT ARTIFACT. p_win is P(graded
side); proj_p_over_line is P(over); 31.4% of settled rows are UNDER-graded,
so comparing them raw measures the ladder backwards on a third of the
sample. Matched + direction-aligned (n=437): 0.252 vs champion 0.352, not
0.108 vs 0.331. The PRODUCT is not making this mistake -- I checked;
projectionChallenger normalises both to the over basis deliberately. The
error was in the measurement.

THE LOSS IS CONCENTRATED. hits (n=245, res 0.060) and total_bases (n=49,
res 0.009) are 67% of rows and carry essentially no signal. Everything else
is fine or better: walks 0.519 vs champion 0.544, runs mean 0.345 vs 0.392,
and on DOUBLES the ladder's mean BEATS the champion's (0.207 vs -0.062).

IT IS THE MEAN, NOT THE SHAPE. On the two failing families the mean itself
carries no signal (0.052, -0.019) against the champion's 0.158 and 0.085.
Where the mean is good the probability is good -- shape follows mean.

A HYPOTHESIS I TESTED AND DISPROVED: prediction compression. I expected
P(>=1 hit) to sit in a narrow band and fail to rank. It does not -- spread
ratio 0.94 overall, 0.80 for hits, 0.94 for total_bases. The ladder has
comparable spread; it is spread in a direction uncorrelated with outcomes.
Recorded because it was a plausible story the data refused.

PRIORS AND PLUMBING CLEAN. proj_factors carries form_rate,
combined_multiplier and breakdown on every row; proj_point 100% populated
with sane centres (hits 0.830 vs line 0.578). Not the environment-style
silent-null failure.

NAMED CAUSE (structural, flagged as hypothesis not finding): the count
model mismatches those two stats. total_bases is a WEIGHTED SUM (1B..HR =
1..4), so an NB treats one home run as four events and mis-states variance
-- and TB has the worst result in the table. hits is BOUNDED BY AT-BATS and
mostly traded at 0.5, so almost everything rides on P(0), the region where
the wrong family hurts most. walks/runs/doubles ARE genuine low-rate counts
and are exactly the ones that work.

FIX BRANCH: targeted per-stat fix for hits and total_bases. THIS REMOVES
THE MLB SIMILARITY BUILD FROM THE CRITICAL PATH -- that branch assumed a
GLOBAL mean weakness, and the mean is fine or better on three of six stat
families. Similarity may be worth building later, on evidence, not on this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 02:50:13 -04:00
builtbykev f5997778a2 Dormant-layer audit: nothing to connect; proj-v1.1 is live and losing
READ-ONLY. Nothing connected, built or wired; the accruing challengers were
not touched. "Dormant" meant three different things and in no case is the
answer "connect it".

DISTRIBUTION LADDER IS NOT DORMANT. projection/distribution.js is consumed
by projectionChallenger (proj-v1.1), live on every snapshot at 94.2%
coverage (276/293) with 437 settled rows since 2026-07-23. It is a FOURTH
accruing challenger, and it is LOSING: resolution 0.108 vs the champion's
0.331. That verdict is no longer thin.

It is also PER-STAT and doctrine-correct -- nine distinct league priors
(hits 0.90, total_bases 1.45, home_runs 0.15, ...) each feeding a
gamma-Poisson posterior into a negative binomial. Correcting the plan:
§10.3's "single additive index across hits/Ks/TB" is engine1's GRADE, not
this ladder, which made a solved problem look open.

SIMILARITY IS WRONG-SPORT. Zero callers, and its weights are NBA
vocabulary: pace 0.15, referee_tendency 0.06, lineup_context 0.12,
score_state_context 0.05, travel_fatigue 0.08. MLB has no pace and no
referees. Connecting it would be the sport-stubbed-in-on-another-sport's-
template breach, and it would fail QUIETLY -- missing factors are skipped,
so the score would silently collapse onto whatever few dimensions happened
to exist. CONSTRUCT, not connect.

BAYESIAN WOULD REGRESS THE MODEL. Zero callers, and DISTRIBUTION_SHAPES
keys on rbis / runs_scored / strikeouts_batter / outs_recorded /
pitcher_strikeouts / walks_allowed / pitches_thrown -- NONE of which are
live stat keys (S41: they are rbi / runs / outs / strikeouts).
getDistributionShape defaults to 'normal' on an unknown key, so wiring it
as-is would model COUNT stats as Gaussian, silently, on most MLB props. It
is also superseded by distribution.js. Do not connect; retire or rewrite.

DEPENDENCY, inverted: a better mean would help the ladder, but the ladder
is already connected and both would-be foundations are unusable -- so this
is not "connect similarity first", it is "the ladder is live and
underperforming, and strengthening its mean requires BUILDING an MLB
similarity layer that does not exist".

Next-order pointer moved to diagnosing proj-v1.1: the only candidate
already carrying settled evidence, and its diagnosis decides whether the
similarity build is worth doing at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 01:51:00 -04:00
builtbykev ec815b0e37 Matchup axis report + plan reconciled: three challengers now accruing
Records the verification that matters: firing measured on a real prod
snapshot rather than inferred. environment 248/293 (84.6%) -- also its
FIRST confirmed ledger write, which the previous session could only infer
-- and matchup 243/293 (82.9%) on tier batter_own_split. Both were 0/634.

Collinearity guard passed at n=243: r = -0.003 vs the projection, +0.074 vs
p_win, +0.003 vs line, -0.068 vs environment, -0.150 vs opportunity. The
axis is not re-encoding recent form. The nudge distribution is also the
right SHAPE -- mean +0.0007, 123 positive / 120 negative -- a balanced
two-sided signal; a one-sided distribution would have suggested a sign or
baseline error.

Plan reconciled in place: arch-v1 condition axes marked firing, three
challengers listed with coverage and their own holdout queries, and the
next-order pointer moved to connecting the still-dormant layers
(similarity, Bayesian, distribution ladder) with archetype_x_archetype as
the named alternative.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 01:15:51 -04:00
builtbykev 9ebd77b68e Build the matchup/platoon axis: three joins fixed, axis now FIRES
The axis was already wired and firing on 0/634 prod rows. Three separate
absences kept it silent, and all three are now joined:

1. oppPitcherByTeam 0 -> the self-origin /api/schedule/mlb/pitchers route
   returned nothing in prod. Added the statsapi probable-pitcher hydrate as
   a fallback, mirroring the one the schedule step already uses. 29/30
   team-sides, one free request.
2. handById 0 -> follows from (1); the batched people call now has ids.
3. bats 0/120 -> batter hand rode ONLY on statcast aggregate rows, which do
   not cover the slate. The season player list we ALREADY fetch and cache
   carries batSide on 1342/1342, so this is a join, not a fetch.
   Switch-hitters ('S') are preserved as-is; platoonSplits decides what to
   do with them, not the map.

Verified end-to-end against the live API: opp_declared 29,
pitchers_with_hand 29, batters_with_hand 1342, and a real read --
multiplier 0.966, L vs R, 287 observed PA, weight 0.324 -- composing
alongside environment in one challenger.

FALLBACK LADDER, and a deliberate deviation from the order. Shipped tier:
`batter_own_split` (the hitter's OWN vs-L/vs-R line, regressed toward HIS
OWN overall rate), labelled on every adjustment.

`league_generic` is deliberately NOT implemented. platoonSplits already
handles thin evidence by regressing toward the hitter's own rate, which
covers the thin case per-player; its own doc-comment argues a hitter with
no split evidence should get NO adjustment. A league split applied to such
a hitter models the LEAGUE, not the player -- the doctrine breach the order
itself names in the same step. Adding it would have produced more firing
rows and a weaker signal.

`archetype_x_archetype` is scoped, not built: it needs the opposing
starter classified per game, which is real work and a separate order. The
tier vocabulary is in place for it.

Honest-absent on every join: no starter, no pitcher hand, or no batter hand
-> NO matchup adjustment, never a fabricated neutral. A neutral multiplier
produces no adjustment row at all.

Holdout committed (scripts/matchup-axis-holdout.sql), filtered to
matchup-carrying rows, and it keeps MATCHUP'S OWN nudge visible rather than
only the combined challenger -- arch-v1 composes four axes into one
p_win_challenger, so a combined-only view could not tell which axis earned
the movement, or which one is dragging.

Champion p_win, ranking, calibration, the armed invariant and the two
accruing verdicts are untouched.

Gates: 4,093 tests / 328 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-02 01:11:05 -04:00
builtbykev 9fc17a4689 Fix the team resolve properly: backfill the name BEFORE confirmation
My first attempt did not work in prod -- team stayed 0/323 after deploy.
I resolved the team name AFTER the hint-confirmation check, but the check
itself reads hit.currentTeam.name, which is undefined because
/sports/1/players returns { id, link }. With a FULL-NAME hint (what
snapshotService passes) neither branch of teamRecordMatchesHint could
match: the name branch had no name, and the abbr branch cannot resolve a
full name to an abbr. Confirmation failed, the team was nulled, and my
later backfill ran on an already-null value.

withTeamName() now backfills the name from the cached /teams list BEFORE
any comparison, and is used at all three confirmation sites plus the
return. Verified against the live API on all four cases: no hint, FULL-NAME
hint, abbr hint -> "Philadelphia Phillies"; WRONG hint -> null.

That last case matters most: a wrong hint must still REFUSE. The
confirmation exists so a namesake collision cannot tag a player to a team
he is not on, which would fabricate opponents downstream. Making the match
succeed must not make it succeed wrongly, and a test locks it.

Gates: 4,087 tests / 327 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 23:34:16 -04:00
builtbykev cfda597fb5 Reconcile MASTER-PLAN to true state; next order = matchup axis
Reconciled in place, not regenerated. Next-order pointer now MATCHUP AXIS
with its verified sourcing table, and an explicit note that
SOURCE-LINEUPS-first is NOT needed.

Marked DONE with their evidence: p_win ranking + edge retirement,
calibration DECIDED, MLB isotonic DECIDED (provisional label retracted),
grade cap 25->500 (board 7->365+), book widening, S59 invariant armed,
environment axis repaired.

Records the honest shape of Phase 1: it is further along than the phase
table implied, but mostly because the work turned out to be CONNECTION AND
REPAIR rather than construction -- the ladder question dissolved, the cap
was discarding 95.7% of the slate, and two condition axes were wired but
firing on zero rows.

Carried forward without softening: WNBA is NOT BUILT rather than failed,
and the ruler is MARKET-not-SHARP with PENDING-RECOVERY status until
PropLine answers the Pinnacle question -- not to be enshrined as permanent.

Remaining ~19 orders, ~9 unblocked. The two accruing verdicts are time,
not code.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 23:30:03 -04:00
builtbykev 03efdda33c Arm the S59 invariant by fixing its input; matchup sourcing = BUILDABLE
PART 1 -- PREMISE CORRECTION, then the real fix.

The order said the invariant's blocker was removed because "team is now
populated 416/416". It is not: what became 416/416 is home_team/away_team.
`team` (the PLAYER'S roster team) is still 0/416. Arming the guard off
home_team would compare the prop's game to itself -- always a match, a
permanent no-op that LOOKS armed. That would be worse than leaving it
disarmed, because it would read as a working guard.

The guard is also ALREADY fail-safe by construction (`if (knownTeam &&
gameTeams && ...)`), so Part 1's requirement was met in code all along.
What was missing was the data.

ROOT CAUSE: /sports/1/players returns currentTeam as { id, link } with NO
name, so searchPlayer's `hit.currentTeam?.name` was ALWAYS undefined and
every resolve returned team: null. The id is present on 1342/1342 and the
/teams list (already cached 24h) maps id -> name, so resolving it costs no
new request. Verified: Schwarber -> Philadelphia Phillies, Ohtani -> Los
Angeles Dodgers, Judge -> New York Yankees.

Five tests lock the fail-safe: drops only on a positive not-in-game;
abstains on unknown player team; abstains on unknown game participants;
and a row carrying only home_team/away_team does NOT satisfy the guard --
so the tautology can never be reintroduced.

PART 2 -- MATCHUP SOURCING: BUILDABLE. Measured on tonight's real board
against the free feeds, by VALUE not endpoint presence (the environment
trap: wired and null 634/634):

  opposing starter   29/30 team-sides (home 14/15, away 15/15)
  pitcher hand       1342/1342 (pitchHand.code)
  batter hand        1342/1342 (batSide.code; L 416 / R 848 / S 78)

SHARED DEPENDENCY, and it is the finding: /sports/1/players -- a list we
ALREADY fetch and cache -- carries currentTeam.id, batSide AND pitchHand.
One join unlocks the invariant's input and two of the three matchup inputs
at once. The third (probable starter) comes from the schedule hydrate that
already exists.

So matchup is BUILDABLE and is the next order; SOURCE-LINEUPS-first is NOT
needed. Archetype-level reach on the opposing starter is available too
(the SP resolves to a player id, so the existing classifier applies) --
noted, not built.

Champion p_win, ranking, calibration and both accruing verdicts untouched.

Gates: 4,082 tests / 327 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 23:29:00 -04:00
builtbykev 0d43fb7db8 arch-v1 axis audit: env/matchup were dead; environment fixed
Report for the audit + fix already committed. Records the two things worth
carrying forward:

1. The environment axis has NOT yet been observed writing to the ledger,
   and I am not claiming it has. recordPipelineGrades upserts with
   ignoreDuplicates and dedupes on (user_id, player_key, stat, line, side,
   game_id) -- correctly, so a re-run never overwrites the original lock.
   Today's 429 rows predate the fix, so the axis cannot backfill onto them;
   first ledger observation is tomorrow's slate. What IS directly verified
   is the resolver (105/120) and the join key (416/416) -- the two things
   that were actually broken.

2. Matchup is not fixed and is not claimed as fixed. It needs the opposing
   starter and BOTH hands, and the audit shows three separate absences:
   oppPitcherByTeam 0, handById 0, bats 0/120. Fixing the pitcher feed
   without the hands, or the hands without the feed, still produces an axis
   that fires on zero rows.

Also noted: the S59 slate JOIN INVARIANT keys off the same null `team`
field, so it is currently inert too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 04:12:28 -04:00
builtbykev 4435856f46 Audit finds env/matchup axes DEAD in prod; fix the environment join
STEP 0 AUDIT -- the "already partly live" premise was half true: the CODE
is wired, the axes are NOT firing. Across 634 graded prod rows the
environment and matchup axes fired on ZERO rows, while 13 archetype axes
fired normally (power 80, swing_miss 69, contact 56, launch 51,
line_drive 43, ...) plus opportunity 142. Ledger confirms it from the
other side: env_multiplier, env_park_base, env_weather_mod, wx_forecast
and env_weather_state are ALL null on 634/634.

ROOT CAUSE, located rather than inferred. A drop-off audit against the
live snapshot: with_team_field 0/120, with_bats 0/120, with_playerId
120/120, oppPitcherByTeam 0, handById 0. `team` is a KEY on every stored
grade and NULL on 416/416 -- so an environment resolver keyed off the
player's roster team could never find a venue, while buildContext sat
there with all 30 teams mapped and 14 weather forecasts resolved and
unused. Coors composes to 1.241 the moment it gets a key.

FIX -- and it is the more correct join, not just a workaround. The park
and the weather belong to the GAME, not to the player's roster team, and
the game rides on the prop from the odds feed. gradeBestSide now carries
home_team/away_team onto the graded row (the legacy grade shape dropped
them), and contextFor joins on the game first, keeping the roster team as
a fallback. This no longer depends on a stats-resolve that can
legitimately fail.

MATCHUP/PLATOON IS NOT FIXED HERE and is not claimed as fixed: it needs
the opposing starter and both hands, and the audit shows
oppPitcherByTeam=0, handById=0 and bats=0 on the slate -- three separate
absences. Per "one axis at a time" that is its own order with its own
diagnosis, not a second fix smuggled into this one.

Champion p_win, ranking, calibration and opportunity_drift's accruing
verdict are all untouched.

Gates: 4,077 tests / 326 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 04:05:43 -04:00
builtbykev c9d56d5668 Audit endpoint: where the env/matchup context chain drops off
The arch-v1 environment and matchup axes fired on ZERO prod rows across
634 graded props while archetype axes fired normally, and buildContext
works locally (15 games, 14 with weather, Coors composing to 1.241). So
the failure is downstream of buildContext and has to be located, not
inferred from an absence.

Replays buildContext + contextFor against the CURRENT cached snapshot
grades and counts the drop-off at each hop: team field present -> resolves
to an abbr -> abbr matches a game -> environment produced; and bats /
playerId / opposing-pitcher known -> matchup produced.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 04:01:57 -04:00
builtbykev 48a2f764ac opportunity_drift: coverage 94%, collinearity PASSES, holdout n-blocked
STEP 1 -- input mapped and measured. opportunity_drift 94% coverage on 100
real props: 100% for batters (total_bases, hits, home_runs), 40-67% for
pitchers, which is correct -- pitchers accumulate few at-bats so the ratio
is genuinely undefined and ABSTAINS rather than being invented.

STEP 2 -- THE COLLINEARITY GUARD PASSES DECISIVELY. Pearson r on n=94:
drift vs l20_avg -0.020, vs l5_avg +0.027, vs ab_per_game -0.029. All
essentially zero, so the axis is orthogonal to every existing projection
input and carries information the projection does not already contain.

That also validates the ratio-over-level decision EMPIRICALLY: ab_per_game
is the same quantity over the same denominator as l20_avg, so the level
would have been redundant. Dividing by the player's own baseline removed
the collinearity -- r = -0.029 against the very quantity it is built from.

STEP 3 -- live as a challenger, verified on prod over an induced 416-grade
snapshot: 142 of 276 rows (51.4%) carry the opportunity axis, the
challenger moved on 190 rows, mean |delta| 0.034, range -0.089..+0.108.
Champion p_win and the live grade path are unchanged.

STEP 4 -- HOLDOUT IS n-BLOCKED BY CONSTRUCTION and I am not manufacturing
one. Settled rows carrying the axis: 0. Its first rows carry game_date
2026-08-01 -- games that have not been played. Running the test on rows the
axis never touched would dilute the comparison with rows where challenger
=== champion by construction, making a null result look like a small
positive one. Query committed for when n arrives; it filters to
axis-carrying rows for exactly that reason, buckets before measuring
reliability, and splits time-forward. BOTH metrics must improve or the axis
is shelved.

A MEASUREMENT TRAP RECORDED: the first prod run showed drift at 0% while
ab_per_game read 94% -- indistinguishable from "the feature does not
compute". It was the 120-second feature-vector cache serving payloads
written by the previous image. A new feature field is invisible for one
cache generation after deploy. I nearly reported it absent, having already
confirmed atBats is present in the live statsapi payload and that the code
produced drift = 1.05 locally on that exact data; the contradiction
between those two facts is what saved it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 03:23:24 -04:00
builtbykev 092f8f09cd Build opportunity_drift axis on challengerProjection (arch-v1)
Champion p_win and the live grade path are BYTE-IDENTICAL: the axis writes
only to p_win_challenger / challenger_adjustments in the ledger.

STEP 1 -- MAP THE INPUT. MLB_LOG_FIELD now maps at_bats -> 'atBats'.
Deliberately NOT added to outcomeService's map or liveTracking's
LIVE_BOX_FIELD: those exist to SETTLE and TRACK graded props, and nothing
grades at-bats, so adding it there would imply a settlement path for a
market we do not carry. A test asserts the settle map still lacks it.

STEP 2 -- DRIFT, NOT LEVEL. opportunity_drift = mean(last-5 atBats) /
(season atBats / games). The LEVEL is collinear with l20_avg (same
games denominator; hits/game ~= (hits/AB) x (AB/game)), so the projection
already embeds it multiplicatively and adding it would double-count. A
deviation from the player's own baseline is the part the projection does
not contain.

HONEST ABSENCE throughout: fewer than 3 at-bat rows, no at-bats in the
logs, or no season baseline all leave drift UNDEFINED -- never 1.0 by
default and never 0. Number(null) === 0 here would read as "zero at-bats",
the strongest possible fade, invented from missing data. Four tests cover
the absent paths.

STEP 3 -- THE AXIS. opportunityNudge composes in the same log-odds space
as park and platoon (log of a ratio), with two guards the measured axes do
not need: a +/-10% DEADBAND (a rest day or a blowout can move a 5-game
window without any role change) and a tighter cap (0.15 vs the
environment's 0.30) so a noisy PROXY cannot outvote measured signals.
Every adjustment carries is_proxy: true and
proxy_for: 'confirmed_batting_order' so nothing downstream can mistake it
for a lineup feed.

The axis can stand ALONE -- without it the early return would gate
opportunity off on exactly the thin-classification rows it is most likely
to help.

Zero extra I/O: analyzeViaEngine1 attaches drift from the feature vector
it has already built, and attachChallenger reads it off the grade. Nothing
re-fetches in a loop that runs over hundreds of props.

COLLINEARITY GUARD added to the coverage probe: Pearson r of drift against
l20_avg / l5_avg / ab_per_game, returning null under n=8 rather than
reporting a correlation on a handful of rows. If drift just re-encodes the
projection, the axis is dead signal and gets shelved.

Gates: 4,073 tests / 326 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 03:15:47 -04:00
builtbykev 8a02c75aec Step 0 input check: stop before wiring opportunity, and why
READ-ONLY. Live grade path byte-identical -- no layer wired, no threshold
moved, no challenger added, no holdout run.

INPUTS ARE 100% POPULATED (n=80 real MLB props, through the grader's own
path): ab_per_game, rest_days, l5_avg, l20_avg, l10_stddev and
game_count_in_7d all 100%; opp_rank_stat 65% overall and 0% on
stolen_bases. So there is no honest-degradation problem to solve.

FOUR FINDINGS THAT STOP THE WIRING, three of which would have made the
work unmeasurable or wrong:

1. THE PREMISE IS WRONG. There is no built opportunity layer to connect.
   ab_per_game is consumed in exactly one place -- analyzeViaEngine1:379,
   which renders "4.3 AB/G" on the grade card. engine1 has NO opportunity
   or usage factor at all. A projected opportunity was never built;
   building one is construction, not connection.

2. THE INPUT IS THE WRONG SHAPE. ab_per_game = season atBats/games. It is
   a per-player CONSTANT (measured: varies for 3 of 20 players, and those
   cannot be legitimate since the value can't depend on stat_type), so it
   can only move all of a player's props together, never separate them.
   And it is collinear with the projection: l20_avg = seasonTotal/games,
   the SAME denominator, so l20_avg already embeds opportunity
   multiplicatively. Adding it additively double-counts.

3. THE REAL INPUT DOES NOT EXIST. depthChartService returns battingOrder:
   null for MLB ("the one lineup slot the free schedule feed exposes") and
   PropLine /context carries lineup_confirmed as a BOOLEAN, not the order.

4. ARCHITECTURE: wiring it into engine1 would be unmeasurable BY THIS
   ORDER'S OWN TEST. Step 2 proves reliability and resolution, both
   measured on p_win. engine1 factors move the grade LETTER and never
   touch p_win. The layer belongs in probabilityEstimator, which already
   adjusts on opp_rank_stat, home_away and a consistency pull.

SEQUENCING IS ALSO STALE: challengerProjection (arch-v1) is already live
with archetype, matchup (platoon) and environment (park) axes, writing
p_win_challenger to the ledger. Step 2 of the order's sequence is partly
done -- and the harness this order needed already exists.

RECOMMENDED INSTEAD, as its own order: an `opportunity` axis on that
harness driven by DRIFT, not level -- recent AB/G (last 5) over season
AB/G. A deviation is not collinear the way the level is. Per-game atBats
is present in the statsapi log rows but MLB_LOG_FIELD never maps it, so it
is a small contained BUILD, which is why it gets its own order. Honest
caveat carried forward: it is still a proxy, not tonight's opportunity.

PROBE BUG RECORDED: the first run reported 0% for every feature including
l5_avg, on a pipeline that had just graded 365 props -- impossible, so the
probe was wrong. getFeatures takes camelCase and returns { features: {} };
I passed snake_case and read the top level. Fixed to call
computeFeaturesForProp. Same class as the earlier silent-false harness: a
measurement that makes working code look broken invites you to "fix"
something that was never broken.

Gates: 4,059 tests / 325 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:40:46 -04:00
builtbykev 212c08b11f Fix the Step 0 probe: it was measuring itself, not the pipeline
The first run reported 0% coverage for EVERY feature including l5_avg --
which projectionFor requires, on a pipeline that had just graded 365
props. That is impossible, so the probe was wrong, not the pipeline.

Two bugs, both in my probe: featureCache.getFeatures takes camelCase
(playerName/statType) and I passed the prop's snake_case shape, and it
returns { features: {...} } while I read the top level. Either alone
yields all-zeros.

Now calls computeFeaturesForProp -- the grader's own entry point -- so it
measures what the grade path actually sees. Same class as the earlier
harness that returned a silent false: a measurement that makes working
code look broken is more dangerous than no measurement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:36:00 -04:00
builtbykev 3e78217678 Step 0 input check: read-only feature-coverage probe
Before wiring any layer into the grade, measure whether its inputs are
actually populated on real props. A layer wired onto sparse inputs does not
degrade gracefully by default -- Number(null) === 0 turns a missing
opportunity into 'zero opportunity', a fabricated input rather than an
absent one.

Reports population per feature, SPLIT BY stat_type, because a feature can
be 100% present for batters and 0% for pitchers and a pooled number would
hide exactly that. Also reports whether ab_per_game varies across a
player's own props -- a per-player constant can only move all of a
player's props together, which is a very different thing from a per-prop
opportunity signal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:34:33 -04:00
builtbykev ecdc644621 Fix superseded assertion after the ?limit= bisect hook
runSnapshot now takes an opts object, so the route call is ('mlb', {}).
Asserted as EMPTY rather than loosened to any-object: a stray limit
reaching production would silently cap every run, which is the exact bug
the hook exists to diagnose.

I pushed the previous commit without reading the suite result -- the
failure was already on screen. Caught and fixed immediately after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:23:24 -04:00
builtbykev 11b0139481 Verify the cap raise on prod: 7 -> 365 graded props
Induced, not projected. DEFAULT_LIMIT=500 produced 365 graded props in
114s (was 7 in 16s) -- 52x the board. All 365 carry a unique forecast_rank
and ZERO leak p_win to anonymous callers, so the tier gating holds at 50x
the volume. Anon payload 220KB in 0.44s. Stat mix went from three stats to
ten. Health green.

Measured cost curve via the ?limit= bisect hook: 1->42s, 25->58s, 60->42s,
120->66s, 500->114s. About 42s of that is FIXED overhead (odds fetch,
roster logs, archetype classify, retention), paid whether we grade 1 prop
or 500 -- grading is the cheap part.

MY PRE-FLIGHT ESTIMATE WAS WRONG. I predicted ~72s from per-prop latency
measured in isolation, which ignored the fixed cost. Real figure 114s.

A FALSE ALARM RECORDED because acting on it would have meant reverting a
fix that works: the first induced run 502'd at 13.4s and I hypothesised
load -- memory or a proxy timeout under 20x the work. Wrong. A limit=25 run
then 502'd in 2 seconds, which no amount of load explains, and both
recovered on retry. The 502s were the deploy rolling, not the cap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:22:30 -04:00
builtbykev 8d052131c5 Add a bisect hook (?limit=) to the internal snapshot trigger
The cap raise 25 -> 500 made an induced snapshot 502 at 13.4s and the run
did not complete in background either, while a 25-prop run had completed
in 16.3s. That rules out a simple duration timeout and means the cause has
to be measured, not guessed. ?limit= bounds one run so the regression can
be bisected without a prod env change; omitted, the real DEFAULT_LIMIT
applies.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:12:07 -04:00
builtbykev a7d6cf8e36 Raise the grade cap 25 -> 500 on measured cost; refusals are correct
PART 1 (read-only, measured on a live prod slate, n=80) OVERTURNS THE
PREMISE. The refusal rate is not a data problem -- it is 98% correct
behaviour. The cap is the entire problem, and it is worse than "25 of 546".

Composition: GRADED 44 (55.0%) | POLICY-SUPPRESSION 35 (43.8%) |
FETCHABLE-GAP 1 (1.3%) | FALSE-THRESHOLD 0 | ARCHETYPE-GAP 0 |
GENUINE-ABSENCE 0.

THE FIFTH BUCKET the order did not anticipate: all 35 "refusals" are
rare_event_over_below_line -- the 2026-07-19 betting-logic audit
deliberately refusing 0.5-line rare events, setting the SAME
insufficient_data flag as a real data gap, which is why they read as one.
They are entirely doubles (18) and stolen_bases (17), while hits (19/19),
rbi (19/19) and total_bases (5/5) grade at ~100%. Had we "fixed" this we
would have re-introduced exactly the bets a previous audit removed, and the
count would have looked like progress.

THE CAP: 585 unique gradeable props, cap 25 -> 560 discarded (95.7%).
Traced to Session 32 (f0c8b4f), commented "bound the herd" -- a guard
written before anyone measured what a grade costs. So I measured it:
721ms mean / 666ms median / 1024ms p90 per grade => ~72s for 500 props at
concurrency 5. Both callers tolerate that: the cron runs 5x/day and
recordDownstream is fire-and-forget.

PART 2 -- item 3 ONLY, because that is what the diagnosis supports.
DEFAULT_LIMIT 25 -> 500, env-tunable via GRADE_SLATE_LIMIT. Concurrency
stays 5 deliberately: the cap raise already multiplies load ~20x, and
concurrency decides how hard we hit statsapi at once. One variable at a
time.

Items 4/5/6 have nothing to act on and I am not manufacturing work for
them: 0 false thresholds to loosen (loosening would be manufacturing
grades); /context wiring is worth doing for grade QUALITY but would not
have graded one extra prop here, so it is not claimed as a coverage win;
archetypes are display-side and do not gate grading at all.

THE REFUSAL RATE DOES NOT DROP, AND THAT IS CORRECT. No threshold lowered,
no grade forced. The board grows because the cap stops discarding 95.7% of
the slate.

Flagged in advance rather than discovered later: snapshot payload and
ledger volume both scale with the same multiple. If the response gets
unwieldy the fix is a response-side cap on what the BOARD returns, never a
re-cap on what gets graded -- grading everything and serving a slice is
honest; grading a slice and calling it the slate is what this fixes.

Gates: 4,052 tests / 324 suites green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:06:25 -04:00
builtbykev d18a19f6aa Part 1 diagnostic: read-only refusal categoriser (25-cap + 72% refusal)
READ-ONLY. Runs the REAL grade path over a REAL slate and categorises
every refusal; writes nothing. Reproduces gradeSlateService.dedupeProps
exactly (MODEL_BOOKS, first-row-wins) and calls analyzeViaEngine1 the same
way, so it measures what the pipeline does rather than a re-implementation.

Adds a FIFTH bucket the order did not anticipate, and it is likely to
change how the 72% is read: (e) POLICY-SUPPRESSION. The 2026-07-19
betting-logic audit deliberately refuses rare-event 0.5 markets (doubles/
triples/HR/SB) on the juiced under, plus any over-juiced price -- and it
sets the SAME insufficient_data flag as a genuine data gap. Counting those
as a data problem would send us hunting for data that is not missing, and
"fixing" them would re-introduce bets we removed on purpose.

Separates (b) FETCHABLE-GAP from (d) GENUINE-ABSENCE by asking the stats
layer directly whether the player has ANY game log, rather than assuming:
no log -> genuine absence, keep refusing; a log that exists while the grade
path found no projection -> a wiring gap with something to fix.

Also measures per-grade latency (mean/median/p90/max, serial and at
concurrency) so Part 2 can decide the cap on cost rather than on taste.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 02:02:53 -04:00
builtbykev 6c97f59546 WNBA truth correction + THE p_win FLIP (live, rollback armed)
PART A -- WNBA TRUTH CORRECTION (no behaviour change).
WNBA does not "abstain" and is not "anti-predictive". The -0.12 that
produced those words was NBA-template machinery run on WNBA data -- WNBA
has never had its own archetypes, variables, conditions or calibration,
which is precisely the "sport stubbed in on another sport's template"
CLAUDE.md forbids. That is an UNBUILT MODEL'S EXPECTED FAILURE, not a
verdict on the sport; reading it as a verdict would quietly retire a sport
we never actually attempted. Its own build is QUEUED, after MLB.

The guard CODE is unchanged -- FORECAST_RANKED_SPORTS = {'mlb'} and the
inheritance test are correct live safety either way. Only the meaning is
corrected, and generalised into the doctrine-as-a-gate: a sport ranks on
p_win ONLY once its OWN model is built and shown to predict (calibration
AND resolution on its own holdout). Others are held out as NOT-BUILT,
never as failed. Re-labelled across gradeRanking, snapshot route, tests,
MASTER-PLAN and the challenger report.

PART B -- THE FLIP, gated on a full-slate re-run.

The re-run found something better than a bigger sample. An induced
snapshot graded 7 props: gradeAndCacheSlate runs with DEFAULT_LIMIT = 25
and ~72% of those refuse for insufficient_data, while 546 props are
gradeable. So 8 props IS the board, structurally -- not a small sample of
it. Logged as its own finding; the cap is a separate order.

For a statistically meaningful delta I used 11 real historical boards
(n=328, board sizes 14-57): 79.9% of rows move, mean 5.16 places per
board, TOP READ CHANGES ON 9 OF 11 BOARDS. The re-ordering holds at real
board size. Query committed.

FLIPPED:
- rankGrades drops its edge key (safe for every sport: removes a
  non-predictive tiebreak without putting p_win in front).
- selectTopGrades leads on forecast_rank, edge key removed.
- flattenToEdgeBoard sorts on forecastRank, not edge -- this board had
  edge as its PRIMARY key, so the whole mobile board was ordered by a
  quantity measured not to predict.
- forecast_rank threaded onto strip props.

Sports whose model is not built supply no forecast_rank, so their boards
fall through to the unchanged grade chain -- the fallback is the guard.

ROLLBACK ARMED: boards sort by forecast_rank WHEN PRESENT, so
FORECAST_RANK=0 reverts every surface on the next response -- no deploy,
no client release.

Edge is still computed, stored, carried and displayed as a labelled
diagnostic. Retired from ranking, not deleted.

Eight superseded tests updated to strictly stronger INVERSE properties --
they now fail if edge is ever re-introduced as a ranking key, which the
originals could not detect.

Gates: 4,045 tests / 323 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 01:55:43 -04:00
builtbykev ef4ac60b81 Per-sport rank guard + edge diagnostic-only display + delta report
DELTA MEASURED on live prod grades (live ordering unchanged): MLB 7/8
props move (87.5%), mean 2.5 places, TOP READ CHANGES (corey seager hits
1.5 under -> jake burger hits 0.5 over). WNBA 25/25 move, mean 4.1, max 12.
This is a large re-ordering, not a tweak.

Caveat recorded rather than buried: MLB had only 8 graded props at
measurement time. The percentages are real; the sample is one small slate.
Re-run before the flip -- it is one call.

PER-SPORT DOCTRINE ENFORCED IN CODE. WNBA moves the most and must NOT
adopt this: its p_win is anti-predictive, so ranking that board by p_win
would sort it by a signal measured to point the WRONG WAY -- worse than
the incumbent, not better. A comment would not have stopped a future flip
from going global, so FORECAST_RANKED_SPORTS = Set(['mlb']) gates the
forecast_rank stamp, with tests asserting no sport inherits MLB's result.
A sport joins only by passing its own holdout.

EDGE IS NOW DIAGNOSTIC-ONLY IN DISPLAY. MobileEdgeBoard.EdgeCell rendered
green (--g-a) for positive edge and red (--miss) for negative. Two things
were wrong: green/red IS a quality claim on a quantity that does not
predict, and ROW-GRAMMAR reserves red for settled-negative ONLY -- a
negative diagnostic is not a settled loss. Now neutral mono with a
diagnostic tooltip; header reads "MKT GAP · DIAGNOSTIC". The number is
still shown -- no display went blank. DeskShowcase neutralised likewise.

PINNACLE LOGGED, NOT ENSHRINED. Per the order, "market-not-sharp" is
PENDING-RECOVERY rather than a confirmed permanent limitation. The single
question for PropLine is in BLOCKERS.md with its evidence, and MASTER-PLAN
now carries the pending status instead of the permanent claim.

Live sorts remain byte-identical: selectTopGrades, flattenToEdgeBoard and
topGradedService all still call the incumbent.

Gates: 4,041 tests / 323 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 01:29:20 -04:00
builtbykev 86d123945c Rank on p_win: challenger instrument + retire edge from decisions
MEASURED BASIS (n=200 settled MLB rows): corr(p_win, outcome) = +0.26;
corr(edge, outcome) = -0.010 incumbent ruler / -0.022 consensus ruler.
Subtracting the market destroys the signal under BOTH rulers, so a
quantity that does not predict must not rank, gate or decide.

CHALLENGER-FIRST -- live ordering is byte-identical. rankGrades (the
incumbent, grade-first with edge as its 4th key) is untouched and tested
as untouched.

NEW: rankByForecast -- takeable-gated p_win -> grade -> confidence -> stable
order, with NO edge term anywhere. p_win LEADS and the letter follows,
deliberately: the letter measured r ~ 0.005 and is inverted (B 52.4% <
C 56.9%) while p_win measures +0.26, so leading with the letter would sort
by the weaker signal and use the stronger one only to break ties.

Recorded in the code: isotonic calibration is a MONOTONE transform, so
ranking on raw vs calibrated p_win gives the SAME ORDER. Calibration
matters when p_win is displayed or thresholded; it cannot change a
ranking. Nothing here needs the calibrated value.

rankingDelta + GET /api/internal/ranking-delta measure how far the board
would move before any flip. The endpoint reports p_win coverage alongside
the delta -- if p_win is absent the challenger degrades to grade order and
the delta UNDERSTATES, which is worth saying rather than reporting a clean
zero.

forecast_rank is stamped on snapshot grades BEFORE stripModelPrice, so
every tier gets the correct order without the paid values (the
topGradedService precedent -- an ordinal can travel where the magnitude
cannot). Additive only: nothing sorts by it yet.

RETIRED AS DECISIONS (not rankings, so done now):
- altLineScanner.compareToBookImplied no longer returns value_detected:
  edge > 0. Edge is still COMPUTED and returned -- losing the record would
  be worse than mis-using it -- but the verdict is an honest null with
  value_basis: 'retired:edge_does_not_predict'.
- scanAltLines no longer filters to edge>0 or calls the survivor "optimal".
  The whole ladder is returned ranked and labelled
  'price_gap_diagnostic_unvalidated'. The module has ZERO callers (verified)
  -- unwired like mlbGrader.js, left in place and made honest.

An honest asymmetry recorded there: ranking props AGAINST EACH OTHER must
not use edge, but choosing between RUNGS OF THE SAME PROP is inherently
price-relative -- ranking rungs by model probability alone would always
pick the lowest line, since P(over 0.5) > P(over 2.5) by construction. So
the gap stays the rung key, explicitly labelled unvalidated.

Two superseded tests updated to stronger properties.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 01:24:55 -04:00
builtbykev 7140e62b65 MLB re-run vs consensus ruler: premise dissolved, isotonic DECIDED
MEASURE-ONLY. No promotion, no flip, no tier spend. Live path
byte-identical: CURRENT_RULER_VERSION still v1_first_book, model still
consumes MODEL_BOOKS only.

MANDATE 1'S PREMISE DOES NOT HOLD. The p_win calibration is
RULER-INDEPENDENT, confirmed two ways: estimateProbability takes
{gameLogs, line, statType, features} and never sees a market price, and
the calibration fits p_win against OUTCOMES. Reliability and resolution
are both p_win-vs-outcome measures, so fair_prob cannot enter either.
There is nothing to re-fit -- the ruler changes edge, CLV and takeable,
not calibration.

I RETRACT MY OWN LABEL. I declared the MLB isotonic result PROVISIONAL
"because it was measured against the bent ruler". That over-applied the
ruler caveat to a measurement the ruler never touched. The result was
never contaminated; it moves PROVISIONAL -> DECIDED, not by re-running but
because the gate I attached does not apply.

RAN THE GENUINELY RULER-DEPENDENT QUESTION INSTEAD -- does a median
consensus rescue EDGE? Timing held constant (both rulers at close; a
lock-time reconstruction joins only 43 rows, and mixing lock-incumbent
with close-consensus would confound WHEN with WHAT).

n=200 MLB settled rows: mean |ruler gap| 0.0085. corr(edge_v1, outcome)
-0.0101; corr(edge_v2, outcome) -0.0220; corr(p_win, outcome) +0.2598.

THE HEADLINE: p_win predicts outcomes at +0.26 while p_win minus the
market predicts nothing under EITHER ruler. Subtracting the market price
destroys the signal -- a direct empirical vindication of the identity now
at the top of CLAUDE.md. Market edge is not merely a poor criterion here;
it is a strictly worse instrument than the raw forecast.

CALIBRATION REFRESH (ruler-independent, but n grew 119 -> 250):
time-forward holdout n=125, reliability 0.0846 (was 0.0939), resolution
0.190 (was 0.123). Both hold and both improved on a fresh later window
the earlier fit never saw. Independent replication.

THE LIMITATION THAT BLOCKS A FULL VERDICT: closing_captures holds only
MODEL books -- exchange quotes were never stored, because normalizeProps
discarded them until yesterday. Mean 1.97 books in the historical join. So
this tested a US-books-median ruler, not the exchange-inclusive consensus
whose live delta showed p90 +10 points. That ruler is UNTESTABLE on
existing data at any n. Per Mandate 4's third outcome: inconclusive, not
forced.

SEPARATE FINDING -- LIVE FEED REGRESSION: pinnacle MLB captures went 4,022
-> 0 on 2026-07-31 and have not returned, while every other book continued
(103,940 captures in the prior 10 days). This also corrects an Order Zero
claim of mine: "no sharp anchor exists in our feed" was accurate for the
day measured but wrong generally -- pinnacle was there until 07-30 with
17,090 two-sided captures. line_type='sharp' is a label in closingCapture
via SHARP_BOOKS, not a separate provider. We had a sharp anchor and lost
it two days ago; not caused by anything in this session.

Both queries committed: scripts/ruler-comparison.sql,
scripts/pwin-timeforward.sql.

Gates: 4,028 tests / 322 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 01:12:47 -04:00
builtbykev c79528abae Order Zero: tier-reality report + widening fingerprint
PHASE 1 resolved on our real keys, and a bad source was discarded on the
way: a fetched rendering of PropLine's docs "tier matrix" claimed
/odds/closing is 403 on free and that /odds returns prices nulled on free.
Both are contradicted by direct observation (200-redacted, and 6,196
two-sided PRICED groups on MLB). Not cited. The report uses only the
machine-readable OpenAPI contract and the verbatim detail bodies our keys
received.

Verdict: every one of the six endpoints behaves exactly as the Free tier's
published contract says. error:"upgrade_required" with an explicit
required_tier is unambiguous -- NOT a key-permission problem, NOT a plan
problem. $9/mo Hobby buys /results + /odds/closing (the CLV instrument) +
/movement (steam across 18 books); $19/mo Pro adds the 90-day settlement
export. Priced and evidenced; not recommended here -- it is a decision.

PHASE 2 fingerprint on the SERVED feed: 5 books -> 13, props rendered
546 -> 2,780 (5.1x), mean 4.22 books/prop.

The unflattering half, stated up front: of 2,234 newly-visible props only
698 (31.2%) carry a real non-DFS market price; 1,536 (68.8%) are DFS-only
pick'em rows. The honest headline is not "80% of the slate unlocked" --
the board is 5x fuller, about a third of the new depth is real market
data, and the rest is pick'em inventory now shown but tagged.

PHASE 3 verified: 546 gradeable props, unchanged. CURRENT_RULER_VERSION
still v1_first_book.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 00:57:12 -04:00
builtbykev 68c5b65427 Thread book_role through the odds route grouping
The route regroups flat props into lines[] and was dropping the role tag,
so the widened feed reached the browser untagged. That is not cosmetic:
on a live prop, PrizePicks prices both sides at even money (+100/+100)
while BetMGM has +450/-750. Rendered side by side without a tag, the
pick'em row reads as a dramatically better price when it is a different
product entirely -- exactly the confusion the three-way split exists to
prevent. Consumers gate on book_role !== 'dfs' before treating a row as a
market price.

The ?book= filter now accepts any DISPLAY book, since shopping a real
book against an exchange is the point of the widening. Grading still only
ever consumes MODEL_BOOKS.

One superseded integration test updated to a stronger pair: an unknown
book still 400s, and a newly-visible one no longer does.

Gates: 4,028 tests / 322 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 00:55:04 -04:00
builtbykev f0543b57a4 Product identity + widen books for DISPLAY, model input byte-identical
IDENTITY (CLAUDE.md top + MASTER-PLAN header). VYNDR is a PREDICTIVE MODEL:
it projects what a player will DO and picks accurately. Market edge is a
BYPRODUCT of a good prediction, never the success criterion. Success =
the forecast is honest about its own confidence AND still ranks --
calibration and resolution, both. No edge/CLV term belongs in a pass/fail
gate; they are diagnostics we report, not thresholds a model must clear.
A model tuned to beat a closing line has been fitted to the market instead
of to the game.

Per-sport doctrine (Phillips 2022, classify by what players DO not by
position): each sport is its own model -- own variables, archetypes,
conditions, calibration, honest ceiling. Shared across sports: ONLY the
Bayesian inference math.

Truth Law: no fabricated data; honest-absent over invented; label
limitations in-band; provisional stays provisional until re-run;
documented is not verified.

PHASE 2 -- AGGREGATOR WIDENING (live). normalizeProps now emits every
DISPLAY book instead of 5 of 18. Before this we discarded 13 books of our
own accord and 64.8% of the MLB slate was invisible to users. Every prop
carries book_role (both/takeable/reference/dfs/offshore) so the display
layer can say WHAT a price is -- a fixed-payout DFS number and a two-way
sportsbook price are not interchangeable objects. Unknown books are still
dropped.

PHASE 3 -- MODEL GATE (the model does not move). bookRoles splits
MODEL_BOOKS (the legacy allow-list, character for character) from
DISPLAY_BOOKS. Both model paths re-filter before they pick a line:
gradeSlateService.dedupeProps (before first-row-wins AND before the limit)
and intradayRefreshService.indexOddsProps (which RE-GRADES at the current
line -- without the gate, widening would have silently moved locked lines
onto books the model has never been calibrated against). A test asserts
the graded set is byte-identical through the widening.

CURRENT_RULER_VERSION stays v1_first_book. The gate lifts only when the
MLB calibration is re-run on the consensus ruler and v2 is promoted.

HONEST FRAMING, recorded in the plan: this is an AGGREGATOR win and it
does NOT fix the model. WNBA still abstains -- a model problem, not a
coverage problem; it is better covered than MLB. MLB isotonic still
provisional. The consensus is MARKET, not SHARP: pinnacle, matchbook and
polymarket are 0% on both sports, so no sharp anchor exists in our feed.

Two superseded tests updated to stronger properties rather than deleted:
roleOf now names the KIND of book, and the normalizer test asserts the
display set widens WHILE the model set does not.

Gates: 4,027 tests / 322 suites green; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-08-01 00:50:54 -04:00
builtbykev 1372e6bcf7 Order Zero Phases 1-3: keyed verification, ruler_version boundary, report
PHASE 1 (measured on the live prod feed with the real key):

- WNBA is NOT thin at the feed -- 4.21 books/prop vs MLB's 3.61. It was
  allow-list-starved exactly as MLB was. This removes one candidate
  explanation for its anti-predictive result; it does not explain it, and
  WNBA stays abstaining.
- We cannot see 64.8% of the MLB slate at all (zero admitted books).
- Exchanges are real (smarkets 27%, novig 22%, kalshi 15% on MLB) but
  pinnacle, matchbook and polymarket measured 0% on BOTH sports. There is
  no sharp anchor for player props. The consensus is a MARKET consensus,
  not a SHARP one -- recorded as a permanent limitation, not a milestone.
- DFS is the trap, quantified: prizepicks covers 82% of MLB props, the
  highest in the feed. Admitting it "for breadth" would have looked like
  the biggest available win. Permanently excluded.
- Endpoints: /context WORKS and is FREE (umpire, roof, pitcher handedness,
  lineup confirmation -- richer than what we hand-built). /odds/closing and
  /movement are REDACTED (full structure, zero prices). /results and
  /exports/resolved-props are 403.
- The $19/mo question is answered: soccer IS graded, ~15 competitions in 30
  days (MLS 41k, Liga MX 15k, Brasileirao 12k, UCL/Europa/Conference). Our
  "soccer grades into a void" is a Pro-tier problem, not a data problem.
  NBA is absent because it is July -- seasonal, not inferable either way.

PHASE 2 delta, corrected: MLB mean +1.50 pts, median 0, p90 +10.0, 17.0%
of comparable props move >=5 pts, one-directional (the incumbent prices
the over below the exchange-inclusive consensus). WNBA symmetric and
tight. The median prop does not move -- the change is a right-skewed
minority. That the rulers DIFFER is established; that the new one is
BETTER is not, and that is the re-run.

PHASE 2 item 6: ledger_entries.ruler_version applied to prod, 1,384
existing rows backfilled to v1_first_book (a statement of fact -- every
row to date was produced by the first-book rule). ledgerService stamps
CURRENT_RULER_VERSION on new rows. Never pool edge or CLV across it.

Repo migration numbering lags prod; 025_ledger_ruler_version.sql records
the DDL for review.

PHASE 3: MLB isotonic p_win remains PROVISIONAL -- calibrated against
v1_first_book, does not promote until re-run on the consensus ruler.

NOT LIVE, deliberately: ALLOWED_BOOKS unchanged, served slate
byte-identical, CURRENT_RULER_VERSION still v1_first_book, no live path
calls consensusRuler.

Gates: 4,022 tests passed / 322 suites; next build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 23:52:28 -04:00
builtbykev c38db1ad65 Fix: the incumbent ruler respects the allow-list (correcting my own model)
My first delta run modelled the incumbent as first-row-wins over the RAW
feed and reported that an EXCLUDED book was "the market" on 69% of MLB
prop-lines, with prizepicks alone at 47%. That is WRONG and I caught it
before it went anywhere.

normalizeProps applies ALLOWED_BOOKS BEFORE gradeSlateService.dedupeProps
runs, so DFS books never reach the incumbent. The allow-list, for all the
coverage it costs, does keep DFS out of the ruler.

incumbentFairProb now takes the allow-list (defaulting to the live
ALLOWED_BOOKS) and reproduces the real chain. Two tests lock it, including
that a prop with no admitted book has NO incumbent -- it is never graded
at all, which is the real loss and is already measured as invisible_props.

Overstating the incumbent's badness would have been as dishonest as
understating it, and more persuasive.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 23:47:13 -04:00
builtbykev a55dd2a6a0 Order Zero Phase 2: three-way book split + challenger consensus ruler
CHALLENGER-FIRST. The live ruler is byte-identical: CURRENT_RULER_VERSION
is still v1_first_book, nothing here writes a cache, a grade or a ledger
row, and no live code path calls consensusRuler yet.

bookRoles.js splits one allow-list into three, because it was answering
two different questions -- "can we show this?" and "can we price against
this?" -- with the same list, which is what bent the ruler.

  TAKEABLE  the user can actually bet here (drives best price / shopping)
  REFERENCE may price the fair-prob ruler; never surfaced as a place to bet
  EXCLUDED  DFS pick'em + offshore, permanently barred from all pricing

Two deliberate calls, both evidence-based:

- The six PropLine-phantom books (caesars/fanatics/bet365/hardrockbet/
  pointsbet/thescore) are KEPT despite the order saying remove. They
  returned zero PropLine quotes, but PropLine is not our only provider and
  the odds-api backup path may carry them. A book that never appears is
  never matched, which costs nothing; deleting them risks silently
  dropping real books on the backup with no upside. Recorded in
  PHANTOM_ON_PROPLINE rather than enacted as a deletion.

- REFERENCE = exchanges + pinnacle + bovada + the four US majors, chosen
  off the measured coverage curve rather than theory. exchange_only is
  cleanest (order-book, ~zero vig) but covers 14.3% of MLB and 5.6% of
  WNBA; adding the US majors gives 28.1% / 46.3%. pinnacle, matchbook and
  polymarket measured 0% on both sports and add nothing. The honest
  limitation is recorded in the config: this is a MARKET consensus, not a
  SHARP one.

consensusRuler.js: median de-vigged fair_prob across >=2 reference books
posting BOTH sides at the SAME line. Median so one stale exchange cannot
drag it. Different lines are never averaged, one-sided quotes never rule,
and n<2 falls back to single-book LABELLED as such with the v1 stamp --
never silently mixed, because a column holding both is two rulers wearing
one name.

The challenger delta runs over the live feed and reports incumbent_book_
roles, which is the real headline: the incumbent is literally first-row-
wins, so it reports what KIND of book has been acting as "the market".
DFS pick'em has the highest coverage in the feed, so a DFS book can be it.

18 ruler tests + 37 total in the two new suites. Full suite 4021 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 23:44:37 -04:00
builtbykev c3bcfaba94 Order Zero Phase 1c: return the aggregate-only bodies in full
/sports and /markets/resolution-summary carry no per-prop data and no
credentials, and the shape summary alone cannot answer the question they
exist to answer -- whether PropLine actually GRADES the sports we cannot
settle. A shape is not a number. Both bodies are scrubbed on the way out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 23:39:38 -04:00
builtbykev 3c466d79cb Order Zero Phase 1b: redaction detection + reference-policy curve
Two corrections to the first pass, both of which would have produced a
false positive.

1) A non-empty body is NOT proof of access. PropLine's free tier returns
   the full STRUCTURE of tier-gated endpoints with values stripped plus an
   upgrade_url -- and the first pass classified /odds/closing and /movement
   as "works" on structure alone. detectRedaction() now counts actual
   prices and downgrades works -> partial when a body advertises an upgrade
   or carries outcomes with zero prices. Same class as the harness that
   returned a silent false, inverted.

2) One hard-coded reference set forces a yes/no on a question that is
   really a curve. reference_policy_curve reports strict eligibility
   (>=2 books, both sides, same line) under exchange_only /
   exchange_plus_sharp / exchange_plus_us / takeable_only, so the ruler
   decision is made on coverage-vs-quality rather than on a guess. DFS is
   absent from every policy by construction and a test asserts it.

Also probes /markets/resolution-summary: /exports/resolved-props being 403
tells us we cannot PULL settlements; resolution-summary tells us whether
they EXIST to be bought. Different questions.

19 unit tests, still hermetic.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 23:37:23 -04:00
builtbykev 2071b79456 Order Zero Phase 1: keyed read-only PropLine verification endpoint
Adds GET /api/internal/propline-verify (internal-key gated, read-only) so
Phase 1 can run WHERE THE KEY LIVES. Touches no cache, no ledger, no
grade; the live adapter and the live ruler are untouched. Breadth reuses
proplineAdapter.fetchRaw -- the exact live request -- so what it measures
is what the pipeline actually receives.

Reports per sport (never pooled): books/prop from the feed vs after our
own ALLOWED_BOOKS, props made INVISIBLE by that filter, reference-book
presence, DFS presence reported separately, and consensus eligibility.

Consensus eligibility is deliberately strict: >=2 REFERENCE books posting
BOTH sides at the SAME line. A one-sided quote cannot be de-vigged, and
two books at different lines are not the same market -- counting either
would overstate how much of the slate can carry a real ruler.

Probes the documented-but-unverified endpoints (/sports, /context,
/odds/closing, /movement, /results, /exports/resolved-props for four sport
keys) and classifies works/partial/no, with 403 = tier-gated and 200-but-
empty = partial rather than works.

Key safety is the other locked property: the key goes via axios params,
never string-interpolated, and every emitted string passes scrubKeys()
which removes the literal key AND any surviving apiKey= query value. A
test asserts a thrown transport error carrying the key cannot escape.

13 unit tests, hermetic (no network, no key).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 23:34:03 -04:00
builtbykev 293367917c Order Zero: book-breadth test + accrual clock correction (measure-only)
STEP 0 disproved the premise before any request was fired. PropLine's
OpenAPI contract states verbatim that `bookmakers` omitted = ALL books,
so proplineAdapter omitting it is correct and always was. Firing a
guessed param would have RESTRICTED the response and produced exactly
the false negative the order warned about.

The real cause is ours: PropLine sends 18 books; oddsNormalizer
ALLOWED_BOOKS intersects them at exactly 5 -- which is precisely the
"5 MLB books" the 2.18 audit measured. Measured on real public data
(no key, no quota): 4.41 books/prop from the feed, 1.50 after our
filter, and 12 of 34 props go invisible entirely.

Also corrected: "73% single-book" is the long tail of deep props
sole-posted by DraftKings or Bovada. On the core props we grade, the
market is 10-12 books wide. pinnacle appears on 0 of 40 MLB props --
the independent low-vig references present on 100% of core props are
exchanges (novig/smarkets/kalshi). DFS pick'em also covers 100% but is
not a market price and must never enter a consensus.

Verdict is outcome (d) ALREADY OPEN, not (a)/(b)/(c) -- all three
assumed the feed was the constraint. Ruler change scoped (not built):
split one allow-list into takeable/reference/excluded, fair_prob_lock
becomes a median consensus with n>=2 or a labelled fallback. Gated on
exchange price validation + the WNBA measurement, which needs the
PropLine key (prod-only, absent locally). MLB isotonic p_win declared
PROVISIONAL until re-run on the real ruler.

Side finding: we use 1 of 29 endpoints. /odds/closing, /movement,
/odds/history, /best-line, /ev, /results, /exports/resolved-props,
/context (free) map directly onto documented gaps -- and resolution
across 33 sports suggests "no free settled feed for NBA/soccer" may be
a $19/mo problem, not a data problem. Documented, not verified.

Plan edits: §10.1 rewritten, §10.2/§10.5 corrected, and §11 adds the
sequential post-completion accrual clock -- pre-completion data does
not count, no pooling across the completion boundary, two clocks
stated separately, per-sport clocks, verification gate before any
accrual, users onboarded to a complete product only. §9.1's "6-10
weeks out" corrected: that is accrual duration, not distance to the
answer. The ruler change independently forces the same no-pooling
boundary by arithmetic.

No API key was used, printed, or committed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 23:19:23 -04:00
builtbykev c98338ef23 plan: add §10 — aggregator + paid-model gaps, and the one root cause behind both
Answers "what makes this the top product, not just a finished one."

THE REFRAME: the aggregator gap and the model gap are the SAME gap in two places.
Our "market" is often ONE book — MLB props are 73% single-book, and
proplineAdapter sends only {apiKey, markets} with NO regions/bookmakers param
(:152), so we take PropLine's default response. That single fact causes four
problems we had been treating as unrelated: no line shopping (the category's #1
free hook), a fair_prob_lock that is a de-vigged single soft book rather than a
consensus (the bent ruler the model is judged against), weak CLV (cannot measure
beat-the-close against one book), and no steam/disagreement detection (needs >=2
books to exist).

So the highest-leverage unblocked action in the whole plan is a cheap API test:
does PropLine return more books with a regions/bookmakers param on our tier? One
request, and if it works it upgrades the free product, the model's denominator and
the CLV instrument simultaneously.

Aggregator gaps catalogued: book breadth, true consensus, historical odds archive
(started — closing_captures 844k rows, lock_lines new, but in-grade history capped
at 24 points, so no full open->close series), market breadth (11 live vs the
category's 50+), ingested alt-line ladders, injury/lineup wire, player news.

Paid-model gaps catalogued: distribution instead of a point (distribution.js
already computes survival probabilities and rungs but is proj-v1.1, ledger-only
and lost to the champion); opportunity/playing-time projected FIRST with its own
uncertainty (the single biggest available modelling gain); per-stat models instead
of one additive index; matchup granularity that actually reaches the grade;
applied calibration; a backtest harness (blocked by the archive gap — you cannot
backtest a price you never stored); CLV as north star.

THE PATTERN: almost every model capability is ALREADY BUILT AND DISCONNECTED.
VYNDR does not have a building problem, it has a connection-and-proof problem plus
one genuine ingestion gap that starves both halves. The expensive part is largely
done, but no new feature fixes it.

Ordering principle recorded: get MLB genuinely good BEFORE replicating across six
sports — a copied-six-times thin model is six times the maintenance for the same
absent edge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 22:35:16 -04:00
builtbykev 37ee952e26 plan: add §9 — what is actually missing for the product to work, not just be built
The phases counted unbuilt code. This section names what is missing for VYNDR to
do what it claims, including the parts that are not builds.

THE CENTRAL GAP: there is no demonstrated edge yet. Every measurement this session
returned null, negative or unproven — served grade r~0.005 and inverted; all three
p_win-vs-fair_prob formulations negative on both sports and both splits; p_win
alone on MLB holdout p~0.07; WNBA negative; CLV null by guard; ROI-by-grade likely
an artifact. The product's core claim is not currently supported by our own data,
and building all 23 orders without closing this leaves a well-built product that
does not do the thing it sells. What closes it is sample and honest iteration, not
code — roughly 6-10 weeks at the current accrual, a clock engineering cannot
shorten and that must not be faked.

Also named: the projection is thin (l5/l20 + opponent rank + rest + usage, with
similarity/archetypes/conditions/Bayesian all built and disconnected, so
connecting them is a hypothesis not a guarantee); it is a one-sport product today
(NBA and soccer do not even settle); there are 3 users and 0 paid so nothing is
validated by usage; there is NO distribution path at all, which appears in no
phase and belongs on the board as its own track; the last mile is unclosed
(push-to-book is a teaser, no affiliate live); and operational fragility remains
(single box, two-sport settlement, three credentials flagged including a Stripe
live key that transited a transcript, no staging).

The honest summary: the truth infrastructure is genuinely well built and this
codebase does not lie about what it knows. What is not yet true is that the model
beats the market — not disproven, unmeasured at adequate n. The finish line is 23
orders PLUS a verdict from accrued data we cannot rush, and the discipline to
report that verdict honestly if it says the edge is not there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 22:16:11 -04:00
builtbykev e3ca1650d9 plan: specs/MASTER-PLAN.md — single source of truth, 7 phases, ~23 orders, defined END
Consolidation only. Nothing built, wired or promoted.

NOTHING WAS RE-VERIFIED and no query was run — all 22 artifacts produced this
session plus the completion matrix were taken as KNOWN, per the order's own clause.
The verification ledger at the top of the plan lists exactly what was taken as
known and which four items remain genuinely open (sport order, board-reasoning
gating, the CLV flag, team colours) — each open because it needs a decision or a
build, not a query.

The plan captures all six tracks in one document: per-sport models (MLB's 8-layer
stack with each layer marked BUILT/PARTIAL/NOT-WIRED, plus the sport order),
design implementation (61 catalogued items), surfaces, the resolution tail, the
sport boundary, and the Chrome audit.

The through-line it makes visible: MLB's layers 2, 3, 5 and 6 are BUILT AND NOT
CONNECTED, while layer 8 (the grade ladder) is connected and meaningless
(r~0.005, inverted). MLB's fix is connection, not construction.

Phasing is by dependency: MLB model truth -> resolution tail -> surfaces/design
(parallel lane) -> sport boundary -> sport rollout (one order per sport) ->
monetization finish -> Chrome audit and hardening. ~23 orders total, ~11 unblocked
today, so "how many sessions left" now has a real answer.

DEFINITION OF DONE is explicit and countable: MLB layers 1-8 connected with a
monotone held-out-proven ladder; every listed sport finished on the same template
or explicitly abstaining with its reason recorded; all 61 design items built; every
surface reachable and honest; the resolution pipeline firing end-to-end; the sport
boundary a registry; the Chrome audit passed; and the record publishable on its own
terms with no claim outrunning its evidence.

STATE.md now points at the plan and is demoted to history.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 22:15:10 -04:00
builtbykev 6d87d7a33c report: p_win recalibration holdout — MLB qualifies on isotonic, WNBA abstains
Measure-only. p_win not flipped live, no grade rebuilt, no calibrator deployed.
Per the doctrine, MLB and WNBA were fitted, selected and judged as SEPARATE
models — and they reach opposite verdicts. No global instrument was fitted.

METHOD: time-forward split per sport (earlier fits, later proves). Both
instruments fitted on TRAIN only — single-parameter Platt and
isotonic-with-pooling. Inputs p_win + outcome only; no market field, no closing
value, no lookahead. Nothing about edge/CLV/beat-the-close enters any pass/fail
line.

MEASUREMENT CORRECTION made mid-run: the first pass reported mean|p - outcome|
(~0.46-0.51), which is NOT calibration — it is noise-dominated individual error
on 0/1 rows and would have made every instrument look identical. Reliability is
only meaningful on BUCKETS (bucket mean predicted vs bucket actual rate,
n-weighted), the metric T0 used. All reported numbers use the corrected metric.

HOLDOUT RELIABILITY (lower better): MLB n=119/4 buckets — raw 0.1038, Platt
0.1120, ISOTONIC 0.0939. WNBA n=93/3 buckets — raw 0.1322, Platt 0.0491,
isotonic 0.0667.
HOLDOUT RESOLUTION: MLB raw 0.1388 -> Platt 0.1284 -> isotonic 0.1225.
WNBA raw -0.1201 -> Platt +0.1269 -> isotonic +0.0322.
Fitted Platt: MLB a=-0.381 b=+0.705; WNBA a=+0.040 b=-0.081.

MLB QUALIFIES, MODESTLY — instrument selected BY HOLDOUT, not assumed: isotonic
beats both raw and Platt, and Platt actually made MLB worse. Reliability improves
0.1038 -> 0.0939 (~10% relative, real but modest) and resolution SURVIVES
(0.1388 -> 0.1225, not crushed). Both Mandate-3 conditions hold.

WNBA ABSTAINS — its Platt result is the best number in the report and is REJECTED
as a fake win. The fitted slope is b = -0.081, negative and near zero, so
sigmoid(0.040 - 0.081*logit p) is nearly constant at ~0.51 for every input: it
"calibrates" by discarding the prediction and emitting the base rate, which is
exactly the failure Mandate 3 pre-registered. Its apparent resolution gain
(-0.120 -> +0.127) is the sign flip, not skill — it would serve the opposite of
its own forecast, fitted on n~96 of anti-signal. Isotonic says the same quietly
(resolution collapses to +0.032).

HONEST CEILING: MLB is a usable-but-unimpressive forecaster (holdout resolution
~0.12, reliability ~0.094, n=119); WNBA has no honest forecast today. Holdout n
and bucket counts (4 and 3) suffice to reject WNBA and prefer isotonic for MLB,
NOT to certify a letter ladder, and the T0 pathology is reduced rather than cured.

CANNOT DETERMINE: per-archetype calibration (Mandate 3d) — bucket n falls below
the reporting floor once split by sport AND archetype on 442 rows.

Queries committed at scripts/pwin-calibration-holdout.sql.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 21:26:21 -04:00
builtbykev 249b3e8235 report: grade diagnostic T0 — p_win is MISCALIBRATED, and it explains the inversion
STOPPED at the T0 gate as instructed. Nothing fixed, no recalibration applied,
no grade touched. T1-T4 deliberately not run.

T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS.

Condition 1 (mean |predicted-actual| > 0.05): MLB ~0.094, WNBA ~0.139.
Condition 2 (monotonic slope): over-confidence GROWS with the prediction —
MLB +0.034 -> +0.043 -> +0.084 -> +0.190 -> +0.189; WNBA +0.044 -> +0.109 ->
+0.349. Worst cases: MLB predicted 0.842 actual 0.652 (n=23), predicted 0.917
actual 0.727 (n=11); WNBA predicted 0.730 actual 0.381 (n=21).

WHY THIS EXPLAINS THE INVERSION, mechanically: p_win is over-stated and the
overstatement SCALES with p_win, so p_win - fair_prob_lock is largest exactly
where p_win is most inflated. Those props hit less than claimed, so the edge
measure correlates negatively. The market was never the problem —
fair_prob_lock is not a bent ruler, the thing subtracted from it is. It also
explains why p_win ALONE still carries signal (+0.23 MLB): rank survives
miscalibration, differences do not.

This independently reconfirms the 2026-07-26 calibration finding (+0.02 at p<.5
-> +0.19 at p>=.8) on a newer, larger population, so it is structural rather
than sampling noise.

PART 0: P0a — only the GRADED side's fair prob is stored (fair_prob_lock;
no opposite-side field), so T1's two-side-sum check cannot run and must use the
stated no-vig recompute fallback. P0b — projection_locked_at exists as a
timestamptz so T2 is potentially runnable, but distinctness from lock time was
NOT verified because T0 gated it.

Two cautions recorded before Part 2 runs: the top MLB buckets where the error is
worst hold n=23 and n=11, so a flexible per-bucket correction would fit noise —
isotonic with pooling or single-parameter Platt is safer; and calibration fixes
magnitudes, so if the market is genuinely better the repaired edge may still land
at ~0, which would be the honest ceiling and gets reported rather than graded
around.

Query committed at scripts/grade-calibration-t0.sql.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 20:12:20 -04:00
builtbykev ea1157d709 report: grade fix Part 1 — the p_win-vs-fair_prob rebuild is REFUTED by the data
STOPPED at the Part 1 gate. Nothing rebuilt, no grade changed, no cutover.

THE FINDING: grading on p_win vs fair_prob does not work. All three candidate
edge formulations correlate NEGATIVELY with outcomes, on both sports, overall,
and in both time splits (n=432 decided rows carrying p_win AND fair_prob_lock):

  ALL  n=432  champ -0.0016  p_win ALONE +0.1221  additive -0.0615  ratio -0.1161  logodds -0.0438
  MLB  n=240  champ +0.0984  p_win ALONE +0.2278  additive -0.0336  ratio -0.1350  logodds -0.0124
  WNBA n=192  champ -0.1143  p_win ALONE -0.0842  additive -0.1326  ratio -0.1281  logodds -0.1243

Subtracting the market's lock-time fair probability destroys and inverts the
signal. The plain reading: props where the model most disagrees with the market
are LESS likely to hit — the market is better than the model, so "edge vs market"
is anti-predictive here, while the raw probability retains some skill alone.

WHAT DOES CARRY SIGNAL: p_win alone, MLB only, and it is modest. Time-forward
split — TRAIN (07-21..07-26, n=120) r=0.2770; HOLDOUT (07-26..07-30, n=120)
r=0.1647, with the additive edge negative in BOTH halves. So p_win survives
forward validation directionally but the holdout is NOT significant (t~1.81,
p~0.07). Suggestive, not proven.

WNBA MUST ABSTAIN: every measure negative including p_win itself (-0.084). Forcing
one threshold across both sports would make a coin-flip sport look sharp, which the
order forbids.

LOOKAHEAD GUARD SATISFIED: fair_prob_lock is the lock-time field, populated on 432
decided rows, range 0.145-0.713. closing_prob (415 rows) is the CLOSE and was NOT
used in any correlation — using it would have manufactured a correlation.

SAMPLE REALITY: 1103 decided rows but only 432 carry both instrument fields, so a
per-sport train/holdout split leaves ~120 per half — enough to show direction, not
to certify a letter ladder.

I did not tune toward a win: three pre-registered candidates were tested and all
three failed; picking a fourth because the first three lost is the overfitting the
order guards against. Recommended instead: grade MLB on p_win alone with WNBA
abstaining and label it modest/accruing (A-RATED hold stays); or wait ~6 weeks for
n~500 MLB; or investigate WHY the market-relative edge inverts, which is the more
valuable question.

Both queries committed at scripts/grade-correlation-proof.sql so no number here
has to be taken on trust. Working settlement untouched; dead resolve endpoint not
wired.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 19:57:21 -04:00
builtbykev 40c61fbb0b report: resolution + CLV investigation — Part 1 premise false, Part 2 is an env flag
Nothing built. No poller wired, no capture change, no env flipped.

PART 1 — GRADES ALREADY AUTO-SETTLE. snapshotScheduler resolves settleAllOutcomes
(:64) and settleAllLedgers (:67) and runs them FIRST at every snapshot slot before
grading (its own comment at :393, Session 61). The record is self-populating: 937
settled rows, growing daily (07-24 through 07-30: 20, 25, 44, 26, 98, 62, 91), and
/api/accuracy reads it live at 937 @ 58% (MLB 526 @62%, WNBA 411 @54%).

/api/grading/resolve is a separate unreferenced legacy path, not the settlement
path. Wiring an ESPN poller to it would create a SECOND settlement path racing the
working one and double-count an append-only ledger — so nothing was built.

The DNP/VOID requirement is already satisfied: outcome carries void and
unrecoverable as terminal states, and getModelAggregate excludes both from the
record denominator, so a DNP is never counted as a loss (105 void rows exist).
Idempotency is enforced too — settleLedger guards on .is('outcome', null) and
outcomeService dedupes on nameKey|stat|line|side|date.

THE REAL GAP is smaller and different: settlement covers MLB + WNBA only. NBA and
soccer grade but never settle because no free settled-result feed is wired. That
is a per-sport feed problem, not a missing poller.

PART 2 — clvCaptureReliable() is ONE LINE:
  return process.env.CLV_CAPTURE_RELIABLE === '1';
It measures nothing. It fails because the operator has not set the flag, not
because the capture is unreliable. So there is no capture code to repair for the
guard to pass — flipping one env var publishes beat_close_pct immediately, which
makes this a judgement call and precisely the "make a number appear" move the
honesty guard forbids.

The guard itself works: beat_close_pct and clv_distribution publish only when the
flag AND settled>=20 AND clv_sample>0; with it off /record shows NOT PUBLISHED YET
and the computable 34/937 = 3.6% is never the publishing path (clvPanel returns
null and a test forbids the fallback).

CANNOT DETERMINE (Supabase MCP upstream-auth outage): the close-vs-locked
distribution, which is the direct test for the old silent-overwrite bug. The exact
query is in the report. A decision rule is stated BEFORE seeing the number so it
cannot be fitted to it: set the flag only if close_moved is a clear majority of
rows carrying a close AND coverage of settled rows is high enough that the
percentage describes the record rather than the captured subset. If either fails,
leave it off — a CLV near zero because close==locked is the fabrication to avoid
and it would look like success.

PART 3 — full outstanding board included in the report, covering model work
(A-flood grade fix on p_win vs fair_prob, the collapsed-output re-adjudication
list, calibration/time-series with no honest source, price-triplet MODEL leg),
surfaces (D1 mount, share cards, notifications, Offseason, /system, S3 media,
45 unwired glyphs, /record has no nav link) and infra (NBA/soccer never settle,
three credentials still flagged for rotation, migration drift 023-029).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 19:38:19 -04:00
builtbykev bedbb8c008 Build 2 Phase B: checkout claims atomically, webhook finalizes, bypass retired
Stripe wired to the Phase-A mechanism. Live prices verified READ-ONLY; no
Stripe object was created and no payment was run.

B1 PRICE KEY -> ID + BOOT ASSERTION (src/config/stripePrices.js). claim_founder_slot
returns a price KEY; this module is the only place a key becomes a Stripe id, and
it reads env (legacy STRIPE_PRICE_ANALYST/DESK accepted as fallbacks so an existing
deploy keeps working). assertPricesConfigured() is wired into server.js and FAILS
BOOT when any of the four is unset — verified by deleting one: it throws
"BOOT FAILED - unset Stripe price env for: desk_founder". A blank price can no
longer sell at the wrong rate or 503 a customer at checkout.

B2 CHECKOUT CLAIMS BEFORE CREATING THE SESSION. resolveCheckoutPrice previously
called founderSeatsAvailable() — a COUNT read, which WAS the race (two checkouts
at seat 99 both read 99, both got founder). It now calls claim_founder_slot and
uses the returned key. The promo-code bypass is retired: founderCode no longer
influences price or metadata, and getPriceId THROWS if handed a code rather than
silently granting a founder rate. metadata.is_founder is renamed is_founder_audit
and the webhook no longer reads it — caller-supplied metadata must never decide
who pays the lifetime founder price.

TRANSIENT-FAILURE POLICY (a real design call, not a default): if the claim RPC
errors we now fail RETRYABLY (503 claim_failed) instead of silently selling at
standing. Both silent options are irreversible — standing permanently overcharges
someone who was entitled to founder, and granting founder without a slot pushes
past the 100 cap at permanent prices. A full cap is NOT an error and still returns
standing normally, per "never error to the customer": a full cap is a real answer,
a DB blip is not.

B3 WEBHOOK FINALIZES THROUGH THE SINGLE WRITER. checkout.session.completed calls
finalize_founder_slot, which flips user_profiles.founder_pricing (canonical) and
mirrors users.founder_status in the SAME txn, so they cannot drift again (they
already had, 1 vs 0). Verify-after-write re-reads the profile and logs the end
state. If finalize errors, the tier is still set so a PAID customer is never left
unentitled, but no founder flag is guessed.

B4 SIGNATURE VERIFICATION was already present (constructEvent with
STRIPE_WEBHOOK_SECRET + express.raw). The live endpoint exists and is enabled:
https://api.vyndr.app/api/stripe/webhook subscribing checkout.session.completed,
customer.subscription.created/updated/deleted, invoice.payment_succeeded/failed.

VERIFICATION: V1 boot assertion proven by simulation. V2 all four prices retrieved
live and confirmed active with correct amounts and monthly recurrence (14.99 /
24.99 / 44.99 / 59.99) — read-only, nothing created. V3 no code path grants founder
except the claim (greps clean; the legacy helper now throws). V4 the handler reads
customer/subscription/metadata.user_id and calls finalize with signature
verification in place. V5 reset to a pristine 100 free / 0 claimed baseline with
both founder flags at 0.

Secrets live only in .env (0600, gitignored, untracked). A pre-commit scan
confirmed NO tracked file contains the key material.

Floor: 320 suites / 3984 passed, 3 skipped (superseded founder-code tests), web
build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 18:42:44 -04:00
builtbykev 7c6fd95e68 Build 2 Phase A: real founder cap — atomic claim, race PROVEN, flags collapsed
DB only. No Stripe call, no checkout/webhook rewire (Phase B). Migrations 035,
036, 037 applied to prod and tracked; repo files added.

035 SCHEMA TRUTH — user_profiles gains stripe_customer_id and
stripe_subscription_id (G3 proved the webhook stores neither today, yet
finalize and grandfather reconciliation both key off the subscription id), plus
a partial unique index so a subscription id resolves to exactly one profile.

036 THE MECHANISM — founder_slots is a real TABLE replacing the decorative view.
The claim is a single UPDATE whose target row is chosen FOR UPDATE SKIP LOCKED;
no count is read in the decision path. UNIQUE(slot_number) plus a PARTIAL
UNIQUE(user_id) WHERE status <> 'free' (one live slot per user). Seeded 100 free.
Q1 global pool: the slot travels with the user, so analyst->desk keeps founder
with no second claim. Q2: release_expired_slots handles TTL abandonment ONLY —
cancelled slots retire, so the counter only rises. A6 redirects
founder_pricing_seats to count claimed slots, capped 100.

PRICE IDS ARE NOT IN SQL. claim_founder_slot returns a price KEY
(analyst_founder / analyst_standing / desk_founder / desk_standing) and the Node
layer maps it to STRIPE_PRICE_* env with a boot assertion — adopted over
hardcoding so a typo fails at boot instead of becoming a permanent mis-charge.

A7 FLAG COLLAPSE — finalize_founder_slot is now the SINGLE writer of both
founder flags in ONE transaction: user_profiles.founder_pricing is canonical and
users.founder_status mirrors it. founder_status is NOT dropped (G5 proved it
live: written at stripeService:163, served at routes/stripe:95, loaded in
middleware/auth:24 PROFILE_COLUMNS). Only the independent write is retired —
the two flags had already drifted in prod (1 vs 0).

A9 RACE TEST, run in Supabase before any Stripe:
  - pool squeezed to ONE free slot; three distinct users claimed concurrently
    -> EXACTLY ONE is_founder=true on slot 100, two returned analyst_standing,
    zero double-allocation.
  - idempotency: the winner claiming again returned the SAME slot 100 and still
    held exactly 1 live slot (two tabs cannot take two seats).
  - constraint layer proven directly: a raw UPDATE granting that user a SECOND
    live slot was REJECTED by the partial unique index, and verify-after-write
    confirmed state unchanged (1 live slot, target row untouched).
  HONEST LIMIT: the three claims contend within one transaction via LATERAL, so
  this proves the claim logic, the SKIP LOCKED path and the constraint that makes
  parallel safe — but it is not N genuinely parallel backend sessions. True
  multi-session concurrency is not drivable through this SQL interface and should
  be exercised once in Phase B against the test key.

037 NEXAPAY DROP — own migration, evidence-led (G4: zero code refs, column
empty). VYNDR is Stripe-only.

CLEAN BASELINE (Q3) verified after the test: 100 free slots, 0 non-free, counter
0/100, and BOTH founder flags cleared to 0 across user_profiles and users — the
inconsistent test record is no longer enshrined as a founder.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 17:19:39 -04:00
builtbykev e970ab1ef3 report: Build 2 Review Zero — G1-G6 + DB verified; two order expectations wrong
No build, no migration, no Stripe object touched. Awaiting Kev on Q1-Q3.

TWO EXPECTATIONS IN THE ORDER ARE WRONG:

1. G5 — users.founder_status is LIVE, not dead. Written by the webhook
   (stripeService.js:163), read and served by routes/stripe.js:95 as is_founder,
   and present in middleware/auth.js:24 PROFILE_COLUMNS so it loads on EVERY
   authenticated request. The guardrail says don't write it unless G5 proves it
   live — G5 proves it live, so A5 must NOT drop it.

2. THE TWO FOUNDER FLAGS ALREADY DISAGREE IN PROD: user_profiles.founder_pricing
   is true on 1 of 3 profiles while users.founder_status is true on 0 of 3. The
   webhook writes both from the same isFounder, so this is a dual-write that has
   already drifted. The build must pick one canonical flag and derive or retire
   the other; two independently-writable founder flags is how a founder loses
   their rate on one code path.

GREPS: G1 founder_pricing has exactly one writer (the webhook mirror) and four
readers (partners MRR attribution, the profile API, the profile badge). G2 the
promo-code bypass is the ONLY founder gate today — getPriceId(tier, founderCode)
against VALID_FOUNDER_CODES, stamped into metadata.is_founder, which the webhook
then trusts, so a code alone mints a founder at any seat number. G3 the webhook
DOES set tier + subscription_status=active + founder_pricing (closing an earlier
CANNOT DETERMINE: a paid sub does flip the Build-1 gate) but stores NO
stripe_subscription_id, confirming A1. G4 nexapay has ZERO code references and
the column is empty, so A5's drop is evidence-supported as its own migration.
G6 price selection is getPriceId -> line_items.

DB VERIFIED: user_profiles has nexapay_customer_id and NO stripe_customer_id /
stripe_subscription_id (A1 needed); users already carries stripe_customer_id;
founder_pricing_seats is a VIEW; 3 profiles, 1 flagged founder.

CANNOT DETERMINE: the four Stripe price IDs — no STRIPE_SECRET_KEY or
STRIPE_PRICE_* in this environment, so I could not independently re-verify that
the IDs in the order are what prod will charge. Since A3 would hardcode them, a
typo becomes a permanent mis-charge; recommend reading them from env (already the
pattern) with a boot assertion that all four resolve.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 17:07:54 -04:00
builtbykev 14f3ce95b9 report: Build 2 Review Zero — payment mechanism BLOCKED, no Stripe credentials
Nothing built. No Stripe object created or changed, no price logic touched.

STOPPED because there is no STRIPE_SECRET_KEY in this environment (.env holds
only ODDS/SUPABASE/INTERNAL keys). The order's standing floor requires
founder/standing/grandfather/race all verified server-side; none of that is
verifiable here, the standing price objects cannot be created, and the
concurrent-checkout race cannot be exercised. On a payment path the failure modes
are permanent and customer-facing — a race bug mis-prices a subscriber forever,
a grandfather bug overcharges one every month — so it must not ship unverified.

VERIFIED ANYWAY:
- Stripe IS live and FOUNDER price objects DO exist. /api/founders/count returns
  {available:true, claimed:0, total:100}, and routes/founders.js returns
  {available:false} whenever countFounderSeats() is null, which it is when
  !STRIPE_SECRET_KEY || founderPrices.length === 0. So available:true proves the
  secret key and at least one founder price ID are configured in prod, and
  claimed:0 is a real count rather than a fallback.
- THE COUNTER IS NOT A GATE. It is a cached (300s) READ, not a claim; founder
  pricing is gated by CODE + EXPIRY, not by the count, so anyone holding
  FOUNDER2026 gets the founder rate at any seat number and the cap is decorative.
  Two simultaneous checkouts at slot 99 would both read 99 and both get founder —
  there is no lock or unique constraint anywhere in the path.
- The gate reads users.tier via config/tiers.js reasoning_visible, so a
  successful subscription must set users.tier for Build 1's gate to open.

CANNOT DETERMINE: whether the STANDING price objects exist (env unreadable, and
getPriceId falls back SILENTLY to a PRICE_UNCONFIGURED sentinel, so a missing
standing object would not surface until the first post-cap checkout 400s in front
of a paying customer); whether the webhook writes users.tier on
checkout.session.completed.

DESIGN IS SETTLED for when it unblocks: a founder_slots table with a unique
constraint on (tier, slot_number) claimed before the Stripe call — the unique
index, not a count read, is what makes the race impossible; price selection from
the claim rather than a code, with the code+expiry bypass retired; grandfathering
by simply never calling Stripe price-migration on a founder sub;
founder-follows-upgrade by claiming on the target tier and releasing the slot on
cancellation; honest display that shows no number when the count is unavailable
(the existing route already sets that precedent).

PREREQUISITES, all needing Kev and none of them code: confirm/create the two
standing price objects and set STRIPE_PRICE_ANALYST / STRIPE_PRICE_DESK; confirm
the webhook sets users.tier; provide a Stripe test-mode key so the race,
grandfather and end-to-end unlock can be exercised rather than asserted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 08:30:32 -04:00
builtbykev e8b15c705a docs: free proof surface recorded (/record, hollow-preserving, CLV honest-absent)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:56:08 -04:00
builtbykev 5930f18d81 Free proof surface: /record — tier-record-forward, honest CLV building panel
Presentation over existing endpoints. src/ untouched (git diff empty): no grade,
model or ledger change. Pricing/migration are Builds 2/3.

TIER-RECORD-FORWARD. /record reads the canonical public aggregates (/api/accuracy
+ /api/ledger/model) and prints them as-is. B 60% n512 and C 57% n413 ship with C
honestly BELOW B; A (n2), D (n5) and F (n5) render HOLLOW with their real sample
instead of a rate. Sport slicing (all/mlb/wnba) is client-side because the
endpoints ignore ?sport= — mlb 526 @62%, wnba 411 @54% come from the sports map.

THE LOAD-BEARING RULE, enforced in lib/proofRecord.js and locked by tests: where
the source withholds a percentage it stays null. A is 1/2 and therefore 50% is
derivable — a test asserts we do NOT derive it, because the API withheld it on
purpose (n < 20).

CLV IS AN HONEST ABSENCE, NOT A NUMBER. beat_close_pct is null because
clvCaptureReliable() has not passed. The panel says "NOT PUBLISHED YET" and
explains that any percentage printed today would be measuring our collection gaps
as much as our edge; it surfaces the accruing sample (937) but no rate. Tests
assert the panel never falls back to clv_beat/clv_sample (34/937 = 3.6%) and that
the serialized panel contains no "3.6" — that number is computable and would be
wrong, which is the exact fabrication this surface exists to refuse. The panel is
built to receive a real number later without a redesign.

HELD, and named on the page rather than faked: calibration and accuracy-over-time
are absent because there is no honest source (no claimed-vs-actual endpoint;
window_days fixed at 30 with no series). The page says so, and says it is not
because they are unflattering.

A page-level test asserts no hard-coded percentage exists in the markup, so no
figure can drift from the aggregate, and that the page never touches /api/snapshot
or itemized rows — the Build-1 gate holds and the exploit stays dead.

Floor: 320 suites / 3986 tests green (16 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:53:44 -04:00
builtbykev 4c302b5722 report: free proof surface Review Zero — three of four leads have no data source
Nothing built. Docs only. The premise was that this is cheap assembly over
existing aggregates; verified, it is not.

0.1 FILTERABILITY — the endpoints are NOT filterable. Probed live:
    /api/ledger/accuracy?sport=mlb -> total 937
    ?sport=wnba -> total 937
    ?window=7 -> total 937
    identical payloads; the params are ignored (the req.query reads at
    routes/ledger.js:61-63 belong to a different route than /accuracy at :68).
    Sport and tier CAN be sliced client-side from /api/accuracy's sports map and
    /api/ledger/model's by_tier. TIME WINDOW CANNOT — window_days is fixed at 30
    inside getModelAggregate with no param and no stored series, so
    "accuracy over time" has no data source.

0.3 CLV CANNOT LEAD WITH A NUMBER. /api/ledger/model exposes the aggregate, and
    live it returns beat_close_pct = null and clv_distribution = null despite
    clv_sample 937. They are null BY DESIGN: ledgerService publishes them only
    when clvCaptureReliable() passes, and it does not — the capture is still the
    starved instrument the 07-28 repair improved but did not finish. The trap to
    avoid is exact: clv_beat/clv_sample = 34/937 = 3.6% is computable and would
    be WRONG, because the value is null due to instrument distrust, not a missing
    division. Publishing it would be the marketing fabrication this order most
    forbids. CLV can only lead with an honest absence.

0.2 The honest-record laws are ALREADY enforced at source: buckets return
    A pct:null (n=2), B 60% (512), C 57% (413), D pct:null, F pct:null — thin
    tiers already refuse to round. C genuinely sits below B, which is the
    unflattering truth and must be shown as-is.

0.4 CALIBRATION CURVE has no data source — clv_distribution is null and there is
    no claimed-vs-actual endpoint; the 07-26 calibration work was a one-off
    read-only measurement, never wired to a served surface.

BUILDABLE NOW: tier hit-rates by sport with existing hollows preserved,
client-side sport/tier filtering, the capped 3-call sample, and honest state copy
including a CLV not-yet-publishable panel that names the reliability guard.

NEEDS ITS OWN ORDER FIRST: CLV as a leading number (blocked on capture
reliability, not presentation), accuracy over time (needs a param or daily
series), the calibration curve (needs a claimed-vs-actual endpoint).

RECOMMENDS shipping tier-record-forward with an honest CLV building panel rather
than CLV-forward — CLV-forward with a null cannot lead, and with 3.6% would be a
lie. That preserves the premise's strongest claim (a real thin honest record
out-credibilizes a fake fat one) without inventing a number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:43:12 -04:00
builtbykev 7cf3892e76 docs: corrected gate recorded + cache-busting verification lesson
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:31:45 -04:00
builtbykev 713f90183f Build 1 CORRECTED: itemized grades are PAID (live AND settled) — exploit killed
Serving/gating change only. src/services/ untouched: no grade, model or
settlement-logic change. Pricing = Build 2, migration = Build 3.

WHY THE PRIOR GATE WAS WRONG: freeing grades at resolution made the free tier a
ONE-DAY-DELAYED FEED OF THE WHOLE PRODUCT — settlement is nightly, so a bettor
watching one cycle behind got the entire method free. There is now NO
per-grade resolution flip: an itemized grade, tonight's or last week's, is
Analyst+.

FREE now gets, none of it itemizing the nightly slate:
  1. the full data aggregator (unchanged — schedule, per-book lines, stats,
     streaks, hubs)
  2. the AGGREGATE track record, which ALREADY EXISTS and is public:
     /api/accuracy (sample 937, byGrade tiers, per-sport mlb+wnba, min_sample 20)
     and /api/ledger/accuracy (per-grade buckets). The honest-record laws are
     already honored there — A/D/F return pct:null under the n>=20 threshold
     rather than a fake percentage.
  3. a CAPPED, day-rotated sample of resolved calls for texture: cap 3, stable
     within a day, rotates across days, and only RESOLVED rows are eligible so a
     live read can never be sampled. The cap is what kills the exploit — three
     rotating past calls cannot reconstruct a nightly slate, whereas the full
     settled list is the feed one cycle late.
  4. the locked shell of tonight's reads: they exist, and their shape.

EVERY itemized grade for an unentitled tier now loses grade, confidence,
confidence_basis, reasoning, kill_conditions_triggered, projection, edge_pct,
matchup_grade, form, alt_lines and kelly, and is stamped locked. Free-side DATA
survives so the board still reads as real: player, market, line, book_odds,
fair_odds (the de-vigged fair number is the free hook and is never the paywall),
season/last10 stats, archetype — and `outcome`, because a RESULT is a fact
rather than a judgment.

The tease stays aggregate-only (live_locked {count, tiers}) computed from the
ungated rows and never joined back to one, and no gated row carries a grade, so
nobody can work out which prop is the A.

Floor: 319 suites / 3970 tests green, web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:29:24 -04:00
builtbykev 7ddf159e4a docs: Build 1 settled/live gate recorded + live anonymous fingerprint
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:21:36 -04:00
builtbykev 6d36e05bfe Build 1: the settled/live gate — unresolved is paid, resolved is free
Serving/gating change only. src/services/ untouched (git diff empty): no grade,
model or settlement-logic change. Pricing and migration are Builds 2 and 3.
Push scoring untouched.

THE RULE: a grade is PAID while its outcome is unknown and becomes FREE the
moment it resolves.

Resolution is read ONLY from a written outcome — never from time, game status or
gradedAt. A game can be final long before the settle pass runs, so treating
"probably over" as settled is exactly how a live edge would leak; a test asserts
an hours-old gradedAt with no outcome is still LIVE. void and unrecoverable ARE
resolutions (terminal results, no live edge left). isResolved FAILS CLOSED:
null outcome, {} with no result, and empty-string result all read as LIVE, so a
settlement failure withholds content rather than exposing it — the same
direction resolveTierFromRequest fails.

FREE/ANON: settled grades pass through IN FULL, reasoning and kill conditions
included — settled reads are the proof product and cost nothing once the outcome
is known. That also converts the previously-unenforced board reasoning leak into
a deliberate rule rather than an oversight.

LIVE grades for unentitled tiers are reduced to a shell: every piece of model
JUDGMENT is dropped (grade, confidence, confidence_basis, reasoning,
kill_conditions_triggered, projection, edge_pct, matchup_grade, form, alt_lines,
kelly) and `locked: true` is stamped so the card renders the unlock prompt. The
free-side DATA stays so the tease is real rather than empty: player, market,
line, book_odds, fair_odds, season/last10 stats, archetype, gradedAt, history.
fair_odds deliberately survives — the de-vigged fair number is the free hook and
is never the paywall. A test asserts the serialized free row carries no trace of
the withheld judgment.

THE TEASE IS AGGREGATE ONLY: live_locked = {count, tiers} computed from the
ungated rows and never joined back to one, and no gated row carries a grade — so
a free viewer learns that N reads exist and their tier shape without being able
to work out WHICH prop is the A.

Gate order in the route: stripModelPrice (S67) first, then gateLiveGrades.
Entitled tiers get the array back by reference — zero cost, zero change.

Floor: 319 suites / 3971 tests green (10 new), web build exit 0.

One test note: the route-level supertest case was removed deliberately — it
needs a live Redis and hangs on ioredis' reconnect timer in a single-suite local
run (known behaviour, CLAUDE.md). The gate contract is fully covered by pure
tests; the wire is verified against prod anonymously in the fingerprint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:19:51 -04:00
builtbykev dbc1416485 docs: tier redesign spec recorded (gate discriminator exists; counter is display-only; base is 3 users)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:07:13 -04:00
builtbykev 4a4a3428d8 spec: tier redesign (Option 2, settled-free / live-paid) — design + build order
Report-first. Nothing built; no tier, price, gate or Stripe object changed.

REVIEW ZERO findings that shape the design:

0.2 The ladder is HALF-EXPRESSIBLE already — PRICE_MAP separates founder from
    standing objects, so lifetime grandfathering is native (a sub created against
    a founder price stays on it). BUT founder access is gated by CODE + EXPIRY
    (FOUNDER2026/VYNDR/BETONBLK/EARLYBIRD, expiry 2026-12-31), NOT by seat count:
    anyone with a code gets founder pricing at any seat number. A real
    Stripe-derived counter exists (/api/founders/count, live 0 of 100) but only
    DISPLAYS — and it is cached 300s, so it cannot enforce "slot 100 and 101
    differ permanently". Making the counter the gate, transactionally and
    uncached at checkout-session creation, is a real build.

0.3 The paid->free flip point already exists ON THE SERVED PAYLOAD: settlement
    writes ledger_entries.outcome + settled_at, and /api/snapshot already merges
    per-grade results — live WNBA returns 25 grades, 5 carrying
    outcome {result:'hit', actual:1}. So the gate discriminator (outcome != null)
    is present on the exact object to be gated; no new pipeline needed.

0.4 THE MIGRATION IS NOT WHAT THE ORDER ASSUMES: the users table holds 3 users,
    all free, created Jun 12-19, and ZERO paid. There is no warm mass base — the
    "founder launch to existing users" is a courtesy note to 3 people, and the
    launch's real audience is people who have not signed up yet.

DESIGN: free = full data aggregator + the COMPLETE settled record (letter,
reasoning, edge, outcome — browsable and filterable), which is the proof hook.
Analyst = tonight's live grades + reasoning + edge, unlimited. Desk = + alt
ladder, Kelly, portfolio, engine2. Reasoning/grade/edge are ONE paid unit while
live and become free together at resolution — which also converts today's
unenforced board-reasoning leak into a deliberate rule.

GATE: outcome == null => live => Analyst+; outcome != null => settled => free.
Filter whole grades server-side (not field-strips) so a live grade cannot leak
partially; never infer resolution from time or game status, only from a written
outcome; fail closed to LIVE so a settle failure withholds rather than exposes;
void/unrecoverable are terminal and therefore free.

BUILD ORDER: (1) the settled/live gate, (2) the free settled-record surface —
noted as arguably shipping WITH (1), since gating live grades without it leaves
free users no graded content at all, (3) Stripe ladder + transactional counter +
grandfather rule + retire the code gate, (4) the founder note to the 3,
(5) pricing visuals (already designed in the package).

CANNOT DETERMINE: whether the four Stripe price objects exist in the dashboard
(env not readable here) — flagged as a prerequisite for build 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 07:06:33 -04:00
builtbykev c6ef4cfb2b docs: tier structure recorded — free board uncapped, board reasoning ungated vs config intent
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 06:41:15 -04:00
builtbykev 844ab96f21 report: tier structure pull — the declared free-tier gate is unenforced on the board
Read-only. Nothing changed.

FREE TIER, EXACTLY:
  - Board /api/snapshot: NO count limit. The only gate is
    stripModelPrice(grades, tier) at routes/snapshot.js:106-107 — no slice, no
    volume branch. Live anonymous right now: MLB 5, WNBA 25 = the full board.
    The "3 scans/day" cap rations the SCAN path only.
  - Grade letter: fully visible on every tier (grade_visible: true). Anon also
    receives confidence, edge_pct and VYNDR's own projection.
  - Edge fields: correctly stripped. p_win/ev_pct/model_odds/value/takeable are
    ALL absent from the anonymous payload, with model_price_locked stamped so the
    card shows a lock teaser rather than an absent leg. This half works as designed.

THE HEADLINE — the two paths disagree on reasoning:
  - Scan REDACTS it: tierGating.js lockReasoning + lockKillConditions +
    tier_gated + upgrade hint, driven by free.reasoning_visible = false.
  - Board SERVES IT IN FULL: snapshotGating MODEL_FIELDS is
    [model_odds, p_win, ev_pct, value, takeable] — reasoning is not in the list.
    Verified live anonymously: full reasoning.summary plus a kill condition WITH
    its reason.

Intent: config/tiers.js declares free: { reasoning_visible: false } with the
comment "blurred — frontend renders tier-locked". One of the two paths does not
enforce the product's own declared line, so the evidence reads as oversight
rather than funnel — a funnel would be declared in config, not contradicted by
it. Flagged with the counterweight: board reasoning is good marketing and the
data layer is already free, so closing it is a monetization tightening (Kev's
call), not a fabrication fix.

FREE DATA IS A REAL AGGREGATOR, not just a limited graded view: schedule,
per-book lines, player stats, streaks, hot lists, team hubs, public record — all
public and uncapped (probed live).

PAID (config/tiers.js, checkout.js:4): analyst $14.99 / desk $44.99. Analyst is
unlimited reads; Desk differentiates on capability (alt ladder, Kelly, portfolio,
engine2). africa tier is defined but activation is blocked on a DB CHECK
constraint. api_access is false on every tier. book_odds/fair_odds deliberately
pass through on all tiers — the de-vigged fair number is the hook and is never
the paywall; only model_odds gates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 06:40:43 -04:00
builtbykev 43281bb885 report: D1-close Review Zero — mount not performed, three findings
Nothing changed: no mount, no row edit, no data threading. Docs only.

1. THE RATIONALE DOES NOT REACH THE ROW. StripProp carries stat/line/side/grade/
   gradedAt/delta/awaiting/outcome/movement/revisedFrom/book/bestBook/dead/
   history — no reasoning, no kill_conditions_triggered — and
   buildPlayerStripsFromProps never threads them. Mounting the hover needs a new
   field on the strip contract threaded through the slate adapter: additive, but
   a data-path change rather than a mount.

2. THE 0.3 PREMISE INVERTS — THE RATIONALE IS ALREADY PUBLIC. Verified live and
   anonymously against prod: /api/snapshot/wnba returns reasoning.summary with no
   locked flag plus kill_conditions_triggered. stripModelPrice removes
   model_odds/p_win/ev_pct/value/takeable but NOT reasoning. So the full model
   rationale already ships to every anonymous browser on the main board, while
   the same content IS tier-gated on the scan path (tierGating.js). Mounting the
   hover would leak nothing new, but would surface content that is currently
   shipped-but-unrendered, and the product gates it in one place while serving it
   openly in another. That is a monetization/consistency decision, so it is
   reported with three options rather than resolved unilaterally.

3. ROW-GRAMMAR IS LAW AND LOCKS StatStrip's SOURCE ORDER. rowGrammar.test.js
   asserts element order via src.indexOf on the component source; adding a
   rationale affordance or a team chip moves those offsets, so specs/ROW-GRAMMAR.md
   and the test must be amended in the same commit. That makes this spec-amending
   work needing its own slot decisions, not an additive mount.

Safely mountable with no blockers: reveal.js (wraps the row list, no StatStrip
internals, no new data, no grammar slot). teamChips needs a grammar slot;
rowRationale needs the data threading AND the gating decision AND a slot.

Recommends splitting D1-close into: mount reveal now; a ROW-GRAMMAR amendment
order for the chip + rationale slots; then the rationale mount once the gating
decision is made.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 06:25:05 -04:00
builtbykev a0501f99c0 docs: D1 finish recorded (rationale real-or-absent, reveal once-on-view, 10/80 chip coverage)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 06:18:42 -04:00
builtbykev 085e8a3a63 D1 finish: row-hover rationale, IntersectionObserver reveal, team-gradient chips
Additive frontend. Backend untouched (git diff src/ = empty): no grade, model,
classifier or ledger change. Scope held to the row anatomy these three items
need — no System-artboard-wide rebuild. Push scoring untouched.

REVIEW ZERO — the two checks that decided whether these could be honest:

0.2 RATIONALE SOURCE — VERIFIED REAL. Live snapshot grades carry `reasoning`
    and `kill_conditions_triggered`. The summary is built by analyzeViaEngine1
    from the actual feature vector (l5/l20 averages, gap to the line, home/away,
    opponent defensive rank, rest days) and kills carry real codes + reasons.
    So the hover shows genuine grade truth, not a placeholder.

0.3 TEAM COLOURS — PARTIAL, and deliberately left partial. The System artboard
    defines a colour pair for only 10 teams (BOS CHC CHI DEN LAD MIL MIN NYY PIT
    SD), lifted verbatim; lib/teams.js holds ~80. The other ~70 are NOT invented
    — a wrong team colour is a recognition error the user reads as fact. Unknown
    teams get the honest-neutral chip (muted border, no colour claim), never a
    guess and never a blank gap. Coverage is reported by coverage(), not hidden.

SHIPPED:
- web/src/lib/rowRationale.js — rationaleFor() returns real summary + kills, or
  NULL. No generic fallback: an empty hover is honest, a manufactured "why" is a
  fabricated model explanation. A locked/tier-gated reasoning is treated as
  ABSENT rather than paraphrased or leaked, and a kill condition with no reason
  explains nothing so it is dropped.
- web/src/lib/reveal.js — IntersectionObserver reveal that fires ONCE then
  unobserves ("react to truth, then rest"), reuses D1-A's bootDelayMs for the
  60ms stagger so there is ONE source of truth for the timing, and reveals
  IMMEDIATELY when IntersectionObserver is absent (SSR/test) so a missing API can
  never hide real content. Reduced motion is handled by the existing CSS, so the
  row is visible either way.
- web/src/lib/teamChips.js — Rev-3 geometry (10px, 135deg, before the abbr,
  inside the row) plus the ranked opacity ramp 1/.86/.64/.48 so chips dim with
  their row. Swap-ready for licensed logos at the same size.

Floor: 318 suites / 3961 tests green (15 new), web build exit 0.

The three modules are pure and unit-locked; mounting them into the live row
components is a follow-up, and the visual result belongs in the Chrome audit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 06:18:12 -04:00
builtbykev e34e99c426 docs: D1-A recorded (glyph buckets, boundary channel completed, primitives)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 06:04:50 -04:00
builtbykev 3d1a3c7794 D1-A: combat glyphs, boundary-channel blue, reaction primitives, READ-FAB
Additive frontend/visual. Backend untouched (git diff src/ = empty): no grade,
model, classifier or ledger change. The 41->74 registry expansion is HELD for
D1-B. Push scoring untouched.

REVIEW ZERO — classifier coverage bounded the glyph wiring. Three buckets, and
the computation was redone three times before it was right (the frontend keys
GLYPHS by ARCHETYPE NAME while the backend keys `glyph:` by SHAPE NAME, and most
registry keys are unquoted identifiers — the first two passes mis-parsed both):
  (a) classifier-backed, already wired: 38
  (b) classifier-backed, package SVG exists, NOT wired -> WIRED HERE: 6
      striker, grappler, pressure, counter, grinder, finisher — all MMA/combat
      archetypes in archetypeService.js that were rendering EMOJI fallbacks
      ('*', 'x', '>', '<>') where the package ships real 24-grid duotone marks.
  (c) package SVG with no classifier -> HELD for D1-B: 39 (wiring them would
      render nothing)
  (d) classifier-backed but NO package SVG: 2 ('dual threat', 'paint boss') —
      a DESIGN gap, not a build gap; flagged for D1-B.
GLYPHS map 38 -> 44 keys, deliberately far short of the package's 83.

BOUNDARY CHANNEL — the blue tokens already existed (--priced-out set) and were
applied on NoMarketState and the scan void box, but PriceTriplet's NO_MODEL
("line not priced") still rendered in neutral text, so the channel was applied
inconsistently. NO_MODEL now renders in the channel, completing "every boundary
state or none". Token-only (no hex fallback and no hex in prose — PriceTriplet's
own test forbids literal hex, and it caught both).

REACTION PRIMITIVES — new web/src/lib/reactions.js + globals.css keyframes at the
exact HANDOFF timings: flash .75s ease-out, boot stagger 60ms steps, reactions
gated 1.5s, WIRE hold 6s. nudge() REFUSES a no-op (null/absent direction -> no
flash) so the primitive cannot be attached to an idle loop — a flash without a
new datum is the UI lying about the feed. Reduced-motion honoured.

READ-FAB — aligned to the exact package geometry: 50px circle, translateY(-14px),
6px void ring (was 46px, marginTop -16, 3px ring).

CARD TOKEN — audit correction: #0E0E14 was already tokenised as --bg-1/--card;
the audit's "1 file" was counting the raw hex, not the token. No change needed.

Floor: 317 suites / 3946 tests green (16 new), web build exit 0.

NOT DONE THIS ORDER (reported, not silently dropped): row-hover rationale and
IntersectionObserver reveal (Phase 3 item 8) and team-gradient chips (Phase 4
item 10) are not implemented — they need the System artboard's row anatomy,
which is a larger port than the rest of D1-A.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 06:04:19 -04:00
builtbykev 49565b5f02 docs: design-vs-build gap audit recorded (61 items, 6 build waves)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 05:46:32 -04:00
builtbykev e257474cc8 report: design-vs-build gap audit — 61 items enumerated, 18 absent
Package specs/design-reference (Jul 22) audited against the CURRENT repo
(bf7c0a3, ~9 days later). No ~/vyndr_design exists; the in-repo copy is the
package. Every claim is a direct file/grep/count check, not the harness that
returned a silent false in Wave 3.

61 implementable items enumerated: BUILT-TO-SPEC 20, BUILT-BUT-DRIFTED 7,
PARTIAL 16, ABSENT 18 (+1 CANNOT DETERMINE: 19-screen mobile parity needs a
visual pass).

Largest single gap: the glyph library — 38 of 83 designed SVGs are wired (46%),
and the design implies 74 display archetypes against a 41-entry backend
registry, so the archetype system is roughly half the designed scope.

Drift found on surfaces built recently: the book comparison wired 07-29 renders
per-book lines but has NO crown, NO disagreement axis, NO SPLIT chip and NO
movement strip — a simpler version than the S2 design. The mobile tab bar has 5
tabs but not the designed READ-FAB. Calibration gating disagrees with the design
(our n>=20 vs designed N30).

Wave-2 reclassification: Newsletter DESIGN EXISTS (S5 The Report is fully
designed) — the earlier status pull was wrong to call it a design gap. Live
tracking and Slip reader remain genuinely design-missing.

Model linkages named: Price Triplet waits on the EV layer producing
p_win/ev_pct/model_odds; the S4 calibration curve waits on the n-threshold
decision plus accrued buckets, while the CLV chips can build on the repaired
instrument now.

Ordered build list in six dependency waves: self-contained first (glyphs,
primitives, boundary-channel blue), then scanner-nudge-gated, model-gated,
resolution-pipeline-gated (share-card masters cannot ship — the tail has no
generation step and no trigger), licensing-gated (book logos, push-to-book),
then the large surface builds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 05:45:56 -04:00
builtbykev 91911cfb1c docs: Wave 3 recorded — /compare live-verified; resolution tail scoped
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 05:35:10 -04:00
builtbykev bf7c0a3c08 Wave 3: /compare built (real head-to-head); resolution tail scoped, not shipped
No grade, ledger or scoring change. Push scoring untouched.

REVIEW ZERO 0.3/0.4 — THE RESOLUTION TAIL DOES NOT FIRE. The resolver is
POST /api/grading/resolve (routes/grading.js:208), and its fanout at :356-371
covers webPush, telegram and discord — but:

  - share-card generation: SPEC'D-NOT-BUILT. Not in the fanout at all (grep
    shareCard in grading.js = 0). shareCards/renderer.js exists with ZERO
    callers, so the component is built but no step would ever invoke it.
  - push notifications: BUILT-NOT-FIRING. In the fanout but gated on
    webPush.configured() (VAPID). push_subscriptions = 0 rows and
    user_notifications = 0 rows — nothing ever subscribed or delivered.
  - Telegram result posts: BUILT-NOT-FIRING (gated on BOT_TOKEN + CHANNEL_ID).
  - Discord result posts: BUILT-NOT-FIRING (gated on webhookFor('results')).
  - recap (all-Final trigger): SPEC'D-NOT-BUILT. No recap file exists in src/.

AND THE WHOLE TAIL IS UNREACHABLE: nothing calls /api/grading/resolve — there is
no ESPN poller in the repo. The live settlement path is the scheduler's
settleAllOutcomes + settleAllLedgers, which fans out to opsNotify only (ops
alerts), with no user-facing output. So even the built channels have no trigger.

Per the order's own rule, ShareCard, /notifications, result posts and recap are
therefore ALL SCOPED, none shipped — no dead shells over a silent pipeline.

BUILT — /compare. Semantics (0.2): a same-market head-to-head, two players with
every row a measure BOTH sides are scored on, aligned via alignRows so the
numbers are comparable — deliberately not two disconnected graded props. Reads
the live /api/stats/player/:name?sport= aggregate. Honest-absent three ways: an
unresolved side reads NO DATA while the other still renders; a measure only one
side has renders a dash, never 0; if neither resolves the page refuses to
compare. NO VERDICT — it shows measures and says the reader draws the call.

Two pre-existing tests (vyndrPhaseE, vyndrParityQA) asserted the in-development
placeholder; both superseded rather than deleted — they now assert the stronger
properties against the real page (live fetch, no sample players, NO VERDICT,
NO DATA, "not a zero").

Floor: 316 suites / 3930 tests green (10 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 05:31:56 -04:00
builtbykev 831d09bdde docs: Wave 1 wiring recorded + honest fingerprint limitation (nav entries -> Chrome audit)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 05:01:33 -04:00
builtbykev ff7f5d8d2d Wave 1: wire /intelligence, /slip, /parlay + /marketplace honesty pass
Wiring + one copy pass. No grade, ledger, model or scoring change (diff empty
across intelligence/, ledgerService, outcomeService, gradeSlateService).

REVIEW ZERO — each surface proven with real data BEFORE wiring:
  0.1 /intelligence vs /system are NOT duplicates. System.dc.html is a
      multi-surface artboard (TERMINAL + INTELLIGENCE + WIRE sections), not the
      design for a distinct /system route; its INTELLIGENCE section is already
      realised as the live app/intelligence/page.tsx. No /system page exists and
      none should be built as a second copy — the prod 404 is correct.
  0.2 /intelligence renders live and gates SERVER-side, not by blur: the proxy
      requires auth and limits by tier (desk 50 signals / non-desk 8), and
      returns 401 to an anonymous caller (verified live). No leak.
  0.3 /slip parses a real DraftKings slip end to end: 3/3 legs,
      needs_review false, Aaron Judge total_bases over 1.5 @ -115. Honest limit
      recorded: parsers are layout-rigid, an unsupported layout yields ZERO legs
      rather than wrong ones (never-guess), so real-world OCR hit-rate across
      layouts is CANNOT DETERMINE until user slips arrive.
  0.4 /parlay direct route hits the real correlation builder on the same
      ParlayContext the drawer uses.
  0.5 /marketplace advertised four unbuilt things but made NO performance or
      profit claim, and its capture was already real (/api/waitlist upserts to a
      waitlist table). The gap was tense, not fabrication.

WIRED: Nav MORE gains Intelligence, Slip Reader and Marketplace; Parlay Lab
re-pointed from the drawer hash to /parlay (the drawer is unaffected —
ParlayPanel stays mounted with its floating badge).

GATING: /intelligence added to GATED_ROUTES because its feed 401s signed-out, so
an ungated link would land visitors on a permanently empty page. /parlay stays
OPEN deliberately — it is the free parlay funnel and gating it would be a
monetization regression.

/marketplace honesty pass: every item body now opens "Not built yet." /
"Not written yet." / "Not produced yet." with what is planned; the subhead states
it is not a purchase, not a pre-order and not a promise of a ship date; the
playbook item carries "No profit claim, no promised return". The capture stays
real — no fake button. Unit-locked.

Floor: 315 suites / 3920 tests green (12 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 04:51:42 -04:00
builtbykev 6328a5ce92 docs: boundary not written (no promotion); build triage waves recorded
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 01:50:31 -04:00
builtbykev e70c02aca0 report: incomplete-surface build triage; boundary tag NOT applied (no promotion exists)
Report-only. Nothing built, wired, tagged, or removed.

PART 1 — the recalibration boundary was NOT written. The build order asks to tag
grades pre/post edge-shading at "the true promotion timestamp"; there is no such
timestamp, and writing the marker would insert a fabricated model transition
into an append-only public record — the exact corruption the order exists to
prevent. Three independent production proofs:

  A. model_snapshots.code_sha — every sha that ran the pipeline in the last 5
     days is a documented commit from this session (f3bf300 currently live,
     then f310608, 9b5235c, b8ee216, afb56b1, 3592aba, 914a057), all stamped
     model_version engine1@2026-07-20. No promotion commit exists.
  B. Daily A-family share 07-24..07-30: 0.0, 0.0, 0.0, 0.0, 1.0, 1.6, 0.0 —
     flat at zero, no step change on any date. A 92.9% re-letter under shading
     would have driven the board to ~79-93% A overnight.
  C. efficiencyShading has zero production importers; HEAD d54eca0, clean tree.

Phase 2 delta: neither 92.9% nor 43.6% is attested in any measurement here, and
no re-lettering occurred at any scale, so partial-slate-vs-full-board cannot
explain a gap that does not exist. The only measured numbers, on all 1250 rows
(the full board): 97.4% would change, 79.8% up, 79.0% A-family, and 0 of 1250
rows actually shaded — the hypothetical re-letter would have come entirely from
an unapproved grading-basis switch, which is why the flip was refused.

Measurement integrity needs nothing new right now: model_version already
separates the S64 eras and takeable tags are complete (1245/1250). The boundary
becomes a hard prerequisite the day a promotion actually ships.

PART 2 — build triage for 14 incomplete surfaces, classified from code with each
one's real dependency and honest size, sequenced into five waves: pure wiring
(/intelligence, /slip, /parlay, /marketplace after a copy honesty pass);
design-only gaps (Live tracking, Slip reader, Newsletter artboards);
self-contained builds (/compare, ShareCard host, /notifications); model-gated
(price triplet MODEL leg and the calibration board both need p_win vs
fair_prob — building either on edge_pct would re-ship the retired 620% lie);
and sport-boundary/quota-gated (/soccer blocked on odds-api 0/500, /system,
Offseason).

Matrix corrections: Live tracking, Newsletter and Slip reader are built/live/
honest and marked incomplete only on the DESIGN column; /parlay is
drawer-reachable, not unreachable; /compare is already honest. No live surface
is showing fabricated data today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 01:49:55 -04:00
builtbykev d54eca0afc report: pre-audit status pull — no promotion occurred; matrix re-derived (15/26)
Verified three ways that nothing was promoted and no re-lettering happened:
HEAD is the no-flip commit with a clean tree, efficiencyShading is imported by
zero production files, and live grades carry 0.0% A-family (MLB {B:1,C:4},
WNBA {B:15,C:10}). Neither 92.9% nor 43.6% is a figure measured here - the
challenger's real numbers were 97.4% would-change / 79.8% up / 79.0% A, with
0 of 1250 rows actually shaded.

Takeable tags landed (1245/1250 tagged, takeable_floor on all, one floor -160;
the 5 untagged have no locked price). The model-version boundary is absent and
correctly so - there was no recalibration to mark. Not a pre-audit gap.

Matrix re-derived from live prod probes, importer counts and nav-link counts:
15 of 26 fully done. Book comparison RESOLVED (BookComparisonPanel routed to
the grade card). Grade card honesty improved by the edge_pct retirement.
ShareCard / MobileEdgeBoard / DemoScan still dead code. Seven live routes
remain orphaned with zero nav links; /system and /offseason are 404. No
LIVE-but-not-HONEST surface found.

Design is NOT complete: Live tracking, Slip reader and Newsletter are shipped
with no design artboard - design is the gap, not build. Offseason is the
reverse (designed, never built).

Chrome audit manifest assembled: 11 items with per-item session state, four
requiring an entitled Desk session that only Kev can drive.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 01:40:04 -04:00
builtbykev 10aaaeb1da docs: promotion gate not passed (no flip); Order B edge_pct display retired
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:58:34 -04:00
builtbykev 0997334f8b Order B: retire edge_pct display. Promotion gate NOT passed — no flip.
THE PROMOTION WAS NOT PERFORMED. Champion grade path byte-identical (diff empty
across intelligence/, gradeSlateService, snapshotService). Projection, p_win and
the CLV instrument untouched.

REVIEW ZERO IS A GATE AND THREE OF FOUR PREREQUISITES FAIL:
  0.1 scores are ESTIMATED priors from the founding spec, not measured. The
      premise's cited values are not in the code either — the module holds
      nba:points .80 and mlb:total_bases .55; there is no NBA 0.72 and no WNBA
      score at all.
  0.2 VERSION-BOUNDARY TAG DID NOT LAND — config/modelEras.js has zero shading
      references. It was deliberately not applied twice (nothing had been
      promoted) and reported both times. The order's own rule says STOP.
  0.3 NO ROLLBACK FLAG EXISTS — zero occurrences of SHADING_ENABLED /
      EDGE_SHADING / shadingEnabled anywhere in src/.
  0.4 takeable tags DID land (migration 034, 1246/1254 rows). PASS.

AND THE APPROVED DELTA DOES NOT MATCH THE MEASURED ONE. Approved: 43.6% of
grades re-letter, efficient markets tighten and soft hold. Measured on all 1250
live rows: 97.4% change (1217), 79.8% move UP, 17.6% down, resulting in 79.0%
A-family (MLB 93.4%) against the champion's 0.2%. And rows_actually_shaded = 0
of 1250 — 96.5% of markets are unscored (f=1) and the one scored market present
is the anchor (f=1.0 by construction). The entire re-letter comes from switching
to edge-vs-fixed-bar grading, NOT from efficiency shading, which is inert on
this board. That is an unapproved grading-basis change riding along, which the
order's own "no new scaling changes riding along" guardrail forbids.

Flipping would re-letter 97.4% of an append-only public record, move 79.8% of
grades UP and mint A's on 79% of the board, on a letter whose measured
correlation with outcomes is r ~ 0.005 — the exact scenario the permanent
founder ruling forbids.

SHIPPED — ORDER B (independent of the promotion, and a live falsehood):
edge_pct display retired from GradeResultCard (confidence strip, EDGE stat cell
now honest-absent, alt-ladder rung) and SoccerGradeResult. DeskShowcase kept
(already honest). Computation and the board's signed-edge sort fallback SURVIVE
— deleting them would re-break the sort fixed on 2026-07-29; a test asserts all
three survive and the sort still orders agrees -> disagrees -> absent.

Fixed two build-breakers the retirement caused (orphaned edgeColor import,
orphaned edge_pct destructure; edge_pct stays on the props contract). Two
pre-existing tests superseded rather than deleted: they asserted the edge figure
is sign-coloured, and now assert the stronger property that no edge percentage
renders at all.

Floor: 314 suites / 3908 tests green (9 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:57:57 -04:00
builtbykev 1e9c808d99 docs: edge-shading challenger measured — flooding persists, input scale is the bug
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:39:51 -04:00
builtbykev c2c7abbb65 Edge-shading challenger: built + measured. Flooding NOT fixed — input scale is the bug
Challenger only. Champion grade byte-identical (verified by diff). Nothing
promoted, no live grade re-lettered, no ledger row deleted or re-settled.

BUILT src/services/challengers/efficiencyShading.js (measured-never-served):
  adjusted_edge = raw_edge * f(efficiency); grade = band(adjusted_edge) against
  ONE fixed bar (A+>=10, A>=5, B>=3, C>=1, D>=0, F<0) that never moves.
  f(e) = E_SOFTEST/e bounded to (0,1] — soft markets intact (never amplified),
  sharp shaded toward but not past zero, unscored -> f=1 and FLAGGED.
  A fence test asserts no production grade path imports it.
  Cross-market behaviour is unit-proven: the same raw 6% edge grades A in soft
  mlb:total_bases and B in sharp nba:points.

MEASURED on 1250 live ledger rows — Phase 2.5's answer is NO, the flooding is
not gone: challenger 79.0% A and 80.9% A/B (MLB 93.4% A) vs champion 0.2% A.

TWO findings explain why, and they are the point of the order:

1. The shading is a NO-OP on the live board: rows_actually_shaded = 0 of 1250.
   96.5% of rows are UNSCORED (f=1), and the one scored market present
   (mlb:total_bases) is the anchor so its f is 1.0 by construction.
   mlb:strikeouts and nba:points do not appear in the ledger at all (our
   basketball is wnba, not nba). Challenger vs baseline: 0 rows changed.

2. Placement was never the bug — the INPUT SCALE is. Against a fixed 5% bar the
   RAW edge already clears A on 100% of MLB doubles, 89.6% of hits, before any
   shading. MLB median raw edge is 60%, twelve times the bar. Decisive test:
   apply the sharpest score in the spec (f=0.647) to EVERY row — the maximum
   the design permits — and 75.8% still clear A (MLB 91.7%). Since f is bounded
   <= 1, no achievable shading can close a 12x overshoot. Moving the multiply
   from the threshold to the edge does not change the outcome.

This is edge_pct behaving as the 2026-07-29 diagnosis described: a price-free
(proj-line)/line gap whose scale is a function of line size. It is not a
betting edge, so no fixed betting-edge bar is meaningful against it.

2.6 efficient-market over-suppression: CANNOT DETERMINE — zero live rows are
shaded, so there is no efficient market in the data to over-suppress.

Phase 3: takeable tagging was completed in the previous order (migration 034,
1246/1254 rows) and is not repeated. The model-version boundary is again NOT
applied: nothing promoted, so no boundary exists.

Unblocking needs the input replaced, not the multiply moved: p_win vs
fair_prob (both already computed) instead of edge_pct, plus scores FIT from our
own record for the markets we actually grade.

Floor: 313 suites / 3899 tests green (9 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:39:13 -04:00
builtbykev 896e6e1a00 docs: takeable tagging shipped, efficiency challenger blocked (board + state)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:20:18 -04:00
builtbykev 2bfaeff572 Ledger takeable tagging (deferred C2); efficiency challenger BLOCKED
Champion grade UNCHANGED. Push scoring untouched. Additive tags only — nothing
deleted, nothing re-settled.

PART A — THE EFFICIENCY CHALLENGER: BLOCKED, NOT BUILT.
Review Zero came back ABSENT on all three inputs:
  0.1 efficiency scores DO NOT EXIST (zero occurrences of market_efficiency /
      marketEfficiency / efficiency_score in src/ or web/src/).
  0.2 base thresholds DO NOT EXIST (engine1.js has zero `edge` references — the
      grade is not an edge-vs-threshold comparison; grade_thresholds.json holds
      PROBABILITY bands).
  0.3 the +/-0.05 additive efficiency nudge DOES NOT EXIST. The only 0.05s on
      the grade path are featureCache.teammate_absence_bump, a bvp_advantage
      cutoff, and p*0.9+0.05 inside probabilityEstimator (the 0.5*0.1 term of
      the shrink-toward-0.5). There is no additive scaling to replace.

So a challenger differing from the champion in EXACTLY ONE thing cannot be
constructed: there is no additive scaling to swap, no base threshold to
multiply, and engine1.js has zero `sport` references so market cannot reach the
grade. A threshold must exist first — that is R1 of
specs/full-output-grade-mapping.md, an explicitly held separate order. Shipping
R1+R4 together would make the Phase-3 delta report misleading: the re-letter
would be driven mostly by switching to probability grading while being
presented as the efficiency fix.

0.4 coverage: the spec names 5 scores; the live ledger has 11 markets and only
MLB total_bases maps to one. 9 of 11 have no score, so "all scored markets"
cannot be satisfied without inventing 9 numbers.

PART B — LEDGER TAKEABLE TAGGING: BUILT (the deferred C2).
New src/config/takeableStandard.js: floor on the minus side, UNCAPPED plus.
Deliberately NOT valueEngine.isTakeable (the -160..+200 PROMOTION band) — a
+400 prop is not promotable but IS takeable; a test asserts the two diverge on
the plus side and agree at the floor so they can never quietly merge. Absent
price returns null, never false (Number(null) === 0 would tag a missing price
takeable). The floor is POLICY not derived (C1 could not derive one) and is
labelled so; each row records takeable_floor so a re-derivation can re-tag.

Migration 034 (applied + tracked): ledger_entries.takeable boolean +
takeable_floor numeric, nullable, partial index. Forward tagging in
ledgerService at row build; backfill in one statement.
Result: 1254 rows, 1246 tagged (781 takeable / 465 below floor), 8 NULL with
null_despite_price = 0 (the NULLs are genuinely priceless rows). Settled 1163
and graded 1254 unchanged.

PART C — the model-version boundary tag is DELIBERATELY NOT APPLIED: no scaling
change shipped, so no boundary exists, and stamping one would mark a model
transition that never happened. modelEras.js is its home when a real one lands.

Floor: 312 suites / 3890 tests green (8 new), web build exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 00:19:13 -04:00
builtbykev a6afac43cc report: market-efficiency scaling check — FLAT, mechanism structurally absent
Report-only. No threshold, grade, or efficiency value changed.

VERDICT: FLAT. marketEfficiency.js does not exist (zero occurrences of
market_efficiency / efficiency_score in src/ or web/src/). The spec's
0.85/0.60/0.55 values appear in grade_thresholds.json only as PROBABILITY
BANDS - a coincidental numeric overlap, not efficiency scores.

The base edge thresholds (MLB A:5%, NBA A:7%) do not exist either: engine1.js
has zero `edge` references, so the live grade is not an edge-vs-threshold
comparison at all. The specced rule threshold = base x efficiency has no host.

DISPOSITIVE: engine1.js contains ZERO `sport` references. computeFactors
receives no sport or market, so per-market OR per-sport scaling is structurally
impossible in the live grader - not merely unwired.

Phase 2: the matched-edge test is confounded (edge is not the grading input -
the same market emits both B and C at one edge). The aggregate that
discriminates: mean grade index wnba points 4.71 at mean edge 10.2 vs mlb hits
4.58 at 69.5 vs mlb total_bases 4.32 at 84.9 - the efficient market earns the
highest grades on one-eighth the edge, the opposite of spec.

PREMISE CORRECTION (measured): this order's opening claim that full-output and
collapsed grades "agree 100%" does not hold - on 512 rows carrying both they
agree 17.8%, with 33.8% differing by 3+ tiers. The prior discrimination result
stands (champion r=0.0050 null vs probability r=0.1313; MLB 0.0686 n.s. vs
0.2356 p~0.0004). Repo unchanged between orders. The collapse was not a
phantom and the re-adjudication list stays open.

Scope: flat thresholds are a grade-CALIBRATION gap only - the projection and
the CLV edge (which measured p_win, never the letter) are untouched, so this is
not a third shadow-model alarm. But the fix is NOT independently bounded: with
no threshold step to multiply, efficiency scaling presupposes probability
grading. It is rule R4 of specs/full-output-grade-mapping.md and belongs to
that MLB-first challenger.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-30 23:56:48 -04:00
builtbykev 708f0fde5c report: full-output grade mapping spec + collapse cost measured
Report-only. Nothing built, reconnected, or promoted.

Premise corrected again: the three-layer engine is BUILT but NOT WIRED and NOT
DEPLOYED (0 python refs in every grade-path file, 0 python in Dockerfile; there
is no engine1Adapter). So no posterior/CI/similarity prior exists to inventory
or diff. Measured against the collapse that actually exists instead.

THREE collapses, not one: (A) estimateProbability's components discarded at
analyzeViaEngine1:521-524; (B) THE SEVERE ONE - p_win never reaches the grade
at all (engine1.js has zero probability references), so the probability is
excluded from grading rather than collapsed into it; (C) grade_thresholds.json
(probability->grade) read backwards to manufacture confidence.
Market-efficiency scaling is never computed - a gap, not a collapse.

MEASURED on 354 settled rows carrying the served letter and the locked pre-game
p_win (forward, not lookahead). Grade->outcome point-biserial r: champion letter
0.0050 (p~0.93, null) vs probability letter 0.1313 (p~0.013). Per sport: MLB
champ 0.0686 n.s. vs prob 0.2356 (p~0.0004); WNBA champ -0.0986 vs prob -0.1258
- BOTH INVERSE. The served letter is inverted between its only two populated
tiers (B 52.4% n=168 vs C 56.9% n=174).

Verdict: costly on MLB, and un-collapsing does NOT help WNBA -> the challenger
must be MLB-FIRST. Five falsifiable mapping rules specced, incl. R2
(uncertainty grades down) stated explicitly and droppable if it fails.

Hard requirement on the next order: persist per-row n, SE and pre-adjustment p,
or R2/R4 can never be adjudicated (not stored today).

Re-adjudication list flagged incl. proj-v1.1's NOT PROVEN verdict (judged
against the collapsed champion, so not final) and ROI-by-grade (with B/C
inverted, the MLB-C +4.57% segment is likely an artifact of a meaningless
letter).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-30 23:47:11 -04:00
builtbykev dd98b0b614 report: model architecture recovery map — the live grade uses 0 of 3 specced layers
Archaeology only; nothing built, reconnected, or promoted.

The champion is two DISCONNECTED estimates: the letter is engine1's additive
factor index (zero references to p_win or any probability in engine1.js), and
p_win is probabilityEstimator's frequencyOver + 5 heuristic layers, computed
after and merely attached. The live grade path never calls the Python service.

The Python three-layer engine is NOT DEPLOYED — no python/pip in the
Dockerfile; app.js only health-checks it. So Layers 1-2 never shipped.
Layer 3 is wired BACKWARDS: grade_thresholds.json maps PROBABILITY->GRADE and
the live JS reads it in reverse to manufacture confidence from an
already-chosen letter. Per-sport market-efficiency scaling is specced-absent.

Consequence stated plainly: every metric audited to date is on the shadow
model, not the specced engine, which has never been measured.

Sport boundary TESTED not asserted: a new sport on the live path is a ~10-file
core edit with four documented silent-failure modes. Per-sport records DO
exist (sports.mlb n=526/62% vs pooled overall n=937/58%, each n>=20 gated),
but /api/accuracy ignores ?sport= and the pooled overall would absorb a new
sport. Park x weather confirmed challenger-only; xwOBA and leash absent.

Recovery map is dependency-ordered with MLB as the reference module.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-30 23:16:34 -04:00
builtbykev f3bf300b36 report: C1 takeable-floor derivation — CANNOT DERIVE (every bucket CI spans zero)
MLB decided overs n=296, 100% with locked_odds. ROI by locked-price bucket
shows every 95% CI containing zero; the curve is NON-MONOTONE and runs opposite
to the premise (deepest buckets positive, the -111..-160 middle most negative);
and price bucket is confounded with market (+200up = doubles/HR longshots).
Rows needed per bucket to resolve a 5-pt edge: 661-2285 vs actual 8-71 (~187
days for one bucket at current accrual). The inherited -160 is neither
confirmed nor refuted. The no-ceiling call is not supported by this data either
(+200up is the worst bucket) though it is not refuted - it stays a design
choice, not a data-backed one.

Recommends C2 proceed with -160 as an explicitly-labelled POLICY floor plus a
re-derivation trigger (any negative bucket n>=300, or end of MLB regular
season; adopt a derived floor only when a bucket CI excludes zero). Enumerates
all 9 takeable sites, incl. the live drift hazard (backend env-tunable,
frontend hardcoded) and the user-visible band copy in PriceTriplet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-30 16:05:09 -04:00
builtbykev f310608ca4 docs: top-graded selector fingerprinted — strip-after-rank boundary proven in prod
The anonymous live order (Brionna Jones edge 29.4 ahead of Rhyne Howard edge
42.9) is only explicable by server-side p_win ranking (.90 vs .745, both
takeable) while the payload carries no paid fields — the free caller got the
paid RANKING without the paid SIGNAL, on live data.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-29 22:08:15 -04:00
builtbykev 72a14dc4cd Build /api/props/top-graded server selector: rank with p_win, serve without it
New READ endpoint. No grade, ledger row, lock_line, or scoring write. Push
scoring untouched.

REVIEW ZERO CORRECTED THE PREMISE: the handler NEVER EXISTED in any commit
(searched git rev-list --all for a /top-graded definition in src/ — zero hits).
Not "removed" — the three axios callers (cheatsheetGenerator, gradeOfTheDay,
widget) and the Next proxy were written against a phantom endpoint, so those
three content generators have silently received [] for their entire life.
Contract recovered from the four consumers, not guessed: {props:[...]},
?sport=UPPERCASE (absent = all sports, which gradeOfTheDay relies on) + ?limit,
rows carrying player/stat/line/direction/sport/grade/confidence? plus the
player_name/stat_type aliases and game_id.

POPULATED-PATH RISK FOUND: the board's populated branch had never run in prod,
and dashboard/page.tsx:463 calls g.stat.replace(/_/g,' ') UNGUARDED (g.player
also feeds the row key, /scan URL and heading; sport must be UPPERCASE for
SportPill). toRow requires non-empty string player+stat and a finite line,
uppercases sport, and DROPS unrenderable rows — a shorter board beats a broken
one.

THE LEAK BOUNDARY (why this is server-side): the browser cannot rank on p_win
for all tiers because stripModelPrice deliberately withholds it from unentitled
tiers. Order of operations is
  read cache -> RANK with p_win (every tier) -> map rows incl. model fields
    -> stripModelPrice(rows, tier) -> serialize
so a free caller receives the paid RANKING without the paid VALUES. Tier comes
from resolveTierFromRequest, which FAILS CLOSED to 'free'. Cache-Control is
private under a bearer token, public otherwise (the /api/snapshot precedent).

ONE SHARED DEFINITION, no drift: new src/utils/gradeRanking.js
(takeablePWin/descNullsLast/rankGrades). heroPropService now imports
takeablePWin instead of its inline copy (behaviour unchanged — it was that
logic verbatim); the selector imports rankGrades; web/src/lib/slateAdapter
keeps its mirror (the browser cannot import src/, S25) and a test cross-checks
the two on identical fixtures (playerName.js precedent). Board is grade-first
("top GRADES"), hero is p_win-first ("top read") — they differ BY DESIGN and
agree within the leading tier.

HONEST LIMIT: the Next proxy (cachedBackendJson) sends no Authorization header
and caches under a shared key, so via the dashboard every viewer gets the
free-tier payload — correct order, no paid values. That is the SAFE behaviour;
forwarding auth into a shared cache is exactly how a paid payload leaks to
anonymous viewers. Per-tier delivery through the proxy needs a tier-keyed cache
and is not done here.

Verified on real prod snapshot data (anonymous path): MLB 8 props, WNBA 10,
0 paid-field leaks, render-contract safe on every row, sport uppercase.

Floor: 311 suites / 3882 tests green (18 new — leak test uses POPULATED p_win,
not today's nulls: entitled gets p_win and it drove the order, unentitled gets
a byte-identical order with all five MODEL_FIELDS absent and no trace in
JSON.stringify, while book/fair market facts survive). Web build exit 0.
Dashboard visual is auth-gated -> tagged for the Chrome audit, not faked.

Held: edge_pct rescale/retirement (Order B); board columns/contract unchanged;
tier-keyed proxy caching.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-29 22:05:30 -04:00
builtbykev 69feab4d25 docs: grade-board sort fingerprinted (deploy boundary captured; ladder induced both directions)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-29 21:18:25 -04:00
builtbykev b85b351993 Grade-board sort: signed signal, takeable-gated p_win, missing sorts LAST
Display ORDERING only. No grade, ledger row, lock_line, scoring, or edge_pct
scale/display change. Push scoring untouched.

Two defects removed from selectTopGrades (wrong at ANY scale, independent of
edge_pct's separate retirement):
  1. edge: Math.abs(numOr(g.edge, -Infinity)) — abs() on an already-
     direction-signed value ranked the model's strongest DISAGREEMENTS level
     with its strongest agreements (177 public ledger rows carry a negative
     edge; positive = the model AGREES with the graded side).
  2. Math.abs(-Infinity) === Infinity, so a row with NO edge sorted FIRST —
     absent data presented as the top pick (the Number(null) class).

New key: grade -> confidence -> takeable-gated p_win (nulls LAST) -> SIGNED
edge (nulls LAST) -> input order. Scales are never mixed in one comparator.
Takeable band = web valueState.isTakeable, asserted byte-equal to the hero's
config/valueEngine.isTakeable (-160..+200) incl. strict-null.

Alt-line ladder (analyzeViaEngine1:506) no longer sorts by edge_pct: ordered
highest-p_win-first derived analytically at zero added compute — P(stat >= k)
is monotone non-increasing in k, so p_win-desc is line-ASC for an over and
line-DESC for an under. base stays marked; no consumer depends on
alt_lines[0]; deskShowcaseService.rungsOf already re-sorted by line.

THREE PREMISE BREAKS found report-first, before code:
  - /api/props/top-graded 404s in prod (absent from src/) so the dashboard
    board renders receipts/empty — the edge sort orders nothing there today.
    The prior order's "97.3% of rows tie" was a LEDGER measurement wrongly
    extrapolated to that board. Fix is correct-in-itself and lands when the
    feed is restored.
  - p_win cannot be a client-side key for all tiers: snapshotGating strips it
    for unentitled tiers ("shipping p_win is shipping the model price").
    Verified live: prod /api/snapshot carries p_win on 0/8 MLB, 0/25 WNBA.
  - Ladder rungs carry no per-rung price, so the takeable gate is inapplicable.

Verified on real data, both sports, both paths: unentitled — WNBA (n=25)
ordering CHANGED, MLB (n=8) unchanged, signed edge non-increasing in every
(grade,confidence) tie group (20 pairs, 0 violations); entitled — 40 real
ledger rows with p_win+locked_odds, p_win-descending, untakeable chalk NOT
promoted (Trea Turner .757 @-275 does not beat Rhyne Howard .745 @-120)
(36 pairs, 0 violations).

Hero consistency, stated honestly: same signal + same gate, different
precedence BY CONTRACT (board = grade-tier-first "top GRADES"; hero =
p_win-first "top read"). Identical within the leading tier (verified); across
tiers the board may lead with an A the hero doesn't pick. Not a contradiction.

Floor: 310 suites / 3864 tests green, web build exit 0. Dashboard + Desk
visuals are auth/feed-gated -> tagged for the Chrome audit, no visual faked.

Held: edge_pct rescale/display retirement (Order B); building the missing
/api/props/top-graded selector; exposing p_win to unentitled tiers.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-29 21:13:51 -04:00
builtbykev 9b5235cf99 docs: hero ranking fix verified (new code serving — max-p_win takeable, chalk excluded)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 04:53:15 -04:00
builtbykev 41b86e3874 Hero ranking fix: rank by champion p_win among takeable, honest empty state
Review Zero found the hero's ACTUAL behavior was worse than "unknown": it ranks
on ev_pct (heroPropService v2), but ev_pct is NULL on served grades and
Number(null)===0 made Number.isFinite(Number(null)) TRUE — so every prop tied at
EV 0 and the "top read" was really the FIRST takeable A/B prop in cache order
(arbitrary, dressed as ranked).

v3: rank by the CHAMPION's p_win (the only signal with a promising, not proven,
edge — its takeable-MLB-over CLV survived the skew audit) among A/B, TAKEABLE-
priced reads (isTakeable band -160..+200, same as the proof/audit). Strict
null guard kills the Number(null)=0 bug. Takeable filter is mandatory (raw p_win
crowns -300 chalk). NO backfill: nothing qualifies → honest empty state
(available:false, reason:'no_qualifying_read'), never a weak recent read.

p_win is RANKING-ONLY, server-side — toHero never exposes it and the route strips
it. Framing unchanged in substance (model number vs book number, grade,
timestamp) — no proven-edge / +EV / best-bet claim, no CLV/ROI/edge number.
Display-only: reads snapshot caches, writes to nothing (no grade/ledger/lock_lines).
Full suite 3852 green, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 04:43:21 -04:00
builtbykev e7ec501054 docs: book-comparison verify note (live-card screenshot auth-blocked; data feed + tests verify)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 04:13:47 -04:00
builtbykev 36014c7c30 docs: Book Comparison wired to the prop card (matrix row 18 DONE, 16/26)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 04:10:58 -04:00
builtbykev 2ab2eeaa7d Wire BookComparison to the prop card (display-only, honest states)
BookComparison.tsx was built but UNROUTED (dead). Route it to the GradeResultCard
via a new self-fetching BookComparisonPanel that reads the live /api/books feed
(source:'bookprices' — the snapshot-locked, fenced, byte-identical store).

Contract fix (Review Zero 0.1): books frequently sit at DIFFERENT lines (WNBA DK
21.5 / FD 18.5; MLB 2/3), so BookComparison now renders EACH book's own line
per-row — never one shared header line implying a false same-number comparison.

Honest states: single-book (the common case for MLB) → one book, "One book
posting this prop.", NO crown/second row; multi-book → all books' own line+price,
NONE crowned (BOOK_CROWN_ENABLED=false — no best-price claim, verified live
crowned:false); no books → renders NULL (panel self-hides), never a placeholder.

No regression: only the always-empty inline d.books section was replaced; grade,
projection, PropLine line, and PriceTriplet price are untouched (wiring test
asserts them). Freshness (0.4): bookprices is written in the SAME snapshot that
locks the grade (intraday refresh touches neither gradedAt.line nor bookprices) —
same fresh, no stale-label needed. Web-only → grade byte-identical trivially.
Full suite 3851 green, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 04:09:34 -04:00
builtbykev 3a05447f77 docs: lock-line persistence unblocks the staleness audit (migration 033, lock_lines)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 02:49:02 -04:00
builtbykev c7067c80c4 Persist lock-time multi-book lines to lock_lines (unblocks the staleness audit)
The over-side skew audit's confirming check — was our locked line stale-high vs
consensus AT LOCK — was BLOCKED because multi-book lines at lock were never
persisted (bookprices is Redis current-only). This persists them.

- migration 033: lock_lines table (tracked + applied to prod). One row per
  (graded prop × book) with both odds + a lock timestamp. RLS enabled, NO
  policies -> service-role only (fence). UNIQUE key -> idempotent re-runs.
- lockLineCapture.js: buildLockRows (pure, graded-props only, honest-absent
  single-book) + idempotent upsert persist. Built from the in-memory props at
  the LOCK moment (ts) -> no Redis re-read, no TTL race.
- snapshotService: persist right after `enriched` (the lock moment; gradedAt
  uses the same ts). Best-effort + fenced.

FENCE (measurement-only): lock_lines is read by NOTHING on the grade path
(gradeSlateService, snapshot dedup/indexOdds, challengers, selector, ledger) —
a grep test asserts it, and RLS locks it to the service role. Grade byte-
identical proven: runSnapshot grades are identical with persist on/off (test).

Volume ~1.5-3k rows/day (graded props x books x 5 snapshots); weeks retained,
no pruning needed short-term. Does NOT retroactively fix the existing 62 rows —
future accrual only; confirmation still needs weeks of settled rows. Full suite
3842 green, web build exit 0. No grade/locked_odds/outcome/served surface changed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 02:47:47 -04:00
builtbykev 37261260d1 report: over-side skew audit — champion over-CLV SURVIVES baseline (promising, not confirmed)
Read-only. Takeable MLB overs n=62. Mechanical baseline (no-edge) CLV +1.51pt
(n=20); high-edge +8.64pt (n=37); difference +7.14pt = real edge. Champion
p_win->CLV partial r=0.375 (SIG p~0.003) survives price control. Skew one-sided
(unders -7, over baseline +1.5). De-vig clean (same-book pairing, analyzeViaEngine1:539);
close well-defined (DK/MGM r=0.92). Greenlights building the takeable-edge grade
ON THE CHAMPION, not proj-v1.1. Flagged promising-not-confirmed: thin n, lock-time
multi-book staleness check BLOCKED (not retained), pinnacle ref n=8. No fix, no
promotion — diagnosis only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 02:28:21 -04:00
builtbykev 59b77cdef2 report: proj-v1.1 takeable-edge proof — NOT PROVEN (n=45 MLB overs)
Read-only proof. N-gate passed (overlap 45). proj-v1.1 edge-CLV partial
correlation controlling for price = 0.245 (n.s.); ~half the raw 0.455 is the
shared -fair_prob_lock term (mechanical). Champion out-predicts proj on the
same rows (champ partial-CLV 0.380 sig; champ-edge->hit 0.25 vs 0.12). Unders
contaminated (CLV -9.3); WNBA proj-v1.1 doesn't run. Promotion HELD; the
under-audit is moot since proj loses to the champion first.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-29 00:47:56 -04:00
builtbykev b8ee216c62 docs: CLV instrument repair + the straight finding (unders lag the close)
closing_prob 59 -> 406 (MLB 248, WNBA 158). Root cause was attachClosingProb's
truncated read + write-once market_unavailable, not capture or the join. CLV
measured: MLB unders lag the close (mean -9.1 prob-pts, 74% lose), MLB overs
+2.0, WNBA flat -> the +4.57% MLB-C and over/under asymmetry are substantially
stale-line artifacts. Unblocks the proj-v1.1 proof order.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-28 19:04:50 -04:00
builtbykev 6552281661 CLV instrument repair: fix attachClosingProb read + recoverable market_unavailable
The closing_prob funnel collapsed 100k priced captures -> 59 usable. Root cause
(VERIFIED against prod, join key is PERFECT with 0 mismatches):
- attachClosingProb read closing_captures with .limit(50000) and NO ORDER BY on
  a 730k-row table that is 86% refusal rows -> saw ~7% for MLB, missed most
  priced closes and declared 200+ rows closeless that HAD a capture.
- market_unavailable_reason was write-once/terminal, so a row wrongly declared
  (truncated read / premature declaration before the capture was visible) could
  never recover even once its genuine capture existed. 298 rows (204 MLB + 94
  WNBA) were stuck this way.

Fix (CLV computation only — no grade/locked_odds/outcome touched):
- Read ONLY priced captures (missed_reason IS NULL, both odds NOT NULL), scoped
  to the candidate rows' game_dates -> small AND complete, no arbitrary truncation.
- Drop the market_unavailable exclusion from candidates; make it a re-checkable
  absence: a genuine close now UPGRADES the row (writes closing_prob, clears the
  verdict). closing_prob stays write-once (first true close wins). No capture +
  past game -> still declared absent (honest). No churn on already-absent rows.
- New internal trigger POST /api/internal/ledger/attach-closing[/:sport] for
  backfill + verification (scheduler already runs attach per tick).

Recovers ~312 usable closes (59 -> ~371), MLB included. Capture itself was
healthy all along (94.9% MLB / 95.8% WNBA per-prop coverage). Full suite 3835
green (17/17 instrument tests incl. 2 new recovery cases), web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-28 18:59:12 -04:00
builtbykev afb56b144b NexaPay purge: VYNDR is Stripe-only — remove all NexaPay traces
NexaPay was cross-project contamination (from another venture) — never a real
VYNDR payment path. Purged; Stripe path untouched.

Removed:
- web/src/services/nexapay.ts (createPaymentLink/getTransaction/HMAC verify)
- web/src/app/api/webhook/nexapay/route.ts (the only importer; Next-registered,
  reachable — now gone)
- NexaPay comments in email.ts + checkout/route.ts
- Active NexaPay entries in docs/SYSTEM-MANIFEST.md (route list, NEXAPAY_* env
  table, service row) + stale claim in wiring-data-train.md
- sw.js precache entry for the deleted webhook chunk

Verified: ZERO NexaPay in code (web/src, src, tests). Full suite 3833 green
(count unchanged — nothing depended on it, confirming it was dead). Web build
exit 0. sw.js parses clean. Stripe checkout untouched (Next→Express→Stripe).

FLAGGED FOR KEV (a repo delete cannot close these):
- Coolify env: remove NEXAPAY_API_KEY / NEXAPAY_WEBHOOK_SECRET / NEXAPAY_API_URL
- Revoke the NexaPay API key + webhook secret at NexaPay's dashboard; de-register
  the webhook if an account was ever configured
- DB column user_profiles.nexapay_customer_id is orphaned (no reader/writer) —
  drop via a follow-up migration (migration 011 left as history)

Cross-project check: ZERO Noctem-Supabase refs; VYNDR references only its own
Supabase (zmdnczhtdxcddsxzttub). NexaPay was the sole contamination found.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-27 16:51:59 -04:00
builtbykev 3592aba8d5 docs: log honesty pass in completion matrix + STATE (rows 19/20/24 HONEST fixed)
Records the six live fabrications removed/hidden, keeps media/newsletter/WIRE on
the board as real work, logs the news/line-movement signal as a future model
input, and logs the known honesty gaps (hit-rate-without-ROI, CLV starved).
Honest state: "no KNOWN live fabrications," not "provably none."

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-27 06:07:21 -04:00
builtbykev 6bc18d823c Honesty pass: remove every live fabrication (REMOVE/HIDE only, no feature cut)
Six live untruths corrected — no grade/snapshot/scorer/pipeline touched:
1. /compare — hardcoded Jokic A+/Wembanyama A + fake VERDICT replaced with an
   honest in-development state; removed from Nav + BottomTabBar (route still
   resolves, never the sample). Real two-player build is later.
2. Pricing — founder Desk $34.99→$44.99 (matches lib/checkout.js), Analyst
   $14.99; removed the struck $19.99/$44.99 "regular" numbers and DeskShowcase's
   stale $34.99. First-100 counter is real (ClaimMeter → Stripe countFounderSeats);
   no fake "first 50" desk claim added (no such counter exists).
3. FAQ "NexaPay" → Stripe (verified: live checkout is Next→Express→checkout.stripe.com).
4. FAQ + Features "Brier/CLV published from day one" removed (not surfaced yet) —
   returns when real. Backend Brier compute untouched.
5. MobileEdgeBoard removed from the Slate — its edge% feed was a miscalibrated
   placeholder (masked >40% as "—"); phones now show the real game cards.
6. Price triplet — never-computed model/EV now derives NO_MODEL (honest absent,
   MODEL "—" / "NOT PRICED", no verdict) instead of QUARANTINE's false "we
   suppressed our price / a leg is poisoned" copy. Fixes grade card + LiveHeroProp.

Full suite 3833 green, web build exit 0. Tests updated to the new honest contracts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-27 06:04:00 -04:00
builtbykev be6b4a2849 inventory: VYNDR-COMPLETION-MATRIX (5-part DONE map of every surface + model)
Read-only inventory order — no code built/wired/fixed/deployed. Maps every
user-facing surface and model component against DESIGNED·BUILT·WIRED·LIVE·HONEST,
re-derived from repo b0a51c8 + prod + design bundle (not STATE.md narrative).
15/26 surfaces fully done. Names the graveyard (BookComparison, ShareCard,
proj_ladder, arch-v1/contact-v1 ledger-only, /intelligence orphan) and the live
honesty gaps (/compare hardcoded, FAQ NexaPay/Brier, MobileEdgeBoard placeholder,
EV fields NULL on served grades). STATE.md now points to the matrix as canonical.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-27 05:12:59 -04:00
builtbykev b0a51c8a0d docs: /api/books contract + crown gate (Book Comparison order)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-27 00:49:10 -04:00
builtbykev e81c9b8c51 Book Comparison Phase 1-3(backend): fenced per-book store + honest gated crown
Per-book prices existed only transiently (odds cache, ~1h, raw names, grade-path
input); every grade-path persistence point collapses to one book. The
/api/books feature was built+mounted but non-functional (fed FLAT rows to a
GROUPED comparator -> always empty).

Phase 1: bookPriceStore captures per-book prices from `props` BEFORE dedupeProps,
keyed nameKey|stat, into bookprices:{sport} (SNAP_TTL) in snapshotService. Fenced:
reads props, writes its own key, read by nothing on the grade path. Grade proven
byte-identical (test + no-grade-path-reference grep test).

Phase 2: scripts/measure-book-spread.js reports same-line best-vs-worst spread
(cents + implied-prob pts), per sport, never pooled. Pre-registered crown
threshold: median >=8c OR >=2pp. Runs post-deploy on real data.

Phase 3 (backend): compareProp is honest-absent (single-book/flat -> no crown)
and the crown is gated (BOOK_CROWN_ENABLED, default OFF until Phase 2 clears).
/api/books repointed to the snapshot-locked store (fallback odds cache),
nameKey-matched; `source` field is the deploy fingerprint.

HELD unchanged: dedupeProps, snapshot dedup, selector, grade, champion,
challengers, ranking, edge_pct/ev_pct. UI routing of BookComparison + crown
treatment deferred to post-measurement (gated on Phase 2). Full suite 3834 green,
web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
2026-07-27 00:45:01 -04:00
builtbykev 914a057611 proj-v1 book-implied: raw odds → de-vigged FAIR (fix self-flattering basis)
proj_book_implied derived from raw book_odds — VIG-INCLUSIVE. A -110/-110 market
implies 52.4%/side (104.8% sum); fair is 50%. Comparing our P against raw book
overstates the book on both sides, biasing the handicapper test IN OUR FAVOR; on
juiced longshots (the Judge HR -18.5pt case) much of that "edge" was vig, not
disagreement.

Fix (fenced to proj-v1's stored comparison basis): proj_book_implied now derives
from DE-VIGGED FAIR via the grade's g.fair_prob — the SAME multiplicative de-vig
the triplet uses (utils/devig.js), so the basis matches the product's shown fair.
Expressed on the OVER basis (under props → 1 - fair) to match our stored P(≥rung);
traded-rung ladder book_implied likewise. HONEST-NULL where fair is uncomputable
(one-sided market, ~14%) — NEVER a raw-book fallback (that would recreate the vig
bias on a subset and mix two bases in one ledger). proj_factors records
book_implied_basis ('fair_multiplicative'|'none').

Phase 0 (prod-verified): fair reachable at store point (g.fair_prob on the grade,
no threading); 86% batting coverage; method = multiplicative/proportional.
Phase 2 FLAG: multiplicative de-vig mis-splits vig on juiced longshots (favorite-
longshot bias), so a longshot fair still carries known method bias — flagged
per-row (longshot_devig_caveat); a better de-vig (Shin/power) is a separate item.
Phase 3: version bumped proj-v1 → proj-v1.1 so pre-fix (raw-book) and post-fix
(fair) rows never silently mix — the projection model is byte-identical, only the
basis changed; pre-fix rows can't be recomputed (only the graded side's odds were
stored). Champion + arch-v1 + contact-v1 + proj-v1's other columns untouched.
proj suites 26/26.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-25 19:09:02 -04:00
builtbykev e96b0dbb6d proj-v1: book-implied from book_odds (grades carry odds, not fair_prob)
The live fingerprint showed proj_book_implied null on every real row: grades
carry book_odds/locked_odds (e.g. -264) but NOT a de-vigged fair_prob, so keying
the book comparison off fair_prob yielded null. The book ODDS are exactly "the
book's implied probability" the handicapper test needs. Now proj_book_implied +
the traded rung's book_implied derive from americanToImplied(book_odds),
expressed on the OVER basis (under props → 1 - implied) so it's directly
comparable to our P(≥rung). Vigged (a known offset the ledger measures both
sides of). proj-v1 suites 24/24.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-23 04:14:04 -04:00
builtbykev 316b79733e proj-v1 sanity fixes (caught in the real-data induction)
1. matchupRead fly-ball signal: the batter metrics `gb_pct_bb`/`fb_ld_pct` are
   MISLABELED — they're exit velocities by batted-ball type (Judge fb_ld_pct =
   100.3 mph, not a rate), not ground/fly RATES. Switched fly-ball lean to
   avg_launch_angle (league p10/p50/p90 = 7.1/13.9/20.1°), the correct signal.
2. Absolute rate now fits the FULL season (recency-weighted), not a 20-game
   window: the window under-sampled rare stats — Judge HR projected 0.11 vs his
   0.28 season rate (a fake -32pt edge). Now point=0.27 (matches season); the
   last-5-2x recency lean is preserved.

Post-fix induction (real statsapi logs + real statcast): Judge HR 0.27 (P>=1
0.235 vs book 0.42 -> flags the juiced over), Judge TB P>=2 0.548 vs 0.48
(+6.8pt), thin-hot 3-game P>=1 0.726 / P>=3 0.164 (credible low, thin high),
.300 hitter != 3.0. proj-v1 suites 23/23.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-23 04:05:23 -04:00
builtbykev 6386e737b9 proj-v1: absolute matchup projection challenger (distribution + full ladder)
A THIRD challenger (after arch-v1, contact-v1), MLB batting v1. Champion is
market-relative P(stat>LINE); proj-v1 is ABSOLUTE — what the hitter will DO —
emitted as a full distribution from which the WHOLE LADDER (P≥1,P≥2,P≥3) derives.
Champion untouched; nothing claimed; the ledger decides per rung, per stat.

- projection/distribution.js — Bayesian Gamma-Poisson → negative-binomial
  predictive. Admits over-dispersion; under-dispersion → Poisson approx
  (conservative, documented). Uncertainty scales with sample by construction
  (r=α): thin → WIDE (real mass on P≥1, honestly thin P≥3), thick → tight.
  NEVER abstains — width carries the honesty.
- projection/matchupRead.js — the input the book doesn't use. HONEST FIDELITY:
  pitcher repertoire is rich (97% pitch-mix) but hitters have NO pitch-type
  performance, so TRUE repertoire-vs-profile is impossible today. This is the
  COARSE version (arsenal buckets fastball/sinker/breaking + whiff/hard-hit
  tendency × hitter whiff/chase/gb-fb/hard-hit) — beats generic L/R, derived +
  documented + TESTED two-sided. A hitter pitch-type feed unlocks the true form.
- projectionChallenger.js — park RELATIVE to the player's own log exposure
  (isHome→own park, away→opp park; Phase B's raw-multiply bug solved), recency-
  weighted fit, per-factor breakdown (form/park/weather/platoon/matchup — show
  your work), full rung set + book-implied per rung. Combined non-form
  multiplier bounded.
- Wired after contact-v1, own try, flag PROJ_V1_ENABLED, reusing arch-v1's
  already-computed park/weather/platoon (no duplicate env I/O). Own ledger
  columns (migration 032, applied to prod): distribution, ladder, point, line,
  our-P, book-implied, factor breakdown — measurable per rung/stat after settle.

Phase 0 (prod-verified): venue join via isHome; NB family; uncertainty-as-width;
coarse matchup honest fidelity; no lineup-slot (per-game rate, volume implicit).
Sanity: thin-hot → wide (credible low rung, thin high rung); .300 hitter ≠ 3.0;
matchup two-sided; champion byte-identical. proj-v1 suites 23/23; snapshot/
ledger/siblings 74 green. Forward-only, version-stamped, PROJ_V1_ENABLED kill.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-23 03:59:36 -04:00
builtbykev b6f12daa98 Contact-quality challenger (contact-v1) — nominate, don't swap
Phase A #2: the champion grade (l5/l20 result-based form) is a HYPOTHESIS that
contact quality predicts better — unmeasured on our props, with zero settled
p_win yet. Swapping l5/l20 (the champion's two heaviest ±1.0 factors) blind
could degrade the core grade undetectably for weeks. So this NOMINATES contact
quality as a second challenger, records what it WOULD project per prop, and lets
the settled ledger decide. Nothing users see changes; the champion is untouched.

- src/services/contactChallenger.js — pure, mirrors challengerProjection. Log-
  odds lean (capped, never a re-forecast) from SEASON contact quality vs league
  percentiles. Metric→prop mapping is the whole game: barrel_pct→HR,
  hard_hit_pct→TB/doubles, k_pct-INVERSE→hits (singles resolve on contact
  frequency, not barrels), k_pct→batter K. rbi/runs/walks ABSTAIN (opportunity/
  discipline — no clean contact predictor). Honest-absent: thin (<50 PA)/absent/
  unmapped/non-batter → p_win_contact NULL (no projection), never a fallback;
  "measured but unremarkable" is distinct (equals champion, delta 0).
- Wired in snapshotService AFTER arch-v1, reusing the already-loaded statcast
  rows; its own try so a second challenger can't break the pipeline. Reads
  g.p_win, never writes it.
- Retained SEPARATELY on the ledger (p_win_contact/contact_delta/
  contact_adjustments/contact_version='contact-v1') so each challenger's marginal
  contribution is measured independently; ledger_entries.stat gives per-prop-type
  segmentation. Migration 031 (applied to prod).

Phase 0 (prod-verified): statcast_aggregates is SEASON cumulative (not rolling),
48h stale now but season-scoped so ~8 PA/600 is negligible; 100% of graded
hitters covered, 92% at ≥50 PA; no xBA/xwOBA in the feed. Forward-only,
version-stamped (contact_version null on pre-nomination rows). Promotion is a
LATER decision on settled evidence, per prop type — never asserted here.

contactChallenger 14/14; snapshot/ledger/arch-v1 suites 80 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 22:42:55 -04:00
builtbykev 125919f86a Scanner reskin: amber → blue boundary channel + build never-built S6/S7 states
Semantic COLOR fix, not cosmetic. The S6/S7 build predated the current Scanner
States spec: it used AMBER (the quarantine / model-suppressed channel) for the
"no market / line not priced" case, telling users "model suppressed" when the
truth is "the board never priced this." Corrected to the HANDOFF Session-3
blue-boundary law: BLUE (--priced-out #8FB2DE) = no-market boundary; amber stays
QUARANTINE; red stays REFUSAL.

Phase 1 (reskin): NoMarketState → dashed BLUE void box + blue header/copy; the
S7 rows → spec format (o 27.5 · BK −114 · ◆ −105 · OPEN READ ▸) at 44px,
390-legible. Input-area surfacer pills reskinned to the blue channel too.

Phase 2 (never-built states, only those Phase 0 confirmed against live data):
- GREEN CTA with LIVE player count ("PLAYER · N PRICED PROPS ▸"), degrading
  honestly to the board path ("N PROPS LIVE · TONIGHT'S BOARD ▸") at 0 — count
  from the SAME fresh index as the rows (pricedCountForPlayer), can't disagree.
- CASE A none-priced DEFAULT: "WE PRICE THESE FOR [player]" — the player's other
  priced stats (pricedStatsForPlayer, filter by nameKey).
- Typed-line-mismatch blue fact line ("o X ISN'T PRICED · NEAREST ↓").

FLAGGED / not built (no shells): Case C off-slate quiet-stop needs schedule/
roster membership the pricedLines index doesn't carry (out of the presentation
fence). Spec CONTRADICTION: Case B says "fair previews amber," but the law
reserves amber for quarantine — fair renders NEUTRAL ◆ (blue-dim), not amber, to
avoid blurring the channel.

Free-tier gate VERIFIED before rendering FAIR: fair_odds is the de-vigged MARKET
price (valueState: "never hide the honest fair number"), NOT the gated
model_odds — no paid leak. Carried through indexPricedLines (additive; keying/
refresh/onPick/stale-tap all unchanged — the proven S7 data path is untouched).
Scan A byte-identical; PRICED_NUDGE_ENABLED still the kill switch; reversible.
Build exit 0; priced + parity suites green (67).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 21:52:41 -04:00
builtbykev 83e9da3663 Consistency classifier: CV → index of dispersion for low-mean counts
The A/D investigation found CV (std/mean) is scale-broken on count data —
for a Poisson-ish stat cv ≈ 1/sqrt(mean), so EVERY stat with mean < 4 blew
past the boom_bust cutoff regardless of behavior. The S63 stopgap made those
return 'unknown', which silently ate a real +1.0 consistency signal on every
MLB batting prop — steady low-mean hitters never got their earned factor.

Fix, fenced to the low-mean branch of consistencyScore (the only branch that
was returning 'unknown'): classify with the index of dispersion (variance/mean,
Poisson baseline 1.0) — the scale-appropriate, UNBIASED statistic for counts.
mean ≥ 4 keeps the NBA-calibrated CV path BYTE-IDENTICAL (zero NBA blast
radius). This is a bug CORRECTION, not threshold loosening: the CV thresholds
and the engine1 ±1.0 delta are unchanged.

Bands (asymmetric around Poisson 1.0, since counts are naturally mildly
over-dispersed): iod<0.60 elite / <0.85 reliable (+1.0) / ≤1.30 volatile
(neutral) / >1.30 boom_bust (−1.0). Sample floor MIN_GAMES_FOR_IOD=8 so a
thin sample abstains ('unknown') — no small-sample guess.

Validated on real 10-game logs (two-sided): Kwan hits 0.67 / Alonso hits
0.78 → reliable (RECOVERED); Alonso TB 2.57 / Henderson hits 1.33 → boom_bust
(no false consistency); HR mean 0.1 → 1.0 → neutral. Direct engine1 proof: a
strong steady prop that grades B+ today reaches A- once the +1.0 fires; a
boom-bust bat stays B (no inflation). A- now emerges NATURALLY from a real
recovered factor. Standing two-sided test pins all three directions.

Forward-only (settled grades are locked in the ledger, never re-graded).
Emitting A- ≠ proving A- — the A-tier record accrues from emission, still
measurement-gated. Full unit suite green (4 pre-existing redis/timing flakes
pass in isolation); web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 21:05:16 -04:00
builtbykev 3440738d9e Landing hero → bundle visual system; claims filtered to earned only
Migrate the Landing hero (Landing-only component; LiveHeroProp/triplet
untouched) to the design bundle's visual language: mono "SPORTS
INTELLIGENCE TERMINAL" eyebrow, tightened 56px/800 headline "Every line,
graded before you bet it.", dual mono CTAs (OPEN THE TERMINAL / SEE THE
PUBLIC RECORD), proof chips.

Claims audited against the held list — the bundle's marketing copy carries
claims we have not earned; those did NOT ship:
- "CLV-VERIFIED" chip + "verified against closing lines" — HELD (C4 broken)
- "312 props / 47 games" count — fabricated demo → ABSENT (no count)
Shipped copy is earned only: pre-graded, edge-ranked, settled in public,
misses included; PUBLIC LEDGER / MISSES INCLUDED / 5 FREE READS chips.

HELD + FLAGGED for a design pass (fabrication-backed, no honest designed
state — check-don't-freelance): the edge-board demo (A+/A don't emit), the
"grades calibrate · A+ hit 80%" claim (calibration unmeasured), and the
tier-record band (real data is B/C-only, no ROI). Not built with fake data.

Tokens via var(--g-a/--void/--border/…) with bundle-hex fallbacks
(documented-intentional pattern). Build exit 0; parity QA green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 19:28:02 -04:00
builtbykev d879a10dd3 Design bundle (2): current source of truth — Scanner States + blue-boundary law
Reference material only, ZERO product code. Adds Vyndr Scanner States.dc.html
(S6/S7 spec) + HANDOFF Session 3 blue-boundary-channel law (#8FB2DE = the honesty
channel: priced-out / no-market / line-not-priced). This is the diff baseline for
the design-migration arc.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 19:10:54 -04:00
builtbykev 9a91837f63 STATE: Session 80 — S7 nudge freshness completed; premise corrected report-first
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 12:20:49 -04:00
builtbykev 4f3f433aae Complete S7 priced-line nudge: freshness on a long-open page + reversible gate
REPORT-FIRST correction. This order's premise — "S7 is a shell that doesn't
update per selection" — is not what the code does. pricedForSelection is a
useMemo on [pricedIndex, selectedPlayer, stat] and setSelectedPlayer/setStat
fire on every user pick, so the chips already update per selection, and the
prior session's verification of that stands. The genuine gap was FRESHNESS: the
snapshot fetch depended on [sport] only, so pricedIndex was fetched once per
sport-change and never refreshed. The pricing cron re-prices at five UTC hours,
so a scanner left open across a cron boundary surfaced hour-stale priced lines.
That is the real defect, and the only one fixed.

FRESHNESS. The fetch is now a refreshPriced callback re-run when the held
snapshot is older than PRICED_STALE_MS (30s, matching the /api/snapshot cache)
at the moment of use — on selection change and on window focus — so a long-open
page never shows a stale line. Sport change still clears the index first, so the
old sport's lines never flash.

STALE-TAP was already safe and is unchanged: the scan submit re-fetches the live
snapshot server-side, so a chip that's gone stale between render and tap either
lands on a real triplet (still priced) or degrades to the honest empty state
(rotated away) — proven in the prior session and re-confirmed here (an
off-snapshot line returns no market and shows the empty state).

REVERSIBLE GATE. The whole nudge sits behind one PRICED_NUDGE_ENABLED flag: false
empties the surfaced set, so the scanner falls back to S6's link-only empty
state with the chips gone. Shipping enabled only after the cases are proven this
session; the flag is the instant revert lever.

DISPLAY-LAYER ONLY. Only scan/page.tsx changed. GradeResultCard, PriceTriplet,
gradeAdapter, valueState, both scan routes and the pure pricedLines helper are
byte-identical — Scan A and the scan-submit resolution are untouched, and the
change is independently revertible.

Tests 3765 passed / 303 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 12:09:33 -04:00
builtbykev 33d72e38f9 STATE: Session 79 — read-card no-market empty state + priced-line surfacer, live
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 11:24:39 -04:00
builtbykev e311f53738 Read-card scan: honest no-market empty state + surface real priced lines
The Session-78 diagnosis stands: the join works, and a marketless scan rightly
shows no triplet. This makes that absence legible and points the user at what IS
priced, without fabricating a market.

PHASE 0 gate — design-check, reachability, timing, all clear. Design-check: the
bundle has the triplet's own REFUSAL language ("we'd rather show nothing than a
number we can't stand behind") as the honesty precedent, and a designed
EmptyState component whose actions give a path forward — so the empty state is
built in the established visual language, not freelanced. Reachability: the
scanner already fetches games/odds/search per selection; the snapshot is one
more public, 30s-cached fetch per sport, re-run when the sport changes.
Staleness: the snapshot rotates 5x/day and every scan re-validates the market
server-side at submit time, so a surfaced line that goes stale degrades to the
empty state on tap rather than a vanishing triplet — the stale-tap guard is
inherent, not bolted on.

Reversibility was the design constraint. The working card, price triplet, grade
adapter, valueState and the scan route are BYTE-IDENTICAL — a test asserts none
of them even reference the new empty state. Everything new lives in two added
files (lib/pricedLines.js, components/vyndr/NoMarketState.tsx) and additive
blocks in the scan page. Removing them leaves the Scan-A path untouched.

Non-fabricating by construction: indexPricedLines only keeps snapshot rows that
carry a real book price, keyed by exact player+stat via nameKey. A different
stat priced for the same player surfaces nothing for the picked stat; an
off-slate player surfaces nothing; nothing is suggested, interpolated, or
rounded to a nearest line. The empty state shows no market numbers of its own —
only real priced lines as one-tap chips, or a link to the live board when there
are none.

Framing is help, not restriction: a "PRICED TONIGHT" chip row sits under the
free-typed line input, and the scanner still accepts any player, stat and line.
Tapping a chip pre-fills the priced line and re-scans it — the market is
re-resolved server-side, so the tap either yields a real triplet or degrades to
the honest empty state.

Path forward, not a wall: a marketless scan no longer dead-ends in blank space.
It states truthfully that the board didn't price that line, keeps the grade, and
routes the user to the priced lines for that exact player+stat or to tonight's
board.

Tests 3760 passed / 303 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-22 11:21:29 -04:00
builtbykev 4c9707ffbb Platoon polarity: VERIFIED correct, pinned by a standing test
Report-only verification of the two unproven claims from the Session-77 wiring,
which fired platoon on synthetic split-less hitters where the K=600 regression
zeroed the multiplier and masked both direction and resolution. No product code
changed — this adds one standing regression test.

FIXTURE — Yordan Alvarez (LHB), real 2026 splits, deep both hands: 115 PA vs LHP
at .529 slg, 327 PA vs RHP at .695, overall .652. The regression leaves a
material multiplier both ways (0.97 vs LHP, 1.023 vs RHP), so unlike the last
test this fixture can actually reveal direction.

RESOLUTION — verified on the live slate that the pitcher-hand attached to a
hitter is the OPPOSING team's probable, not his own. CLE (home) resolved to
Minnesota's away starter 696070; MIN (away) resolved to Cleveland's home starter
800048. The chain — hitter's team, the game, the other team, that team's
probable, that pitcher's hand — is correct, and it is pinned independently of
direction because a backwards resolution is invisible on a neutral hitter.

DIRECTION — deterministic L-vs-R on the frozen Alvarez fixture. Facing RHP nudges
UP (1.023) because he slugs .695 there, above his .652 overall — a favorable
opposite-hand matchup, exactly what platoon theory predicts for a left-handed
bat. Facing LHP nudges DOWN (0.97). The two move opposite directions, and
crucially the SPECIFIC sides are asserted, not merely "opposite" — a
mirrored-but-inverted implementation would put RHP below 1 and fails here. Both-
backwards is ruled out.

The test also pins a REVERSE-split hitter, Brandon Nimmo, who hits better vs LHP
than RHP. His multiplier goes up vs LHP, following his real numbers rather than a
hardcoded LHB-vs-RHP assumption — proof the sign is data-driven, which is the
correct design.

VERDICT: PASS. Resolution correct, direction correctly signed against both the
real split and platoon theory. Pinned by tests/unit/platoonPolarity.test.js so
the polarity cannot silently regress — the opp_rank_stat lesson applied.

Tests 3750 passed / 302 suites.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 23:08:11 -04:00
builtbykev 535a7b70ea STATE: Session 77 — dormant adjusters wired; env non-null proven, ledger row pending
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 21:41:21 -04:00
builtbykev 7b25d97891 Wire the four dormant adjusters live — pure input-wiring
Verified state going in: parkBase, weatherMod and platoonSplits were called by
nothing, and env_multiplier was non-null on zero rows across four orders. The
adjusters were correct in isolation and starved of inputs. This gives them their
inputs and changes none of their internal logic — the five adjuster files are
byte-identical after this commit.

PHASE 0 GATE — all three inputs are available at snapshot build, and the two
join keys already existed. Venue: always, on every schedule game object.
First-pitch: always, gameTime on the same object. Opposing-pitcher hand:
present once the probable is declared, via the pitchers endpoint's pitcherId
joined to statsapi handedness — 15 of 15 games declared this afternoon, though
morning locks precede declaration and those props honest-absent on platoon,
correctly. The batter-handedness join (statcast bats) and the MLBAM id were
already on each grade from earlier sessions.

environmentContext.js is the wiring, kept separate from the adjusters so they
stay pure. It fetches once per snapshot: the schedule (team to venue, gameTime),
probable pitchers (team to opposing pitcher id), one batched handedness call,
one Open-Meteo forecast per home park, and batter splits per graded hitter. Park
coordinates for 30 parks live here as public geometry, the same class as the
dome list and centre-field bearings already in weatherMod, rather than inside an
adjuster. Everything is best-effort: a missing venue drops park and weather, an
undeclared pitcher drops platoon, and any fetch failure degrades that prop to
archetype-only rather than breaking the pipeline the adjusters are measured
inside.

attachChallenger becomes async and takes a per-grade contextFor that returns the
environment coefficient (park_base x weather_mod, composed) and the matchup
(platoon). Point-in-time holds: the weather is a forecast for first pitch fetched
now, and the split is the hitter's line entering the game — neither reads a
settle-time value.

Attribution is independent. env_multiplier, env_park_base, env_weather_mod and
env_weather_state land in their own ledger columns, and challenger_adjustments
keeps every axis — archetype, environment, matchup — as a separate entry, so
when volume accrues each of the four can be measured for its own marginal
contribution rather than as one blended delta.

The combined move stays bounded, tested on the worst case: a Coors slugger with
wind out and a favourable platoon, all at once, still moves under 12 percent,
because every layer is capped and the total nudge is clamped. Stacking leans, it
does not compound into a re-forecast.

Non-MLB honest-absents entirely — park, weather and platoon are MLB-only today,
so a WNBA prop gets no environment and no matchup.

The champion is untouched throughout: p_win is read, never written, the served
snapshot payload is still the enriched object, and a test confirms p_win passes
through byte-for-byte while the challenger moves.

Tests 3741 passed / 301 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 21:38:51 -04:00
builtbykev 2adf192d98 STATE: Session 76 — platoon regressed; plumbing gap now four orders old
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:58:04 -04:00
builtbykev 6dd6f59481 Layer 3 Step 6: platoon splits, regressed hard
The highest-value adjuster and the thinnest sample in baseball. The regression
is not a refinement here, it is the entire feature: applying raw splits would
adjust projections on noise, which is worse than not building it.

PHASE 0 — both gates clear, and one was already closed. Splits are a statsapi
pull, one call per hitter (statSplits with sitCodes vl,vr). The batter-handedness
join that Session 69 recorded as pending is in fact DONE: statcast_aggregates
carries bats for 604 of 604 batters, 210 left, 327 right, 67 switch. STATE said
pending; the data says otherwise, and the note is corrected. Point-in-time holds
as long as the split is fetched before first pitch, since a season split queried
this afternoon cannot contain tonight — but a historical backtest would use
season-final numbers and leak, so clean measurement is forward-accruing.

THE SPINE — regressed = (PA x observed + K x prior) / (PA + K), with K = 600 PA
and the prior being the hitter's OWN blended rate rather than the league's. The
question a platoon adjustment answers is whether he is DIFFERENT against this
hand than he normally is, so his own line is the correct null and a hitter with
no evidence of a split correctly gets nothing. K is deliberately conservative:
platoon skill is famously slow to stabilise, with the half-signal point for
right-handed batters near a thousand PA.

THE MAKE-OR-BREAK TEST, both halves. A .310 average against left-handed pitching
on 30 PA gets 4.8% weight and moves the projection by 0.003 — essentially
nothing, which is the correct answer rather than a limitation. The SAME .310 on
400 PA gets 40% weight and moves it by 0.023, eight times as far. A test asserts
that ratio stays above five, so if the regression ever breaks the suite says so
instead of the projections quietly drifting onto noise.

Real data behaves exactly as the mechanism predicts and is worth recording:
Josh Bell hits .259 against lefties and .248 against righties, which looks like
a platoon split until the sample speaks — 126 PA earns 17% weight and the
adjustment lands at 1.005. Aaron Judge, 76 PA against lefties, comes out at
0.999. Neither is material. Most hitters will get nothing from this adjuster,
and that is the honest output, not a failure.

Honest-absent has five distinct routes, all returning exactly 1.0: no batter
handedness, no pitcher handedness, no splits, a stat platoon says nothing about,
and a missing side falling back to the prior rather than to zero.

INDEPENDENT of the environment. Park and weather compose into one coefficient
because they both describe the stadium; platoon describes this hitter against
this pitcher's hand, so it rides its own slot with its own label. Entangling
them would make both harder to attribute when the instrument scores them.
Directional, mirrored on the under, capped at 15%, and inverted for strikeouts
where a higher rate means a higher prop rather than a better hitter.

Tests 3729 passed / 300 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:57:33 -04:00
builtbykev 9086951852 STATE: Session 75 — weather modulation composed onto park; venue gap still open
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:43:41 -04:00
builtbykev 9f60ceba10 Layer 3 Step 5: weather modulation composed onto the park base
Completes the coupled environment: effective = park_base x weather_mod. Weather
tilts the park, it never overrides it — a wind-out night at Oracle Park is still
Oracle Park.

PHASE 0 — both feeds are free and keyless. statsapi /venues gives every park's
coordinates in one call; Open-Meteo returns hourly temperature, wind speed and
wind direction for those coordinates hours before first pitch, which is when we
project. Verified live.

THE SPINE — two weather values, two purposes, never crossed. The FORECAST we
held at projection time drives the live adjustment AND is what the instrument
measures, because it is what we actually knew. It lands on the ledger row beside
p_win. The ACTUAL goes only to game_context as raw material for future
self-derived weather factors, and is read by nothing that scores a projection.
Using the actual to measure tonight would be scoring ourselves on information we
did not have. The actual is also pulled from Open-Meteo's ARCHIVE endpoint
rather than the forecast endpoint, because asking a forecaster after the fact
returns a re-forecast, not what happened.

WIND IS PARK-ORIENTATION CONDITIONED. Wind direction is meteorological — the
direction it comes FROM — so blowing out to centre means arriving from the
opposite bearing. Getting that backwards would invert every wind adjustment in
the system, so the 180-degree rotation is commented at the site and pinned by a
test on all three cases: straight out, straight in, and crosswind. Centre-field
bearings are public geometry, in the same class as the dome list; a park missing
from the table gets no wind effect at all rather than a guessed one, and keeps
its temperature effect.

THREE HONEST DO-NOTHING STATES, all multiplier 1.0, none fabricating an effect.
Dome: weather does not apply, and the PARK factor still does — verified that a
domed venue keeps its sub-1.0 park base while weather stands down. Forecast
absent: none available for this park and time. Sub-threshold: a real forecast
below a meaningful bar, because manufacturing a 0.3% nudge on a light breeze is
false precision. Weather also says nothing about a strikeout prop and returns
not-applicable rather than a neutral it might later be tempted to fill.

Conservative and ledger-tunable: every magnitude is an env var, the total is
capped at 12%, and nothing here is asserted. This is a nominated challenger that
earns its place on the instrument or is cut.

Induced at Wrigley, whose centre field bears 32 degrees: wind from 212 at 15 mph
computes as 15 mph straight out, weather 1.12 composed with park 1.06 for an
effective 1.187 and a +0.043 nudge; the under mirrors exactly; the pitcher's
home-runs-allowed prop moves with the hitter's, since both are P(over) on a ball
leaving the park. Wind in drops the coefficient to 0.955. A calm 72-degree
evening, a dome, and a missing forecast all return 1.0 by three different
honest routes, with the park base still applying in each.

One correction to the order worth recording: it describes a wind-out night as
helping the hitter and hurting "the pitcher there's HR-allowed" as opposite
sides. In prop terms both go the same way — the HR-allowed OVER is more likely
too. The sign lives in the stat, exactly as established for park factors, and
the implementation follows that rather than the phrasing.

Migration 036. Tests 3707 passed / 299 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:43:13 -04:00
builtbykev 5b4af67d93 STATE: Session 74 — public park base pluggable, game-context capture live
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:23:28 -04:00
builtbykev f33091ddb8 Layer 3 Step 4b: public park base, source-pluggable, plus game-level capture
PHASE 0 — the settle path sees a player's game-log line, not the game. It knows
date and teams, never venue or final totals. But the grain is far cheaper than
per-prop or even per-game: ONE statsapi schedule call per game DATE returns
every game that day with venue, linescore and scoring plays. Fifteen games, one
call, verified live.

PUBLIC BASE — the ingestion was already done. The static FanGraphs table from
Session 15 is the public base; this converts its 100-indexed values into the
multipliers the composable architecture wants (Coors 128 becomes 1.28) rather
than ingesting a second copy of a number we already hold.

It is labelled COMMODITY in the code, not just in a comment. Every resolution
carries a provenance record, and the public one reads proprietary: false with
the note "Commodity: a public number. Not a VYNDR derivation." The proprietary
label exists but belongs only to the self-derived version, and only once it
beats this base on the instrument. A surface rendering a park effect can state
which it is rather than implying the flattering one.

Honest-absent where even the PUBLIC number is thin: a relocated club in a
temporary venue gets no factor, because a public number for a park with one
season behind it is no more trustworthy than ours would be.

SOURCE-PLUGGABLE is the architectural point. resolveParkBase() is the only
accessor, public and derived return identical shapes, and callers never branch
on source — so when self-derived factors clear their floor they swap into the
same slot with nothing downstream to rewrite. A derived source with no factor
available returns absent rather than silently falling back to public, because a
silent fallback would make a proprietary claim out of a commodity number.

GAME-LEVEL CAPTURE starts now because it cannot start retroactively. Game grain,
deduped on game_id, never copied onto prop rows — a game's totals belong to the
game, and duplicating them per prop is how one fact starts disagreeing with
itself. Every field is tied to a named future derivation: venue for park
factors, runs for the run environment, HR totals for HR factors. Nothing else is
stored. Only Final games are captured, since an in-progress total is not a
result, and a game with no scoring plays reports HR as absent rather than zero.
HR totals come from scoring plays, which is complete because every home run
scores at least the batter.

The accrual target is stated rather than promised: 150 home games per venue at
roughly 81 per season means about two seasons before a self-derived factor can
be nominated, and accrualStatus() reports live progress per venue so the wait is
measurable.

Induced: Coors home runs +0.061 for the hitter and identically +0.061 for the
pitcher's home-runs-allowed at the same park, mirrored on the under; San
Francisco negative; Tampa flagged weather-N/A with its factor still applying;
the Athletics' temporary venue absent; strikeouts untouched. A real 2025-07-19
capture produced 15 games across 15 venues, 12 with HR totals, zero duplicate
game ids.

Migration 035. Induce with POST /api/internal/gamectx/:date, progress at
/gamectx/accrual. Tests 3688 passed / 298 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:20:29 -04:00
builtbykev 63e4bce858 STATE: Session 73 — park factors derived + composable; pipeline gap documented
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:06:49 -04:00
builtbykev 3ac91c3d96 Layer 3 Step 4: derived park factors, composable for weather
PHASE 0 GATE — the answer is BOTH, and the important half was already here.
A STATIC FanGraphs park-factor table has existed since Session 15
(src/data/parkFactors.js) and computeFeatures already consumes it, so park is
not a new idea in this codebase. What was missing is OUR derivation. I nearly
built a second source of truth before finding it; the new service lives at
src/services/parkFactors.js and the two are deliberately distinct.

That discovery changes the point of this order rather than just its scope. If
the champion already sees a park factor, adding one to the challenger risks
double-counting — which is exactly the redundancy the Session-72 harness exists
to catch. So park ships as a NOMINATED CHALLENGER whose job is to be tested for
marginal contribution, not as an assumed improvement. Checked and worth noting:
the static table reaches computeFeatures but NOT probabilityEstimator, so it
does not currently touch p_win at all.

DERIVATION, not ingestion. statsapi gives every game with venue, linescore and
scoringPlays in one call per date range — and since every home run scores at
least the batter, HR totals are fully recoverable from scoring plays. Derived
from 5,055 real games across 2022-2025: Coors tops the run environment at
1.099, Dodger Stadium tops home runs at 1.106, Oracle Park and PNC suppress
them at 0.923 and 0.917. Eighteen parks cleared the floor, eighteen did not and
are honestly absent.

COMPOSABLE BY CONSTRUCTION — the architectural point. Park emits a multiplier
around 1.0, never an additive nudge, because weather has to modulate it next
order: effective = park_base x weather_mod. Additive terms do not compose
correctly (a 5% park and an 8% wind are 1.05 x 1.08, not +13%), and the
challenger converts the multiplier to log-odds so stacking stays correct. A
test multiplies a placeholder weather term onto the park base to prove the shape
composes with no rearchitecting.

DIRECTIONAL BY PROP-OWNER: home_runs and home_runs_allowed both key off hr_base
in the same direction, because the sign lives in the STAT, not the park. Coors
inflates the hitter's home run prop and the pitcher's home-runs-allowed prop
identically.

THREE HONEST STATES, deliberately distinct. Absent (thin sample, adjust
nothing), present (adjust), and weather_na for domes — where the park factor
STILL APPLIES because a dome has a real run environment, and the flag exists so
next order's weather modulation correctly does nothing there. N/A is not absent;
conflating them would either drop a valid park factor or apply wind indoors.

Structural breaks: a season deviating past the threshold starts a new regime
only if the FOLLOWING season confirms it — one odd year is noise, two
consecutive years on the same side is a rebuilt park. Only post-break seasons
are used, so a humidor or moved wall cannot be diluted by the stadium that
preceded it. Factors regress toward neutral by sample size, so a two-season park
cannot assert a Coors-sized coefficient, and fine conditioning stays unavailable
until its own larger floor.

Tests 3669 passed / 297 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 01:06:17 -04:00
builtbykev beac1816d2 STATE: Session 72 — Tier-1 live, Tier-2 harness forward-accrual (no historical OOS)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 00:30:46 -04:00
builtbykev 927e867a23 Layer 3 Step 3: Tier-1 mappings live; Tier-2 nomination harness
PHASE 0 GATE — historical out-of-sample testing is NOT available, and the reason
matters. statcast_aggregates is overwritten nightly by design (Layer 1 is a full
re-pull upsert), so it holds season-TO-DATE numbers with no point-in-time
history. Classifying a player for a 15 July game using today's aggregate would
feed the model games from 15-21 July — look-ahead leakage, and the resulting
"out-of-sample" verdict would be worthless. The harness therefore reads the
archetype vector RETAINED at grade time (Session 70's instrument) and runs
FORWARD-ACCRUAL, not historical. Reported rather than worked around.

CANONICAL NAMES ASSERTED. Every mapping references the axis keys the classifier
actually emits, and a test walks both maps against BATTER_AXES / PITCHER_AXES.
A key that does not exist would look active and never fire — a mapping that
appears wired while silently doing nothing is the exact failure this guards.

TIER 1 IS LIVE, tautological and directional: PUNCHOUT/WHIFF raises strikeouts;
SINKER/SEAM lowers home runs allowed and FLY BALL/ELEVATOR raises them (a ball
on the ground cannot leave the park); SURGEON ARM/PINPOINT lowers walks allowed;
SLUGGER/BOMBER raises total bases and home runs; TECHNICIAN/SURGEON raises hits
and lowers strikeouts; GRINDER/SNIPER raises walks. Each adjusts only its named
stat, mirrors exactly on the under side, and leaves an average player untouched.

SPEED IS HONESTLY ABSENT. BURNER/stolen-bases has no axis to key on — SB is a
statsapi field that never reached the aggregate store, so Layer 2 shelved it.
The mapping is an empty object rather than an invented one.

THE TIER-2 HARNESS tests MARGINAL CONTRIBUTION, not correlation. A ground-ball
arm obviously correlates with fewer home runs; the question is whether the
archetype explains the PROJECTION'S RESIDUAL (outcome minus p_win). If the
projection already knows it, the residual carries no signal and the mapping is
rejected as redundant — that hurdle is what catches double-counting. The split
is by DATE, never random, because rows from one game share a pitcher, a park and
a lineup and would leak across a random split. Direction is validated from the
held-out data and a contradicted sign is REJECTED, never silently flipped to
whatever the data says, which would be fitting noise.

LIFECYCLE ENCODED — nominated, live, claimed. A mapping that survives runs live
and is measured; only the quantified public claim waits for the ledger. Nothing
sits dark.

One fixture bug worth recording: my first synthetic generator aliased the
carrier selector against the outcome draw and manufactured a 0.038 effect where
the generator had put zero. The harness rejected it correctly — it just gave the
sign reason instead of the redundancy reason, which is how I found it. The draw
now uses a coprime modulus.

Real candidate run end to end, GROUND-BALL to hits-allowed: INSUFFICIENT, 0 of
200 settled rows, because no settled row carries p_win yet (Session 70's
instrument starts recording at the next new lock). That is the correct verdict
and the expected one.

Tests 3654 passed / 296 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 00:30:16 -04:00
builtbykev 56fee267c9 STATE: Session 71 — challenger deployed; verification gap documented honestly
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-21 00:04:24 -04:00
builtbykev f2da9dd7e8 Layer 3 Step 2: archetype-aware CHALLENGER, measured not claimed
The champion (probabilityEstimator -> p_win) keeps serving and grading users,
completely unchanged. The challenger is a second probability computed from the
same inputs at the same instant, landing on the same ledger row so it joins to
the same outcome and the same close. Identical conditions, one difference —
the only clean A/B.

NOTHING IS CLAIMED. Running a challenger is honest beta; asserting it is better
before the settled ledger says so is not. Promotion stays a later decision gated
on Brier and calibration over sufficient segmented volume.

INTERPRETABLE, NOT A RE-ESTIMATION. The challenger is the champion's probability
adjusted by the Layer-2 axes, applied in log-odds space so a nudge cannot push
past 0 or 1 and means the same thing at p=0.5 as at p=0.9. Every deviation is
attributable to a named axis and a signed nudge, stored as
challenger_adjustments, and the total is capped at 0.45 log-odds — a lean on a
real signal, never a re-forecast. Only mechanically obvious stat/axis
relationships are mapped; a speculative mapping would be the same guessing this
layer exists to replace.

IDENTICAL WHERE THERE IS NO SIGNAL, by construction. An unremarkable player, a
thin sample, an unmapped stat or a missing classification all return the
champion's probability byte-for-byte with an empty adjustment list and a stated
reason. The experiment therefore differs only where archetype-awareness could
possibly help or hurt, with no dilution from rows the treatment never touched.

Induced on real players. Judge home runs over: 0.42 -> 0.447, via BOMBER +0.22
and WHIFF RISK -0.11 — two real opposing signals netting positive. The same prop
under mirrors it exactly to -0.027. Judge strikeouts: delta exactly 0, because
WHIFF RISK and GRINDER cancel — an honest "no lean" with both signals still
recorded. Skubal strikeouts over: 0.60 -> 0.702 via WHIFF, TRAPDOOR and CANNON
all aligned; his hits-allowed goes the other way, 0.50 -> 0.392, because a
strikeout arm makes hits less likely. Josh Bell and a 12-PA sample are
untouched.

Isolation is structural: adjust() is pure, the champion field is read and never
written, the served snapshot payload is still the untouched champion object, and
a challenger failure is caught so it can never break the pipeline it is measured
inside. Statcast aggregates load once per snapshot run rather than per prop, so
grade-time I/O stays at zero.

Migration 034. Tests 3634 passed / 295 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 23:53:32 -04:00
builtbykev 80f7100fc3 STATE: Session 70 — measurement instrument wired; the baseline accrues forward
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 23:12:27 -04:00
builtbykev 474ebc5d3a Fix the close-attach: de-vig raw prices, not a column that does not exist
Caught by inducing on real rows. The first attach ran and marked 642 rows
market-unavailable while attaching ZERO closes — because it selected a
`fair_prob` column from closing_captures, which has none. That table stores
over_odds and under_odds deliberately (Session 64) so the de-vig can run later
against the same engine the grade-time fair price uses; asking it for a
probability returns nothing and makes every row look closeless.

The de-vig now runs here, via devigTwoWay, which is what makes lock and close
comparable at all. A one-sided capture yields no fair probability and is
correctly not a close.

Repair checked rather than assumed: the 642 markings turn out to be CORRECT —
every one is a game from before closing capture existed on 2026-07-20, so those
rows genuinely have no close and the absence is true. Zero capture-era rows were
wrongly marked. The bug would have mis-marked every future row, which is what
the fix prevents.

Two tests added: the de-vig path with real prices, and a source assertion that
the query never again asks closing_captures for a column it does not have.

Tests 3616 passed / 294 suites.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 23:11:52 -04:00
builtbykev c5580f333e Layer 3 Step 1: wire the measurement instrument
Step 0 found we have been flying without one. p_win lives only in
model_snapshots, which has 1,000 rows and ZERO settled outcomes; the closing
line lives only in closing_captures, which carries no link to a result; and
ledger_entries, the row that actually settles, carries no probability at all.
So "is the projection calibrated" and "does it beat the market" have never been
answerable — the entire measurable universe was 35 rows recovered by a lossy
in-memory join.

PHASE 0 — closing coverage verified BEFORE reuse, because an instrument built
on a partial close measures a biased subset. closing_captures holds 70,254 rows
of which 13,364 are usable, and the 56,890 refusals are candidates we never
graded plus one-sided prices — not refusals of our props. Coverage on graded
props since capture started is 83/83, 100%. Safe to reuse, with the honest
caveat that capture only began 2026-07-20.

THE FOUR-TUPLE NOW LANDS ON ONE ROW. ledger_entries gains p_win, fair_prob_lock,
archetype_vector and projection_locked_at at LOCK time, and closing_prob plus
closing_captured_at from the append-only capture store. The join is the whole
point: calibration is p_win against outcome, market-comparison is p_win against
the close, and both become plain SQL on one record instead of a join that
silently drops 90% of the rows.

p_win and the archetype vector are IMMUTABLE — written once at lock via the
existing ignoreDuplicates upsert, never re-derived at settle. A re-derivation
would measure a projection we never made.

The archetype is stored as the VECTOR, not the label. "Did archetype-awareness
help?" can only be answered against the axes that were live at grade time, and
a single text column cannot express a blend. A grade with no archetype stores
null rather than an empty object.

HONEST-ABSENT BOTH WAYS. A past game with no usable capture is marked
market_unavailable_reason and never given an imputed line; calibration still
scores on those rows, only market-comparison is absent. And a game that has not
started yet is NOT declared closeless — a close can still arrive, and premature
absence is as dishonest as imputation in the other direction.

One bug caught before it shipped: the scheduler hook iterated a SPORTS
identifier that does not exist in that scope. Inside its try/catch it would have
thrown ReferenceError every tick and silently never run — the instrument would
have looked wired and captured nothing. Now iterates cadence.ALL_SPORTS.

The baseline accrues FORWARD. Historical p_win and closes are gone, discarded
before this existed. Calibration and market-comparison stay honest-absent until
volume accrues.

Tests 3614 passed / 294 suites, web build exit 0. Migration 033 applied.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 22:58:30 -04:00
builtbykev 063e9fb3f7 STATE: Session 69 — multi-axis archetypes live, FLEX fabrication deleted
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 22:40:26 -04:00
builtbykev 7ac6aa73e3 Layer 2: multi-axis archetype classifier; the FLEX fallback is gone
A player is a blend across independent axes, not one label. Skubal is a STARTER
and a strikeout arm and a ground-ball arm and a control arm — four true things
at once, and single-label classification threw three of them away.

AXIS INDEPENDENCE WAS MEASURED, NOT ASSUMED. Correlations over the live store
(467 batters, 531 pitchers); anything |r| >= 0.70 is one underlying trait and
was collapsed so we never show one trait as two archetypes. Batter k% ~ whiff%
+0.89, hard-hit% ~ exit velo +0.88, chase% ~ swing% +0.87, chase% ~ bb% -0.72;
pitcher k% ~ whiff% +0.76, gb% ~ fb% -0.73 — all collapsed.

The survivors are genuinely orthogonal, and one result is worth stating: pitcher
velocity correlates +0.14 with K%, +0.07 with whiff% and +0.07 with GB%.
Velocity is NOT a proxy for missing bats — a hard thrower who misses no bats is
a real distinct type, so CANNON earns its own axis rather than being folded into
STRIKEOUT. Pitcher K% ~ GB% is -0.10, so PUNCHOUT and SINKER are independent,
which is exactly the multi-axis thesis.

Cut-lines are the measured p75 (distinctive) and p90 (elite), per role where the
tails differ even when the medians agree: reliever GB% p90 is 54.1 against a
starter's 48.9, both with a median of 42.5.

THE FALLBACK IS DELETED. classify() used to return FLEX (mlb) / SHIELD (wnba) /
CONNECTOR (nba) at weight 1.0 when nothing scored — "could not classify"
rendered as a fully-confident classification of a real archetype, with
descriptive education copy attached. 8 of 18 MLB players carried it, and FLEX
could never be earned because its only scoring input had zero writers. Every
sport now does what MMA already did: unclassified is absent.

Induced on real players. Skubal: STARTER, throws L, WHIFF + SEAM + PINPOINT, all
elite. Judge: BOMBER + GRINDER + WHIFF RISK — elite power, patient, strikes out,
three true things. Kwan: SURGEON + SNIPER + SLASH with NO power claimed (0.4
barrel% is absent, not "low power"). Josh Bell, who used to classify as DRIVER:
empty blend, "No standout profile — league-average across every measured axis."
Alan Roden, who was FLEX at weight 1.0 on 21 PA: every axis absent, "Not enough
plate appearances yet — no profile claimed."

Per-axis honest-absence holds: a velo-less pitcher keeps every other axis, and
NO DATA is distinguishable from LEAGUE-AVERAGE rather than collapsing into one
shrug. The full vector is stored for Layer 3; only the top three distinctive
traits surface.

Three existing tests asserted the fallback and were updated to assert absence.
One of them surfaced a real robustness gap: classify(sport, null) threw, because
an explicit null does not trigger a default parameter and every scorer
dereferences its argument. Guarded.

Every baseball name is accounted for in docs/ARCHETYPE-AXES.md — built, alias,
tier, or shelved with its unlock condition. Zero orphans; cross-sport names left
for their sport.

Tests 3601 passed / 293 suites.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 22:38:32 -04:00
builtbykev 1265c23305 Tier-A joins: handedness, true role, and the velo fix
Three Layer-2 prerequisites, each one free call.

HANDEDNESS — statsapi /sports/1/players carries batSide, pitchHand AND
primaryPosition for every player: 1,316/1,316 in the live probe. Batter
handedness was 100% absent, which made every platoon or switch-flavour
archetype unbuildable; it is now populated from the same call that gives
pitchers theirs, with statsapi as the authority and the movement feed as the
fallback.

ROLE — statsapi season pitching with playerPool=ALL returns 751 rows (the
default returns only the ~57 qualified). Real usage: gamesStarted, gamesPitched,
gamesFinished, saves, holds. roleDetail derives starter/closer/setup/reliever
from that instead of the season-IP proxy, which drifts all year as innings
accumulate and left 32 pitchers in a 60-80 IP trough.

VELO — recovered from 53% to 99% (721/729 pitchers). The movement feed carries
only each pitcher's PRIMARY pitch, so matching by position could never do
better than one pitch each. The wide pitch-arsenals feed has one column per
pitch type, matched BY TYPE: Skubal now has velo on all 5 of his pitches. Velo
archetypes are therefore buildable rather than shelved.

Migration 032 applied.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 22:30:12 -04:00
builtbykev 27aabd078f STATE: Session 68 — Layer 1 mechanism data live (1,354 rows, 100% join)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 21:55:30 -04:00
builtbykev a49959867d Statcast: take the full arsenal, not each pitcher's primary pitch
Caught by spot-checking a real row after the backfill landed: Skubal stored
with one pitch. The pitch-movement endpoint with an empty pitch_type returns
ONE row per pitcher — their primary offering — so 677 rows for ~700 pitchers,
and a five-pitch arsenal was being recorded as a one-pitch one. Not a
fabrication, but a silent under-representation of the single most important
pitcher-mechanism field, which is worse than useless for Layer 2: it would have
classified every pitcher as a one-pitch arm.

Mix now comes from pitch-arsenal-stats (3,205 rows = pitcher x pitch type)
carrying usage%, whiff%, K%, put-away% and run value per 100 for every pitch.
Movement still supplies velo, break and handedness, folded onto the primary
pitch; a pitcher present only in the movement feed keeps his handedness and his
one measured pitch rather than being dropped. Velo on non-primary pitches is
null — absent, not guessed.

Skubal now stores 5 pitches, throws L, FF first by usage with velo 96.7.

Tests 3583 passed / 292 suites.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 21:53:32 -04:00
builtbykev a011ae79fe Statcast: role belongs in the key (two-way players)
Found by inducing the real job on the server, not by review: the first chunk
wrote, the second failed with 'ON CONFLICT DO UPDATE command cannot affect row
a second time'. A player can legitimately appear in BOTH the batter and the
pitcher feeds — two-way players, position players who pitch, pitchers who bat —
so (sport, season, source_id) collapsed two real profiles into one key and a
single batch hit the same row twice.

Ohtani has a real batter profile and a real pitcher profile. Merging them would
invent one player out of two genuinely different sets of measurements, so role
goes in the primary key rather than one profile winning. Migration 031 applied;
conflict target updated; a two-way case is now a test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 21:49:53 -04:00
builtbykev 528cb1a6d0 Layer 1: Statcast mechanism-data ingestion (backfill + nightly refresh)
The data foundation for the archetype and projection layers, built as the
pattern every sport inherits. Layers 2 and 3 are not touched.

PHASE 0 GATE — both match rates measured live, both 100%. Batters 40/40;
PITCHERS 66/66 across five real rosters (CLE, DET, MIN, NYY, LAD) joined by
MLBAM id against the 713-pitcher Savant feed. Zero honest-absent on identity,
because the join is an integer both systems use natively — and the snapshot
pipeline already stores it per graded row.

SOURCE — five Baseball Savant CSV leaderboards, free and public, pulled with
axios and the CSV parser savantAdapter already runs in prod. pybaseball is
deliberately NOT used: it is an MIT wrapper over these same URLs, and adding it
would reintroduce a Python runtime in a stack where the existing Python service
is already offline. min=1 on every feed, not Savant's default min=q, so the
long tail arrives and OUR minimum-sample gate decides what is thin — explicit
and testable rather than silently dropped upstream.

Measured: 1,354 rows per season (604 batters, 750 pitchers), all five feeds in
about five seconds. Pitcher mechanism includes arm angle, GB/FB/LD, chase and
whiff; batters get exit velo, launch angle, barrel and hard-hit, chase and
z-swing. Handedness rides in free on the movement feed (677 pitchers); batter
handedness stays absent pending a roster join rather than being guessed.

BACKFILL AND REFRESH ARE THE SAME CALL — a full re-pull upserted on
(sport, season, source_id). Idempotent and self-healing: a missed night
self-corrects on the next run, with no incremental who-played bookkeeping to
drift out of sync. At 1,354 rows the simple thing is also the robust one.

HONESTY RULES, each with a test: a metric the feed did not carry is null and
never 0; a thin sample is STORED and flagged rather than dropped or inflated,
because thin and missing are different claims; an unjoined player is stored
with a null player_key and joins later; and if every feed comes back empty the
job REFUSES to write, so a bad night can never blank a good table.

Freshness is treated as a truth property. updated_at on every row, and the
scheduler pages on a failed run AND on silent staleness — a job that stops
being scheduled never produces a failure, so staleness has to alarm on its own.
Never-built is deliberately not stale: different condition, different fix, and
paging on a fresh install teaches the operator to ignore the alarm.

Nightly at STATCAST_HOUR_UTC (default 11 UTC, after every game is final), kill
switch STATCAST=0, and induce-able at POST /api/internal/statcast/refresh with
a freshness probe at /statcast/status — we verify a refresh by running it, not
by waiting for the slot.

Migration 030 applied. Promoted columns for the classification-critical metrics
plus a metrics JSONB carrying every raw field, so Layer 2 can reach something we
did not promote without a re-ingest. Raw per-pitch stays out of Postgres on
purpose: one season is ~0.85 GB against a 500 MB plan ceiling, and it is
re-pullable from the free source if Layer 3 ever needs it.

Pattern documented in docs/MECHANISM-DATA.md for NBA tracking and NFL Next Gen.

Tests 3581 passed / 292 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 21:37:30 -04:00
builtbykev 0264486bf9 STATE: Session 67 — price layer gated at two doors, read card wired
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 19:59:33 -04:00
builtbykev 4e2f488341 Forward model_price_locked to the hero — a paywall was wearing poison's copy
Live induction on the deployed landing page caught this; markup review and the
unit suite both passed it. With the model leg stripped for anonymous visitors,
LiveHeroProp forwarded book/fair/model/ev/quarantine to PriceTriplet but NOT
model_price_locked, so deriveValueState fell through to the missing-model-price
branch and the card rendered "MODEL READ WITHHELD" — the quarantine state,
whose copy says we suppressed our own price because a leg is poisoned. Nothing
was poisoned. The real reason was the paywall, and the two must never share a
face: one says our data is untrustworthy, the other says you don't have this
tier.

Now forwarded, with a test asserting both the separation in deriveValueState
and the forwarding at the call site.

Tests 3559 passed / 291 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 19:56:21 -04:00
builtbykev aaa41134d4 Close the second leak path: /api/hero-prop bypassed the snapshot gate
The landing hero reads the snapshot from Redis DIRECTLY via heroPropService,
so it never passed through routes/snapshot.js and was still serving model_odds,
ev_pct, value and takeable to anonymous visitors after the first fix. Same
strip, same tier resolution, same private-cache rule for authenticated callers;
the Next proxy now forwards the bearer token.

PRODUCT CONSEQUENCE, FLAGGED RATHER THAN BURIED: the landing hero is served to
anonymous visitors, so it now renders BOOK and FAIR with the model leg LOCKED
instead of the full triplet it showed this morning. That follows the stated
free-tier rule exactly, but it does trade a strong marketing moment (VALUE
+21.1% VS FAIR on the shop window) for consistency of the gate. Reversing is
one line — add 'model_price' to the free tier in src/config/tiers.js, or
special-case the hero route — and is a product call, not a correctness one.

Tests 3557 passed / 291 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 19:52:00 -04:00
builtbykev fbcb00b7b1 Close the public model-price leak; wire the read card's price layer
PHASE 0.5 GATE — the three checks, and one correction.

`fairLine` does not exist. Zero hits across src/ and web/src. Option A as
written had no referent, but it resolves better than feared: `fair_odds` is
already a real de-vigged American price on every graded snapshot row, so there
is nothing to derive.

  Gate 1 (is it a price): PASS. fair_odds is American odds from
  impliedProbToAmerican inside devigTwoWay; fair_prob is the probability. Both
  distinct from `line`, the stat threshold.

  Gate 2 (numeric match): PASS, 8/8 exact. Recomputed fair_odds and fair_prob
  independently from the stored raw over/under prices; every value matched the
  stored one to the integer and to 3dp. Same de-vig, same numbers the component
  was proven against.

  Gate 3 (poison independence): PASS, and proven on the quarantined cohort
  itself. devigTwoWay's inputs are (over_odds, under_odds) — market prices
  only, no model term is reachable. The 8 rows recomputed above are all
  wrong_opponent_grade rows, and their fair prices reproduce exactly from the
  market. The poison is in the grade, not the price. Quarantine therefore
  suppresses the MODEL leg only; the fair leg stands, as designed.

THE LEAK WAS REAL AND ALREADY LIVE. GET /api/snapshot/:sport is public and
unauthenticated, and it was serving model_odds, p_win, ev_pct, value and
takeable to anonymous callers on every graded row — 25 of 25 on the live wnba
board. The Session-66 gate on /api/analyze was bypassed entirely by this
endpoint.

The strip covers more than model_odds, because model_odds is not the only way
to read the model price: p_win IS the price in another base, and ev_pct is
INVERTIBLE — ev is a function of p_win and book_odds, and book_odds is public,
so leaving ev behind hands the price over. All five model-derived fields go.
book_odds, fair_odds, fair_prob, overround and devig_method stay on every tier:
the fair leg is never the paywall. Rows that keep a book+fair pair are stamped
model_price_locked so a gated price is never mistaken for a missing one.

Tier comes from resolveTierFromRequest, which reads a bearer token when one is
present and otherwise returns 'free'. It FAILS CLOSED on every error path, so a
resolution failure can only ever withhold the price. The response now varies by
entitlement, so the /:sport handler downgrades Cache-Control to private for
authenticated callers and the browser proxy forwards the bearer token —
otherwise a CDN could hand a paid payload to an anonymous viewer, or every
request would look anonymous and paid users would lose the leg.

READ CARD — a manual scan carries no market. The request is {player, stat,
line, direction}, so the engine has no over/under prices to de-vig and
book_odds/fair_odds are legitimately absent from its response; that is why the
triplet was hidden there. lookupSnapshotPrices recovers them from the
pre-graded snapshot via the same cache-only read this route already performs
for locked odds and team. The join is exact on player + stat + line + side
(fair_odds is side-specific), and returns nothing unless book and fair are BOTH
present — a user-chosen line the board never graded has no market attached, so
the triplet stays hidden rather than borrowing another line's price.

FAIR-LEG ABSENCE, measured before shipping: 636 graded rows, 636 with book,
636 with fair, 0 one-sided. Absence rate 0.0%. The hero number is not a
sometimes-number on current data.

Tests 3556 passed / 291 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 19:47:42 -04:00
builtbykev 16697a4b90 STATE: Session 66 — price-layer tokens + the triplet, with induction proof
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 18:33:02 -04:00
builtbykev 8cf9cbfc26 Price-layer token foundation + the price triplet, built natively on it
PHASE 0 finding, reported before building: the token layer this order asked
me to establish ALREADY EXISTS and already matches HANDOFF.md exactly.
web/src/app/globals.css :root carries the design's surfaces, borders, text
ramp, fonts and grade colours byte-for-byte (aligned 2026-07-16), and
lib/colorContract.js already encodes the green-is-edge-only and
glow-is-A-tier-only laws with a test enforcing them. The stack is Tailwind v4
CSS-first (no config file) with components styled by inline style={{}} reading
var(--x) — 1,916 such reads — so CSS custom properties are the only vehicle
the stack natively consumes. Creating a second parallel layer would have meant
two competing sources of truth, so this EXTENDS the existing one.

PHASE 1 — additive only. globals.css gains one colour the system did not have,
the priced-out blue (#8fb2de + tints), plus a tokenized A-tier glow and the
fair-leg tints. The block writes the LAWS into the token layer itself — green
= takeable edge only, glow = A-tier only, amber = caution + the fair leg, red
= miss/negative only, blue = edge priced out, JetBrains Mono = all data — and
a test asserts every newly-declared name is new (zero collisions, zero
overrides). No existing hardcoded style was touched and no live surface was
migrated: the diff over existing files is 149 insertions, 0 deletions.

PHASE 2 — lib/valueState.js is the single verdict function; the component
renders what it returns and never re-derives one. VALUE fires only on
ev >= 2 AND a takeable price, mirroring src/config/valueEngine.js with a test
that cross-checks both files and fails on drift. Five states: VALUE (green),
EDGE-NOT-TAKEABLE (blue), NO EDGE (grey, stated at full voice), QUARANTINE
(model leg withheld, book+fair stand), REFUSAL (nothing rendered). Free tier
is gated at the wire — tierGating strips model_odds and sets
model_price_locked, so the lock is real rather than a blur over data already
sent; book and fair pass through on every tier because the fair leg is never
the paywall. Wired into the landing hero (data was already on /api/hero-prop)
and the read card, where the projection block reads first and the triplet sits
beside it, not in place of it. Ledger and public profile are out of scope —
no fair-odds columns exist there.

PHASE 3 — induced all six states in a real browser and read back computed
styles, not just markup. Green resolves on VALUE alone: rgb(0,212,160) on the
model leg and verdict; the +11.7%-EV-at-210 row renders rgb(143,178,222) blue
and a white model leg; quarantine shows MODEL "—" with book and fair intact;
refusal renders no legs at all; free tier renders a lock bar with book -120 and
fair -104 still honest. Landing hero on live data: book -153, fair -129, model
-343, VALUE +21.1% vs fair at +28.1% EV. At a real 390px column the three legs
hold at 117px each with no horizontal overflow and fair no smaller than its
neighbours.

Induction caught a real bug that markup review would not have: the "VS FAIR"
figure compared BOOK to fair, printing "VALUE · -6.5% VS FAIR" — a
contradiction on screen. The design's own two worked examples pin the formula
as MODEL minus FAIR in implied-probability percentage points; modelVsFair now
reproduces both exactly (+2.9 and -1.8) and a test locks them. The figure is
shown only when its sign agrees with the verdict, so a row that clears the EV
bar on the book price while our price sits level with fair leads with the EV
instead of a number that reads as a contradiction.

Tests 3539 passed / 290 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 18:22:00 -04:00
builtbykev f549422758 Design bundle: Price Triplet (ACT 01) + Offseason/Intelligence + 83 glyphs
Drops the updated claude.ai/design bundle into specs/design-reference/ as the
build reference. New since the Jul-18 export: Vyndr Price Triplet.dc.html (the
reality-corrected ACT 01 — five honesty states incl. EDGE-NOT-TAKEABLE, and the
projection-then-price read-card hierarchy), Vyndr Offseason.dc.html,
Vyndr Intelligence.dc.html, PNG export masters, and 83 archetype glyphs
(74 display + 9 classifier-legacy).

Reference material only. No product code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 17:53:23 -04:00
builtbykev 38ee83d0be STATE: Session 65 — CLV un-claim + FORM un-fabricate, with live fingerprint
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 17:09:01 -04:00
builtbykev f5156dd16d Un-claim CLV on the public profile; un-fabricate player-page FORM
Two Truth-Law fixes found by auditing the product logged-out.

FIX 1 — /u/[handle] claimed a "CLV-verified record" with "closing-line
value included" while ZERO closing-line value renders there. Verified
live: GET /api/profiles/vyndr returns beat_close_pct null (gated behind
CLV_CAPTURE_RELIABLE, unset while C4 is open). Eight instances found —
two of them (the OG + portrait "CLV-VERIFIED RECORD · 30D" eyebrows)
only by the post-removal residual sweep; two more printed the claim in
exactly the no-record branch.

Copy now describes what the page shows. The gated CLV-VERIFIED badge and
the BEAT CLOSE figure are removed from the public profile, OG card and
portrait card. DISPLAY ONLY: beat_close_pct, clvCaptureReliable() and
the whole CLV data path are untouched, and the earned directional badge
stays Analyst+Desk. The claim returns when CLV genuinely renders here.

Also fixes the doubled "· VYNDR · VYNDR" title (layout's '%s · VYNDR'
template already supplies the suffix); verified on composed output by
serving the build and reading the real HTML, not on source.

FIX 2 — the player page's FORM was `70 + 4 × (count of tonight's graded
props)`. Nothing on the HTTP path ever sets stats.form, so that fallback
WAS the live number: Josh Bell's "74" is 70 + 4×1 prop, confirmed
against his live payload. MATCHUP was gradeFromForm(that number), with a
hardcoded 'B' on the no-archetype branch — both fabricated letters with
no opponent input on the path. Systemic: buildIntel is the unconditional
path for every player and sport.

FORM and MATCHUP now render "—" (kind 'plain', so no bar width or colour
is computed off a null). gradeFromForm is deleted and the prop count is
no longer passed into buildIntel. computeFormScore's hardcoded 75 now
returns undefined. Induced across MLB/NBA/WNBA: all render cleanly, and
real values (USAGE 3.6 AB/G, REST B2B) still render.

Neither form value feeds the grade — engine1 reads raw l5_avg/l20_avg
against the line and never a form key; buildIntelFields decorates the
already-graded object. Grade inputs are byte-identical.

Held (needs a per-sport headline-stat design call): a real player-level
form metric + label disambiguation.

Tests 3491 passed / 289 suites, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
2026-07-20 17:05:04 -04:00
builtbykev ca9ca34cbb Remove legacy C4 CLV from FIVE surfaces; install directional badge on ledger
TRUTH FIX FIRST. row.clv/clv_result derive from the overwritable
closing_line — 637/699 rows had closing == locked — so of the 578 rendered
chips, 522 "flat"s encoded a CAPTURE FAILURE as a held line. That number
was live on FIVE surfaces, not the one the order assumed:

  1. ledger card            ClvChip@319            (authed)
  2. public profile /u/     duplicate chip@260     (PUBLIC, shareable)
  3. dashboard              "· CLV BEAT"@508       (authed)
  4. board / Slate          "CLV BEAT/FADED"@422   (FREE surface)
  5. board / Slate          "✓ HIT · CLV BEAT"@517 (FREE surface)

Removing it from the ledger alone would have left the lie live on three
surfaces including two public ones, defeating the stated PRIMARY GOAL, so
the removal covers all five. That is a deliberate extension beyond the
"scoped to ledger card" guardrail and is flagged as such — the guardrail
protected against feature creep, and this is the same defect at four more
addresses.

DIRECTIONAL BADGE installed in the old ledger slot. Three time states now
read distinctly on the row: entry (locked_odds, at-grade) · close (badge,
at-close) · outcome (hit/miss, final). The RESULT stays the row hero.
The PUBLIC profile deliberately gets NO badge — it never receives dclv
data (Analyst+Desk, server-gated), so that surface now shows the settled
result alone.

HONEST ABSENCE, tested: the 578 formerly-chipped rows now render NOTHING —
not a "flat", which is the old chip's lie in subtler form. Rows settled
before capture existed will never get dclv, and permanent silence is the
correct output. All six states verified in place; a 40-row dense ledger
stays legible (27 badges, max 40 chars, one line, consistent slot after
OutcomeChip); no CLV sort or filter exists.

FIELD REMOVAL DEFERRED as a separate scoped cleanup: /api/ledger/mine and
/api/profiles still SEND clv/clv_result, and dashboard/Slate still type
them. Display removal is local and safe; stripping the fields mid-swap
could break a response consumer.

Suite 289/3485 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 13:58:47 -04:00
builtbykev da8bfdf1db Render directional-CLV badge — Analyst+Desk, server-gated, receipt-bearing
CARD BADGE ONLY. Ticker-CLV explicitly DEFERRED (named, not lost).

PHASE 1 — SERVER-SIDE GATE AT THE DATA LAYER. A Free request never
RECEIVES dclv data: the CLV columns are appended to the SELECT only behind
canAccess(tier,'clv_badge') (new capability, analyst+desk), and responses
are ALSO stripped as defence in depth so a future SELECT change cannot
quietly leak. No CSS/client gate — data that reaches the browser has left
the building. dclv_fair_lock/fair_close are de-vig internals and are never
sent at all.

SURFACE AUDIT, all six channels, each test-locked to contain no CLV:
public profile (share link), snapshot/card feed, ticker feed, share
card/OG, embeddable widget, newsletter. A test also asserts no
ledger_entries read uses select('*') — a star would auto-leak every new
column, which is exactly how a gate becomes theatre.

PHASE 2 — IMMUTABLE ONCE COMPUTED. A settle can re-run (stat correction,
protested game) and a badge that flips positive->negative AFTER a user saw
or screenshotted it is a credibility failure. First computation wins: dclv
is only computed when dclv_computed_at is null, so a re-settle can never
rewrite a shown badge. Same discipline as the locked grade.

PHASE 3 — RENDER, test-first, ABSENCE IS HONEST. unknown / flat / null /
missing-receipt all render NOTHING — no element, no placeholder, no
"pending". Proven on an ALL-NULL board (today: 0 badges) and a MIXED board
(tomorrow: 1 of 4 badged, badge-less cards clean). Binary states only:
  positive -> MOVED TOWARD US  "graded -110 · closed -145"
  negative -> MOVED AWAY       "graded -110 · closed +120"
The RECEIPT is the persuasive part, so a badge with no numbers is
suppressed rather than shown as a bare claim. Negative is neutral context
and NEVER touches the locked grade — no back-door re-grading.

NO aggregate, count or rollup exists by construction: the module exports
exactly {clvBadge, fmtPrice} and a badge payload carries exactly
{tone,label,receipt} — asserted by test, because an on-screen tally would
be the held aggregate claim through the side door.

Build gotcha hit and fixed: clvBadge is CommonJS (allowJs) with no TS
types, so the .tsx needed an explicit cast at the call site — the build
worker exits 1 on type errors even though compilation "succeeds".

Suite 288/3473 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 13:37:52 -04:00
builtbykev dcdad60896 Directional CLV — per-read signal, side-bound, with a real compute trigger
PER-READ ONLY. No aggregate CLV stat, no CLV marketing un-held.

PHASE 0 FINDING THAT SHAPED THE BUILD: the LOCK end must come from
model_snapshots, NOT ledger_entries. The ledger stores only the graded
side's locked_odds (694 rows, single-side) which CANNOT be de-vigged.
model_snapshots retains BOTH side prices on 520/520 graded rows AND an
already-de-vigged fair_prob on 520/520 — produced by the same
devig.devigTwoWay the close uses, so "same method both ends" holds by
construction rather than by convention.

THE COMPUTE TRIGGER is the SETTLE PASS (ledgerService.settleLedger). At
settle the game is final, so the close has landed and the read is final —
the only moment both ends of the comparison exist. Grade and locked prices
are written hours earlier and the close at lock, so without this trigger a
correct CLV function would simply never populate.

JOIN INHERITS THE PROVEN KEY: (sport, player_key, stat, side, game_date),
WITHOUT line — a close that moved off the graded line is the entire point.
Verified clean earlier: 164 identity groups, zero ambiguity. Rows whose
capture refused (missed/ambiguous/one-sided) are UNKNOWN for CLV, matching
the capture layer's own honesty.

SIGN IS SIDE-BOUND and proven by test before the logic existed — the
badge-inverting trap. Same market move:
  OVER-graded  -> positive  clv +0.0800  (fair .500 -> .580)
  UNDER-graded -> negative  clv -0.0800  (fair .500 -> .420)
exact mirrors. FLAT is a PROBABILITY-space threshold always (1.5pp): a
40-cent price move on a deep favourite reads flat, correctly, because
price space lies about magnitude.

UNKNOWN is a first-class state, never 0 — zero asserts "the market did not
move", which is a claim; a missing close asserts nothing. describe()
returns null for unknown so a badge can never render for it.

migration 030 adds dclv/dclv_state/dclv_fair_lock/dclv_fair_close/
dclv_computed_at as NEW columns rather than reusing the C4 clv fields —
conflating a verified per-read signal with a known-broken one would be the
worst kind of quiet lie.

Caught pre-deploy: the trigger call passed `deps`, which is not in scope in
settleLedger (it uses `opts`) — a ReferenceError at the call site, outside
the helper's try/catch, which broke two settlement suites.

Suite 286/3447 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 13:17:34 -04:00
builtbykev b1ed675500 STATE: MLB opp_rank_stat live — consumption proven, coverage 8/14, factor inert this slate 2026-07-20 12:53:39 -04:00
builtbykev 63302d194e Wire MLB opp_rank_stat into live features (consumption path verified)
featureCache.teamFeatures now derives MLB opp_rank_stat from statsapi team
pitching splits when the ESPN path yields nothing — which for MLB is
always, because ESPN's MLB team endpoint carries no defensive metric at
all. mlbStatsAdapter.getTeamPitchingStats fetches all 30 teams in one free
unauthenticated call, cached at the season TTL.

CONSUMPTION PATH VERIFIED before wiring, not assumed:
  featureCache.teamFeatures sets out.opp_rank_stat (line 338)
    -> engine1.computeFactors READS features.opp_rank_stat (lines 96-102)
    -> fires weak_opponent_defense (>=0.70) / top_opponent_defense (<=0.30)
So teamFeatures is the correct insertion point: the grader reads exactly
the field we populate. A value written anywhere else would have been a
dead end — computed, retained, and still not affecting the grade.

Contract preserved: the derived value goes into the SAME field with the
SAME 0-1 scale and the SAME high=weak polarity WNBA uses, so engine1 reads
one field with one meaning across sports. Isolated and best-effort — a
derivation failure leaves the field ABSENT (honest null), never a guessed
rank. Only fills when the ESPN path produced nothing, so WNBA behaviour is
untouched.

Suite 285/3435 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 12:49:13 -04:00
builtbykev 55157b3288 Close-capture retry (lock-walled) + MLB opp_rank_stat derivation
PHASE 1 — CLOSE-CAPTURE RETRY, test-first. The closing capture gets a
retry the snapshot path deliberately does not: a snapshot re-runs at the
next slot, but a MISSED CLOSE IS PERMANENT, and the feed flaked once on a
dry induce. Three hard rules, each driven by a test written before the
logic:
  - BOUNDED attempts (default 3) with short backoff so every attempt fits
    inside the window. Never infinite.
  - HARD LOCK-WALL: inside lockWallMinutes of first pitch (or past it) it
    stops and records missed_close. A price captured AT or AFTER lock is
    NOT a close; storing one would fabricate the CLV baseline.
  - NO BOUND LOCK TIME -> refuse immediately, never burn retries on a prop
    whose close cannot be timed.
On exhaustion it records missed_close with NO price — never a stale,
mid-day or post-lock line.

PHASE 3 — MLB opp_rank_stat DERIVED, contract-locked. MLB previously had
no opponent metric at all (ESPN's MLB team endpoint carries none), so
engine1's +/-1.0 opponent factor never fired for the sport carrying most
of our volume. Derived from data we already ingest: statsapi team pitching
splits, all 30 teams in ONE free unauthenticated call.

THE SHARED CONTRACT is documented and TESTED, not assumed: 0-1 scale,
HIGH (>=0.70) = WEAK opponent, LOW (<=0.30) = TOUGH — identical to WNBA's
live semantics. Polarity is the highest-risk part: backwards polarity does
not fail loudly, it silently adjusts every MLB grade the wrong way. A test
asserts MLB polarity EQUALS WNBA polarity using engine1's own thresholds.

PROVEN AGAINST THE LIVE FEED:
  Colorado Rockies  BAA .286 -> opp_rank 0.983  (weak, fires weak_opponent)
  LA Dodgers        BAA .215 -> opp_rank 0.017  (tough, fires top_opponent)
  POLARITY HOLDS: true

HONEST NULLS, tested: thin league baseline, thin opponent sample, unmapped
stat, unknown opponent, or a missing field all return NULL with a reason —
we are FIXING a silent null, so it is never replaced by a confident guess
off three games. opponentStrengthHealth pages on an empty source AND on
derived-null-for-a-sport-we-expect-to-derive.

Suite 285/3435 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 12:27:47 -04:00
builtbykev 2321267346 STATE: harness armed + closing capture started; reverted my own non-close writes 2026-07-20 12:09:12 -04:00
builtbykev 77a58e4113 Arm harness on our scheduler + start closing-line capture (capture only)
PART A — HARNESS ARMED ON OUR OWN INFRA. snapshotScheduler now runs the
nightly backtest at HARNESS_HOUR_UTC (default 14), appends to
harness_results, and pages via opsWatch.harnessStaleAlarm — a validator
that stops running looks exactly like one that keeps passing. No external
dependency: the join is plain SQL through the service client and the
harness is a pure function. POST /api/internal/harness/run induces the
same code path on demand, because a scheduled mechanism is verified by
inducing it, never by waiting for a slot.

PART B PHASE 0 — GATE PASSED for what is capturable:
- C4 diagnosed: closing_line is ONE overwritable field with no timestamp
  and no provenance. captureClosing writes the current line and, when a
  prop fails to match, silently leaves the earlier value (= the lock) in
  place — so "captured a real close" is indistinguishable from "never
  updated". It is 92% equal, not 100%: 56 rows DID record movement, so
  the defect is provenance, not the value.
- Feeds: normalized props already carry BOTH raw side prices per book,
  with game_time, and the intraday refresh polls every ~20 min during
  slate hours — so the last observable pre-lock line is available.
- SHARP close: pinnacle is in ALLOWED_BOOKS -> a no-vig reference is
  capturable ("beat the market").
- ODAWA: NOT capturable. 'odawa' exists only as a UI preference option in
  onboarding/settings; it is in no adapter, no ALLOWED_BOOKS, no feed. An
  un-capturable source is a finding, not a gap to paper over.
- JOIN: must drop `line` from the natural key, because a close that MOVED
  off the graded line is the entire point of CLV. Verified safe — all 164
  current identity groups have exactly ONE line per
  (sport, player_key, stat, side, game_date). Zero ambiguity.

PART B PHASE 1 — CAPTURE ONLY, built test-first. The refusal was proven
before the capture logic existed: unbound game_time, doubleheader
ambiguity, a missed pre-lock window, or a one-sided price all record
missed_reason with NO price. A stale or mid-day line substituted for a
close would manufacture a CLV proof from a number that was never the
close.

migration 029 closing_captures: append-only, never overwritten (that is
the provenance C4 lacked), BOTH raw side prices so the existing de-vig
engine can compute a fair closing probability later, sharp vs book line
types kept distinct. Wired into the intraday refresh with a capture-rate
alarm — a missed close is unrecoverable.

NO CLV metric built, as ordered. This starts the clock.

Suite 284/3417 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 12:05:16 -04:00
builtbykev e809a0eb3c Backtest harness — the validator, built refusal-first
Phase 0 gate PASSED: the join is clean. No FK exists; the natural key
(sport, player_key, stat, line, side, game_date) yields 283 clean 1:1
joins with ZERO ambiguity. game_id is NOT usable — 400/550 snapshot rows
carry UNK@UNK because home/away names weren't threaded into the grader
until Order 1.6. Non-joining rows are EXPECTED, not errors: retention
stores both sides plus refusals; the ledger keeps only the graded side.
Outcomes are NOT denormalized — ledger_entries stays the source of truth.

BUILT TEST-FIRST, and the first property proven is the REFUSAL, not the
math. Below threshold the harness emits INSUFFICIENT with n and the
shortfall and NO rate anywhere in the payload, so a downstream renderer
cannot surface one by accident. A test asserts the payload contains no
hit_rate number at all.

- Wilson intervals (correct at the n we actually have, unlike the normal
  approximation which emits negative lower bounds).
- Strata NEVER mix sport or model_version.
- Denominator excludes quarantined, void, unrecoverable, pending, push —
  asserted by test.
- Monotonicity refuses to RANK buckets whose intervals overlap; it reports
  "not distinguishable on this sample".
- Probability calibration (Brier + reliability) also respects the
  threshold: a thin sample returns status INSUFFICIENT and a NULL score.
- Replay seam reads the STORED feature vector only. A row whose input was
  never retained is UN-BACKTESTABLE, never scored with substituted current
  data. Identity replay reproduces the live prediction exactly.

The tests caught a real bug in my own code: `Number(null) === 0` let a
null p_win through as a confident 0% forecast — this codebase's signature
fabrication bug, inside the harness whose entire purpose is refusing
invented numbers. Fixed with a strict null guard.

FIRST LIVE RUN — the correct, passing output:
  VERDICT: INSUFFICIENT_HISTORY (can_validate=false)
  283 joined -> 35 scored (120 quarantined, 124 pending, 4 terminal)
  C n=18 (short by 2), B n=17 (short by 3)
  strata: mlb 7, wnba 28 — never mixed

migration 028 adds harness_results (append-only trend log; INSUFFICIENT
rows are expected and correct) and opsWatch.harnessStaleAlarm pages if the
harness stops running — a validator that isn't running looks exactly like
one that keeps passing.

Suite 283/3403 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 11:31:48 -04:00
builtbykev f73fb64a43 STATE: heal executed — 64 re-voided, 25+200 quarantined, scopes separated 2026-07-20 05:45:27 -04:00
builtbykev b33612675d Heal execute: quarantine markers, re-enable DNP voiding, two exclusion scopes
Order 2 Phases 2 + 4. Pre-heal rollback point secured first:
vyndr-20260720-093821.dump (856,890 bytes) VERIFIED ON THE BOX, not just
exit 0.

MIGRATION 027 — two DISTINCT exclusion scopes, deliberately separate:
- quarantine_reason: the row's GRADE is untrustworthy (wrong_opponent_grade).
  The row REMAINS a real public settled result — the bet happened, the
  outcome is real — but it must never train or validate, so
  getModelAggregate now excludes it from the denominator alongside
  void/unrecoverable.
- analysis_flags: the row is VALID for settlement and the record but
  unattributable for PER-GAME analysis (doubleheader dates). Explicitly NOT
  filtered from aggregates.
Collapsing these would either wrongly drop 166 doubleheader rows from the
record or wrongly keep 25 wrong-opponent grades inside model validation.
Tests assert both directions, including that analysis_flags is NOT filtered.
Also adds re_settled_at + settlement_source to model_snapshots.

DNP VOIDING RE-ENABLED — reversing my own Order 1.5 disable, with scrutiny,
because its premise was FALSE. Order 1.5 assumed a missing player row meant
the row's DATE was wrong. The Phase 0 dry-run disproved it: across every
bindable row the stored date matched a real game (MIS-DATED: 0), and the
players I had cited as counter-evidence were genuine DNPs on their true
dates (Freeman 07-18; Kwan/Hedges/Davis 07-17 — their teams played, they
did not). The evidence is positive: games FINAL + no line in a full-season
log = no bet existed.

I got this wrong twice tonight in opposite directions; the dry-run is what
caught it. Recording the reasoning in the code so the next reader sees why
the flag flipped back.

Suite 282/3386 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 05:41:41 -04:00
builtbykev b76e35f575 STATE: grading date binding fixed + blast radius quantified 2026-07-20 03:52:34 -04:00
builtbykev 6415751f2e Grading binds opponent features to the REAL game, not ESPN's "today"
Order 1.6 Phase 1. This is a MODEL-OUTPUT fix, not bookkeeping.

computeFeatures.lookupTodayGame called the ESPN scoreboard with NO date
param and took whatever ESPN calls "today". Renamed to lookupGameOnDate
and now sends ?dates=YYYYMMDD from the prop's BOUND game — the same game
the ledger, retention and settlement use, so all four finally agree.

PROVEN against live ESPN (before/after, same instant):
  dateless "today"      CLE->PIT  NYY->LAD  LAD->NYY   (Jul 19 card)
  bound to 2026-07-20   CLE->MIN  NYY->PIT  LAD->PHI   (the real games)
  bound to 2026-07-19   CLE->PIT  NYY->LAD  LAD->NYY   (reproduces OLD)
Every opponent was wrong. opponentAbbr feeds opp_rank_stat (a +/-1.0
factor) and isHome feeds home_away (+0.5), so late-slot grades were
scored against the wrong matchup.

Note the window is WIDER than the 01:00/03:00 UTC slots: this ran at
07:5x UTC = 03:5x ET and ESPN's dateless scoreboard was STILL returning
the previous day's card.

HONEST DEGRADATION: with no bound game date the grader does NOT fall back
to a dateless lookup — it records 'no_bound_game_date' and leaves
opponentAbbr/isHome/gameId null, so engine1 simply omits the opponent and
home/away factors rather than scoring a wrong matchup. Tests assert both
directions.

Same class of bug fixed alongside: the Tank01 augmentation used TODAY's
UTC date for its cache key; it now uses the bound game date.

gradeSlateService threads game_date/game_time/home_team/away_team into the
grader so the binding reaches computeFeatures at all.

Audited the rest of the feature path for dateless/"today" lookups — none
remain (weather is current-conditions by venue, park/pace are static).

Suite 282/3383 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 03:50:07 -04:00
builtbykev 8e17de4010 STATE: game_date root fixed, 64 voids reverted, grading blast radius flagged 2026-07-20 03:27:42 -04:00
builtbykev 1bcdd8b305 Fix the ROOT: bind props to their real game, never date by the grade clock
Order 1.5 Phase 1. PropLine emits NO commence_time (grep-verified: zero
hits in proplineAdapter), so ledgerService's
`dateET(prop.game_time) || dateET(gradedTs)` always fell through to the
GRADE timestamp — and a 01:00/03:00 UTC snapshot is 21:00/23:00 ET the
PREVIOUS day. Tonight's props were filed under yesterday, settlement
correctly found no game there, and Order 1's void logic turned that into
64 destroyed results.

gameBinder.attachGameTimes() now matches every prop to a scheduled game by
TEAMS across the plausible ET window (grade date, +1, -1) and attaches the
GAME'S OWN time/date/id. It runs in snapshotService before grading and
before the ledger write, so ledger, retention and settlement all inherit
the correct date from one place.

HARD CONTRACT: an unbindable prop returns NOTHING. ledgerService no longer
has a grade-clock fallback — a row with no real game time is SKIPPED and
counted, because a mis-dated row is fabricated data and the ledger holds
real values or nothing. A slate that binds nothing pages.

DOUBLEHEADERS are reported, never guessed: two games with the same teams
on one date mark the binding `ambiguous` so settlement can decline rather
than attribute a prop to the wrong game. (Real example already in the
data: mlb:2026-07-11:MilwaukeeBrewers@PittsburghPirates(Game1).)

Also fixes retention, which had the SAME bug from last night — I had
dated model_snapshots rows with the snapshot clock. Rows now take the ET
date of the bound game_time.

Suite 282/3381 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 03:25:01 -04:00
builtbykev 270c4db47a STOP voiding on player-absence — it destroyed real results
Correctness fix to code I shipped minutes ago. The induced live settle
pass voided 64 rows as 'player_dnp' and a large share of them are WRONG:
the Jul 18 set is everyday starters (Freeman, Bellinger, Tucker, Chisholm,
Conforto). They played.

ROOT CAUSE — and my Phase 0 diagnosis was wrong. It is not DNP. The
ledger row's game_date is WRONG. ledgerService derives game_date from the
GRADE timestamp when the feed carries no game_time, and a 01:00/03:00 UTC
snapshot is 21:00/23:00 ET the PREVIOUS day, so rows get labelled with the
previous ET date. Verified against fresh season logs (cache disabled, so
not staleness; found:true, so not name resolution):
  Freddie Freeman  played Jul 17 and Jul 19 (x2, doubleheader) — NOT Jul 18
  Steven Kwan      played Jul 18 (x2) and Jul 19               — NOT Jul 17
Settlement was correct to find no game on the labelled date. My void logic
then converted a data-labelling bug into destroyed results.

FIX: never void on player-absence alone. Voiding now requires POSITIVE
evidence — the games themselves postponed/cancelled. Absence returns
'unknown' (reason player_absent_unconfirmed), so the row retries and ages
out to 'unrecoverable' at the cap. We cannot distinguish "did not play"
from "mislabelled date", so we must not claim DNP. Both terminal states
are excluded from the record denominator either way.

Window-decay remains genuinely fixed (full season log vs a rolling
window), and terminal states still prevent immortal rows.

NOT DONE HERE: the 64 wrong voids are still in the table, and the
game_date derivation is still wrong at the source. Both are reported for
the table — no healing in this order.

Suite green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 03:10:34 -04:00
builtbykev d4a6170ffa Settlement fix forward: date-targeted resolution + terminal states
Order 1 of 2. Push scoring UNTOUCHED — it is correct. No healing here.

PHASE 1 — DATE-TARGETED FETCH replaces the rolling window for settlement.
settleSource.resolveOutcome() resolves the SPECIFIC DATE and, when the
player is absent, reads GAME STATE to learn what the absence MEANS:
  game final + player has a line  -> SETTLE (a partial game is a real
                                     result, never a void)
  game final + player absent      -> VOID (confirmed DNP)
  postponed / cancelled           -> VOID
  scheduled / in progress / SUSPENDED -> PENDING (a suspended game resumes;
                                     voiding it would destroy a real bet)
  player played, stat missing     -> unknown, NEVER void a real appearance
This is FREE for MLB: mlbStatsAdapter.getPlayerGameLog already returned the
full season log and getPlayerStats was discarding it with .slice(-10).
Settlement now reads fullLog — same request, same cache — which removes
window-decay entirely (the verified failure was a Jul 12 game outside a
last10 starting Jul 6). Projections keep using last10, unchanged.

PHASE 2 — TERMINAL STATES (migration 026 applied). outcome CHECK widened to
hit/miss/push/void/unrecoverable; added settle_attempts, settlement_source,
settlement_version, model_version. A row that cannot be resolved after
SETTLE_ATTEMPT_CAP (4) date-targeted attempts becomes 'unrecoverable'
rather than pending forever. CRITICAL: getModelAggregate now EXCLUDES void
and unrecoverable from the settled selection — it used
.not('outcome','is',null), so without this a void would have counted as a
settled row and silently moved the public record. Verified in the record
calc, not just the settle path.

PHASE 3 — SETTLEMENT-RATE ALARM. zeroSettleAlarm only caught a TOTAL zero
while ~30% of a slate failed quietly (Jul 17: 57/86). opsWatch
.settlementRateAlarm pages when resolved/attempted falls below
SETTLE_RATE_FLOOR (0.8). Voids count as RESOLVED — a void is a legitimate
terminal state — so healthy voiding never pages. Third silent-failure
surface of the night, now closed.

PHASE 4 — VERSION STAMPING. src/config/modelEras.js defines the cutoff
ONCE (2026-07-19T22:50:00Z); migration 026 backfilled pre-cutoff rows as
'pre-retention-unknown' (naming the uncertainty, not implying knowledge);
new rows carry model_version.

Regression caught pre-deploy: getScheduleFn was not injectable, so the
ledger suite hit the real network and HUNG. Now injectable via opts and a
no-op under NODE_ENV=test. The "no row -> pending" test was updated to the
new behaviour deliberately: a missing row on a FINAL game now voids.

Suite 281/3373 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 03:04:55 -04:00
builtbykev e46c88e364 STATE: retention clock proven ticking via induced cron entrypoint 2026-07-20 02:06:43 -04:00
builtbykev 5a5e37e32e Retention: fill enrichment fields + page on a zero-write slot
PHASE 1 — cron capture needed NO wiring. Verified in code: the scheduler
tick calls runAll = snapshotService.runAllSnapshots, which loops
runSnapshot per sport, which already carries the onGraded -> retention
hook. The scheduled path and the manual path are the SAME function. The
reason no cron cycle had been captured is simply that no slot has fired
since retention deployed (slots are 14/19/22/1/3 UTC; retention landed
~02:55). Induced proof follows the deploy.

PHASE 2 — archetype/team/opponent were permanently null because retention
persisted at GRADE time, before enrichment attaches them. Retention still
COLLECTS at grade time (the only moment the feature vector exists) but now
PERSISTS after enrichment, merging those three fields via
retentionService.mergeEnrichment. The merge is pure and fills ONLY those
three fields — features and every model output are grade-time values and
must never be rewritten by enrichment; a test asserts that. Unmatched rows
(refusals not in the enriched slate) keep nulls rather than guesses. The
empty-slate early return now persists too: a refusal-only slate is still
history worth keeping.

PHASE 3 — ZERO-WRITE ALARM. opsWatch.retentionZeroWriteAlarm pages at
missed-snapshot severity when a slot GRADED props but retention wrote
fewer rows than the slate (or nothing). runSnapshot now returns
retentionRows so the scheduler can evaluate it. Retention is best-effort
by design so it can never break a snapshot — which means a broken write is
silent by construction. This is the counterweight. A slot that graded
nothing never false-pages; an absent count reads as NOTHING and still
pages, distinct from a reported 0.

Suite 280/3349 green, build exit 0. Outcome stamping deliberately NOT
implemented (depends on the settlement fix).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 02:00:40 -04:00
builtbykev c2f6041406 STATE: off-box round trip CLOSED — 645 restored from the box copy
Phase 3 complete. The insurance chain is proven end to end rather than
assumed: dump -> validated -> pushed off-box -> verified on the box ->
pulled back down -> rebuilt into a live database.

Pulled vyndr-20260720-051158.dump FROM the Storage Box (not the local
copy) with the in-session key through the pinned host key, never
bypassing StrictHostKeyChecking. Restored into scratch Postgres 17: 715
archive objects, 42 public tables, ledger_entries with all 27 columns and
real spot-checked rows.

ASSERTION PASSED: ledger_entries restored 645 == live 645 (target >= 645).
model_snapshots restored 100/100, so the retention store shipped yesterday
is covered by backups from day one.

Records the operational gotcha the restore surfaced: the dump is written
by pg_dump 17 (Supabase 17.6) and pg_restore 16 CANNOT read it —
'unsupported version (1.16) in file header'. The first attempt failed on
exactly this. Any DR runbook must use PG17+ tooling. Restoring into
vanilla Postgres also logs 12 ignored errors (Supabase roles/extensions
absent locally) which are harmless.

Scratch DB torn down, pulled copy deleted, both dumps still on the box,
nightly cron untouched. No private key material echoed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 01:28:29 -04:00
builtbykev 491e636b7f STATE: off-box backup working + verified on the box
Records off-box as WORKING with the root cause (key was only in Hetzner's
project store, never in the box's authorized_keys — the box previously
offered an EMPTY auth list) and the proof: offbox_ok:true, and the file
independently VERIFIED on the box via rsync --list-only through the pinned
host key (vyndr-20260720-051158.dump, 833,917 bytes, 05:12:28 UTC,
byte-identical to the local dump).

Env truth captured from the run output: key is correctly base64-decoded,
destination has no leading-slash bug.

Hardening recorded: host key statically pinned (accept-new gone, missing
pin refuses the push), remote dir guaranteed, failed required push now
pages at urgent with offbox_ok:false while the exit code still tracks
on-box durability.

Flags the ONE outstanding acceptance item honestly: the round-trip restore
is NOT done, because the dev box cannot authenticate to the Storage Box
(the authorized key is Kev's, not the in-session keypair) and the
container has no Postgres server. Lists both unblocks and the assertion
target (>= 645).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 01:16:30 -04:00
builtbykev 2bfae804da Verify off-box presence by reading the remote dir back
Exit 0 from the backup script is deliberately tied to ON-BOX durability,
so it is not proof the off-box copy landed. GET /api/internal/backup/offbox
runs rsync --list-only against BACKUP_REMOTE using the SAME pinned
known_hosts as the push (checking never disabled) and returns the dumps
actually present, with size and timestamp — so off-box presence is a
verified fact rather than an inference from an exit code.

Needed because the dev box cannot authenticate to the Storage Box: the
authorized key installed there is Kev's ~/vyndr-backup-key, not the
keypair generated in-session, so independent verification has to run from
the container that does hold working credentials.

Suite 280/3338 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 01:14:38 -04:00
builtbykev c4c9b97604 Off-box backup: pin the host key, guarantee the remote dir, page on failure
PHASE 1 — HOST KEY STATICALLY PINNED. ssh-keyscan -p 23 returned an
ED25519 key whose fingerprint EQUALS the out-of-band value
SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM, so it is safe to pin.
scripts/storagebox_known_hosts now carries that verified line and ships to
the container (Dockerfile already COPYs scripts/). backup-db.sh uses
StrictHostKeyChecking=yes + UserKnownHostsFile=<pin> instead of
accept-new, which was trust-on-first-use and would have accepted an
impostor on the very first run. A missing pin file REFUSES the push rather
than silently falling back. Never weakened to accept-new/=no//dev/null —
a test asserts that on executable lines.

PHASE 1b — REMOTE DIR GUARANTEED. The box has only .ssh/, and rsyncing a
file into a missing parent either fails or silently writes the dump AS the
directory name — one file, overwritten nightly, reading as "backups exist"
while retaining exactly one. Uses rsync --mkpath when available, else an
explicit remote mkdir -p ahead of the push.

PHASE 2b — FAILED OFF-BOX PUSH IS NOW LOUD. Off-box is required, so the
failed-push path pages at "urgent" (was "low"/deferred) and the script
emits a machine-readable OFFBOX_OK=1/0/deferred that
POST /api/internal/backup/run surfaces as a distinct offbox_ok field.
Exit code deliberately still reflects ON-BOX durability — a good on-box
dump must not raise a false total-failure alarm. Surfacing the truth, not
manufacturing a failure.

No key material is echoed anywhere; only the PUBLIC host key is committed.

Suite 280/3338 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 01:10:23 -04:00
builtbykev 7e150f9342 STATE.md: pin header + point to the orientation block 2026-07-20 00:10:15 -04:00
builtbykev 7ea0af2081 STATE.md: CURRENT STATUS + OPEN ITEMS orientation block
Top-of-file ground-truth block for orienting a fresh session.

Records what shipped tonight (probability layer revived 32/32, value
engine arc 1, grade-range work, backup durable on-box at 643/643,
model_snapshots retention live with 100 rows incl 36 refusals, ESPN
parser fix) WITH two honest qualifiers rather than a clean win: A still
does not emit in production so the A-RATED marketing hold stands, and EV
is overconfident (+62%/+61%/+56.9% captured, p_win clamps at 0.95) while
hero v2 already ranks on it.

OFF-BOX BACKUP recorded as NOT WORKING and deferred — never succeeded
once, every dump lives only on the Hetzner volume. Documents the two real
blockers fixed (missing base64 decode; container had rsync but no ssh
binary) and the remaining one: the Storage Box offers an EMPTY auth-method
list, which is an account refusing all auth rather than a wrong key.
Explicitly marks as UNVERIFIED that no Chrome/UI diagnostic was run — no
data on the SSH-support toggle, external reachability, project-vs-box key
scope, or any Hetzner outage — and lists those as untested hypotheses in
likelihood order rather than implying they were checked. The full
scratch-Postgres restore proof is recorded as still OWED.

Open items with status: settlement zero-pushes bug, ~28 props/day
unsettled, permanent model-version contamination (with the hard-cutoff
rule), A-grade unreachable, EV overconfidence, edge_pct broken scale, C4
CLV, consistency CV stopgap. Plus the three live credentials to rotate:
Storage Box password, VYNDR_INTERNAL_KEY (pasted in a transcript), and the
GitHub PAT still in .git/config.

Next queued: backtest harness (needs ~2wk history, currently holds one
night), settlement audit, opponent-strength sourcing (MLB solved via
statsapi pitching splits; NBA/WNBA open) behind the source-adapter
pattern, then the metrics engine gated on the harness.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 00:10:02 -04:00
builtbykev 971f641d12 STATE: retention live + settlement findings + version contamination
Records model_snapshots as LIVE and verified capturing (100 rows over 2
cycles: MLB 14 graded/36 refused, WNBA 50 graded; features, grade_11,
p_win, ev_pct on 100% of graded rows).

First-ever refusal visibility: juiced_no_edge 18, rare_event_over_below_
line 13, insufficient_data 5 — the MLB gate refused 36 of 50 sides (72%),
now measurable for the first time.

Flags EV as OVERCONFIDENT and not fit to surface: first captured values
include +62.1%/+61%/+56.9%, which real markets do not offer. Cause is the
estimator clamping p_win at PROB_CEIL 0.95 off ~10 games. Hero v2 already
ranks on ev_pct, so it will pick the MOST overconfident read — calibration
must gate this before EV drives anything user-facing.

Logs the two settlement-correctness findings Kev asked to track (zero
pushes across 470 settled rows; ~28 props/day never settling) and the
ledger model-version contamination, with the rule that any backtest off
existing history must treat the 2026-07-19 fix boundary as a hard cutoff.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 23:14:58 -04:00
builtbykev 14f47af74b Dockerfile: install openssh-client — rsync cannot exec ssh without it
The off-box push failed with 'rsync: Failed to exec ssh: No such file or
directory (2)'. The container had rsync and pg_dump from S62 but no ssh
binary, and rsync shells out to ssh for every remote transport. The dump
itself succeeded, so this failed AFTER a good backup and reads like a
network/auth problem when it is a missing package.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 23:02:16 -04:00
builtbykev d3ffa1b8c2 Retention: model_snapshots live + base64 SSH key support
RETENTION (Phase 2, priority zero). History starts compounding tonight.

migration 025 model_snapshots — APPLIED to prod. Append-only, one row per
graded prop PER SIDE PER CYCLE, with a unique index on
(snapshot_id, player_key, stat, line, side) so a retried cycle cannot
duplicate. RLS on, service-role writes only.

What it captures that the ledger never did:
- features jsonb — the model's INPUTS. Without these a backtest can only
  grade our own homework; with them any future model can be replayed
  against the exact conditions this one faced.
- REFUSALS (refused + refusal_reason). The ledger drops them, so a gate
  refusing props that would have WON is invisible — unmeasurable lost
  edge. Captured via a new onGraded hook in gradeSlateService that fires
  with BOTH sides before any filtering.
- grade_11, the pre-collapse grade. The 4-letter map throws away the
  entire live C-/C/C+/B- range.
- model_version + code_sha on every row. ledger_entries mixes pre/post-fix
  grades with no marker and cannot be separated retroactively.
- p_win / ev_pct / fair_odds / takeable / value — none of which any
  permanent store held.

Wiring: analyzeViaEngine1 attaches _features/_grade_11 (underscore =
internal); gradeSlateService fires onGraded then STRIPS them so they never
reach a cache or API payload; snapshotService builds rows and persists
best-effort. Retention reuses the LEDGER's dateET/gameIdFor helpers so
rows share the ledger's natural key exactly — otherwise the settle pass
could never join outcomes onto them. Rows are written BEFORE the empty-
slate early return: a slate that refused everything is exactly the case
worth recording.

CONTRACT HELD: retention is injectable and every path is caught. persist()
returns errors, never throws; a missing Supabase client is SKIPPED, not an
error. A retention failure can never break a snapshot.

BACKUP: backup-db.sh now accepts BACKUP_SSH_KEY as base64 (recommended —
survives env-var newline mangling, which is how injected SSH keys usually
break silently) OR raw PEM, detected by decoding and looking for the PEM
header. Verified both forms detect correctly against a real generated key.

Suite 279/3325 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 23:01:12 -04:00
builtbykev 04a09ec1b2 Phase 2 (a) report + (b) retention design — REPORT-FIRST, nothing built
(a) WHAT REPLAYABLE HISTORY EXISTS — the headline is confirmed and worse
than "6 days".

ledger_entries is the ONLY store of model history in the database: 640
public rows, 6 distinct game days (Jul 11/12/16/17/18/19 — 13/14/15 are
missing entirely), 2 sports, 215 players, 470 settled, 465 settled WITH
odds. Every other candidate is 0 rows: grade_history, line_snapshots,
historical_props, closing_lines, resolution_results, accuracy_tracking,
model_predictions_extended, engine1_weights, prediction_registry and ~30
more. A data warehouse was designed and never filled. Redis holds no
history either (latest/previous at 24h TTL; the outcomes log carries no
odds/confidence/projection).

The blocking gap is not the day count, it is that NO MODEL INPUTS ARE
STORED ANYWHERE. No feature vectors, so we can score the grades we
emitted but cannot ask whether a different model would have done better —
which is the only question a harness exists to answer, and the exact gate
the metrics-engine north star requires. Also missing: p_win/ev_pct/
fair_odds (born tonight, on no column), grade_11 (only the 4-letter
collapse is stored, so the entire live C-/C/C+/B- range is unrecoverable),
and any model_version, so pre- and post-fix rows are already silently
mixed in one table. CLV remains unusable (C4). Settlement gaps surfaced
too: Jul 17 MLB 86 graded/57 settled, Jul 18 103/75, and 0 pushes across
470 settled rows — both feed the settlement-correctness audit.

Verdict: we cannot meaningfully backtest yet. Retention is priority zero;
every night without it is history we can never recover.

(b) DESIGN PROPOSAL — model_snapshots in Postgres (not Redis, which is
what lost us history twice). One append-only row per graded prop PER
CYCLE, capturing market values, model output, outcome (stamped later by
the settle pass), and critically a `features` JSONB — the counterfactual
enabler. Carries model_version + code_sha so eras never mix, grade_11 so
resolution is not thrown away, and refused/refusal_reason because
refusals are training data the ledger currently discards entirely.
Written from snapshotService (the existing chokepoint), best-effort so a
retention failure can never break a snapshot.

Volume: ~800 rows/day ~ 292k/year, ~300-600MB/yr of features, which would
exceed the Supabase free tier alone — so the proposal keeps full features
90 days and scalars forever. Three open questions for Kev before building:
the 90-day policy, whether to backfill the 640 existing rows as
scalars-only with explicit null features, and confirming we store
refusals.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 22:40:43 -04:00
builtbykev 8aafceaa3b STATE.md: pin header to cc8e478 2026-07-19 22:32:59 -04:00
builtbykev cc8e47884d STATE.md: backup is durable on-box and verified by read-back (643==643) 2026-07-19 22:32:59 -04:00
builtbykev 97e4dc72d5 Backup: chown /app/backups in image + report uid/writability
The real backup run failed: pg_dump could not write to /app/backups —
'Permission denied'. Cause: the container runs as the non-root 'vyndr'
user (Dockerfile USER vyndr) and the Coolify-mounted volume is
root-owned, so the mount is present but unwritable.

- Dockerfile now creates AND chowns /app/backups to vyndr alongside the
  existing /app/data + /app/.pm2 line. Docker seeds ownership into a
  NAMED volume on first creation, so this fixes it for a fresh volume;
  a host bind-mount still needs a host-side chown, which is why the
  next change exists.
- GET /api/internal/backup/verify now reports process uid/gid,
  backup_dir_writable and the access errno, so the exact chown target is
  observable instead of guessed. A mounted-but-unwritable volume reads as
  'configured' everywhere else — this makes it loud.

Suite 278/3310 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 22:30:37 -04:00
builtbykev ef7f17610f Backup: durable on-box volume, off-box DEFERRED, and a real read-back check
BACKUP_DIR is now a persistent volume (/app/backups), so the dump already
survives redeploys — the container-ephemeral risk that made this urgent is
closed. Storage Box SSH auth is not sorted yet, so the off-box push is
explicitly DEFERRED rather than failing:

- gated on BACKUP_OFFBOX=1 (plus BACKUP_REMOTE and BACKUP_SSH_KEY); until
  then the script logs "off-box push DEFERRED" and exits clean.
- if an enabled push DOES fail, it is a LOW-priority "deferred" notice, not
  a failure — the durable on-box dump succeeded, and calling that an
  incident would train us to ignore backup alerts.

Adds the read-back check, because a backup nobody has read is a hope:
countRowsInDump() runs `pg_restore --data-only --table=X -f -` and counts
the rows between `FROM stdin;` and the terminating `\.`, proving the
archive CONTAINS the data rather than merely parsing. Needs no Postgres
server, so it runs inside the API container. Validated against a real
pg_dump from a scratch Postgres: counted exactly 604 rows.

GET /api/internal/backup/verify exposes it (newest dump in BACKUP_DIR,
size, table, rows_in_dump). Unit tests inject spawn/fs so CI needs neither
docker nor pg_restore.

Suite 278/3310 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 22:27:58 -04:00
builtbykev 13ca070096 Spec: metrics-engine north star + sourcing scope report (no code)
NORTH STAR (design philosophy, not built): VYNDR measures players by
MODERN FUNCTION, not legacy label — the principle already under the
archetype system, from Rashad Phillips' Basketball Position Metric. The
rule: every proprietary metric is baselined against the player's
functional ARCHETYPE's CURRENT-SEASON behavior, never the position's
inherited standard. The edge is that the market often prices today's
players against yesterday's baselines, so archetype-vs-position baseline
disagreement is a repeatable mispricing. Generalizes across sports. Moat
= proprietary metrics x current-game calibration x our private outcome
data. Metrics ship as VALIDATED FAMILIES: hypothesis, flagged build,
backtest, ship-or-delete with the negative result written down. Nothing
is real until the harness proves it predicts better.

SOURCING SCOPE (report, no code): MLB opponent strength IS derivable from
statsapi, verified live — one free call returns all 30 teams' pitching
splits (era/whip/avg/slg/ops/homeRuns/strikeOuts/HR9), which beats the
ESPN field we were reaching for because it is STAT-SPECIFIC, exactly what
opp_rank_stat wants. NBA/WNBA cannot use ESPN (its team endpoint carries
only a team's own stats, no defensive rating or pace); options are
stats.nba.com dashboards, deriving allowed-points from scoreboard finals
we already fetch, or API-Sports. API-Sports is a fallback tier at best —
100/day will not survive per-team-per-day. ESPN stays last, always behind
an adapter.

Proposed the SOURCE-ADAPTER pattern: one interface per feed, config-driven
primary+fallback per (sport x capability), normalized output so vendor
quirks stay in adapters, fallback announced rather than silent, sources
with zero callers deleted rather than left as corpses, and a health check
that PAGES when a source returns empty or broken — where EMPTY IS A
FAILURE. Tonight's crash (captured 0 / errored 15) and the months-null
opp_rank_stat are both exactly what that check exists to catch.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 21:43:59 -04:00
builtbykev 40aba37f83 Backup: env-injected SSH key, nightly off-box push, triggerable run
Closing the backup for real. Three changes, each fixing something that
would have made the Storage Box target fail or silently rot.

1. SSH KEY COMES FROM ENV, not from the container. Generating a keypair
   inside the API container was the obvious move and it is wrong: the
   container filesystem is ephemeral, so the key dies on the next
   redeploy and the off-box push starts failing silently. backup-db.sh
   now reads BACKUP_SSH_KEY (a Coolify secret), writes it to a 0600 temp
   file per run, and removes it on exit via trap.

2. PORT 23, verified live. Hetzner Storage Box runs full OpenSSH on 23;
   port 22 answers with mod_sftp (SFTP only). Banner-checked both against
   u635423.your-storagebox.de. rsync now uses
   -e "ssh -p ${BACKUP_SSH_PORT:-23} ... -i <key>"; the old invocation had
   no -e at all and would have gone to 22.

3. OFF-BOX PUSH IS NIGHTLY, not Sundays-only. A weekly push meant up to
   six days of dumps existed ONLY inside an ephemeral container, which is
   the same as not existing. Alert copy updated to say exactly that when
   the push fails or is skipped.

Also adds POST /api/internal/backup/run (internal-key gated) so a real
backup can be TRIGGERED and OBSERVED — it returns exit code, duration,
output tail, and whether the remote + ssh key are configured. The backup
can only run where SUPABASE_DB_URL and the Supabase route live (this
container), and there was no way to fire or inspect it without a shell.

Connectivity established this session: Storage Box reachable from the dev
box on 22/23; Supabase :5432 NOT reachable from WSL2 (so the dump must
run in-container, as designed); docker IS available locally, so the
restore-verify can run against a scratch Postgres using the real dump.

Suite 278/3305 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 21:42:27 -04:00
builtbykev 26b276fbfb Fix the ESPN team-stats parser + report: opponent rank is still underivable
Ran the manual regrade with the internal key (thanks). Results are mixed
and the honest half matters more.

CONFIRMED WORKING — the probability layer is fully alive in production.
After POST /api/internal/snapshot/{mlb,wnba}: p_win, ev_pct, model_odds
and value are present on 32/32 live grades (mlb 7/7, wnba 25/25), up from
0/8 before. That fix is done.

NOT WORKING — the grade-range half did not land, and I am not going to
claim it did. The live distribution is unchanged (wnba B17/C8 before AND
after; mlb B4/C3), no A, no D, same four confidence values. Diagnosis:
matchup_grade is 0/25 on the live board, i.e. opp_rank_stat is still
null, so engine1's +/-1.0 opponent factor still never fires and the
ceiling is still +3.0 against the +4.5 an A requires.

Two distinct causes, both verified against the live ESPN feed:
1. refreshTeamStats CRASHED on every team — "buckets is not iterable",
   captured 0 / errored 15. ESPN's current shape is results.stats =
   an OBJECT with categories[], not an array. The old parser did for...of
   on it. This was invisible until S63 gave the function its first
   production caller. FIXED here (now captured 15 / errored 0) with a
   regression test covering the current shape, the legacy array shape,
   and empty payloads.
2. Even parsed correctly, the endpoint does not carry a
   defensive-strength metric at all: defensive_rating, opponent_ppg,
   pace and opponent_fg_pct all normalize to null — it returns only a
   team's OWN stats. So defensive_rank_normalized cannot be computed and
   opp_rank_stat remains underivable from this source. A test documents
   the gap and will fail if that ever changes.

Consequence: A STILL DOES NOT EMIT, so the A-RATED marketing hold STAYS.
Reviving the opponent factor needs a different derivation (opponent
points allowed from scoreboard/schedule, or a different ESPN endpoint) —
logged as the concrete next item, not hand-waved as done.

Suite 278/3305 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 20:31:15 -04:00
builtbykev b742230d94 Phase 1: ship the backup cron as CODE + a manual regrade trigger
FOUNDATION-FIRST re-order, phase 1 (tooling + safety).

BACKUP (highest-severity open item) — INSTALLED, not re-proven.
src/backupScheduler.js runs scripts/backup-db.sh nightly from inside the
API container, armed at boot in server.js. The container already has
SUPABASE_DB_URL, pg_dump and the Supabase route, so deploy == installed:
no host crontab, no Coolify click. Arming is deliberately opt-OUT (armed
whenever SUPABASE_DB_URL exists; BACKUP_CRON=0 kills it) because the S62
design was opt-in and nobody ever opted in — the DB went unbacked every
night for weeks. A failed run pages high-priority ntfy; silence is the
danger with backups.

Durability is the one part still needing a human: the container FS is
ephemeral, so a dump dies on redeploy unless BACKUP_REMOTE (off-box
rsync) or BACKUP_DIR (persistent volume) is set. The scheduler detects
that and pages a WARNING at boot rather than letting an undurable backup
read as "backed up". Runbook rewritten to lead with the code path.

MANUAL REGRADE TRIGGER — scripts/run-snapshot.js, runnable via
docker exec with no VYNDR_INTERNAL_KEY and no new HTTP surface. Runs the
SAME snapshotService.runSnapshot the cron runs (including the team-stats
refresh that powers opp_rank_stat), supports `all` and `--settle`, and
prints the grade/confidence distribution plus p_win/ev_pct presence —
which is the thing you actually want when verifying a grading change.

ACCESS BLOCKER, logged honestly in specs/model-train.md: there is no
VYNDR_INTERNAL_KEY in the local .env and SSH to the box times out from
WSL2, so I can neither curl the internal endpoints (which already exist
from S45) nor docker exec. The trigger is built and correct but only Kev
can run it until a key or SSH access exists. This is the highest-leverage
unblock for phases 2 and 3, which both need on-demand regrade+settle to
verify anything.

Also logged the standing cautions: CLV ledger stays private until
backtest-proven; "self-improving model" is unsupported marketing until
the loop closes; the engine is MLB/WNBA-calibrated and NFL/NBA/soccer
need their own calibration before the hub grades them (scaling gate).

Suite 277/3300 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 19:28:15 -04:00
builtbykev bf8ecb45ad Track U-deg pt2 (edge_pct scale) + the dispersion classifier as open items
Not new work — logging so neither gets lost.

U-deg part 2: edge_pct is on a broken scale and it is the number users
actually see. Live: edge_pct 100 on a single-digit-edge prop; ledger-wide
311/604 rows (51.5%) exceed the frontend's sane cap of 40, 39 exceed 100,
worst 620. Mapped the consumers, and the split is the whole problem: 13
frontend files + deskShowcase/contentTemplate/parlayScan/tierGating read
the BROKEN edge_pct, and ledgerService:199 persists it to the
column of an append-only table right now. NOTHING on the frontend reads
ev_pct; only heroPropService does. Noted that S-b (rank board on EV) is
the real remedy and should be done as one piece with the scale fix, and
that EDGE_BOARD_SANE_MAX is damage control that nulls half the board.

Dispersion classifier: MIN_MEAN=4 is the honest stopgap; it leaves a
+/-1.0 dead for MLB low-count stats. The scale-free fix is variance/mean
vs the Poisson baseline of 1.0. Logged with its explicit validation bar —
backtest harness first (still does not exist), replay settled outcomes,
show no tier degradation, report before flipping, env-gate it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 19:08:04 -04:00
builtbykev 58ce1c3e56 STATE.md: header to deployed+fingerprinted HEAD 2026-07-19 18:58:26 -04:00
builtbykev a80868c4eb S63 fingerprint: probability layer verified live (p_win/ev_pct/model_odds)
POST /api/analyze/prop on prod returns p_win 0.523, ev_pct -10.4,
model_odds -109, confidence_basis grade_band, value false — every one of
which was absent on 100% of grades before this change. The value triplet
is whole (book -140 / fair -125 / model -109) and correctly refuses to
call a -140 price value when the model gives it 52.3%.

A-emission still pending the 01:00 UTC snapshot (opp_rank_stat populates
only when refreshTeamStats runs in a snapshot). MARKETING HOLD on A-RATED
copy stays until that passes. edge_pct scale remains broken (U-deg pt 2).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 18:58:12 -04:00
builtbykev 1a94ef5fcf Revive the dead probability layer + restore grade range ON MERIT
Folds re-sequenced steps 1+2 into one change (Kev's call): same bug
family — features wired to sources that return null.

THE PROBABILITY LAYER WAS DEAD IN PRODUCTION. p_win/ev_pct/kelly/
model_odds/value were absent on 0/8 live grades because
gameLogService.getGameLogs returns null for MLB by construction and
depends on the offline Python service for NBA/WNBA, so meta.gameLogs was
[] for every sport. This was the S46 bug in a second location — that fix
gave featureCache an MLB branch (why grades still worked) but never the
estimator. featureCache.getStatRows now supplies normalized rows
([{date,[statType]:v}], most-recent-first) for every sport, feeding the
estimator AND consistency AND game_count_in_7d from one fetch.
VERIFIED on real props: p_win 25/25 WNBA, 8/8 MLB (was 0).

GRADE RANGE, ON MERIT — never by rescaling (permanent founder ruling:
minting A's without new information is a relabelled B sold as an A and
corrupts an append-only ledger).
- refreshTeamStats wired into runSnapshot — it had ZERO production
  callers, so opp_rank_stat was permanently null and a +/-1.0 factor
  could never fire. Test-env no-op (opsNotify precedent).
- L20 made SYMMETRIC: both branches were delta +1.0, so the season
  baseline could only ever ADD. No negative path was a structural reason
  D was unreachable. New l20_contradicts_* carries -1.0.
- game_count_in_7d derived from real logged dates (heavy_workload_7d).
- NOT wired, deliberately, with reasons inline: teamId (no team_id
  column; getFeatures reads it top-level; factor also needs a starter-id
  list) and season_type (ESPN 2 = REGULAR season; threading it raw would
  fire veteran_in_playoffs in July). Dead code dressed as a fix is the
  thing we are removing, not adding.

CALIBRATION GUARD (found by verifying, not assuming): consistency CV is
NBA-tuned; for a Poisson-ish stat cv ~ 1/sqrt(mean), so any stat with
mean < 4 auto-classifies boom_bust. First verification run showed 8/8 MLB
props boom_bust — a blanket -1.0 that dropped the board to all-C. Floored
at CONSISTENCY_MIN_MEAN=4 -> 'unknown' below. Absent beats wrong. MLB
low-count stats therefore still get no consistency factor: honest, not
fixed. Scale-free index-of-dispersion classifier is the open follow-up.

CONFIDENCE IS NOT A PROBABILITY: payloads carry confidence_basis:
'grade_band'. Corrected mlb-grade-degradation.md — its "25/25
grade<->confidence agreement" is a TAUTOLOGY (confidence is derived FROM
the letter, so it would report 25/25 even if every grade were wrong), not
a validation. Removed dead mlbGrader.js (referenced only by its own test)
and the stale computeFeatures comment claiming a penalty that never ran.

VERIFICATION (scripts/verify-grade-range.js, real props/logs/engine):
WNBA 25 props B 68%->32%, C 32%->64%, D 0->1 (4%); 11-step spread went
from 2 steps to 5 (C/C+/B-/D). The D is earned: Angel Reese assists o2.5,
p_win 0.365. Nothing flooded — grades got HARDER. A did not emit locally
because opp_rank_stat needs the Redis cache only prod populates (local
ceiling +3.0 vs the +4.5 A needs); reachability is proven arithmetically
and locked in tests. Prod A-emission is the outstanding fingerprint.

MARKETING HOLD: "A-RATED" (AccuracyBadge, TopSignals) is unsupported
until that fingerprint. Confirmed honest fallbacks render today —
/api/ledger/accuracy has B and C buckets only, so the badge shows
"MODEL · 63% HIT" and TopSignals self-hides. Nothing fabricated ships.

Suite 276/3286 green, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 18:54:51 -04:00
builtbykev 416639efe4 Grade collapse: mechanism traced — A is mathematically unreachable
Completes the diagnosis. Report only; no grade logic or thresholds changed.

The grade is an integer index (GRADE_SCALE, NEUTRAL_INDEX 3) moved by a
flat sum of +/-1.0 and +/-0.5 factor deltas, then clamped and rounded.
grade_thresholds.json is NOT an input mapper in the JS path — engine1
reads it BACKWARDS, taking the letter the index already produced and
looking up that band's midpoint to manufacture `confidence`. So
confidence is a cosmetic re-encoding of the letter: zero information
beyond it, and it can never disagree with it. There is no
data-sufficiency penalty in the live path (the one CLAUDE.md describes is
in mlbGrader.js, which is dead code).

Six of thirteen factors are wired to features nothing populates —
verified: refreshTeamStats has ZERO production callers (so opp_rank_stat
is permanently null, killing a +/-1.0), teamId/season_type/
game_count_in_7d are never passed (gameContext is built as {home_away}
and nothing else), and MLB consistency starves on the same dead
gameLogService path as Finding 2. Also verified: BOTH l20 branches are
delta +1.0 — there is no negative L20 contribution at all.

Arithmetic: an A needs sum >= +4.5; the live maximum is +3.0 (+2.0 on a
back-to-back, and MLB rest_days is 0 most days). D needs <= -1.51; the
live minimum is -1.5 and Math.round(1.5)=2, so it misses by one rounding
tick. Reachable band is index 2..6 = {C-,C,C+,B-,B}, which the adapter's
FOUR_LETTER_MAP (a 3->1 collapse) renders as exactly {C,B} — the observed
output, derived from first principles. Reachable confidences {42,47,52,
57,63} match the live values {47,52,57,63} exactly; C- is truncated by
gradeSlateService keeping the higher-confidence side.

mlb-grade-degradation.md's "25/25 grade<->confidence agreement" is a
TAUTOLOGY, not a validation — confidence is derived from the letter, so
it would report 25/25 even if every grade were wrong.

Recommends feeding the starving factors (restores A/D on merit) and
explicitly REJECTS re-scaling thresholds, which would mint A's without
adding information — every "A" would be a relabelled B.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 18:15:14 -04:00
builtbykev 3fe840ab83 Diagnose grade collapse + find the DEAD probability layer (report only)
Kev's call: investigate the B/C grade collapse before building. Report
only — no grade logic, thresholds, or engine code touched.

FINDING 1 — the collapse is real, live and structural. Across 604 ledger
rows and both sports the engine has emitted exactly TWO grades (B, C) and
NINE confidence values (63/57/55/52/47/45/35/25/20), ceiling 63. Still
true today on both sports. Confidence does NOT determine the letter:
conf 45 -> B while 47 and 52 -> C (non-monotonic), so the surfaced
confidence is not the quantity the letter came from. Edge scale still
broken: 311/604 rows exceed the frontend's sane cap of 40, 39 exceed 100,
worst 620.

FINDING 2 (bigger) — the entire probability layer is DEAD in production.
Live /api/snapshot/mlb: p_win, kelly, ev_pct, model_odds and value are
absent on 0/8 grades, while alt_lines (Desk-gated) IS present 8/8 —
proving nothing is tier-stripped, they are simply never computed.

Root cause: gameLogService.pythonPath returns null for MLB by
construction and the Python service is offline for NBA/WNBA, so
meta.gameLogs is [] for every sport; estimateProbability returns
p_over null; every field guarded by `if (pWin != null)` is skipped.
This is the S46 bug in a second location — that fix added an MLB branch
to featureCache.gameLogFeatures (which is why grades/projections still
work) but never to the estimator path.

Consequences: EV — the Model Train's whole ranking signal — has never
been computed on a live prop. Hero v2 matches nothing and always falls
through to the recent-read fallback (live /api/hero-prop returns
is_recent:true). Quarter-Kelly, sold on the pricing page and listed BUILT
in PROMISE-AUDIT.md, never runs. The value triplet is a duet live.

Recommend re-sequencing: revive the probability layer BEFORE G-a and
C-led (C-led would persist a column of nulls; G-a's EV_FLEX_THRESHOLD
would gate on a permanently-null value — Kev's EV_FLEX_ENFORCE=0 ruling
accidentally prevented an outage). featureCache:206-226 already has both
adapter branches and is the template.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 18:13:06 -04:00
builtbykev 669479097c Model Train G-b/C-cal: gate simulation + calibration report (docs only)
REPORT-FIRST per the arc order. G-a is HELD — the data changes the
recommended dials. No engine code touched.

Replayed against live ledger_entries (576 rows, 6 game days, 470 settled)
because the "30 days of stored snapshots" does not exist: snapshot Redis
keys are latest/previous only at 24h TTL, and no backtest harness exists
anywhere in the repo.

Findings that change the plan:
- The -400 floor shipped this morning was the whole win: past -400 hit
  80.3% against an 86.9% breakeven = -13.29u / -7.7% ROI on 173 settled.
- Arc 2's incremental cut over the live gate is ~11 props in 6 days. The
  only material change is gating the flex band behind 2x EV.
- The flex band (-161..-250) is our BEST band (+2.2% ROI, n=70) and the
  takeable band is flat (-0.3%, n=209) — the opposite of the assumption
  behind EDGE_FLEX_WALL. Recommend shipping the knob with enforcement
  OFF until EV is persisted and measured.
- ev_pct/p_win are on NO ledger row, so the EV half of the gate cannot be
  replayed at all. C-led (persist EV) is now the highest-leverage item.
- Confidence is monotonic but understates hit rate by ~20-25 points, and
  the entire public ledger contains only B and C grades — zero A/A+.
  That breaks hero v2 (isAB) and undermines "A-RATED" copy. Escalated.
- L-a answered: alt_lines carry NO odds and the feed has no alternate
  markets. L-b is blocked on a data source, not engine work.
- C-led needs no odds backfill (locked_odds 99.1% populated).
- U-deg: the projection==0 leak is already closed (0 since 07-18).
- C4 confirmed in data (359/376 MLB closes == the lock). Stays suppressed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 18:04:29 -04:00
builtbykev 18bf3ecb51 Reconcile the record: STATE.md header + Model Train arc 1 spec (docs only)
Arc 1 (7a925f4) shipped without a spec and without a STATE.md entry — the
plan lived only in a session context that was lost. Both written from the
code on disk, not from memory.

- STATE.md: header was stale at a8e383e; now 7a925f4 (pushed, NOT yet
  deploy-fingerprinted). New top section records de-vig + EV + takeable/
  value gates + hero v2 + the value triplet, the real config values, and
  the 274/3289 -> 276/3306 test baseline.
- specs/model-train.md (NEW, per CLAUDE.md rule #1): every knob and its
  ACTUAL default (TAKEABLE_ODDS_CEILING -160, TAKEABLE_ODDS_MAX +200,
  VALUE_EV_THRESHOLD 2, JUICE_ODDS_FLOOR -400, RARE_EVENT_LINE_MAX 0.5);
  the gate AS BUILT (flat price-only refusal at -400 pre-feature, NOT
  edge-aware, no -250 wall; the -160..+200 band never refuses a grade and
  is enforced on the hero alone); the triplet + hero v2; and a checklist
  of what is NOT built.

Recorded honestly rather than assumed: EDGE_FLEX_WALL, HARD_JUICE_WALL,
LADDER_ODDS_MAX and MIN_RUNG_PROBABILITY do not exist in the codebase; no
frontend reads ev_pct/fair_odds/model_odds/book_odds/suppressed_reason yet;
the value fields are ungated to free tier; a second EV implementation
(processing/EVCalculator.js) still coexists with devig.evPct.

No engine code touched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 17:56:09 -04:00
builtbykev 7a925f43eb Model Train arc 1 (engine): de-vig + EV + takeable/value gates + hero v2 + triplet
Steps 1-6 — make "real opportunities at takeable prices" the engine, not a filter.

1. DE-VIG (src/utils/devig.js): two-way multiplicative de-vig strips the vig and
   returns fair prob + fair price per side + the overround. One side missing →
   fair UNAVAILABLE (null), never faked. Method noted in code + the `devig_method`
   field.
2. EV (devig.evPct): ev_pct = model prob × decimal − 1 at the graded side's
   ACTUAL price. This is the ranking signal now, replacing raw |model−consensus|.
3. TAKEABLE gate (src/config/valueEngine.js, TAKEABLE_ODDS_CEILING −160 .. +200,
   env-tunable): promoted surfaces only (hero/featured/alerts). The full board
   still shows everything; Parlay Lab exempt; JUICE_ODDS_FLOOR (−400) stays the
   absolute backstop underneath. Strict null-guard (Number(null)===0 would have
   made a missing price "takeable").
4. VALUE flag: passes BOTH gates (takeable AND ev_pct ≥ VALUE_EV_THRESHOLD).
   Grade = read quality; value = the price pays you. Shipped in payloads.
5. HERO v2 (heroPropService): highest ev_pct among takeable A/B reads — a huge
   gap on a −900 line is trivia, not an opportunity.
6. VALUE TRIPLET: book_odds · fair_odds · model_odds on every read (snapshot,
   hero, scan — they all spread the grade). Handoff documents the fields; the
   rendering is Session-2 Design's job.

All wired in analyzeViaEngine1's existing p_win/kelly block (real quantile
probability × real book odds, or nothing). 33 new tests; suite 276/3306 green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 02:43:30 -04:00
builtbykev 348a82b4a0 Generalize the no-edge guard: suppress by the BOOK'S PRICE, not a stat whitelist
Follow-up to the rare-event under fix — the whitelist (doubles/triples/HR/SB)
was fragile: the same juiced-under problem exists for steals, blocks, and any
other low-frequency market, and a new stat would slip through.

The real signal is the book's own price. The doubles unders were priced -625 to
-1100 — laying 6-11x to win 1x on an ~82% event, with no value the model could
recover. So the PRIMARY guard is now stat/sport-agnostic: analyzeViaEngine1
refuses any read whose graded-side odds are past the juice floor
(JUICE_ODDS_FLOOR, default -400, env-tunable). That catches every version of
this — steals, blocks, anything — and it also keeps the public record honest
(those -800 "wins" hit ~82% of the time and would inflate the hit rate, the same
class as the projection-0 degradation).

The structural rare-event rules stay as the BACKUP for props with no odds
(list also expanded cross-sport: + steals, blocks). Normal + longshot prices
(-110, -250, +600) are preserved. 16 tests cover both layers.

Reported: the doubles projection was REAL per-player (not a fallback); the fix
is the price guard, not a bigger list.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 02:09:11 -04:00
builtbykev f72f063e6f Suppress rare-event 0.5 unders (juiced, no-edge) — config-driven grade + board fix
Betting-logic audit: the CONSENSUS-vs-MODEL board flooded with fake reads like
"DOUBLES u0.5 · MODEL 0.2 · +edge" — the juiced under side of rare counting-stat
markets (doubles/triples/HR/SB), which is never a takeable edge and violates the
no-unders-default doctrine.

Report finding (item 3/4): the doubles projection is REAL per-player, not a flat
fallback — 'doubles' maps to a real game-log field (MLB_LOG_FIELD doubles→
doubles) and the live values varied (0.03/0.16/0.2/0.22). So no projection-gate
refusal for fakeness; the problem is purely structural (a rare event's real
projection always sits below a 0.5 line, so the under always "wins").

Fix (config-driven — src/config/rareEventMarkets.js, tunable stat list + line
threshold):
- Grade layer (analyzeViaEngine1): a rare-event UNDER at ≤0.5 is always REFUSED
  (grade null + suppressed flag/reason). A rare-event OVER at ≤0.5 is refused
  UNLESS the model genuinely projects the event above the line — because a
  0.2-over-0.5 carries the SAME |edge| as the suppressed under and would just
  take its rank on the board. The over grades normally once projection > line.
- Board layer (marketBreadth.collectBreadth): drops null-model rows so a
  suppressed/ungraded prop can't rank a "MODEL —" placeholder onto the board.

10 suppression tests + config locks; also fixed a settingsPage book assertion
left over from the ESPN→theScore swap. Suite 274/3289 green, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 00:19:23 -04:00
builtbykev b8b954bb96 STATE.md: backup + founder-checkout tasks
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 23:08:57 -04:00
builtbykev c2c43cdc92 Task A — make the container backup-capable + validated dump + mechanism fingerprint
SUPABASE_DB_URL is set in Coolify on the API service, and this WSL2 box can't
reach db.<ref>.supabase.co — so the backup runs INSIDE the API container, which
has the env + Supabase network. Made that real:
- Dockerfile: install postgresql-client (pg_dump/pg_restore) + rsync + bash in
  the runner image.
- backup-db.sh: added an integrity fingerprint on every run — pg_restore --list
  must parse the archive AND find ledger_entries, else the run FAILS + pages
  (stronger than the size check; catches a corrupt/structureless dump).
- BACKUP-RUNBOOK.md: rewritten for the container-exec reality — host cron does
  `docker exec <api> sh /app/scripts/backup-db.sh` (inherits env + network +
  pg_dump), or a Coolify Scheduled Task. Full restore-fingerprint steps included.

MECHANISM FINGERPRINT (run locally, docker + pg16): seeded a ledger_entries
table (137 rows) → ran backup-db.sh (dump + validate: 22 archive objects,
ledger_entries present) → pg_restore into a scratch DB → 137 rows restored,
exact match. The dump/validate/restore path is proven end-to-end; it's the same
pg_dump/pg_restore that run in the container against Supabase.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 23:08:31 -04:00
builtbykev ccb9668f0c Task B — founder checkout is SEAT-GATED; payment-decline grace spans retries
1+2. Checkout price was CODE-gated (founder price only with a valid founder
   code) — so "Claim a Founder Desk" would have charged the $44.99 standard
   price, not the advertised $34.99. Now it's SEAT-gated: resolveCheckoutPrice()
   attaches the founder price while founder seats remain (< FOUNDER_SEATS_TOTAL,
   read from the SAME countFounderSeats() truth as the ClaimMeter), and flips to
   standard at seat 100. createCheckoutSession uses it; the founderCode param is
   kept for back-compat but no longer drives price. The meter flips to "SOLD
   OUT" at capacity. When the count can't be verified we honor the advertised
   founder price (never overcharge).
   - Also hardened countFounderSeats to manual pagination (the for-await form
     broke on non-async-iterable list mocks).

3. Tests: resolveCheckoutPrice at seat 0 → founder, seat 100 → standard, the
   99/100 boundary, null-count → advertised founder price.

4. Grace: invoice.payment_failed now sets a 14-DAY grace (spans Stripe's Smart
   Retry window) instead of 48h — a transient decline no longer revokes access
   mid-retry. Access is revoked only when Stripe actually cancels
   (customer.subscription.deleted keeps its 48h grace). Test updated.

Stripe + founders suites green, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 23:04:19 -04:00
builtbykev 39c07a03b9 STATE.md: security + plumbing follow-up (items 0-7)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-18 13:57:45 -04:00
575 changed files with 70226 additions and 1480 deletions
+4
View File
@@ -27,3 +27,7 @@ out/
# Vercel # Vercel
.vercel/ .vercel/
.seq-cache/
.content-out/
+48
View File
@@ -273,3 +273,51 @@ match `^[a-z0-9_]{3,20}$` (400); a handle owned by another user → 409.
Frontend: `/u/[handle]` (OPEN route, server shell + client record + OG card, Frontend: `/u/[handle]` (OPEN route, server shell + client record + OG card,
Node runtime) via proxies `web/src/app/api/profiles/me|[handle]`. Node runtime) via proxies `web/src/app/api/profiles/me|[handle]`.
## Value Engine fields (Model Train arc 1, 2026-07-19)
Every graded read (snapshot `grades:{sport}`, `/api/snapshot/:sport`, `/api/hero-prop`,
`/api/scan`) now carries — all OPTIONAL + self-hiding when absent:
- `ev_pct` (number) — expected value % at the graded side's ACTUAL price
(model prob × decimal 1). The ranking signal on boards + hero, replacing raw
|modelconsensus|.
- `value` (bool) — the read passes BOTH gates: takeable price AND ev_pct ≥
VALUE_EV_THRESHOLD. Grade = read quality; `value` = "the price pays you." An A
without `value` is honest (right read, price gone).
- `takeable` (bool) — the graded-side price is inside the promotable band
(TAKEABLE_ODDS_CEILING 160 .. +200).
- **The value triplet** — `book_odds` · `fair_odds` · `model_odds` (all American):
the book's price, the de-vigged FAIR price (two-way multiplicative de-vig;
present only when BOTH sides were priced — else absent, never faked), and the
model's implied price. This is the story to render: "book 145 · vig-free 132
· model 110". Also `fair_prob`, `overround`, `devig_method`.
- Suppression reasons ride on refused reads: `suppressed` + `suppressed_reason`
(`juiced_no_edge` / `rare_event_under`) + `reasoning.summary` (branded "No
read" copy) — for the visible-refusal moment (Model Train step 10).
DESIGN: the triplet + VALUE marker + refusal copy are Session-2 design surfaces
(reveal card, board row, hero). Backend ships the fields; rendering is Design's.
## Book Comparison — `/api/books` (per-book price display, snapshot-locked)
Data-only layer for the (still-unrouted) `BookComparison.tsx` grid. Never touches
the grade path.
- `GET /api/books/:sport/:player/:stat?side=over|under` → the honest per-prop
grid: `{ player, stat, line, side, books:[{book,line,over_odds,under_odds,
isBest}], bestBook, bestOdds, bookCount, savings, crowned }`. HONEST-ABSENT:
a single-book prop returns that one book with `isBest:false, bestBook:null,
crowned:false` — never a placeholder second row. `:player` is matched by
`nameKey` (accent/dot/nickname-folded).
- `GET /api/books/:sport?side=&limit=` → "best lines tonight" (`bestLines`, sorted
by savings). This is a CROWN claim, so it returns `[]` while the crown is gated
off. Response `source` = `bookprices` (snapshot-locked store serving) or
`odds-cache` (fallback before first snapshot) — also the deploy fingerprint.
- **The crown is threshold-gated + OFF by default** (`BOOK_CROWN_ENABLED=1` to
enable). Pre-registered gate: median same-line spread ≥8¢ OR ≥2 implied-prob
pts (`scripts/measure-book-spread.js`). Measured 2026-07-27 on the live feed:
MLB median 5¢/0.95pp, WNBA 7¢/1.38pp → NOT met → crown does NOT ship. Render
books WITHOUT crowning one until a future measurement clears the bar.
- Source store: `bookprices:{sport}` (Redis, SNAP_TTL), written by
`snapshotService` from the pre-dedup multi-book `props`. Grouped
`{player, stat_type, books:[...]}`, normalized-name-keyed. DISPLAY-ONLY —
fenced from grade/selector/challenger/ledger.
+19
View File
@@ -24,3 +24,22 @@
**Workaround:** Apply migration via Supabase Dashboard SQL Editor **Workaround:** Apply migration via Supabase Dashboard SQL Editor
**Resolution path:** Either DNS propagates, or add Supabase IP to /etc/hosts, or use `supabase link` with access token via api.supabase.com **Resolution path:** Either DNS propagates, or add Supabase IP to /etc/hosts, or use `supabase link` with access token via api.supabase.com
**Owner:** Kev **Owner:** Kev
## PINNACLE MLB PROP COVERAGE STOPPED — 2026-07-31 (open, external)
**One question to PropLine:** *why did Pinnacle MLB player-prop coverage stop on
2026-07-31?*
**Evidence:** `closing_captures` MLB, pinnacle: 103,940 rows over 07-20 → 07-30,
then **4,022 → 0** on 07-31 and zero since, while every other book continued
normally (07-31: 47,606 non-pinnacle rows). `line_type='sharp'` is a label applied
in `closingCapture.js` via `SHARP_BOOKS` — same feed, not a separate provider.
**Why it matters:** Pinnacle is the only sharp anchor we have ever had for props
(17,090 two-sided captures in that window). Without it the consensus ruler is a
**market** consensus, not a **sharp** one.
**Status:** `market-not-sharp` is **PENDING-RECOVERY, not confirmed permanent.**
Do NOT enshrine it in MASTER-PLAN as a permanent limitation until this is
answered. Not caused by any VYNDR change — the display widening only adds books.
+486 -1
View File
@@ -1,7 +1,455 @@
# VYNDR — Build State # VYNDR — Build State
## Last Updated ## Last Updated
2026-07-12 2026-08-03
## Session 94 (2026-08-04) — Causally-correct platoon + park inputs ✅
4,307 tests / 344 suites green, build exit 0. Counter + frozen clusters
byte-identical.
- **PLATOON SEVERITY built + tested: n=452, 48 short of the gate.**
CANDIDATE_PENDING — not proven, not theatre. Uses each hitter's own vs-LHP/
vs-RHP split, shrunk by the smaller side's PA, refused below 60 PA. It moves
LESS than flat platoon (0.021 vs 0.026), consistent with the pattern.
- **The refusal costs sample honestly** — 452 vs 741 rows is exactly the hitters
whose splits are unreadable.
- **PARK DIMENSIONS INGESTED** (free, statsapi venue endpoint): fence distances,
roof, turf, elevation. Prod-verified 15 venues. Joined by real `venue_id` from
the schedule, never inferred from the home team.
- **RAW WEATHER RETAINED** (temp/wind_mph/wind_dir/precip) — and the old guard
that dropped the environment entirely when the multiplier was 1 is fixed, which
had been discarding the forecast on every ordinary night.
- **PLATOON/PARK ingest prod-verified:** 270 lineups, 265 platoon, 15 park dims.
- **NOT built: park+weather→hit-type.** Its inputs landed this session and carry
ONE as_of date; testing it on total_bases needs accumulated dated rows, so
building it now would be plausibility not proof.
- **Proven for hits: pitcher_contact_profile, defense_by_direction.**
## Session 93 (2026-08-04) — Causally-correct defence atom PROVES ✅
4,297 tests / 342 suites green, build exit 0. Counter + frozen clusters
byte-identical.
- **ATOM 1 BUILT AND PROVEN.** `defense_by_direction` (spray×trajectory ×
positional OAA, joined by handedness) **PROVES** on hits: n=528, Brier 0.0034,
CI [0.0059,0.0009] at 99.9%. Team-average `defense` remains NOT PROVEN.
Kev's insight confirmed: the crude operationalization was the problem.
- **It moves the number LESS (0.013 vs 0.030) and is more reliably right**
bigger movement is often the tell, not the signal.
- **Zero new sourcing**, as predicted: Savant `leaderboard/batted-ball` has
pull/straight/oppo × gb/air (609 hitters, free); per-position OAA is a
regrouping of the fielding feed. Prod-verified: spray 609, team_defense 31.
- **Proven set for hits now: `pitcher_contact_profile` + `defense_by_direction`.**
Both pooled; per-archetype still sample-blocked (BOMBER 225294, GHOST 86125).
- **ATOM 2 INPUT-BLOCKED.** Weather free source PASSES (Open-Meteo wired, exposes
temp/wind_mph/wind_dir/precip) but raw fields are collapsed to a scalar and
`wx_forecast` is 0/1119. **Park dimensions are not ingested at all** — a
hit-type conversion needs them. Not half-built.
- **Next:** retain raw weather fields (cheap), source park dimensions, then build
ATOM 2. Sequence/reliever layer stays gated behind proven atoms.
## Session 92 (2026-08-04) — The two-part factor gate; one factor proves ✅
4,286 tests / 340 suites green, build exit 0. Counter + frozen clusters
byte-identical.
- **`factorGate.js`** — a factor must MOVE the prediction off the base rate AND
improve out-of-sample Brier. Movement alone = **THEATER**, rejected by name.
Baseline is the player's leave-one-out base rate (the literal "he's due" null).
- **Cumulative correction applied to the INTERVAL** (99.9% at 50 tests). This
flipped defense and platoon out of "proves" — a plain 95% CI would have shipped
two unproven factors.
- **NOT_PROVEN_AT_CORRECTED_BAR added** as distinct from THEATER; conflating them
would repeat "insufficient evidence = evidence of absence".
- **RESULT (hits, n=741): `pitcher_contact_profile` PROVES** (Brier 0.0066, CI
[0.0114,0.0016]). defense (0.0043) and platoon (0.0039) NOT_PROVEN at the
corrected bar. park_hits sample-blocked (n=405). **Zero theater.**
- Per-archetype all sample-blocked (BOMBER 252294, GHOST 67125).
- **Spec gaps found:** approach identities (SPRAY/DAMAGE-DEALER/COUNT-WORKER)
don't exist; `parkFactors` has no hits-specific factor (hits → run_base).
- **Grade rescale NOT run** — it was gated on factors proving, and one pooled
factor with a 0.0066 Brier gain is not a factor-informed distribution.
## Session 91 (2026-08-04) — Hits calibrated point-in-time; parlay partially unblocked ✅
4,275 tests / 339 suites green, build exit 0. Counter + frozen clusters
byte-identical (`p_win` untouched).
- **PARTIAL PASS.** Fit <2026-08-02 (n=589) → held-out >= (n=383). Corrected
held-out: 0.477→0.506, 0.587→0.580, 0.667→0.603 (raw was +0.191/+0.279/+0.246).
Ordering preserved. **Certified band 0.400.60 (n=276, err 0.029).**
- **HONEST CEILING 0.667** — no 80%+ hit reads survive calibration. 4-leg ticket
at ceiling = **0.198**, against 0.686 implied by the raw numbers.
- **Banded certification** (`certifyBands`/`inCertifiedBand`) instead of a
blanket flag: middle honest (0.029/+0.007), edges not (0.167/+0.063).
- **`calibrationService`** fits strictly-before-today, splits by TIME, returns
null on thin history (⇒ nothing stackable). Wired into the snapshot: hits
grades carry `p_win_calibrated` + `calibrated`; `p_win` untouched.
- **Parlay surface unblocked for in-band legs only**`chainAcross` compounds
them, cross-game preferred, same-game a labelled approximation.
- Caught: a pass condition that demanded ≥0.70 bins exist would have failed the
map for succeeding (calibration removes that band).
## Session 90 (2026-08-04) — chaining-v1: portable chain + the calibration gate ✅
4,269 tests / 339 suites green, build exit 0. Counter + frozen clusters
byte-identical.
- **HIT-PARLAY BLOCKED (the order's own prerequisite, failed decisively).**
n=972: predicted 0.911 → actual 0.630; flat ~63% above 0.70. A 4-leg 91%
ticket is 0.686 by the model, 0.157 in fact. `chainAcross` refuses
uncalibrated atoms — verified end-to-end on real data.
- **`calibration.js`** — reliability table, `isCalibrated` gate (tolerance 0.05,
high-end weighted), and `fitIsotonic` (monotone: ordering preserved, numbers
corrected). Real map: 0.65→0.594, 0.85→0.639, 0.91→0.639.
- **`chain.js`** — the portable core. Aggregator pluggable (ACROSS=parlay,
UP=score), sport parts as inputs, archetype-`redistribute` hook (dormant in
baseball, live in basketball). Self-check flags internal inconsistency as LOW
CONFIDENCE; market divergence flags a contested script WITHOUT claiming we're
right. `propagate` is shrinkage-weighted by sample.
- **NOT built this order:** the independent game-script projection. It needs
proven team-level atoms + out-of-sample validation vs actual margins, and no
atom has passed the gate yet — building it now would be plausibility, not proof.
- **Next:** fit the point-in-time isotonic map and re-gate hits; that is the
unlock for the parlay surface.
## Session 89 (2026-08-04) — Lineup + baserunner context ingested ✅
Spec: commit history + `src/services/lineupContextService.js`. 4,250 tests / 338
suites green, build exit 0. Counter + frozen clusters byte-identical.
- **THE INPUT RBI/RUNS ALWAYS NEEDED, ingested free from statsapi.**
`lineup_context` (batting order, 153 rows / 10 games) + `hitter_opportunity`
(RISP share, 149 rows, range 0.1700.528). Both DATED in the PK.
- **Rung 2 was cheap:** situational splits give the RISP aggregate in ONE call
per player, not play-by-play reconstruction.
- **Prod-verified** via the new `POST /api/internal/lineup-context/refresh`
(added because the first run wrote 0 while the parser worked locally — a 0 is
a wiring bug until proven otherwise).
- **Coherence check passes:** top RISP-share hitters all bat 4th/5th.
- DRIVER/CATALYST theories now INPUT-READY, sample-blocked. Proofs run later
under native cumulative correction — ingesting is not proving.
## Session 88 (2026-08-04) — Re-adjudication: nothing to demote, hole closed ✅
Spec: `specs/re-adjudication.md`. 4,238 tests / 337 suites green, build exit 0.
Counter byte-identical. Nothing recalibrated — nothing needed to be.
- **PROVEN SET IS EMPTY, verified 3 ways** (proven-status, featureRegistry
summary, validatedSkills). Zero conditioning entries ever reached PROVEN, so
STEP 3 (demote) and STEP 4 (recalibrate) are vacuous — correctly.
- **Correction: the cumulative gate did NOT catch a false positive last session.**
It caught nothing; it tightened α 0.0026 → 0.0013, demonstrating the mechanism.
- **THE REAL HOLE, CLOSED:** `promote()` could bypass cumulative correction.
`isSufficient` now requires `bonferroni_tests`, refuses anything below the
cumulative count, and refuses a p that doesn't clear 0.05/tests. Same guard on
`recordConditioning(PROVEN)`. Verified: no-correction / per-session-8-vs-38 /
weak-p all refused; cumulative-38 with p=0.0005 accepted.
- **Cumulative correction now NATIVE on all analysis paths** — pitcher-prove-k
and tb-solo-and-interactions migrated off per-session counts.
- **`reAblation.js` built** (standing second line): pure/injectable, records both
p-values + both test counts per verdict, `PENDING_RETEST` when there is no
fresh measurement (absence is not evidence).
- **Net effect on the proven set: ZERO.** No demotions, no recalibrations, no
ledger event — announcing a recalibration that changed nothing would itself be
a false signal of rigour.
## Session 87 (2026-08-03) — Defence ingested; cumulative correction locked ✅
Spec: `specs/defense-ingest-and-cumulative-correction.md`. 4,228 tests / 336
suites green, build exit 0. Counter + clusters byte-identical.
- **DEFENCE INGESTED (free):** Statcast OAA feed → 514 fielders → `team_defense`
(31 teams, dated from row one). Prod-verified: fielding rows 514,
team_defense_written 31. Cubs +56 best, Mariners 29 worst.
- **CUMULATIVE BONFERRONI LOCKED** (`testLedger.js` + `mc_test_ledger`): the
denominator is now distinct hypotheses across the programme lifetime, not the
session. Demonstrated 19 → 38, α 0.0026 → 0.0013. Re-tests don't inflate it.
- **THE PREDICTED DIFFERENTIAL APPEARS:** defence solo r = **+0.130 GHOST**
(contact/speed) vs **0.018 BOMBER** (power). Exactly "defence matters, and for
whom". Both UNDERPOWERED (n=104/245, p=0.188 vs α=0.0013) — signal shape only.
- **Bug class recorded:** the feed 404'd on a doubled `/leaderboard` path and,
because feeds degrade to an empty index by design, reported "0 rows" — which
reads like an honest absence. Any feed reporting 0 is suspect.
- **Nothing proved → nothing recalibrated, nothing shipped.** validatedSkills()
is {} everywhere; proven set still EMPTY.
- **Next:** sample only. GHOST×hits needs ~396 more rows, BOMBER×hits ~213.
Prefer re-testing standing candidates — every new hypothesis now tightens α
for everything after it.
## Session 86 (2026-08-03) — Conditioning registry + a probe so "proven" stops drifting ✅
Spec: `specs/conditioning-registry.md`. 4,221 tests / 335 suites green, build exit
0. Counter + batter model + pitcher engine byte-identical.
- **`scripts/proven-status.js`** recomputes the proven set from the ledger.
PROVEN_SET = **EMPTY**. Built because four consecutive orders opened by calling
null results proven; prose decays, a recomputed number does not.
- **COUNTING BUG CAUGHT:** joining model_snapshots to ledger_entries fans out
(one snapshot row per cycle) — BOMBER x hits read 641, true distinct 287.
Fixed in both the analysis and the status probe.
- **NO archetype x stat reaches the gate.** BOMBER x hits 287 (short 213) is
closest; pitcher archetypes untestable (58 settled Ks total).
- **Structured registry built:** `recordConditioning` keys archetype x SKILL x
interaction x status + lift, with the skill tag ENFORCED (untagged refused,
PROVEN-without-evidence refused). `validatedSkills()` = {} everywhere, by design.
- **BOMBER x hits conditioning tested, all UNDERPOWERED:** arsenal (barrel x
breaking share) incr +0.043, batted-ball (launch x pitcher GB) +0.001, contact
quality 0.020/0.015, K x K 0.063. Within BOMBER the counter still leads
(0.218 vs 0.160).
- **Bug fixed mid-run:** `fromStatcastRow` doesn't carry pitch_mix, so the arsenal
category read n=0 — it was measuring nothing, not failing.
- **DEFENSE: genuinely not derivable** from ingested data (no OAA/DRS; pitching
proxies conflate skills). Needs Savant's free fielding feed — not sourced,
because sourcing it to test at n=282 answers nothing.
- **Nothing proved → nothing recalibrated, nothing shipped.**
## Session 85 (2026-08-03) — Rung 1 derived free; the cap fix fingerprinted ✅
Spec: `specs/lineup-k-rate-rung1.md`. 4,221 tests / 335 suites green, build exit
0. Counter + batter cluster + pitcher engine byte-identical.
- **CAP FIX VERIFIED IN PROD: 334 -> 907 grades/snapshot; strikeouts 6 -> 17.**
n>=500 for Ks is ~a week out instead of ~3 months. NOTE: the manual internal
snapshot endpoint now 524s at Cloudflare (>100s) but COMPLETES server-side.
- **RUNG 1 DERIVED, zero new sourcing:** opposing-team K-rate from the roster
joined to batter k_pct we already ingest, 94.7% coverage — now PA-WEIGHTED.
That change flipped its contribution: unweighted HURT (0.174->0.129),
PA-weighted HELPS (0.174->0.195). Head-to-head delta +0.259, CI
[0.0167,+0.5645] — nearly excluding zero, still INCONCLUSIVE at n=57.
- **Within-archetype:** FLAME incremental 0.152, non-FLAME +0.145 — opposite
signs, invisible when pooled (+0.077). But n=20/24 and the direction
contradicts the theory. Structure to re-test, not a finding.
- **Rungs 2/3 NOT triggered** — Rung 1 is n-blocked, not failed. Do not source
confirmed lineups.
- **Nothing proven, nothing calibrated, nothing shipped.** Counter is still
anti-predictive on Ks (0.064); skill model leads by 0.26.
- **Next:** wait ~1 week for n>=500 + a statcast_history window, re-run, re-test
the strata at ~200/stratum, and give arm_angle a registry entry + mechanism.
## Session 84 (2026-08-03) — Pitcher engine built; the cap was eating the board ✅
Spec: `specs/pitcher-engine-strikeouts.md`. 4,221 tests / 335 suites green, build
exit 0. Batter model + counter byte-identical (verified by diff).
- **THE REAL FIND: the grade cap, not pitcher data.** 1,244 unique gradeable
props/slate; the 500 cap graded ~334, and first-row-wins-in-feed-order gave
pitchers 6 props a slate. Raised 500 -> 1500 on measured cost (~179s for the
full board at concurrency 5, cron 5x/day). Unblocks EVERY n-blocked stat.
Pitcher props were never being refused (graded 5, refused 0, suppressed 0).
- **`pitcherEngine.js` — own archetypes (FLAME/SCALPEL/SINKER/DEFAULT), own
inputs (stuff), own projection** (log5 K% vs THIS lineup x batters faced).
Test asserts its weight keys differ from the batter engine's. 17 tests.
- **Strikeouts NOT proven** (n=57 vs 500): pitch-v1 0.1285 vs counter 0.0639,
delta +0.192 CI [0.098,+0.509]. Four solo features clear the |r|>=0.15 bar and
fail only on n — arm_angle 0.250 (largest in the programme), whiff +0.213,
k_pct +0.206, chase +0.195.
- **The counter is ANTI-PREDICTIVE on Ks (0.064)** — recent K counts track
opponent and workload, not skill.
- **Bug caught:** `resolveTeam` needs an abbreviation; the game log gives names,
so lineup coverage was 0% and the theorized carrier was never tested. Fixed via
NAME_TO_ABBR → 94.7%. The carrier still shows no incremental signal (n=54).
- **Calibration not reached** — nothing passed BAR 1.
- **Next:** let the cap accrue (~2 weeks to n>=500), re-run with a point-in-time
window from statcast_history; give arm_angle a registry entry + mechanism.
## Session 83 (2026-08-03) — Batter cluster measured; the proven set is EMPTY ✅
Spec: `specs/batter-cluster-prove.md`. 4,204 tests / 334 suites green, build exit
0. skillProjection byte-identical (TB frozen, verified by diff); counter untouched.
- **PREMISE CORRECTED: total_bases has NOT passed BAR 1.** It is inconclusive at
parity (CI includes zero) and contaminated. Installing it as the "proven
reference" would make the cluster's bar "be inconclusive at parity".
- **HITS CLOSED — well-powered negative.** n=803 CLEARS the gate sample bar, so
features were tested not refused: max |r| 0.053, interactions ≈0, head-to-head
0.096 CI [0.165,0.029].
- **Others n-blocked:** TB 383, rbi 391, HR 228, runs 188. Leads: home_runs
barrel r=0.135 (negative = a correction, not a predictor); runs K×K
incremental +0.132 (largest in cluster).
- **RBI is half-unmodellable** — power × opportunity, and baserunner state is not
ingested at all.
- **statcast_history retention LIVE + verified in prod** (1,387 rows, as_of
2026-08-03). First run failed on a drifted hand-written schema; table now
mirrors the source via LIKE. Usable point-in-time window starts 2026-08-04.
- **Stage B has nothing to calibrate.** Proven set is empty.
- **Next (waiting, not building):** let history accrue a week + TB/rbi reach
n>=500, then re-run `scripts/cluster-prove.js`. Ranked: TB → rbi (needs
baserunner state) → HR → runs. Do not re-run hits.
## Session 82 (2026-08-03) — TB solo+interactions; point-in-time validation unblocked ✅
Spec: `specs/tb-solo-and-interactions.md`. 4,204 tests / 334 suites green, build
exit 0, counter byte-identical.
- **BLOCKER FOUND + FIXED FORWARD:** `statcast_aggregates` keeps ONE as-of date
(upsert in place). Yesterday's backtest was clean only because the refresh was
dead code and the table sat at 2026-07-21; fixing the cron destroyed the
window. New `statcast_history` table + retention on every refresh (best-effort,
never fails the refresh). Until it accrues, all skill results are CONTAMINATED.
- **SOLO (n=383, Bonferroni-12): nothing passes.** hard_hit_pct marginal r=0.135
(p=0.0080) fails both the 0.15 bar and α=0.00417 — and DRIFTED DOWN from 0.153
at n=295. Everything else <0.09.
- **INTERACTIONS: none pass.** barrel×power_archetype is the only one whose
incremental partial (0.101) exceeds its parts (0.019), at n=260. A lead.
- **Caught a fabricated finding:** the archetype proxy was a transform of barrel
itself, so the "interaction" was barrel² — it produced the only positive result
until a scale-free collinearity check + real `model_snapshots.archetype` labels
replaced it.
- **COMBINED vs COUNTER on TB: 0.2718 vs 0.2647, delta +0.0071, INCONCLUSIVE**
the first challenger that did not LOSE (hits was 0.116, CI excluding zero).
- **BUILT: compound TB projection** (per-PA bases convolution, barrel→HR share,
exit velo→XBH share). Replaces the refusal; non-degeneracy locked by test.
- **Next:** let statcast_history accrue a point-in-time window (~a week) while TB
reaches n>=500 (~117 short), then re-run. Do not re-run hits.
## Session 81 (2026-08-03) — The gate, built and run: hits is dead, total bases is the stat ✅
Spec: `specs/stagea-gate-result.md`. 4,200 tests / 334 suites green, build exit 0.
Counter byte-identical (zero diff on probabilityEstimator/analyzeViaEngine1).
- **PREMISE CORRECTED:** statModel.js and correlateValidator.js do NOT exist in
this repo. The spec lived only in an offline Python blueprint, and
supplementSystems.test.js inlines its own validateFactor (requires just
fs/path). Nothing to connect — so the gate was BUILT to spec.
- **`correlateValidator.js`** — n>=500, |r|>=0.15, p<0.05, Bonferroni. Exact
p-value (incomplete beta), unit-verified against known values.
- **GATE RUN, hits (n=570, Bonferroni-8): EVERYTHING FAILS.** Max marginal |r|
0.062 vs the 0.15 bar — an effect-size failure at a well-powered n. Head-to-head
also loses: 0.0499 vs counter 0.166, delta 0.116 CI [0.189,0.043].
- **GATE RUN, total_bases (n=295): CANNOT TEST — and that is the finding.**
hard_hit_pct marginal r=0.153 (above threshold), exit_velo 0.124; refused only
on n. ~205 more settled rows needed. Matches the physics: contact quality
drives extra bases, not singles.
- **Architecture change the run forced:** per-STAT feature verdicts, so a feature
dead for hits stays alive for TB. Gate now reports r/p when underpowered.
- **Next:** build the compound TB value projection (per-hit bases distribution
from launch/barrel — skillProjection still refuses TB by design), accrue to
n>=500, re-run the gate. Leave hits alone. Do not lower the bar.
## Session 80 (2026-08-03) — The skill engine: built, gated, and Stage A honestly lost ✅
Spec: `specs/skill-engine-architecture.md`. 4,182 tests / 333 suites green, build exit 0.
- **BUILT `src/services/model/`:** `featureRegistry` (CANDIDATE/PROVEN/DEAD per
sport; `liveFeatures()` = PROVEN only; promotion needs n>=200 + positive lift +
CI excluding zero, no override) and `skillProjection` (PA outcome tree, log5
odds-ratio K/BB, archetype-selected contact quality, Binomial over a PA
distribution). 22 tests assert the five disciplines as BEHAVIOUR.
- **The gate works by construction:** with only PROVEN features allowed the
projection returns NULL. Registry ships with ONE proven feature (the counter).
- **STAGE A: skill-v1 LOSES → NOT PROMOTED.** Out-of-sample (profiles frozen
07-21, only later games scored), 570 rows, 91.9% pitcher coverage: resolution
0.0499 vs champion 0.166, delta 0.116 CI [0.189,0.043]. Not selective
either (top-8 hit 50%, lift 0.065).
- **Two false starts caught:** (1) units — statcast stores PERCENTAGES, raw rows
made bip negative and refused 568/576; now one chokepoint `fromStatcastRow`.
(2) an INVALID first verdict — ledger team/opponent are NULL, so the pitcher
resolved for 1 of 570 rows and it was silently measuring a batter-only model.
Fixed via each player's statsapi game log.
- **Not exercised yet (so the loss is real but partial):** park (passed 1.0),
handedness, opportunity_drift, and PA projection is season-PA/103. And the
skill profiles carry NO recency while the champion has last-5.
- **Fixed: Statcast nightly refresh was UNREACHABLE CODE** — inside tick() below
the HOURS_UTC guard while testing h===11. Never ran; 13 days stale; both alerts
in the same dead branch. Now its own tick; test rewritten to catch it.
- **Next:** recency into the skill profile, wire park/handedness/opportunity_drift,
real PA from lineup slot, then re-run Stage A.
## Session 79 (2026-08-03) — Reality assessment vs the FORWARD-PROJECTION objective ✅
Spec: `specs/forward-model-reality-assessment.md`. READ-ONLY (src/web untouched).
- **Finding: the forward model's parts all EXIST and are all wired downstream of
the grade.** `probabilityEstimator` (the served p_win) reads 3 features + the
game log. Statcast/arsenal/park/weather/platoon/archetype load in
`snapshotService` AFTER grading, into challenger columns nothing serves.
`mlbContext` has zero consumers.
- **Statcast nightly refresh is DEAD CODE by guard** — tick() returns for hours
not in HOURS_UTC (14,19,22,1,3); the block tests h===11. Data frozen at
2026-07-21 (13 days stale); its own failure alert is in the same dead branch.
- **Inputs are HAVE** — 1,354 statcast rows, handedness complete both sides,
pitch mix/velo/break, GB/FB, barrel, exit velo, launch. MISSING: team defense
(OAA/DRS), catcher framing/umpire. PARTIAL: batter GB/FB (in `metrics` JSONB),
lineup slot (role tables 0 rows).
- **Design shows the COUNTER.** Factor labels are all `l5_hot_vs_line`-family
plus NBA leftovers (refs, coach pace). The card's forward-read slots
(archetypeBlend "Why this grade", vyndrIntel.matchup, propDNA) exist and go
unfilled. Needs feeding, not redesign.
- **STAGED DISTANCE:** Stage A (forward baseball model) = ONE real build, ZERO
data acquisitions — assemble hitter profile × pitcher stuff × conditions as the
SPINE with frequency demoted to a prior; risk is sample, not feasibility.
Stage B (calibrated + scouting surface) = short once A exists (clamp/calibration
already diagnosed + swap the factor vocabulary). Stage C (per sport) = blocked
on mechanism data we do not have for NBA/WNBA (ESPN is box scores, Python
service offline) and soccer is odds-api quota-blocked.
- **Verdict re-checks:** proj-v1.1 + hits-v1 stay refuted AS DISTRIBUTION SWAPS
(neither tested a matchup-fed projection); arch-v1 is market-relative by
construction = the one measured on the wrong axis; "AT CEILING" is provisional.
## Session 78 (2026-08-03) — Champion decomposed: the edge is a hit-rate counter ✅
Spec: `specs/champion-input-diagnosis.md`. READ-ONLY (src/web untouched);
4,159 tests green.
- **The champion is 5 lines.** base = empirical frequency of (stat > THIS line),
0.6/0.4 blend with last-5, ±0.03 opponent, ±0.015 home/away, cv>0.40 pull,
clamp [0.10,0.95]. It reads 3 features; featureCache retains a dozen more that
p_win never touches.
- **Exact analytic ablation, per stat, paired bootstrap.** Removing ALL THREE
adjustments changes resolution by nothing everywhere (hits 0.0059, TB 0.0015,
rbi +0.0106, runs +0.0130, walks +0.0008) — and rbi's home/away is mildly
HARMFUL (+0.0053, CI excludes 0). ~100% of the edge is base+recency.
- **Pooled 0.46 is an artifact** — per stat 0.196 (hits) … 0.499 (rbi). Corrected
last session's reading; paired differences unaffected.
- **BIGGEST LOSS = the clamp.** 20.6% of settled rows pinned to a constant (no
ranking possible there), and `0.900` covers home_runs-under truly 99.5% AND
hits-under truly 51.9%. Global over-prediction +3.5pt (TB +7.6). No new data
needed to fix.
- **One real lead: `opportunity_drift`** (residual +0.156 hits, +0.145 TB —
repeats across stats, unlike the weather hits which sit inside the expected
false-positive count). We ALREADY compute it; arch-v1's opportunity axis
extracts nothing from it. Wrong implementation, not a missing feature.
- **Archetype: UNMEASURABLE** — 2 of 41 labels have testable n. Not refuted.
- **Next order priority:** (1) clamp + calibration, (2) opportunity as a rate
scaler, (3) prune the diluting axes, (4) get archetype coverage. Explicitly NOT
another projection variant.
- Flagged: `model_snapshots.outcome` NULL on all 22,032 rows — retention is
never settled, so replays must join the ledger for labels.
## Session 77 (2026-08-03) — Settlement was dead for two days; scoreboard now readable ✅
Specs: `specs/challenger-scoreboard.md`, `specs/odds-429-diagnosis.md`.
4,159 tests / 332 suites green, web build exit 0.
- **THE FIND.** Three challenger axes read exactly ZERO settled rows. Not low —
zero, on games played days earlier, with `settle_attempts = 0`. `settleLedger`
refetched rows via `.in('id', ids)`; 500 UUIDs = an 18,499-char URL the fetch
layer rejects, and the result was destructured with no error binding, so it
returned all-zeros indistinguishable from a clean "nothing to settle".
Volume-triggered: 2026-08-01 was the first day past the 500-row limit.
The zero-settle ops alarm reads those same return values and was blind to it.
- **FIXED + DRAINED.** One query, all columns, no id list; failed fetches surface.
`captureClosing` chunked at 100 (same defect family). 1,444 rows from 08-01
settled (1,376 hit/miss + 68 void, 0 remaining). Settled n **493 → 1,741**.
- **SCOREBOARD — nothing promoted, nothing earned it.** arch-v1 n=1,741 Δ0.0000
CI[0.0050,+0.0054] (moves 76% of rows by 2.5pp mean = active movement carrying
no information); contact-v1 n=1,055 +0.0008 inconclusive; proj-v1.1 ladder
n=1,664 **0.0301 CI[0.0543,0.0060] = reliably WORSE**. matchup/tb-v1/hits-v1
STILL PENDING (rows dated 08-02+, settle after ET midnight). Champion
byte-identical; all challengers stay wired.
- **429 DIAGNOSED (read-only) — premise refuted with numbers.** PropLine 5/3,000
daily (0.17%); the 429 is **odds-api at 478/500 monthly, blocked at 95%**,
surfacing whenever PropLine returns empty. One snapshot = ONE PropLine call per
sport. Change-based pull is NOT the fix and no tier upgrade is needed. Could
NOT verify PropLine movement endpoints (auth-gated docs, prod-only keys) — not
asserted. Book-breadth invariant recorded: we never discard books; DFS is
excluded from PRICING only.
- **Next:** the silent PropLine fall-through (empty slate must not report the
backup's 429); diagnose the projection family's INPUTS (two independent
measurements now say it trails the champion).
## Session 76 (2026-08-02) — hits-v1: a challenger built, measured, and REFUTED ✅
Spec: `specs/hits-v1-binomial.md`. 4,156 tests / 332 suites green, web build exit 0.
Scope was hits only; champion, ladder, ranking, calibration, reference ruler and
the four accruing challenger verdicts are byte-identical (the diff has ZERO
deleted lines).
- **What was built.** `src/services/projection/binomialHits.js` — hits as a
bounded conversion: `N ~ the player's empirical at-bat distribution`,
`hits | N ~ Binomial(N, q)`. At the 0.5 line (84% of real hits rows) this
states `P(>=1) = 1 E[(1q)^N]` directly instead of inferring P(0) from a
count family. The multiplier scales `q` (conversion), never `N` (opportunity).
Wired in `projectionChallenger` as `proj_hits_p_over` / `proj_hits_meta`
(new ledger columns, migration applied).
- **STEP 0 first — inputs before model.** `scripts/hits-input-coverage.js`:
30/30 real ledger players, 100% combined-input coverage, mean 3.518 AB/G,
mean per-AB rate 0.248.
- **FIRING, on the real board.** `scripts/verify-hits-v1.js` runs the production
`attachProjection` over the live prod snapshot: 158/159 hits props (99.4%), one
honest abstention. 94 of 159 props sit OUTSIDE the promotion band and 93 were
modelled anyway — scoping by book identity kept 59% of the board a price-shape
rule would have deleted.
- **AND IT LOST.** Point-in-time replay (log truncated strictly before each row's
game_date, real grade-time multiplier), hits-only, direction-aligned, n=242:
resolution champion **0.195** / ladder **0.048** / hits-v1 **0.026**. Paired
bootstrap: hits-v1 ladder = 0.022, CI95 [0.046, 0.0003]. NOT PROMOTED.
- **The finding is what it eliminates.** Family was wrong AND mean was not the
constraint (hits-v1 moved the line-0.5 mean 0.554→0.581 toward a 0.598 base
rate while resolution FELL). The hits deficit is per-prop DISCRIMINATION — the
ladder's inputs, not its distribution.
- **A pre-registered branch recorded as WRONG.** The spec's fallback ("hits may
be genuinely low-resolution for anyone") is refuted by the champion scoring
0.276 on the identical 189 rows. Kept in the doc rather than deleted.
- **Next order is a DIAGNOSIS, not a model:** what does the champion's `p_win`
read on a hits prop that the projection ladder does not? Building another
projection variant first would repeat this session's mistake.
- Logged not fixed: local `.env` has a transposed Supabase ref — local scripts
need `SUPABASE_URL=` override; prod unaffected.
## Session S11 (a1 board, 2026-07-12) — Live Tracking: the read locked, the game watched ✅ ## Session S11 (a1 board, 2026-07-12) — Live Tracking: the read locked, the game watched ✅
Spec: `specs/LIVE-TRACKING.md` (+ ROW-GRAMMAR §2/§3 S11 amendment). Spec: `specs/LIVE-TRACKING.md` (+ ROW-GRAMMAR §2/§3 S11 amendment).
@@ -4685,3 +5133,40 @@ Complete frontend overhaul. 18 pages, 22 API routes. `npm run build` passes with
- Timestamped records in evolution_detections table, Evolution Watch content formatter - Timestamped records in evolution_detections table, Evolution Watch content formatter
- **Migration 008:** coaching_tendencies, player_out_history, evolution_detections, unconventional_validations (all with indexes + RLS) - **Migration 008:** coaching_tendencies, player_out_history, evolution_detections, unconventional_validations (all with indexes + RLS)
- **Integration:** 3 new blueprints registered in app.py (coaching_bp, redistribution_bp, unconventional_bp), evolution + odds_scanner extended with new endpoints - **Integration:** 3 new blueprints registered in app.py (coaching_bp, redistribution_bp, unconventional_bp), evolution + odds_scanner extended with new endpoints
---
## Session — Under-querying vs out of data (2026-08-05)
**Shipped**
- `scripts/backfill-context.js` — platoon splits backfilled to all 380 settled
hitters (was 298; ingest had only ever seen tonight's lineups).
- `scripts/reconstruct-game-environment.js` — joins the ledger's game slug to
statsapi, writes `game_context` on the LEDGER's key, pulls actual archived
Open-Meteo weather. 96/101 settled games now carry real weather; park
dimensions 15 -> 30 venues.
- `src/services/model/parkWeather.js` (+ tests) — park geometry + air read onto
HIT TYPE, not P(hit). Wind refused (no park orientation).
- `factorGate` — cluster-aware bootstrap + `effective_n`. Unclustered rows keep
the original path byte-for-byte.
- `specs/under-querying-vs-out-of-data.md` — the full record.
**Decided**
- `pitcher_contact_profile` DEMOTED: point estimate halved on 2x sample, interval
now spans zero.
- `platoon` / `platoon_severity` clear the bar but are NOT promoted — 4.5% median
contamination (season-to-date splits contain the games they predict) on a
-0.0001 bound.
- Park+weather on total_bases: 47 clusters < 500, point estimate WORSE (+0.0011).
**Next**
- Weather is reachable: ~57 days at 7 games/settled-day.
- Park geometry is NOT — 30 ballparks exist, so a venue-constant factor can never
reach 500 independent units. Needs a hierarchical fixed-effect treatment or it
stays unvalidatable.
- Point-in-time platoon splits would settle the two held passes.
- Park orientation is the one column that would unlock wind.
**Blocker**
- `git push` has no credentials in this environment; commit `7b85934` is local.
+751 -3
View File
@@ -1,5 +1,49 @@
# VYNDR — Claude Code Project Context # VYNDR — Claude Code Project Context
---
# 🔷 PRODUCT IDENTITY — READ FIRST, EVERY SESSION
**VYNDR IS A PREDICTIVE MODEL.** It projects what a player will **DO**, and picks
accurately. It reads and pulls the market apart — a student of the game that is
also an aggregator.
**Market edge is a BYPRODUCT of a good prediction. It is NEVER the success
criterion.**
> **SUCCESS = the forecast is honest about its own confidence AND still ranks.**
> Calibration (does 60% mean 60%?) *and* resolution (do higher forecasts actually
> hit more often?). Both, or it isn't working.
**No edge or CLV term belongs in a pass/fail gate.** CLV and market-relative edge
are *diagnostics we report*, never thresholds a model must clear to ship. A model
that forecasts honestly and ranks correctly is working even in a week the market
moved against it; a model tuned to beat a closing line has been fitted to the
market instead of to the game.
## PER-SPORT DOCTRINE
*(Rashad Phillips, "Basketball Position Metric," 2022 — classify players by **what
they do**, not by position labels.)*
**Each sport is its OWN model** — its own variables, archetypes, conditions,
calibration and honest ceiling. The **only** thing shared across sports is the
**Bayesian inference math**. Never one model fit to all sports; never a sport
stubbed in on another sport's template and counted as covered.
## TRUTH LAW
- **No fabricated data anywhere.** If it renders a number, it comes from the
database or it doesn't render. `Number(null) === 0` is the classic breach.
- **Honest-absent beats invented.** An empty state is a valid answer.
- **Label limitations in-band** — e.g. "market consensus, **not sharp**",
"RECORD BUILDING", "MODEL · LEARNING".
- **Provisional results stay provisional until re-run.** A measurement taken
against an instrument that has since changed is not a result; it is a result
*pending*.
- **Verify by inducing the real code path on demand** — never wait on a cron slot
to find out whether something works. Documented ≠ verified.
---
## What This Is ## What This Is
Sports betting intelligence SaaS. Real software product. Sports betting intelligence SaaS. Real software product.
Three tiers: Free (5 scans), Analyst ($19.99 / $14.99 founder), Desk ($49.99 / $34.99 founder). Three tiers: Free (5 scans), Analyst ($19.99 / $14.99 founder), Desk ($49.99 / $34.99 founder).
@@ -672,9 +716,11 @@ snapshot, locked to the line, and read from cache.
props. It's `rbi` now; a normalizer test locks it. props. It's `rbi` now; a normalizer test locks it.
- **The streaks/hotlist path uses its own `rbis` key** built from raw MLB stats — - **The streaks/hotlist path uses its own `rbis` key** built from raw MLB stats —
independent of the odds normalizer. Don't "unify" them; the split is intentional. independent of the odds normalizer. Don't "unify" them; the split is intentional.
- **MLB is the ONLY end-to-end-live sport.** Outcome settlement is MLB-only - **MLB and WNBA both settle end-to-end.** CORRECTED 2026-07-26 (was "MLB-only"):
(WNBA/NBA/soccer grades never settle → `accuracy` reflects MLB only). Fixing `outcomeService.SPORTS` includes wnba, which settles via ESPN box scores
that (ESPN box-score settle path) is roadmap Session 57. (`espnStatsAdapter` box-field map) — verified: 376 WNBA rows settled, settle
logic spot-checked correct. NBA/soccer still do not settle (no free settled
feed wired). So `accuracy` reflects MLB + WNBA, not MLB only.
- **`src/utils/opsNotify.js`** pushes pipeline alerts to ntfy (`vyndr-pipeline- - **`src/utils/opsNotify.js`** pushes pipeline alerts to ntfy (`vyndr-pipeline-
kev2026`). It NEVER throws and is auto-disabled under `NODE_ENV==='test'` / kev2026`). It NEVER throws and is auto-disabled under `NODE_ENV==='test'` /
`PIPELINE_ALERTS=0` (inject `fetchImpl` to test it). `snapshotService` alerts on `PIPELINE_ALERTS=0` (inject `fetchImpl` to test it). `snapshotService` alerts on
@@ -1001,6 +1047,708 @@ phased plan in the Session-57 conversation / BUILD-STATE Next section).
"TRACKING — READ LOCKED PRE-GAME" renders once per live card (GameCard, "TRACKING — READ LOCKED PRE-GAME" renders once per live card (GameCard,
dim — it's meta, not a caution signal). dim — it's meta, not a caution signal).
## Probability Layer + Grade Range (Session 63 — non-obvious)
- **`gameLogService.getGameLogs` is a TRAP: it returns null for MLB by
construction** (`pythonPath` `default: return null`) and depends on the Python
service, which is OFFLINE in prod. Anything wired to it is dead. S46 fixed this
for FEATURES (`featureCache.gameLogFeatures` MLB branch) but NOT for the
estimator — so `meta.gameLogs` was `[]` for every sport and `p_win`, `ev_pct`,
`kelly`, `model_odds`, `value` were absent on 100% of live grades for months.
**`featureCache.getStatRows(player, sport, statType)` is now the one true source
of normalized per-game rows** (`[{date, [statType]: v}]`, MOST-RECENT-FIRST —
the estimator treats `slice(0,5)` as the recency window). Use it; never add a
new caller of `gameLogService` directly.
- **Hero v2 requires a finite `ev_pct`** — when EV was dead it matched nothing and
fell through to the recent-read fallback silently (`is_recent:true` was the
tell). A "working" endpoint returning data is not proof the intended rule ran.
- **`confidence` is NOT a probability.** engine1 picks a letter from an additive
factor index, then reads that letter's band MIDPOINT out of
`grade_thresholds.json` to make the number — so it carries zero information
beyond the letter and can never disagree with it. Payloads carry
`confidence_basis: 'grade_band'`. The real signal is `p_win`. Corollary:
mlb-grade-degradation.md's "25/25 grade<->confidence agreement" is a TAUTOLOGY,
not a validation (corrected in that file) — never cite it as grade quality.
- **`grade_thresholds.json` is NOT an input mapper in the JS path** — only the
Python side compares scores to it. In JS it is a confidence lookup table read
BACKWARDS from the already-chosen letter.
- **The grade is an integer index** (`GRADE_SCALE`, `NEUTRAL_INDEX` 3) moved by
flat +/-1.0 and +/-0.5 deltas. A needs sum >= +4.5, D needs <= -1.51. Six
factors were wired to features nothing populated, pinning the live range to
{C,B} — only TWO letters ever emitted across 604 ledger rows.
**`refreshTeamStats` had ZERO production callers**, so `opp_rank_stat` (a +/-1.0)
was permanently null; it is now called in `runSnapshot` (test-env no-op, the
opsNotify precedent). L20 was asymmetric (both branches +1.0 = no downside path)
and is now symmetric.
- **Consistency CV is scale-dependent — this is a live landmine.** The thresholds
are NBA-tuned (points ~20/gm). For a Poisson-ish stat `cv ~ 1/sqrt(mean)`, so
ANY stat with mean < 4 auto-classifies `boom_bust` (real: Alonso hits mean 0.60
-> cv 1.17). Reviving consistency without a guard stamps a blanket -1.0 on
nearly every MLB prop. Floored at `CONSISTENCY_MIN_MEAN` (4) -> `unknown` below.
The scale-free fix is an index-of-dispersion classifier (open item).
- **NEVER rescale thresholds to make A's appear** (founder ruling, permanent).
Minting A's without new information is a relabelled B sold as an A and it
corrupts an append-only ledger. Fix the grade on MERIT or don't claim the scale.
- **A-RATED copy is on hold** until a prod fingerprint shows real A grades.
`/api/ledger/accuracy` currently returns B and C buckets only, so AccuracyBadge
correctly falls through to "MODEL · X% HIT" and TopSignals self-hides.
- `scripts/verify-grade-range.js` replays live-board props through the real engine
on free feeds. It UNDERSTATES range locally (no Redis -> no `opp_rank_stat`).
Redis runs degraded locally, so the script must `process.exit(0)` — otherwise a
reconnect timer holds the process open and piped output is lost to SIGTERM.
## hits-v1 — a REFUTED challenger, and why it stays (Session 76 — non-obvious)
- **`specs/hits-v1-binomial.md` is the record.** hits-v1 models hits as a
binomial over the player's EMPIRICAL at-bat distribution (P(0) stated directly,
since 84% of hits rows trade at 0.5). It FIRES at 99.4% on the live board and
it DOES NOT WORK: point-in-time replay, hits-only, direction-aligned, n=242 —
resolution champion 0.195 / ladder 0.048 / hits-v1 0.026. Paired bootstrap
(same rows) puts hits-v1 ladder at 0.022, CI95 excluding 0. NOT PROMOTED.
- **Two explanations are now ELIMINATED for hits, which is the useful part.**
The family was wrong (swapping it made things slightly worse) AND the mean was
not the constraint (hits-v1 moved the line-0.5 mean 0.554→0.581 against a 0.598
base rate — closer — while resolution FELL). What remains is per-prop
DISCRIMINATION: the ladder's inputs don't separate hitters. Don't build another
projection variant for hits; diagnose what the champion's `p_win` reads first.
- **"Honest ceiling" needs its control checked before you claim it.** The spec's
own pre-registered fallback ("hits may be genuinely low-resolution for anyone")
was REFUTED by the champion scoring 0.276 on the identical 189 rows. A ceiling
claim is only honest if no instrument on the same rows beats it — check that
BEFORE writing the branch, not after.
- **Backtest ≠ verdict.** The replay truncates each player's game log strictly
BEFORE the row's `game_date` and reuses the row's stored grade-time
`combined_multiplier` (both live on real ledger rows) — without that truncation
it would be scoring predictions with the answer in hand. The verdict of record
is still the forward accrual, so hits-v1 stays wired, writing
`proj_hits_p_over`/`proj_hits_meta` only. Never served.
- **Paired bootstrap, not two independent SEs.** Challengers score the SAME rows;
comparing independent standard errors overstates uncertainty and would have
read a reliable 0.022 regression as noise. `scripts/hits-v1-holdout.js` has
the seeded implementation — reuse it for the next challenger.
- **The takeable axis paid off measurably:** 94 of 159 live hits props are
OUTSIDE the promotion band and 93 were modelled anyway. Scope by
`isTakeableMarket` (book identity); record `isWithinPromotionBand` and never
let it gate a model — a price-shape rule would have deleted 59% of the board.
- **Local `.env` has a transposed Supabase ref** (`zmdnczhtdxcddszxttub`; real is
`zmdnczhtdxcddsxzttub`), so local scripts hitting Supabase need an explicit
`SUPABASE_URL=` override. Prod + the MCP connection are fine.
## Settlement outage + challenger scoreboard (Session 77 — non-obvious)
- **`.in('id', [...])` IS A URL, NOT A QUERY.** PostgREST puts filters in the
URL: 500 UUIDs = an 18,499-char request that the fetch layer rejects with
`TypeError: fetch failed`. This silently killed `settleLedger` for two days
(2026-08-01/02) — it refetched rows by id, destructured `const { data: rows }`
with NO error binding, so rows was null, the loop never ran, and it returned
`{settled:0,voided:0,unrecoverable:0,pending:0}`, byte-identical to a healthy
"nothing to settle". 1,444 rows sat with `settle_attempts=0`. NEVER send an
unbounded id list; `ID_FILTER_CHUNK` (100) is the guard, and settleLedger now
selects every column it needs in ONE query.
- **It was VOLUME-TRIGGERED, which is why it hid.** Daily volume ran 20260 rows
for weeks; 2026-08-01 was the first day past `SETTLE_FETCH_LIMIT` (500). If a
pipeline "works for weeks then stops", suspect a threshold that volume just
crossed, not a code change.
- **The ops alarm was blind to it BY CONSTRUCTION.** The zero-settle watchdog
reads settleLedger's own return values, so `pending: 0` told it the backlog was
empty. An alarm that trusts the return value of the thing it watches cannot see
that thing fail silently — the signal must come from OUTSIDE (a direct
`game_date < today AND outcome IS NULL` count).
- **Settled n went 493 → 1,741 the moment it was fixed.** Anything reading
"n-blocked" across MULTIPLE independent challengers at once is a pipeline
symptom, not a sampling fact. Count settled rows before believing it.
- **`specs/challenger-scoreboard.md` is the board.** Nothing promoted: arch-v1
Δ0.0000 CI[0.005,+0.005] on 1,741 (and it MOVED 76% of rows by 2.5pp mean —
active movement carrying zero information), contact-v1 +0.0008 inconclusive,
proj-v1.1 ladder 0.0301 CI excluding zero = reliably WORSE. matchup/tb-v1/
hits-v1 genuinely pending (rows dated 08-02+, settle after ET midnight).
- **arch-v1/contact-v1 need no replay** — they wrote p_win at grade time into
their own columns, so scoring them is a TRUE prospective holdout. Only a
challenger that did not exist at grade time (hits-v1) needs a point-in-time
replay. Don't conflate the two kinds of evidence.
- **Score a nudge on the rows it MOVED**, not on all rows — otherwise the
unmoved rows are the champion measured against itself and dilute any real
effect toward zero. `scripts/challenger-scoreboard.js` does both slices.
## The 429 is odds-api, NOT PropLine (Session 77 — non-obvious)
- **`specs/odds-429-diagnosis.md`.** MEASURED: PropLine 5/3,000 daily (0.17%);
odds-api 478/500 MONTHLY, `allowed:false` (tracker blocks at 95%). One snapshot
= ONE PropLine call per sport (all markets comma-joined) — there is no
per-prop/per-book fan-out and no request-pattern problem to optimize.
- **The 429 text is the BACKUP's.** `oddsService.getOdds` falls through silently
when PropLine returns null/empty, then odds-api's quota gate throws
`429 "Odds data temporarily unavailable"`. So an EMPTY PropLine slate is
indistinguishable from an outage, and the error names the wrong provider. Open
order — don't read a 429 as "PropLine exhausted" without checking
`GET /api/internal/quota`.
- **BOOK BREADTH INVARIANT: we never discard books.** All books are KEPT and
SHOWN (`DISPLAY_BOOKS = MODEL REFERENCE DFS`). The ONLY selectivity is that
DFS pick'em is excluded from PRICING/consensus (`EXCLUDED_FROM_PRICING`) — a
fixed-payout shaded number is not a market price. Never "clean up" breadth.
## Champion decomposition — the edge is a hit-rate counter (Session 78 — non-obvious)
- **`specs/champion-input-diagnosis.md`.** `probabilityEstimator` IS the champion
and it is five lines: `base` = empirical frequency of (stat > THIS line) over
the game log, blended 0.6/0.4 with the last-5 frequency, then oppAdj(±0.03) +
homeAdj(±0.015) + a cv>0.40 pull toward 0.50, then clamp [0.10, 0.95].
- **EXACT ANALYTIC ABLATION (no refit):** every adjustment is closed-form from
stored features and the consistency step is linear (`f(x)=0.9x+0.05` ⟹
`f(a+b)=f(a)+0.9b`), so layers subtract algebraically out of the stored p_win.
Result: **removing ALL THREE adjustments changes resolution by nothing on every
stat**, and on rbi/runs it IMPROVES it (rbi home/away removal +0.0053, CI
excludes zero = mildly HARMFUL). ~100% of the edge is base+recency.
- **POOLED RESOLUTION IS INFLATED — do not quote 0.46.** Per stat the champion is
0.196 (hits) to 0.499 (rbi); pooling stats with different base rates adds
correlation because p_win tracks the base rate across stats. Paired DIFFERENCES
(the scoreboard) stay valid; the absolute level does not. Always per-stat.
- **THE CLAMP IS THE BIGGEST LOSS, not a missing feature.** 358/1,741 settled rows
(20.6%) sit ON the boundary, so the model emits a CONSTANT there and cannot rank
within a fifth of the book. And `0.900` hides home_runs-under truly 99.5%
(9.5pt under-confident) next to hits-under truly 51.9% (+38.1pt over-confident).
`PROB_CEIL=0.95` makes the 99.5% case inexpressible. Global over-prediction
+3.5pt (total_bases +7.6). Fixable with NO new data.
- **`opportunity_drift` is the ONE real missing-weighting lead** — residual corr
+0.156 (hits) and +0.145 (total_bases), i.e. it REPEATS across independent
stats. Discipline: 14 features × 5 stats = 70 tests, so 34 CI-excludes-zero
results are expected BY CHANCE; a single hit (weather on TB) is noise. And we
already compute it — arch-v1's opportunity axis uses it and extracts NOTHING
(delta +0.0001). Wrong implementation, not a missing feature: opportunity must
scale the RATE, not nudge the probability.
- **ARCHETYPE VERDICT: unmeasurable, not refuted.** Only 2 of 41 archetypes
(BOMBER, GHOST) reach n≥40 settled rows; all mean residuals straddle zero. That
is "we have not measured it", NOT "archetypes carry no signal". Don't act
either way. (Their uniformly negative residuals are the global over-prediction,
not an archetype effect.)
- **Why every challenger has failed:** the ladder/hits-v1 REPLACE the frequency
question with a fitted distribution; arch-v1's env axis adds park/weather the
champion ignores. Asking "how often has he cleared THIS number" directly is the
thing that works — improve its inputs, never substitute it.
- **`model_snapshots.outcome` is NULL on all 22,032 rows.** The retention table
built for exactly this kind of replay was never settled, so ablations must join
outcomes from `ledger_entries` on (player_key, stat, line, side, game_date).
## Forward-model reality check (Session 79 — non-obvious)
- **`specs/forward-model-reality-assessment.md`.** THE OBJECTIVE is a FORWARD
matchup projection (hitter profile × pitcher stuff × park/conditions, read
through archetype), not a market-edge number. Every component that needs
exists AND is loaded in prod — and ALL of it sits DOWNSTREAM of the grade.
- **The served grade sees NONE of it.** `probabilityEstimator` reads exactly
three features (`opp_rank_stat`, `home_away`, `l10_stddev`/`l20_avg`) plus the
game log. statcast rows, arsenal, park, weather, platoon and archetype are all
loaded in `snapshotService` AFTER grading and written to CHALLENGER columns.
`mlbContext` (platoon/handedness) has ZERO consumers — dead code.
- **STATCAST NIGHTLY REFRESH IS UNREACHABLE CODE.** `snapshotScheduler.tick()`
returns at `if (!HOURS_UTC.includes(h)) return` (14,19,22,1,3); the statcast
block then tests `h === STATCAST_HOUR_UTC` (default **11**), which that guard
can never admit. Data frozen at its 2026-07-21 backfill; its failure alert is
inside the same dead branch so it can't warn. Same shape as the settlement
outage — guarded-out code that reports nothing. Set STATCAST_HOUR_UTC to one of
HOURS_UTC or move the block above the guard.
- **Inputs are HAVE, not missing** — `statcast_aggregates` 1,354 rows (750
pitchers / 604 batters): exit velo, launch, barrel, hard-hit, whiff, chase,
pitch_mix, GB/FB, arm angle, and **bats/throws complete on all 1,354**. Gaps
are team DEFENSE (only a coarse `opp_rank_stat`) and catcher framing/umpire.
PARTIAL: batter GB/FB land in the `metrics` JSONB not the typed columns;
lineup-slot tables (`player_role_profiles`, `lineup_role_profiles`) are 0 rows.
- **The card is designed for the forward read; the engine never fills it.** The
factor vocabulary the "SIGNAL BREAKDOWN" renders is entirely counter-restating
(`l5_hot_vs_line`, `l20_over_line`, `back_to_back`, `home_game`) with several
structurally-NBA labels (`ref_foul_high`, `coach_pace_delta`,
`opp_3plus_starters_out`). No signal names a pitcher, pitch type, handedness or
park. Surface needs FEEDING, not redesigning.
- **What the prior verdicts do and don't say.** Resolution = corr(forecast,
outcome) was never a market/edge test — the metric was right, the QUESTION was
narrow ("does challenger out-rank champion?"). proj-v1.1 and hits-v1 remain
correctly refuted AS DISTRIBUTION SWAPS on thin inputs; neither tested a
matchup-fed projection. arch-v1 IS market-relative by construction and is the
one component genuinely measured on the wrong axis — re-test its axes as
forward inputs. "AT CEILING" (runs/walks) is provisional: measured only against
features the champion already reads.
## The skill engine + feature registry (Session 80 — non-obvious)
- **`specs/skill-engine-architecture.md`.** `src/services/model/` is the forward
engine: `featureRegistry.js` (CANDIDATE/PROVEN/DEAD per feature PER SPORT) and
`skillProjection.js` (PA outcome tree: K/BB via log5 odds-ratio vs league, then
archetype-weighted contact quality → Binomial(PA, p_hit) mixed over a PA
distribution). Challenger-only; champion untouched.
- **THE GATE IS STRUCTURAL, not a habit.** `liveFeatures()` returns PROVEN only,
and the registry ships with exactly ONE proven feature (the incumbent counter).
A test asserts that with only PROVEN allowed, `projectSkill` returns NULL — an
unproven model cannot reach a user by accident. `promote()` requires n>=200,
positive lift, CI excluding zero, and has NO override argument.
- **STAGE A RESULT: skill-v1 LOSES, not promoted.** 570 rows, 91.9% pitcher
coverage, resolution 0.0499 vs champion 0.166, delta 0.116 CI [0.189,0.043].
Also NOT selective — its top-8 most confident picks hit 50% (lift 0.065).
- **UNITS: `statcast_aggregates` stores PERCENTAGES (0100), not fractions.**
`k_pct: 29.6` means 29.6%. Feeding raw rows in made `bip = 129.617.1` negative
and refused 568/576 rows. ALWAYS convert via `skillProjection.fromStatcastRow`
(the one chokepoint); it nulls out-of-range values rather than clamping, and
leaves mph/degrees fields alone.
- **`ledger_entries.team`/`opponent` are NULL on ~all rows** — do NOT join a
matchup on them. The first Stage A run resolved a pitcher for 1 of 570 rows and
would have reported a verdict on a batter-only model. Resolve the opponent from
the player's own statsapi game log (`getPlayerGameLog` → `{date, opponent}`),
which is authoritative and point-in-time safe → 91.9% coverage.
- **Archetype = FEATURE SELECTOR, not a nudge.** `ARCHETYPE_MAP` weights decide
which skill inputs drive a hitter (BOMBER barrel 0.50 / gb_speed 0; GHOST
barrel 0.05 / gb_speed 0.60). Locked by test: same hitter read through two
archetypes moves >0.15. Weights are DOCUMENTED, not fitted — fitting on 1,741
rows is curve-fitting; the registry exists so they get measured.
- **total_bases is deliberately REFUSED by skillProjection.** A deterministic
bases-per-hit multiplier made P(TB>=2) exactly equal P(hits>=1) — a relabelled
hits curve carrying no new information. TB needs tb-v1's compound per-hit bases
distribution; refusing beats shipping a relabel.
- **STATCAST REFRESH WAS UNREACHABLE CODE** (fixed): it sat inside `tick()` below
`if (!HOURS_UTC.includes(h)) return` (14,19,22,1,3) while testing `h === 11`.
Never ran once; data 13 days stale; BOTH its alerts were in the same dead
branch. Now its own `statcastTick`. The old test only checked the string
existed — the new one asserts it is not behind the snapshot-hours guard.
## The validation gate + the stat that was wrong (Session 81 — non-obvious)
- **`statModel.js` and `correlateValidator.js` NEVER EXISTED** in this repo. The
spec's only prior form was `src/services/python/blueprints/unconventional.py`
(Flask, in the OFFLINE python service, scoring NBA factors against an empty
warehouse), and `tests/unit/supplementSystems.test.js` INLINES its own
`validateFactor` (line 368; only `fs`/`path` are required). So those tests
passed for months with no implementation to connect — that is the real reason
every challenger was measured ungated.
- **`src/services/model/correlateValidator.js` is the gate now** — n>=500,
|r|>=0.15, p<0.05, Bonferroni. The p-value is EXACT (t-transform via a
regularized incomplete beta, Lentz CF) and unit-verified against known values;
scipy isn't available in Node so don't reach for an approximation. Pairs with
an unknown side are DROPPED — zero-filling a correlation invents a point at
the origin.
- **HITS IS A CLEAN NEGATIVE — stop modelling it.** Gate run at n=570,
Bonferroni-8: EVERY skill feature fails, max marginal |r| = 0.062 vs a 0.15
bar. Not a power problem — an effect-size problem. And the value engine loses
head-to-head (0.0499 vs 0.166, CI [0.189,0.043]). At the 0.5 hits line there
is very little for skill inputs to know.
- **TOTAL BASES IS WHERE THE SIGNAL IS, and it is n-blocked.** Same features:
`hard_hit_pct` marginal r = **0.153** (above threshold), `exit_velo` 0.124,
raw r 0.196/0.167 — refused ONLY because n=295 < 500. Needs ~205 more settled
rows. This is what the physics predicts: contact quality drives EXTRA BASES,
not whether a grounder finds a hole.
- **The gate reports r and p even when underpowered** (`underpowered: true`,
`rows_needed`). "Not enough data yet" and "nothing here" need OPPOSITE
decisions — collapsing them into a bare refusal hid the best signal on the board.
- **Feature verdicts are PER STAT** (`recordStatVerdict` / `statusForStat` /
`candidateFeaturesForStat`). Marking these DEAD sport-wide on hits evidence
would have killed the features most alive on TB. Per-sport doctrine one level
deeper: physics differ per stat.
- **Next is the compound TB projection** — `skillProjection` still REFUSES
total_bases (a deterministic bases-per-hit made P(TB>=2) == P(hits>=1)). Build
the per-hit extra-base distribution off launch/barrel (tb-v1's shape, fed by
skill inputs), accrue to n>=500, re-run this gate. Do NOT lower the bar.
## Point-in-time skill validation + TB solo/interactions (Session 82 — non-obvious)
- **`statcast_aggregates` KEEPS NO HISTORY** — upserted in place on
(sport,season,source_id,role), ONE as-of date, prior versions destroyed. The
first skill backtest was honest only BY ACCIDENT: the nightly refresh was
unreachable code so profiles sat frozen at 2026-07-21, BEFORE the settled
window. Fixing that cron refreshed them to today and made point-in-time
validation impossible from that table. **`statcast_history` (new) retains a
dated snapshot per refresh** — query `where as_of_date < game_date order by
as_of_date desc limit 1`. Retention is best-effort and must NEVER fail the
refresh (unit-tested). Until it accrues a window, ALL skill-feature results are
CONTAMINATED/DIRECTIONAL, never gate verdicts.
- **TB solo pass: NOTHING passes.** n=383, Bonferroni-12 (α=0.00417).
`hard_hit_pct` is closest at marginal r=0.135, p=0.0080 — fails BOTH the 0.15
effect bar and corrected α. **It DRIFTED DOWN from 0.153 (n=295) → 0.135
(n=383)**: an estimate regressing as noise averages out, not an effect firming.
Don't keep quoting the older better number.
- **Interactions: none pass.** Only `barrel × power_archetype` has incremental
(partial, controlling for both components) exceeding its parts — 0.101 vs
0.019 at n=260. A lead, not a finding.
- **INTERACTION-PROXY TRAP:** the archetype conditioner was first
`barrel_pct/LEAGUE.barrel_pct` — a monotone transform of its own component — so
the "interaction" was barrel² measuring NONLINEARITY, and it produced the run's
only positive result (0.132). A Gauss-Jordan pivot test does NOT catch this
(the columns differ by a scale factor); use a **scale-free pairwise correlation
check** on control columns. Real archetype labels come from
`model_snapshots.archetype` (260 labelled TB rows: 140 BOMBER / 120 other).
- **Interactions must be scored by PARTIAL correlation** vs the counter residual,
controlling for both components — raw correlation can't distinguish
PASSES-AND-ADDS from PASSES-BUT-REDUNDANT.
- **TB is at PARITY with the counter** (0.2718 vs 0.2647, CI [0.065,+0.079],
inconclusive) where HITS lost by 0.116 with CI excluding zero. Same engine,
same day — the stat choice was the whole story. Parity under contamination is
NOT a win; nothing promoted.
- **`skillProjection` now models total_bases** as a compound convolution (per-PA
0/1/2/3/4 bases; barrel→HR share, exit velo→2B/3B share). The old
deterministic bases-per-hit made P(TB>=2) EXACTLY P(hits>=1); non-degeneracy is
locked by test.
## Batter cluster + the bar (Session 83 — non-obvious)
- **total_bases has NOT passed BAR 1.** Its head-to-head is INCONCLUSIVE at
parity (delta +0.004..+0.007, CI includes zero) and CONTAMINATED. It is
frozen, but frozen as an *inconclusive* model — do NOT install it as "the
proven reference standard", because then the bar other stats must clear
becomes "be inconclusive at parity", which admits everything on a null result.
**The proven set is EMPTY.**
- **HITS IS CLOSED — a well-powered negative.** At n=803 it CLEARS the gate's
sample bar, so its features were properly TESTED, not refused: max marginal
|r| = 0.053 vs the 0.15 bar, every interaction's incremental ≈ 0, and the
model loses head-to-head 0.096 with CI [0.165,0.029]. Don't re-run hits.
- **Everything else is n-blocked:** TB 383, rbi 391, HR 228, runs 188 (gate needs
500). Two leads worth carrying: `home_runs · barrel_pct` marginal r = **0.135**
(NEGATIVE — higher barrel goes with the counter OVER-predicting, i.e. a
correction not a predictor), and `runs · batterK×pitcherK` incremental **+0.132**
(largest in the cluster; mechanism = strikeouts destroy PA, and a PA that never
happens cannot score).
- **RBI is half-unmodellable today:** it is power × OPPORTUNITY and we ingest NO
baserunner state. A weak RBI result is evidence we model half the stat, not
that skill inputs fail for RBI.
- **`statcast_history` retention is LIVE and verified in prod** (1,387 rows,
as_of 2026-08-03). Two gotchas: the first run failed on a drifted hand-written
schema (`swing_pct` missing) — the table is now `create ... (like
statcast_aggregates)` and the writer passes rows through whole; and the
refresh still succeeded during that failure, confirming the best-effort guard.
A usable point-in-time WINDOW starts 2026-08-04 (as_of < game_date).
- **`scripts/cluster-prove.js`** runs the whole both-ways program for any stat via
`CLUSTER_STAT=`. Per-stat interaction sets are the TB map RE-WEIGHTED, never
copied — reuse speeds the search and grants no pass.
## Pitcher engine + the cap that was eating the board (Session 84 — non-obvious)
- **THE GRADE CAP WAS THE BINDING CONSTRAINT ON EVERY STAT.** `dedupeProps` takes
FIRST-ROW-WINS IN FEED ORDER and stops at `GRADE_SLATE_LIMIT`. Measured via
`GET /api/internal/diagnose-refusals`: **1,244 unique gradeable props/slate**,
a 500 cap graded ~334, and pitchers (2.6% of the feed) got **6 props a slate**
— putting n>=500 three months out. RAISED 500 -> 1500 (measured: 721ms/prop at
concurrency 5 ≈ 179s for the full board; cron runs 5x/day; statsapi is free).
Expect pitcher Ks ~6 -> ~32/slate, so n>=500 in ~2 weeks. Concurrency stays 5.
- **Pitcher props were NEVER being refused** — `strikeouts: graded 5, refused 0,
suppressed 0`. Don't hunt for a data gap here; it was truncation.
- **`src/services/model/pitcherEngine.js` is its OWN engine** (per-role doctrine):
archetypes FLAME (whiff .65) / SCALPEL (chase .40) / SINKER (k_rate .50) /
DEFAULT, and the projection is `K% (log5 vs THIS lineup) x batters faced ->
Binomial(BF, k)`. A test asserts its weight keys are NOT the batter engine's.
Unclassifiable -> DEFAULT map, never a guessed archetype.
- **THE COUNTER IS ANTI-PREDICTIVE ON STRIKEOUTS: resolution 0.064.** Recent K
counts are dominated by which lineups a pitcher drew and how long he was left
in, not by skill. This is the one stat where the incumbent has no defensible
edge — the strongest theoretical case for the skill model in the programme.
- **Strikeouts NOT proven** (n=57 vs 500): pitch-v1 0.1285 vs counter 0.0639,
delta +0.192, CI [0.098,+0.509]. But FOUR solo features exceed the |r|>=0.15
bar and fail only on n: **arm_angle 0.250** (largest in the programme),
whiff +0.213, k_pct +0.206, chase +0.195. Batter cluster's best was 0.135.
- **`resolveTeam` wants an ABBREVIATION, not a team name.** The game log supplies
full names ("Cincinnati Reds"), so the roster join silently resolved nothing
and the first run showed 0% lineup coverage — the theorized stuff x lineup
carrier was never tested, not failing. Use `environmentContext.NAME_TO_ABBR`;
coverage went 0% -> 94.7%.
- The stuff x lineup-K-rate carrier shows NO incremental signal so far (its raw r
is explained by whiff alone), and adding the lineup term LOWERED head-to-head
resolution (0.174 -> 0.1285). n=54, so not a verdict — but recorded, not dropped.
## Lineup K-rate (Rung 1) + the cap fingerprint (Session 85 — non-obvious)
- **THE CAP FIX LANDED: 334 -> 907 grades/snapshot, strikeouts 6 -> 17** (2.7x
across the board). n>=500 for pitcher Ks is now ~a week away, not 3 months.
- **OPERATIONAL: `POST /api/internal/snapshot/:sport` now 524s at Cloudflare** —
grading the full board exceeds the 100s edge timeout. **The run still COMPLETES
server-side** (the 907-grade snapshot was written by a 524'd request), and the
cron is in-process so it is unaffected. Never read that 524 as a failure; check
`/api/internal/snapshot/status`.
- **PA-WEIGHT the team K-rate.** Opposing-lineup K-rate is derived free by joining
the opposing roster to batter `k_pct` we already ingest (94.7% coverage, zero
new sourcing). An UNWEIGHTED roster mean counts a 12-PA callup like an everyday
starter and it HURT the model (0.174 -> 0.129); PA-weighted it HELPS
(0.174 -> 0.195). Same hypothesis, same data — the derivation was the problem.
Always weight a team aggregate by playing time.
- **A conditioner can have ~zero solo signal and still matter.** Lineup K-rate
solo r = +0.004. That is not evidence against it — it is hypothesised as a
CONDITIONER, not a standalone predictor. Judge it by its incremental partial,
stratified.
- **Within-archetype strata have OPPOSITE signs** (FLAME incremental 0.152,
non-FLAME +0.145) and the pooled value (+0.077) sits between them — the shape a
conditional effect makes, and invisible when pooled. But n=20/24 (SE≈0.22) and
the DIRECTION contradicts the theory (predicted stronger for finesse; magnitudes
are near-equal with flipped signs). Structure to re-test, NOT a finding.
- **Rungs 2 and 3 are NOT triggered.** A rung only fails once fairly tested, and
Rung 1 is n-blocked, not failed. Do not source confirmed lineups yet.
- **Pitcher features still have NOT passed the gate** — all refused at n=57.
Four exceed the |r|>=0.15 effect bar (arm_angle 0.250, whiff +0.213, k_pct
+0.206, chase +0.195) but exceeding one of three thresholds is not passing.
## Conditioning registry + the proven-status probe (Session 86 — non-obvious)
- **RUN `node scripts/proven-status.js` BEFORE planning on a "proven" claim.**
Four consecutive orders opened by calling null results proven. The script
recomputes from the ledger: PROVEN_SET is **EMPTY** (hits LOSES 0.096 CI
excluding zero; total_bases +0.004 inconclusive; strikeouts +0.259 inconclusive
at n=57). It deliberately reports SAMPLE READINESS separately from RECORDED
VERDICTS so "n>=500" is never mistaken for "passed".
- **JOINING `model_snapshots` TO `ledger_entries` FANS OUT.** model_snapshots
holds one row per prop PER SNAPSHOT CYCLE, so a naive join counts each ledger
row once per cycle: BOMBER x hits read as **641** when the true distinct figure
is **287**. Always dedupe on `ledger_entries.id`. This is the difference
between "gate-ready" and "short by 213".
- **NO archetype x stat reaches n>=500.** Best: BOMBER x hits 287, BOMBER x TB
142, BOMBER x rbi 128, GHOST x hits 124. Pitcher archetypes are untestable
(58 settled Ks across ALL archetypes).
- **`featureRegistry.recordConditioning`** keys archetype x SKILL x interaction x
status + lift. The skill tag is MANDATORY and enforced (untagged → refused;
PROVEN without sufficient evidence → refused). `validatedSkills()` returns the
coherent profile — currently `{}` for every archetype, by design.
- **`fromStatcastRow` does NOT carry `pitch_mix`** (it maps PCT_FIELDS/RAW_FIELDS
only). Attach it explicitly or arsenal features silently read n=0 — which
would have recorded "arsenal doesn't matter" from a column that was never
populated. pitch_mix shape is `[{type, usage_pct, velo, whiff_pct, ...}]`.
- **DEFENSE IS GENUINELY NOT DERIVABLE from what we ingest.** No OAA/DRS/range
anywhere; opposing pitchers' hits-allowed conflates pitching WITH defense so it
would validate the wrong skill. It needs Savant's fielding endpoint (free, same
host as the five feeds already ingested). Don't proxy it.
- Within BOMBER, the counter still leads on hits (0.218 vs 0.160) — consistent
with the closed pooled hits negative.
## Defence ingest + cumulative Bonferroni (Session 87 — non-obvious)
- **BONFERRONI IS NOW CUMULATIVE ACROSS THE PROGRAMME LIFETIME**
(`src/services/model/testLedger.js` + `mc_test_ledger`). Correcting per-session
(8 tests → /8, forever) let the false-positive rate compound silently; the
denominator is now DISTINCT hypotheses ever tested. Demonstrated: 19 → 38 in
one session, α 0.0026 → **0.0013**. RE-TESTS DO NOT INFLATE IT — re-asking the
same question on more data is not a new shot on goal, and counting it would
punish waiting for sample. The alpha only ever shrinks, so prefer re-testing
standing candidates over inventing new hypotheses — that is now mathematically
the disciplined choice.
- **DEFENCE IS INGESTED** — free Statcast OAA (`FEEDS.fielding_oaa`), 514
fielders → `team_defense` (31 teams, dated). Team-level is the right unit (the
defence behind the pitcher faced). `oaa_sum` + `oaa_mean` (mean because a team
with more measured fielders would else look better for being measured more);
<3 fielders → absent. **OAA 0 is a REAL "exactly average" reading** — coercing
absence to 0 asserts every unmeasured fielder is league-average, the commonest
profile there is. `team_defense` carries `as_of_date` in the PK FROM ROW ONE
(the statcast_aggregates lesson, applied before it was needed).
- **`BASE` in statcastAdapter ALREADY ENDS IN `/leaderboard`** — the new feed
doubled it and 404'd. And because a failing feed degrades to an EMPTY index by
design, it surfaced as "fielding_oaa: 0 rows", which reads exactly like
"Statcast has no fielding data". **Graceful degradation makes a wiring bug look
like an honest absence — treat any feed reporting 0 as suspect until the URL is
fetched by hand.**
- **THE DEFENCE DIFFERENTIAL APPEARS AS THEORY PREDICTS:** solo r vs counter
residual is **+0.130 for GHOST** (contact/speed, n=104) and **0.018 for
BOMBER** (power, n=245). A GHOST's hits depend on fielder range; a BOMBER's
barrels clear the defence. **A flat BOMBER result is the theory working, not
the test failing.** Both UNDERPOWERED (p=0.188 vs corrected α 0.0013) — a
signal shape, not a result.
- `team_defense` keys on Savant's DISPLAY NAME (a nickname, "Cubs") while game
logs give full names ("Chicago Cubs") — match on both.
## Re-adjudication + the promotion bar (Session 88 — non-obvious)
- **NOTHING HAS EVER BEEN PROVEN.** `proven-status.js` = EMPTY; `validatedSkills()`
= {} for all archetypes; 0 conditioning entries. The only PROVEN *feature* is
`recent_frequency_prior` — the incumbent COUNTER itself (S78 ablation showed it
is ~100% of the champion's resolution). It is the baseline, not a conditioning
interaction; demoting it would leave nothing to grade from.
- **The cumulative correction did NOT catch a false positive.** It caught nothing
(empty proven set). It tightened α 0.0026 → 0.0013 in one session — the
mechanism working, not a demotion. Don't restate that as a catch.
- **THE REAL HOLE (now closed): `promote()` could bypass the cumulative
correction.** `isSufficient` now REQUIRES `evidence.bonferroni_tests`, refuses
it if lower than `opts.cumulativeTests`, and refuses a `p_value` that doesn't
clear `0.05 / bonferroni_tests`. Same rule guards
`recordConditioning(status:PROVEN)`. This is what makes a retroactive
re-adjudication pass unnecessary — the bar is applied at promotion time.
- **Cumulative correction is now NATIVE on every analysis path** — `cluster-prove`,
`pitcher-prove-k` and `tb-solo-and-interactions` all use `testLedger`. If you
add a new analysis script, wire it or it silently corrects per-session.
- **`src/services/model/reAblation.js` is the standing second line.** Pure +
injectable (no DB, no measurement) so the decision rule can't drift from the
gate's. Records BOTH p-values and BOTH test counts per verdict so a demotion is
re-derivable. **No fresh measurement = `PENDING_RETEST`, never DEMOTE** —
absence of a re-test is not evidence, and demoting on it would punish whichever
stat is off-season. A feature promoted at α=0.05/20 CAN demote on the same
p-value once the bar is 0.05/60; that is correct, not unfair.
- **Don't emit a public "recalibrated after re-adjudication" ledger event when
nothing changed** — announcing rigour that did no work is itself a false signal.
## Lineup + baserunner context ingest (Session 89 — non-obvious)
- **RBI/runs were INPUT-blocked, and the input is now ingested** — both halves
free from statsapi (already used for game logs/schedules/pitchers).
`src/services/lineupContextService.js` → `lineup_context` (batting order) +
`hitter_opportunity` (RISP share). Both DATED in the PK.
- **`schedule?hydrate=lineups` → `homePlayers`/`awayPlayers` are ORDERED arrays
of 9 — the array order IS the batting order** (index 0 = leadoff). Nothing is
inferred; a short lineup records fewer slots rather than padding to nine.
- **RUNG 2 IS CHEAP, contrary to expectation.** "How often does he bat with
runners on" looked like a play-by-play reconstruction; statsapi serves it via
`people/{id}/stats?stats=statSplits&sitCodes=risp,r0` — ONE call PER PLAYER
(season aggregate), not per game. Real example: 87 PA with RISP → 25 RBI vs
302 PA bases-empty → 17 RBI. Probe for a cheaper aggregate endpoint before
assuming per-event reconstruction.
- **`risp_share` is RISP ÷ (RISP + bases-empty)** — deliberately NOT ÷ season PA,
because runner-on-first-only belongs to neither split. It is "RISP as a
fraction of the PAs we can classify", stated exactly rather than implied.
- **A hitter with no splits is NULL, never a 0 share** — 0 would assert he never
bats with runners on, a strong and usually false claim.
- **`POST /api/internal/lineup-context/refresh`** verifies the ingest in seconds.
Built because the first prod run wrote 0 rows while the parser worked locally
(144 rows / 10 games) — diagnosing that via a full snapshot costs ~3 min and
524s at the edge. Same lesson as the defence feed: **a 0 is a wiring bug until
proven an honest absence.**
- **Coherence check that the data passes:** the top RISP-share hitters all bat
4th/5th — the mechanism showing up the moment both tables joined.
- DRIVER/CATALYST pre-registered theories are now INPUT-READY (were
input-blocked); they are sample-blocked from here. Ingesting is not proving.
## chaining-v1 + the calibration gate (Session 90 — non-obvious)
- **THE HIT-PARLAY SURFACE IS BLOCKED, and the block is structural.** Measured on
972 settled hits props: the model is monotonically over-confident exactly where
a parlay stacks — predicted **0.911 → actual 0.630** (err 0.281), and it is
FLAT at ~63% for everything above 0.70 (no discrimination there at all).
A 4-leg "91%" ticket: model 0.686, reality 0.157 — a **4.4x overstatement that
compounds with every leg**. `chain.chainAcross` REFUSES atoms not marked
`calibrated: true`. Refusing is the feature.
- **Calibration ≠ resolution, and chaining cares about calibration.** A model can
rank fine and be useless compounded. Single props survive a calibration error;
a parlay multiplies it. Never stack a probability that has not passed
`calibration.isCalibrated`.
- **`calibration.fitIsotonic` is the honest repair** — monotone (pool-adjacent-
violators), so the model's ORDERING survives untouched while the NUMBERS move
to what actually happened. Real map: 0.65→0.594, 0.85→0.639, **0.91→0.639**.
Fit it POINT-IN-TIME (outcomes preceding the prop) or it has seen the answer.
- **`src/services/model/chain.js` is the portable core:**
`base_events + context → chainFn → aggregator`, aggregator pluggable —
ACROSS = compound ticket, UP = team score. Sport parts are INPUTS, not code
paths, so basketball is content not a rebuild.
- **`redistribute` is the archetype hook** — DORMANT in baseball (a nine-run lead
doesn't change who bats next), LIVE in basketball (blowout fades the star,
feeds the bench). It exists now so the machine doesn't need rewriting later.
- **Treating same-game legs as independent errs in the FLATTERING direction** —
they share pitcher/park/weather, so the joint is likelier than the product.
`chainAcross` applies a bounded shift toward the weakest leg and labels it an
approximation, not a joint distribution. Prefer cross-game legs.
- **An unreadable atom is DROPPED, never p=0** — a single zero leg would zero an
entire ticket.
- **Market divergence is NOT an error signal** and must not downgrade confidence:
it flags a contested game script (its props are the best or the worst on the
board, unknown which). **Internal inconsistency IS** — per-entity reads not
summing to the team read means one is wrong and we don't know which, so the
honest output is LOW CONFIDENCE, not a correction.
- **`propagate` is shrinkage-weighted by sample** — one game moves a 400-obs atom
barely and a 4-obs atom a lot. That gap is the difference between learning and
noise-chasing.
## Hits calibration + the parlay unblock (Session 91 — non-obvious)
- **PARTIAL PASS: hits are calibrated and stackable ONLY in 0.400.60.** Fit on
game_date < 2026-08-02 (n=589), evaluated on >= (n=383) — the map never saw the
evaluation rows. Held-out after correction: **0.477→0.506 (0.029, n=83),
0.587→0.580 (+0.007, n=193), 0.667→0.603 (+0.063, n=63)**, vs raw errors of
+0.191/+0.279/+0.246. Ordering preserved (verified pairwise, not assumed).
- **THE HONEST CEILING IS 0.667.** Once the numbers are truthful this model has
NO 80%+ hit reads at all. A 4-leg ticket at the ceiling is **0.198**, not the
0.686 the raw numbers implied. The "high-floor parlay" is a ~0.67-per-leg
proposition — say that plainly rather than selling the old number.
- **CERTIFY BY BAND, never a blanket flag.** Held-out error was 0.029/+0.007
through the middle but 0.167 at the bottom and +0.063 at the top. A single
true/false would either discard the 72% that works or ship the edges that
don't. `calibration.certifyBands` + `inCertifiedBand`; only in-band atoms get
`calibrated: true`, which is what `chainAcross` requires.
- **A PASS CONDITION CAN FAIL A MAP FOR SUCCEEDING.** My first gate demanded
honest bins ≥0.70 — but honest calibration REMOVES those bins (ceiling 0.667),
so it failed the repair for working. Test the highest REMAINING band, not a
fixed threshold.
- **`calibrationService.fromLedger` fits STRICTLY before today** and splits by
TIME, not at random — certifying on rows the map was fitted on always looks
perfect, and a random split leaks the future. No calibrator ⇒ NOTHING is
stackable, never "pass raw numbers through".
- **`p_win` is never mutated.** Calibration rides beside it as
`p_win_calibrated` + `calibrated` on hits grades, so the counter stays
byte-identical — a calibration map is a correction TO a forecast, not a
different forecast.
- **Synthetic-data trap in the tests:** front-loading wins makes outcome
correlate with date, so a time-split trains on wins and certifies on losses —
the generator creating the exact leakage the split prevents. Interleave.
## The two-part factor gate (Session 92 — non-obvious)
- **`src/services/model/factorGate.js` asks a question correlation cannot.** A
factor must (a) MOVE the prediction off the player's base rate AND (b) improve
out-of-sample BRIER. Movement alone is **THEATER** — the grade LOOKS like it
read tonight's game while reading nothing, and neither a user nor a
correlation test can see it. arch-v1 was exactly this: moved 76% of rows by
2.5pts, changed resolution by 0.0000, live for months.
- **BRIER, not correlation.** Correlation asks whether the ORDERING improved;
this asks whether the NUMBER got closer to what happened. For a graded
probability the number IS the product, and a factor can improve ordering while
degrading the number.
- **Baseline = the player's LEAVE-ONE-OUT base rate** — literally the "he's due"
null. A factor earns its place only by beating that. Leave-one-out matters: a
row must never contribute to its own baseline.
- **CUMULATIVE CORRECTION APPLIES TO THE INTERVAL ITSELF.** A plain 95% CI is
right for ONE test; at 50 cumulative tests ~2-3 of them exclude zero by chance.
The bootstrap interval now widens to 1 0.05/tests (currently **99.9%**).
Applying it flipped defense and platoon from "proves" to not-proven — a 95% CI
would have shipped two unproven factors.
- **NOT_PROVEN_AT_CORRECTED_BAR ≠ THEATER, and conflating them is the same error
as "insufficient evidence = evidence of absence".** THEATER is reserved for
brier_delta >= 0 (moves, reads nothing). A favourable point estimate whose
corrected CI spans zero is a real candidate held to a rising bar.
- **RESULT for hits (n=741):** `pitcher_contact_profile` **PROVES**
(Brier 0.0066, CI [0.0114,0.0016] at 99.9%). `defense` (0.0043) and
`platoon` (0.0039) are NOT_PROVEN at the corrected bar; `park_hits` is
sample-blocked (n=405). **Zero theater.** Per-archetype all sample-blocked
(BOMBER 252294, GHOST 67125).
- **Two spec gaps found:** approach identities (SPRAY / DAMAGE-DEALER /
COUNT-WORKER) **do not exist** in the registry — MLB batter archetypes are
BOMBER/GHOST/TORCH/BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. And
`parkFactors.STAT_BASE` maps `hits → run_base`, so there is **no hits-specific
park factor**: a park that turns outs into hits without scoring is invisible.
## Causally-correct atoms (Session 93 — non-obvious)
- **KEV'S INSIGHT IS CONFIRMED ON DATA: crude operationalizations under-prove.**
`defense_by_direction` **PROVES** (n=528, Brier 0.0034, CI [0.0059,0.0009]
at 99.9%) where team-average `defense` does NOT (CI spans zero). And the
causally-correct atom **MOVES THE NUMBER LESS THAN HALF AS MUCH** (0.013 vs
0.030) while being reliably right — the crude version was moving more and
knowing less. Bigger movement is not better; it is often the tell.
- **`src/services/model/sprayDefense.js`** = spray×trajectory × positional OAA,
joined by HANDEDNESS. Pull for a RHB is the LEFT side (3B/SS/LF); for a LHB the
RIGHT side (1B/2B/RF). Getting that backwards sends half the league's grounders
to the wrong infielders and **still looks like it's reading defence** — nothing
downstream would catch it. Switch hitters are UNREADABLE (they bat opposite the
pitcher, unresolved here), not guessed.
- **Both halves were already free.** Savant's `leaderboard/batted-ball` carries
pull/straight/oppo × ground/air (609 hitters) — the `statcast` leaderboard does
NOT (19 cols, no direction). And the OAA feed already carries each fielder's
position, so per-position defence is a REGROUPING of last week's ingest.
`team_defense.position_oaa` + `batter_spray` (dated). Prod-verified 609/31.
- **Unmeasured zones are RENORMALISED AWAY, never zero** — a zero asserts an
exactly-average fielder standing there. `coverage` states what share of a
hitter's contact we could actually read; nothing readable → null.
- **ATOM 2 (park+weather→hit-type) is INPUT-BLOCKED, not sample-blocked.**
Weather's free source check PASSES (Open-Meteo already wired via
`weatherService`, exposing temp_f/wind_mph/wind_dir/precip_mm) — but those raw
fields are collapsed into a scalar `env_weather_mod` and `wx_forecast` is
**0/1119** on settled rows. And **park DIMENSIONS are not ingested at all**
(parkFactors holds coefficients, not wall heights or fence distances). A
hit-type conversion needs both; retaining the raw weather fields is the cheap
half, dimensions are the missing one.
## Causally-correct platoon + park inputs (Session 94 — non-obvious)
- **PLATOON SEVERITY built, n=452 — 48 SHORT of the gate.** Not proven, not
theatre. `platoonSeverity.platoonRead` uses each hitter's OWN vs-LHP/vs-RHP
split (statsapi `sitCodes=vl,vr`, one call per hitter), shrunk toward league by
the **SMALLER side's PA** (500-vs-40 is a 40-PA read) and **REFUSED outright
below 60 PA** — a heavily-shrunk severity is indistinguishable from a MEASURED
league-average one, and those are different claims. Without the refusal the
atom would assert a league-typical split about every call-up in the league.
- **The refusal costs sample, honestly:** severity has n=452 where flat platoon
has 741. That gap IS the hitters whose splits we cannot read.
- **Switch hitters are the easy case misread as hard.** He bats opposite by
choice so DIRECTION is never in doubt; the per-side VALUE of his swing is the
unanswerable part. UNREADABLE, never credited with an automatic edge.
- **PARK DIMENSIONS are free from statsapi** —
`/venues/{id}?hydrate=location,fieldInfo` returns leftLine/leftCenter/center/
rightCenter/rightLine + roofType + turfType + elevation (Wrigley: 355/400/353,
595ft). `parkFactors` holds RUN COEFFICIENTS, which structurally cannot express
a park that turns outs into hits without scoring — that is why crude park failed.
- **JOIN THE PARK BY `venue_id` FROM THE SCHEDULE, never by home team.**
Neutral-site and international games break the home-team assumption silently.
`lineup_context.venue_id` carries the real venue.
- **RAW WEATHER IS NOW RETAINED** (`wx_forecast`: temp_f/wind_mph/wind_dir/
precip_mm). Two bugs fixed at once: the scalar `weather_mod` cannot express a
hit-TYPE conversion (wind-out-and-warm vs cold-heavy-air collapse to the same
number), AND the old guard dropped the entire environment when the multiplier
was 1 — discarding the forecast for every ordinary night, which is the majority
of games and exactly the rows a hit-type model must learn the ordinary case from.
- **Proven factors for hits remain: `pitcher_contact_profile`,
`defense_by_direction`.** Crude `defense` and `platoon` both still NOT_PROVEN.
## Active Skills ## Active Skills
- vyndr-voice (all user-facing output) - vyndr-voice (all user-facing output)
- prop-analysis (grading methodology) - prop-analysis (grading methodology)
+11 -4
View File
@@ -26,8 +26,15 @@ RUN npm ci --omit=dev --no-audit --no-fund
FROM node:20-alpine AS runner FROM node:20-alpine AS runner
WORKDIR /app WORKDIR /app
# curl is used by the /api/health smoke check (Coolify HEALTHCHECK). # curl /api/health smoke check (Coolify HEALTHCHECK).
RUN apk add --no-cache curl tini # postgresql-client (pg_dump/pg_restore) + rsync + openssh-client + bash — the
# nightly DB backup. openssh-client (Session 64) is NOT optional: rsync shells
# out to `ssh` for any remote transport, and without it the off-box push dies
# with "Failed to exec ssh: No such file or directory" AFTER a successful dump —
# which reads like a network problem and isn't one.
# (scripts/backup-db.sh) runs INSIDE this container, where SUPABASE_DB_URL and
# the Supabase network are available. See docs/BACKUP-RUNBOOK.md.
RUN apk add --no-cache curl tini bash postgresql-client rsync openssh-client
# PM2 is installed globally so the entrypoint can call `pm2 start` to # PM2 is installed globally so the entrypoint can call `pm2 start` to
# boot all three pollers (NBA / WNBA / MLB) alongside the Express API. # boot all three pollers (NBA / WNBA / MLB) alongside the Express API.
@@ -57,8 +64,8 @@ COPY content ./content
# Persistent volume for JSONL training data (resolutions survive # Persistent volume for JSONL training data (resolutions survive
# redeploys via the Coolify mount). PM2_HOME lives outside it so # redeploys via the Coolify mount). PM2_HOME lives outside it so
# supervisor state is local to the container. # supervisor state is local to the container.
RUN mkdir -p /app/data/training /app/.pm2 \ RUN mkdir -p /app/data/training /app/.pm2 /app/backups \
&& chown -R vyndr:vyndr /app/data /app/.pm2 \ && chown -R vyndr:vyndr /app/data /app/.pm2 /app/backups \
&& chmod +x /app/scripts/docker-entrypoint.sh && chmod +x /app/scripts/docker-entrypoint.sh
USER vyndr USER vyndr
+290
View File
@@ -0,0 +1,290 @@
# VYNDR — CANONICAL STATE FILE
Read-only ground-truth audit. Written 2026-07-26 by Claude Code (repo + deployed DB access).
Every item tagged **VERIFIED** (file/line or measured number), **CANNOT DETERMINE** (with reason),
or **BLOCKED** (with what unblocks it). Where a prior claim conflicts with code/data, the code/data wins.
Nothing was built, changed, deployed, or migrated by this audit.
---
## REVIEW ZERO — PREMISE
- **0.1 What I can read** — VERIFIED. The **repo** (full source), the **deployed Supabase DB**
(read-only via MCP, project `zmdnczhtdxcddsxzttub`), and the **deployed API** (`api.vyndr.app`).
- **0.2 Key reachability** — VERIFIED. `PROPLINE_API_KEY_*` are **NOT on the box** (`.env` absent);
they were pasted in-session earlier this conversation and are usable for probes. `ODDS_API_KEY`
is on the box but **exhausted (0/500)**. `ODDSPAPI_KEY` not on box. DB-dependent items were run
against prod directly, so nothing here is BLOCKED on PropLine keys.
---
## PHASE 0 — THE ROI RECONCILIATION (headline)
- **0.1 ROI computed in code?** — VERIFIED. **Not for the model ledger.** ROI exists only for
**user bet-tracking**: `performanceService.js:42` (`roi: stats.roi`) and `betService.js:172-174`
(`profit = payout - amount`). There is **no ROI computation over `ledger_entries`** (the model's
public record). That is why the model's ROI has been invisible.
- **0.2 Price-at-grade persisted?** — VERIFIED. **YES**`ledger_entries.locked_odds` (text,
American), written by `ledgerService.recordPipelineGrades`. Populated on essentially all settled
rows (5 nulls on B). **ROI IS computable for accrued history.**
- **0.3 11-point index persisted?** — VERIFIED (nuanced). **NOT in `ledger_entries`** — only the
collapsed letter (`grade`); `_grade_11` is explicitly `delete`d at `gradeSlateService.js:97`
before the ledger write. **BUT it IS retained in `model_snapshots.grade_11`** (migration
`025:51`, `retentionService.js:117`) — 2,888 of 4,600 snapshot rows carry it. So sub-tier is
**lost on the settled-outcome table but recoverable** by joining `model_snapshots` to
`ledger_entries` on (player, stat, line, date). STATE.md's "grade_11 is stored" and this audit's
"deleted before ledger write" are BOTH correct — different tables.
- **0.4 Price distribution by grade (American odds, settled)** — VERIFIED:
| sport | grade | n(decided) | median | q1 | q3 |
|---|---|---|---|---|---|
| MLB | B | 285 | 270 | 650 | 155 |
| MLB | C | 176 | 169 | 500 | +109 |
| WNBA | B | 208 | 124 | 145 | 110 |
| WNBA | C | 160 | 120 | 134 | 108 |
Blended B median 160, C median 130. Overall range 10000 … +600. MLB grades skew to **deep
favorites**; WNBA grades cluster tight around 120.
- **0.5 Flat-stake unit ROI by grade (1u/row at `locked_odds`, decided rows hit/miss, voids
excluded)** — VERIFIED:
| sport | grade | n | hit% | **ROI%** |
|---|---|---|---|---|
| MLB | B | 285 | 70.2 | **1.02** |
| MLB | C | 176 | 60.8 | **+4.57** |
| WNBA | B | 208 | 52.4 | **4.79** |
| WNBA | C | 160 | 51.9 | **5.26** |
| A / D / F | (all) | 1 / 2 / 4 | 0 / 0 / 0 | 100 each (negligible n) |
Including voids as net-0 (my first pass) gives blended B 2.31%, C 0.10%.
**CONCLUSION — the premise's binary is a false dichotomy; the truth decomposes:**
- MLB **B** = high-hit (70%) **break-even favorites** — the premise's "favorites at fair prices,
zero edge" hypothesis is CONFIRMED here.
- WNBA **B/C** = **losing** (~52% hit at ~120 where break-even is ~54.5%).
- MLB **C** = **genuinely +EV (+4.57% on 176 decided)** — a real edge the blended "no ROI" masked.
So: ROI is not uniformly zero-edge. It is **not computed in code**, and when computed here it shows
**one profitable segment (MLB-C) hidden under WNBA losses and break-even MLB-B**.
- **0.6 64/56 population** — VERIFIED. The displayed `/api/ledger/accuracy` figures are
`hits/(hits+misses)`, **excluding voids** (88 void rows, 8.8%) and unsettled (67). Distinct
`outcome` values are **{hit, miss, void, null}** — **no `push`** (pushes structurally impossible:
half-number lines). The ROI table above uses the SAME decided (hit/miss) population, so hit-rate
and ROI are comparable. Blended B hit is 63.1% on `hits/(hit+miss)` but 55.9% on all-settled
(the 64 void B rows are the gap).
- **0.7 Hit rate in UI/marketing?** — VERIFIED. Displayed on **≥10 surfaces**: `app/page.tsx`
(landing), `GradeCard`, `TopSignals`, `vyndr/TierRecord`, `vyndr/ModelRecord`, `vyndr/AccuracyBadge`,
`game/[id]/page`, `u/[handle]/portrait`, `ledger/page`, `scan/page`. **ROI / edge is displayed
NOWHERE.** Honesty gap: users see "63% / 70% HIT" with no indication MLB-B is 1% and WNBA is 5%.
- **0.8 CLV computed / closing persisted?** — VERIFIED (with a broken-ness caveat). `closing_odds`
on **962/996** rows, `clv` on **842/996** — so CLV IS computed and closing IS persisted. BUT
`closingCapture.js:7-8` documents it as **effectively broken**: `closing_line == locked_line on
92% of rows` (only ~56 rows show real movement), because `captureClosing` overwrites the field
with the current feed on every snapshot and most props leave the feed near their lock. CLV exists
but is **largely degenerate (≈0)**.
---
## PHASE 1 — THE GRADED LINE
- **1.9 Selector location** — VERIFIED. Two stages: (a) `gradeSlateService.dedupeProps`
(`gradeSlateService.js:34`, called `:144`); (b) the `snapshotService` dedup at
`snapshotService.js:425-434`.
- **1.10 Reads over which set?** — VERIFIED. **One provider at a time** — PropLine-normalized rows
(primary) via `oddsService``recordDownstream``gradeAndCacheSlate`. Falls back to
odds-api / oddspapi only if PropLine fails (`oddsService` fallback chain). Not all providers merged.
- **1.11 dedupeProps exists?** — VERIFIED. **YES, it EXISTS** (`gradeSlateService.js:34`). Logic:
Set on key `` `${player}::${stat_type}::${line}` `` — keeps the **FIRST** row per (player, stat,
line), discards later rows at the same (player, stat, line) (i.e. other books at the same line),
caps at `limit`. Written **Session 32** (`f0c8b4f`, "Grades pipeline + NFL/NHL wiring"). This
settles the asserted/un-asserted question: **it is real, not imagined.**
- **1.12 snapshotService 411-434** — VERIFIED. Confirmed: dedup by `` `${nameKey}|${stat_type}` ``,
keeps the row with the **highest `confidence`**. `confidence` derives from the **grade letter**
(band-midpoint from `grade_thresholds.json`, `confidence_basis:'grade_band'`) — so "highest
confidence" == "highest grade letter". It carries no information beyond the letter.
- **1.13 consensus vs first-book** — VERIFIED. **NO consensus rule exists in code.** The graded line
is chosen by **first-book-at-each-line (dedupeProps) then highest-grade-across-lines
(snapshotService)**. The "book-agnostic consensus rule" description is a prior order's *proposal*,
never built. The "first book" description is CLOSER to correct. **Code wins: no consensus.**
- **1.14 5 books at 3 lines → which line?** — VERIFIED (walked). `normalizeProps` emits one row per
book → `dedupeProps` keeps the first book at each of the 3 distinct lines (3 survivors) →
`gradeAndCacheSlate` grades both sides of each → `snapshotService` keeps the **highest-grade** of
the 3. **The graded line = whichever of the 3 lines grades highest** (a best-grade-for-us
selection, active now that the feed is 27% MLB / 82% WNBA multi-book).
- **1.15 Consumers assuming one book/prop** — VERIFIED (partial list): the grade path
(`dedupeProps` collapses book multiplicity), `analyzeViaEngine1` (`book_odds`/`fair_prob` use the
single surviving row's odds), the snapshot GameCard overlay, `detectBestBook` (no-ops <2 books).
- **1.16 Test pinning graded line vs book-set change?** — VERIFIED: **NONE.** `dedupeProps` and
`snapshotService` have unit tests for dedup mechanics, but **no test pins graded-line stability
against a change in the book set** (a multi-book different-line scenario).
- **1.17 Ledger flag for pre/post book-set change?** — VERIFIED. **No dedicated flag.** `model_version`
stamps the model era and challenger-version columns exist, but **nothing distinguishes grades by
book-set basis.** A silent book-set change would not be visible on the ledger.
---
## PHASE 2 — THE FEED
- **2.18 Books-per-prop TODAY (2026-07-26, live `/api/odds`)** — VERIFIED, and it **CORRECTS the
prior "137/140 single-book, DK129/FD14" measurement — the feed has broadened:**
- **MLB**: n=322 → `{1 book: 234 (73%), 2: 51, 3: 26, 4: 10, 5: 1}`. Books: **draftkings 273,
betmgm 105, betrivers 49, pinnacle 30, fanduel 2.** So 73% single-book (mostly DK), 27%
multi-book, **5 books now present** (not DK-only).
- **WNBA**: n=163 → `{1 book: 29 (18%), 2 books: 134 (82%)}`. Books: **fanduel 149, draftkings 148.**
**WNBA is graded off a genuine 2-book feed (DK + FD)** — better multi-book coverage than MLB.
- **2.19 ALLOWED_BOOKS** — VERIFIED. `oddsNormalizer.js:9` = 11 books (draftkings, fanduel, betmgm,
caesars, fanatics, bet365, hardrockbet, pointsbet, betrivers, pinnacle, thescore), applied at
`:110`/`:198`/`:235`. **Drops nothing today** (every book in the live feed is on the list).
- **2.20 PropLine request** — VERIFIED. `proplineAdapter.js:128` `buildUrl = ${BASE}/${sportKey}/odds`;
`:151-152` params = `{ apiKey, markets }`. **No `regions`/`bookmakers` param is sent.** The feed
broadening (2.18) happens on PropLine's DEFAULT response, not a param we added.
- **2.21 PropLine docs / multi-book param** — CANNOT DETERMINE (not re-tested this order). Prior
research (WebSearch): PropLine advertises "13 books + 5 exchanges; every payload includes a
bookmakers array" and is The-Odds-API-compatible (which uses `regions`/`bookmakers`). Whether a
param unlocks the full set on our tier is **unconfirmed** — would need a keyed test with the param.
---
## PHASE 3 — THE CHALLENGER LEDGER
- **3.22 Row counts** — VERIFIED:
| challenger | total | settled | first row | note |
|---|---|---|---|---|
| arch-v1 | 164 | 128 | 2026-07-21 | 94 with non-zero delta |
| contact-v1 | 122 | 86 | 2026-07-23 | 107 non-null `p_win_contact` (15 abstained) |
| proj-v1 + proj-v1.1 | 76 + 46 = 122 | 86 | 2026-07-23 | 119 with `proj_point` |
**All MLB** (challengers are statcast-gated → MLB only; the 407 WNBA rows carry none). Public
ledger total = **996 rows** (929 settled, 67 unsettled). **The earlier "ZERO settled p_win /
measurement not begun" is now STALE — 86-128 settled per challenger.**
- **3.23 Population rate** — VERIFIED. 872 rows in the last 14 days across **12 active days** (2 days
had no rows). Challengers populate a fraction of rows (arch 164/996) — the rest are pre-deploy or
WNBA.
- **3.24 Idempotent-lock lag** — VERIFIED, **still structural.** The ledger upsert is
`ignoreDuplicates:true`, so a challenger's fields land ONLY on rows first written AFTER that
challenger's code deployed (arch 07-21, contact/proj 07-23); re-running a snapshot never
backfills challenger fields onto an already-locked row. Any challenger added later inherits the
same gap. **Settlement itself is healthy** (`stale_unsettled = 0`).
- **3.25 Silent-failure / abstain path** — VERIFIED: **none producing fake accrual.** contact-v1's
nulls are honest abstentions (thin/absent Statcast); proj-v1 projects 119/122; arch-v1's
non-moved rows are byte-identical-to-champion (no distinctive axis). No caught-throw-writes-null
path masquerading as accrual.
- **3.26 Void/DNP rate** — VERIFIED. **88 void (8.8%)** + 67 unsettled of 996. Matches the ~9% premise.
- **3.27 model_snapshots / archetype / opp_rank** — VERIFIED (partial). `model_snapshots` is LIVE:
**4,600 rows, latest 2026-07-26 22:01**, `grade_11` on 2,888. Archetype + `opp_rank_stat` were
verified live per-sport (MLB + WNBA) in prior orders; **CANNOT re-confirm per-sport freshness here
without additional queries** (not run to keep this pass bounded).
---
## PHASE 4 — WHAT IS ACTUALLY WIRED
- **4.28 SportsGameOdds wired?** — VERIFIED. **NOTHING.** No SGO reference in `src/`, `web/src/`,
or `.env`. Not wired, configured, committed, or deployed. (The audit that qualified it as a source
was report-only.)
- **4.29 Design surfaces live vs designed** — VERIFIED:
- **S2 (book comparison / crown / disagreement):** `BookComparison.tsx` EXISTS in
`web/src/components/` but is **NOT imported/routed anywhere (dead component)**. `BookChip` +
`BookWordmark` exist and are used. `MovementStrip`, `CrownBadge` → **NOT FOUND** (never built).
(Corrects an earlier order that said "BookComparison doesn't exist" — it exists, just unrouted.)
- **THE WIRE:** = the **daily newsletter/content format** (see 4.30). Newsletter engine built
(`newsletterService`); send is unscheduled (internal endpoint only).
- **S3 article media:** **NOT FOUND** (no article-media generator; `mediaEngine` only has wire/
share text templates).
- **S-2 Offseason hub / SeasonBoard:** **NOT FOUND** (never built).
- **System / Intelligence:** design files present in `specs/design-reference/`; live components partial.
- **4.30 What is THE WIRE?** — VERIFIED. It is **VYNDR's daily newsletter / content voice**, not a UI
surface: `newsletterService.js:219-220` (`THE WIRE — {date}`) + `mediaEngine.js:6,121`
(`MORNING WIRE`, `SIGNAL` deterministic templates). It's the editorial format for the daily report.
- **4.31 detectBestBook + LineSparkline** — VERIFIED. `slateAdapter.detectBestBook` returns a book
only when ≥2 books post the SAME line at differing prices (no-ops on 1 book, never marks a lone
price "best"). `StatStrip.LineSparkline` (`StatStrip.tsx:130`) renders only at **≥3 history
points**, returns `null` below. Both degrade honest-absent on single-book/shallow data.
- **4.32 Line history / closing overwrite** — VERIFIED. `history` (`{t,line}`, line-deduped, cap 24)
is shallow — most props sit at 1-2 flat points (line moves are rare; books move odds, not the
half-point line). Two distinct closing mechanisms: the **`closing_captures` TABLE is INSERTED /
appended** (`closingCapture.js:234`), but the **`ledger_entries.closing_line/closing_odds` is
OVERWRITTEN every snapshot** (`ledgerService.js:18-21` doc; overwrite is why CLV is degenerate — 0.8).
- **4.33 Migration drift** — VERIFIED, **still present.** Repo `supabase/migrations/` has
**001-022, 025, 030, 031, 032**. **MISSING: 023, 024, 026, 027, 028, 029** — applied to prod but
untracked (they carry the challenger columns `p_win_challenger`/`challenger_*`, `model_version`,
`env_*`, `archetype_vector` — confirmed live in prod). The repo does NOT reflect prod schema; check
`information_schema.columns` via MCP, not the repo, before schema work.
- **4.34 WNBA grading live?** — VERIFIED. **YES.** 407 WNBA ledger rows, **376 settled**, graded off
a 2-book (DK+FD) feed. WNBA carries no challenger rows (statcast-gated MLB-only).
---
## PHASE 5 — DOCTRINE AND THE BOARD
- **5.35 CLAUDE.md (read in full)** — VERIFIED. Standing rules & findings:
- **NO CODE WITHOUT A SPEC**; 5 quality gates; WSL2 heredoc rule (python3 for >10-line files);
update BUILD-STATE.md / BLOCKERS.md.
- **Data Semantics Rule:** VYNDR never generates lines/odds — market values are REAL captured book
numbers; only model_value/grade/edge are model output. `Number(null)===0` is the recurring
fabrication bug; use strict null guards.
- **Grade internals:** grade is an additive integer index (`engine1`, NEUTRAL_INDEX 3), moved by
flat ±1.0/±0.5 deltas; A needs sum ≥+4.5, D ≤−1.51. `confidence` is NOT a probability (grade-band
midpoint). `p_win` is the real signal. **NEVER rescale thresholds to mint A's (permanent founder
ruling).** **A-RATED marketing on hold** until a prod fingerprint shows real A grades.
- **Three stat_type whitelists must stay in sync** (analyze.js, scan.js, validation.py).
- **Snapshot pipeline** is the product model (scheduled grade → lock to line → read from cache);
on-demand "Read" retired. SNAP_TTL 24h.
- **DO-NOT-WIRE / DO-NOT-TOUCH:** Tank01 player props = empty (do not wire); ParlayAPI host dead;
`gameLogService.getGameLogs` returns null for MLB (a trap — use `featureCache.getStatRows`);
`mlbGrader.js` is dead code; the legacy `--grade-a` token alias block kept until consumers migrate.
- **Three separate MLB stat maps on purpose** (featureCache / outcomeService / liveTrackingService) —
do not merge. Settlement is MLB-only (WNBA/NBA/soccer never settle via that path — but WNBA IS
grading + settling per 4.34, so confirm the settle path).
- **5.36 Running task list / open items** — VERIFIED. Primary: **`specs/STATE.md`** (1,886 lines,
"STATE OF THE WORLD," CURRENT STATUS + OPEN ITEMS block, last dated 2026-07-22). Also
`BUILD-STATE.md`, `BLOCKERS.md`, `DECISIONS.md`, `AUTONOMY.md`, `PROMISE-AUDIT.md`, `ROADMAP.md`.
**Open threads I have been tracking across recent orders that a strategist chat may not have:**
(a) challenger measurement now HAS settled rows (arch 128 / contact 86 / proj 86) — promotion is a
future per-prop-type ledger decision; (b) the design-migration arc (Landing hero migrated; scanner
S6/S7 blue-channel reskin shipped; S2/THE WIRE/Offseason/article-media are GAP/unbuilt);
(c) multi-book: SGO qualified (report-only), nothing wired; (d) **ROI is uncomputed in code and
decomposes to MLB-C +4.57% / MLB-B 1% / WNBA 5%** (this file).
- **5.37 Spec / roadmap files** — VERIFIED (in `specs/`): `STATE.md` (running state),
`VYNDR-NORTH-STAR.md`, `DESIGN-SPEC.md`, `ROW-GRAMMAR.md` (row grammar law), `LIVE-TRACKING.md`,
`VOICE.md`, `model-train.md`, `phase-0-kill-the-lies.md`, `phase-1-truth-infrastructure.md`,
`propline-audit.md`, `feature-1-1…4-1` + `a1-s3/s7/s9/s10` feature specs, `combat-intelligence.md`,
`design-reference/` (the design bundle). Root docs: `CLAUDE.md`, `DECISIONS.md`, `ARCHITECTURE.md`,
`AUTONOMY.md`, `BACKEND_HANDOFF.md`, `BUILD-STATE.md`, `BLOCKERS.md`, `PROMISE-AUDIT.md`.
- **5.38 Standing decisions a new session might contradict** — VERIFIED: `DECISIONS.md`
(DECISION-001+ architecture log); CLAUDE.md's permanent rulings (never mint A's; A-RATED hold;
data-semantics; do-not-wire list); STATE.md open items (A does not emit in prod; EV overconfident/
unvalidated — hero ranks on ev_pct and picks the most overconfident read; grade_11 stored in
model_snapshots). `AUTONOMY.md` = the zero-touch loop trace.
---
## CLAIMS I WAS ASKED ABOUT THAT TURNED OUT TO BE FALSE OR UNSUPPORTED
1. **"the ledger reports … 597 settled with 'no ROI'"** — the population is now **929 settled / 996
total** (597 was an earlier snapshot); and **ROI IS computable** (`locked_odds` persisted) — it is
simply **not computed in the ledger code**. "No ROI" = no computation, not incomputable.
2. **"Either ROI is not computed, or the model is selecting heavy favorites at fair prices"** — the
binary is false; **both are partially true and it decomposes by sport/grade**: MLB-B = high-hit
break-even favorites (the hypothesis), WNBA = losing, **MLB-C = genuinely +4.57% EV**. Not
uniformly zero-edge.
3. **"the graded line is chosen by a book-agnostic consensus rule"** — **no consensus rule exists in
code.** It is first-book-at-line (`dedupeProps`) + highest-grade-across-lines (`snapshotService`).
4. **Prior "137/140 single-book, draftkings 129 / fanduel 14" (~98% single-book)** — **corrected**:
today MLB is 73% single-book with **5 books present** (DK/BetMGM/BetRivers/Pinnacle/FanDuel), and
**WNBA is 82% two-book (DK+FD)**. The feed broadened.
5. **"the 11-point index is lost / unrecoverable"** — **it is retained in `model_snapshots.grade_11`**
(2,888 rows); lost only from `ledger_entries`. Recoverable via join.
6. **"ZERO settled p_win yet / challenger measurement not begun"** (from prior challenger orders) —
**stale**: arch-v1 128, contact-v1 86, proj-v1 86 rows are now settled.
7. **"BookComparison doesn't exist, only BookChip"** (an earlier order) — **BookComparison.tsx
EXISTS**, it is just unrouted/dead.
8. **CLV framed as "held / not computed"****CLV IS computed** (842 rows) and closing IS persisted
(962), but it is **largely degenerate** (closing==locked on 92%) due to the overwrite — a
different problem than "not computed."
9. **WNBA implicitly treated as not-really-grading** (settlement described as MLB-only in CLAUDE.md) —
**WNBA IS grading AND settling** (376 settled). The doctrine note about MLB-only settlement is
contradicted by the data; the WNBA settle path should be confirmed.
---
*End of canonical state. Regenerate the measured numbers before citing them in a later session —
they move as the ledger accrues. Structural facts (file/line, schema, wiring) are stable until code changes.*
+90
View File
@@ -0,0 +1,90 @@
# MLB ARCHETYPE AXES — Layer 2
A player is a **blend across independent axes**, not one label. Skubal is a
STARTER *and* a strikeout arm *and* a ground-ball arm *and* a control arm —
four true things at once. Single-label classification is lossy, and its failure
mode is the FLEX disease: when nothing matches, invent a bucket.
**There is no fallback.** Unremarkable on an axis → absent on that axis. A
genuinely average player surfaces nothing, and says so.
## Axis independence — measured, not assumed
Correlations over the live store (467 batters PA≥50, 531 pitchers IP≥10).
**|r| ≥ 0.70 = one underlying trait → collapsed**, so one trait is never shown
as two archetypes.
| Pair | r | Decision |
|---|---|---|
| batter k% ~ whiff% | **+0.89** | collapsed → one SWING-AND-MISS axis |
| batter hard-hit% ~ avg exit velo | **+0.88** | folded into POWER as intensity |
| batter chase% ~ swing% | **+0.87** | collapsed → one AGGRESSION axis |
| batter chase% ~ bb% | **0.72** | same axis inverted (discipline) |
| pitcher k% ~ whiff% | **+0.76** | collapsed → one STRIKEOUT axis |
| pitcher gb% ~ fb% | **0.73** | collapsed → one signed tilt axis |
| batter barrel% ~ hard-hit% | +0.70 | at the line; barrel leads |
| **pitcher velo ~ k%** | **+0.14** | **INDEPENDENT** — velo earns its own axis |
| pitcher velo ~ whiff% | +0.07 | independent |
| pitcher velo ~ gb% | +0.07 | independent |
| **pitcher k% ~ gb%** | **0.10** | **INDEPENDENT** — PUNCHOUT ⊥ SINKER |
| batter barrel ~ launch | +0.23 | independent |
| batter k% ~ chase% | +0.05 | independent |
The velocity result matters: **velo is not a proxy for missing bats.** A hard
thrower who misses no bats is a real, distinct type.
## Cut-lines
`p75` = distinctive · `p90` = elite, from the measured distribution — **per
role where the tails differ** even when medians agree (reliever GB% p90 = 54.1
vs starter 48.9, both median 42.5). Sample floors: **PA ≥ 50 · IP ≥ 10**.
## Every baseball name accounted for
| Name | Status | Axis / reason |
|---|---|---|
| WORKHORSE | **TIER** | role=starter + high IP; the old hardcoded default is gone |
| PUNCHOUT / WHIFF | **BUILT** | strikeout (p75 / p90) |
| SINKER / SEAM | **BUILT** | ground_ball |
| FLY BALL / ELEVATOR | **BUILT** | fly_ball |
| SURGEON ARM / PINPOINT | **BUILT** | control (low BB%) |
| NIBBLER / SCATTERGUN | **BUILT** | wild (high BB%) |
| BAIT / TRAPDOOR | **BUILT** | chase |
| HITTABLE / BATTING PRACTICE | **BUILT** | contact_allowed |
| CANNON / HOWITZER | **BUILT** | velocity — unlocked by the 53%→99% velo fix |
| SIDEARM / SUBMARINE | **BUILT** | slot (low arm angle) |
| STARTER / RELIEVER / CLOSER / SETUP | **BUILT** | role, from real usage |
| SLUGGER / BOMBER | **BUILT** | power (barrel%) |
| TECHNICIAN / SURGEON | **BUILT** | contact (low K%) |
| WHIFF RISK / WINDMILL | **BUILT** | swing_miss |
| GRINDER / SNIPER | **BUILT** | patience |
| FREE SWINGER / HACKER | **BUILT** | aggression |
| TOPSPIN / LOFT | **BUILT** | launch |
| SLASH / BARREL FINDER | **BUILT** | line_drive |
| HAMMER | **ALIAS** | of chase/strikeout — best-pitch whiff r=0.50 with overall; not independent enough for its own axis |
| BRUSH | **ALIAS** | of contact (TECHNICIAN) |
| DRIVER | **ALIAS** | of power — RBI is lineup context, not a player trait |
| **FLEX** | **RETIRED** | it was the fallback, never earned: `utility` had zero writers |
| MIRROR | **SHELVED** | needs switch-hitter axis; `bats='S'` now joined — unlock when a platoon-split feed lands |
| GHOST / BURNER / LEG | **SHELVED** | speed: SB is statsapi, not yet joined into the aggregate store |
| CATALYST | **SHELVED** | table-setter: needs lineup slot (not ingested) |
| HYBRID | **BUILT-ADJACENT** | two-way = both role profiles exist (position `TWP`) |
| ALPHA | **ALIAS** | of strikeout+control combined; the blend expresses it |
| FASTBREAK · ENFORCER · CHAINMOVER · CONDUCTOR · ARCHITECT · TORCH · LINK · ARTILLERY · BELL COW · etc. | **OTHER SPORT** | NBA/WNBA/NFL names — correctly not built here |
Zero orphans.
## Output shape
```js
{
role, roleDetail, roleLabel, sufficient, sample, bats, throws,
vector: { axisKey: {tier,label,value,strength} | null }, // Layer 3 reads ALL
blend: [ {label, axis, tier, value, strength} ], // top ≤3 to surface
absent: [axisKey], // NO DATA — distinct from unremarkable (vector null)
note // honest copy when the blend is empty
}
```
`absent` vs a `null` vector entry is a real distinction: absent means we could
not measure it, null means we measured it and he is ordinary.
+69 -20
View File
@@ -4,29 +4,68 @@ Supabase free tier has **zero** backups (no scheduled, no PITR). `scripts/backup
is the safety net: a nightly full-database `pg_dump`, 14 days kept locally, a is the safety net: a nightly full-database `pg_dump`, 14 days kept locally, a
weekly copy pushed off-box, ntfy alert on any failure. weekly copy pushed off-box, ntfy alert on any failure.
## What Kev needs to set (one env var) ## Where it runs
**`SUPABASE_DB_URL`** — the Supabase **direct** connection string (session mode). `SUPABASE_DB_URL` is set in Coolify **on the VYNDR API service** — so the backup
Supabase → Project → Settings → Database → **Connection string → URI**, the runs **inside that container**, which already has the env, the Supabase network,
`db.<ref>.supabase.co:5432` one (NOT the `:6543` transaction pooler — `pg_dump` and now `pg_dump`/`pg_restore`/`rsync` (added to the Dockerfile). The Hetzner box
needs a real session). Paste it in Coolify as an env var; never commit it. can't reach `db.<ref>.supabase.co` directly and doesn't hold the connection
string; the container is the right place.
Optional: `BACKUP_REMOTE` (off-box rsync target for the weekly copy — see below). **`SUPABASE_DB_URL`** must be the **direct** connection string (session mode):
Supabase → Settings → Database → **Connection string → URI**, the
`db.<ref>.supabase.co:5432` one (NOT the `:6543` pooler — `pg_dump` needs a real
session). Already set. Optional: `BACKUP_DIR` (mount a Coolify **persistent
volume** here so dumps survive redeploys — e.g. `/var/backups/vyndr`),
`BACKUP_REMOTE` (off-box rsync target, below).
## Install the cron (on the Hetzner box) ## ✅ THE CRON NOW SHIPS AS CODE (Session 64) — no install step
**Read this before following the manual instructions below; they are now the
FALLBACK, not the primary path.**
`src/backupScheduler.js` runs the nightly backup **inside the API container**,
armed from `server.js` at boot. The container already holds `SUPABASE_DB_URL`,
`pg_dump` and the Supabase network route, so **deploy == installed**. Nothing to
add to crontab, nothing to click in Coolify.
- **Arming is opt-OUT:** armed whenever `SUPABASE_DB_URL` is set. The S62 design
was opt-in (a host cron someone had to add) and nobody ever added it — the
database went unbacked every night for weeks. That failure mode is now
impossible.
- **Kill switch:** `BACKUP_CRON=0`.
- **Schedule:** `BACKUP_HOUR_UTC` (default 3) / `BACKUP_MINUTE_UTC` (default 10).
- **Failure pages high-priority ntfy.** Silence is the danger with backups.
- Boot log line: `[backupScheduler] armed — nightly 03:10 UTC ...`
### ⚠️ DURABILITY — the one thing still requiring a human
The container filesystem is **ephemeral**: a dump written inside it is LOST on the
next redeploy. The scheduler detects this and pages a warning at boot when
neither is configured. Set ONE of:
1. **`BACKUP_REMOTE`** — off-box rsync target (Hetzner Storage Box, ~€3/mo). Best.
2. **`BACKUP_DIR`** pointed at a **Coolify persistent volume** (e.g. `/var/backups/vyndr`).
Until one is set, backups run but do not survive a deploy. An undurable backup
that reads as "backed up" is worse than a loud gap — hence the boot-time page.
---
## Install the cron (host → docker exec into the API container)
Find the API container name (`docker ps | grep vyndr`), then a host cron:
```bash ```bash
# 1. Ensure the client is present # Nightly at 03:10 UTC. BACKUP_REMOTE can also be set in Coolify instead.
sudo apt-get install -y postgresql-client rsync
# 2. Nightly at 03:10 UTC. Pass the env the script needs (or source an env file).
sudo crontab -e sudo crontab -e
# add: # add (replace <api-container>):
10 3 * * * SUPABASE_DB_URL='postgresql://...' BACKUP_REMOTE='u123456@u123456.your-storagebox.de:vyndr-backups/' /path/to/vyndr/scripts/backup-db.sh >> /var/log/vyndr-backup.log 2>&1 10 3 * * * docker exec -e BACKUP_REMOTE='u123456@u123456.your-storagebox.de:vyndr-backups/' <api-container> sh /app/scripts/backup-db.sh >> /var/log/vyndr-backup.log 2>&1
``` ```
(If the cron runs inside the Coolify container instead, the env vars are already `docker exec` inherits the container's env (`SUPABASE_DB_URL`) + network + the
present — just schedule `scripts/backup-db.sh`.) newly-installed `pg_dump`. (If you'd rather, Coolify's **Scheduled Tasks** can run
`sh /app/scripts/backup-db.sh` on the API service on the same cadence.)
## Off-box target (simplest reliable pick) ## Off-box target (simplest reliable pick)
@@ -38,14 +77,24 @@ rather keep it off Hetzner entirely — swap the `rsync` line for `rclone copy`.
## Restore / FINGERPRINT (proves it's a real backup, not just a file) ## Restore / FINGERPRINT (proves it's a real backup, not just a file)
A dump only counts once a restore of it succeeds. Load one into a scratch DB: **Built-in (runs on every backup):** the script validates each dump with
`pg_restore --list` — a dump that isn't a valid archive, or that doesn't contain
`ledger_entries`, is treated as a FAILURE and paged. So every successful run has
already proven the archive parses and holds the ledger.
**Full restore proof (run once to fingerprint):** from the host, restore the
newest dump into a throwaway postgres and count the ledger:
```bash ```bash
# spin a throwaway local postgres, restore the newest dump, count a known table # 1. Produce a dump on demand (writes into the container's BACKUP_DIR)
docker exec <api-container> sh /app/scripts/backup-db.sh
# 2. Copy the newest dump out of the container
newest=$(docker exec <api-container> sh -lc 'ls -t /var/backups/vyndr/vyndr-*.dump | head -1')
docker cp "<api-container>:${newest}" /tmp/vyndr-latest.dump
# 3. Restore into a scratch postgres + count a known table
docker run -d --name vyndr-restore-test -e POSTGRES_PASSWORD=x -p 55432:5432 postgres:15 docker run -d --name vyndr-restore-test -e POSTGRES_PASSWORD=x -p 55432:5432 postgres:15
sleep 5 sleep 6
newest=$(ls -t /var/backups/vyndr/vyndr-*.dump | head -1) pg_restore --no-owner --no-privileges -d "postgresql://postgres:x@localhost:55432/postgres" /tmp/vyndr-latest.dump
pg_restore --no-owner --no-privileges -d "postgresql://postgres:x@localhost:55432/postgres" "$newest"
psql "postgresql://postgres:x@localhost:55432/postgres" -c "select count(*) from public.ledger_entries;" psql "postgresql://postgres:x@localhost:55432/postgres" -c "select count(*) from public.ledger_entries;"
docker rm -f vyndr-restore-test docker rm -f vyndr-restore-test
``` ```
+65
View File
@@ -0,0 +1,65 @@
# `/api/content-studio` — the agent-ready contract
The endpoint Kev's `/studio` page reads today and an autonomous poster reads
later. **The page is a thin client**: no posting logic, no fact handling. Wiring
a bot means pointing it here — nothing on this side changes.
Distinct from `/api/content` (Session 29), which serves structured content
*objects* by data level. This serves finished **posts**.
## Auth
`x-internal-key: $VYNDR_INTERNAL_KEY` — private, never public. The browser never
holds the key; the Next proxy at `web/src/app/api/content-studio/[...path]`
attaches it server-side.
## `GET /api/content-studio/:date?`
`:date` optional, defaults to today ET.
```jsonc
{
"date": "2026-08-07",
"count": 3,
"posts": [{
"id": "honesty_flex",
"label": "The Honesty Flex",
"sport": "mlb",
"status": "pending", // pending | approved | skipped | regenerate_requested
"ok": true,
"skipped": false,
"reason": null, // why it was skipped, when it was
"honest_absence": false, // a real "nothing tonight" post, not a failure
"copy": "WE GRADED 2140 PROPS TONIGHT...",
"card": { "title": "...", "lines": [...] },
"card_svg": "<svg ...>", // ready to render or rasterise
"fact_contract": ["graded", "ceiling_letter", "..."], // REQUIRED fields
"facts": { "graded": 2140, "...": "..." } // what backed it
}]
}
```
**`fact_contract` + `facts` are the point.** An agent (or a reviewer) can check
what a claim rests on instead of trusting the sentence. A post whose contract
could not be met never appears with invented values — it arrives `skipped` with
a `reason`, or as an `honest_absence`.
## `POST /api/content-studio/:date/:id/status`
```jsonc
{ "status": "approved" } // approved | skipped | regenerate_requested
```
Editorial state only, stored in Redis for 14 days. **Approving a post changes
nothing about the model** — status never touches a serving, model or ledger
table.
## For the agent build
1. `GET` the date → filter `ok && !skipped`.
2. Post `copy`; rasterise or attach `card_svg`.
3. `POST` status `approved` on success.
4. **Never** synthesise a claim not present in `facts`. The engine refuses to
render an unbacked token; an agent must not reintroduce one downstream.
Adding a template changes the payload not at all — a new `id` simply appears.
+64
View File
@@ -0,0 +1,64 @@
# git push in this environment — it was never the firewall, and never missing credentials
## What actually happened
Every push attempt this session used `git push origin main` and failed with:
```
fatal: could not read Username for 'https://github.com'
```
That error names the cause exactly, and it was misread all session as "no git
credentials on this machine." Two things were true instead:
1. **`origin` is GitHub** (`github.com/kev3109/betonblk.git`) and has **no stored
credential.**
2. **`gitea` is the working remote** (`git.builtbykev.com/builtbykev/vyndr.git`)
and **a valid credential for it was on the machine the entire time.**
The habit of typing `origin` is what kept twenty commits local. Nothing was
blocked.
## The GATE-0 firewall theory — tested and REJECTED for VYNDR
The theory was that GATE-0 (Hetzner `mastermind-core-fw`, inbound deny-by-default
except 80/443, SSH/22 restricted to Tailscale + Kev's IP) was blocking an
SSH-based push, as it did for COLYRA.
**It does not apply here. Measured:**
| check | result |
|---|---|
| `git remote -v` | **both remotes are already HTTPS** — no `git@…:…` URL anywhere |
| `curl -I https://git.builtbykev.com` | **HTTP 200** in 0.64s |
| `curl -I https://github.com` | **HTTP 200** in 0.17s |
| Gitea git endpoint over 443 | **HTTP 200** |
| GitHub git endpoint over 443 | HTTP 401 (auth required, reachable) |
There was no SSH remote to be blocked and no connectivity failure of any kind.
COLYRA's HTTPS-remote fix was the right fix for COLYRA's problem; **VYNDR was
already in the state that fix produces.**
Applying it here would have meant creating a new Gitea token to solve a problem
that did not exist — and the pre-existing credential would have made the new
token look like the cure.
## The fix
```
git push gitea main # not origin
```
Result: `6452926..ecf78b9`, 21 commits, verified by `git ls-remote` matching
local `HEAD`.
## Standing note
- **`gitea` is VYNDR's push remote.** `origin` (GitHub) is unauthenticated on this
machine and will always fail.
- Read the error text before reaching for an infrastructure theory. `could not
read Username for 'https://github.com'` is a *credential* message naming a
*specific host* — it is not a connectivity message, and it named the wrong
remote, not a wrong protocol.
- The `~/vyndr-full-history-2026-08-07.bundle` and patch series stay as
belt-and-braces. They are no longer the only copy.
+98
View File
@@ -0,0 +1,98 @@
# MECHANISM DATA — the Layer-1 pattern every sport inherits
Layer 1 of the archetype+projection build. It lands **mechanism data** — how a
player actually does what he does — and keeps it current. It does **not**
classify (Layer 2) or project (Layer 3).
MLB is the first instance. NBA tracking and NFL Next Gen slot into the same
four boxes with a different adapter and a different source id.
## The pattern
```
SOURCE ADAPTER free/public first · absent-not-zero · injectable fetch
│ no new runtime deps · one file per sport
IDENTITY BRIDGE source-native id ↔ our player_key
│ the match rate is MEASURED and REPORTED, never assumed —
│ the unmatched rate IS the honest-absent rate
AGGREGATE pure, re-runnable, explicit MIN-SAMPLE gates
DERIVATION thin ≠ missing: both stored, distinguishable
HOT STORE small, indexed, upserted on the natural key
(Supabase) + updated_at, because freshness is a truth property
```
### The five rules that make it a pattern
1. **Backfill and refresh are the same call.** A full re-pull upserted on the
natural key is idempotent and self-healing: a missed night self-corrects on
the next run. No incremental "who played today" bookkeeping to drift out of
sync. Only viable because the aggregate grain is small — which is the point
of aggregating.
2. **Aggregate grain, not raw.** One season of raw MLB per-pitch is ~0.85 GB in
Postgres against a 500 MB plan ceiling; the aggregate set is ~1,350 rows.
Raw stays retrievable from the free source if Layer 3 ever needs it.
3. **Absent is absent.** A metric the feed did not carry is `null`, never `0`.
A player below the minimum sample is stored and flagged, not dropped and not
inflated. A player we cannot join is stored with a null `player_key` and
joins later — storing him is not a claim about him.
4. **Refuse to write nothing.** If every feed returns empty that is a source
failure, not "there is no mechanism data". The job refuses the write so a
bad night can never blank a good table.
5. **Freshness is monitored.** `updated_at` on every row; the scheduler pages on
a failed run AND on silent staleness — a job that stops being scheduled never
produces a failure, so staleness must alarm on its own. **Never-built is not
stale**: different condition, different fix, and paging on a fresh install
teaches the operator to ignore the alarm.
## MLB instance
- **Adapter** `src/services/adapters/statcastAdapter.js` — five Baseball Savant
CSV leaderboards (free, public, no key). Direct `axios` + the existing CSV
parser; **pybaseball is deliberately not used** — it is a Python wrapper over
these same URLs, and the stack's Python service is already offline in prod.
- **Service** `src/services/statcastAggregateService.js``buildRows` (pure),
`refreshSeason` (the job), `getFreshness` / `isStale` (the alarm predicates).
- **Store** `statcast_aggregates` (migration 030), PK `(sport, season, source_id)`.
Promoted columns for the classification-critical metrics + a `metrics` JSONB
carrying every raw field, so Layer 2/3 can reach something we did not promote
**without a re-ingest**.
- **Schedule** nightly at `STATCAST_HOUR_UTC` (default 11 UTC ≈ 7 AM ET, after
every game is final). Kill switch `STATCAST=0`.
- **Induce** `POST /api/internal/statcast/refresh` · **probe**
`GET /api/internal/statcast/status` (internal key). We verify a refresh by
running it, never by waiting for the slot.
### Measured (2026-07-20)
| | |
|---|---|
| Batter match rate | **100%** (40/40 real players) |
| Pitcher match rate | **100%** (66/66 real roster pitchers) |
| Rows per season | **1,354** (604 batters + 750 pitchers) |
| Join rate | **1,354 / 1,354** |
| Handedness present | 677 pitchers (from the movement feed) |
| Sufficient / thin | 998 / 356 at PA≥50, IP≥10 |
| Pull time | ~5 s for all five feeds |
## Adding a sport
1. Write `src/services/adapters/{sport}Adapter.js` returning the same shape:
indexes keyed by source id, metrics strictly parsed.
2. Confirm and **report** the identity match rate before building on it.
3. Extend `buildRows` with the sport's role vocabulary and its minimum-sample
gates.
4. Add the scheduler hour and the induce endpoint.
Nothing else changes: the store, the alarm and the upsert semantics are shared.
## Commodity, not moat
Raw Statcast is public — every competitor can pull the same numbers in about a
second. Ingesting it is table stakes. The edge is Layer 2 (which mechanism
signals define an archetype, and where the boundaries sit), Layer 3
(projections built on them), and the settled ledger that proves whether any of
it predicts anything. **Having the data is not having an edge. Having it plus
an attributed record is.**
+8 -10
View File
@@ -101,7 +101,7 @@ Mounted in `src/app.js`. Auth column meanings:
These are proxies or thin wrappers; they hit Express via `BACKEND_URL` These are proxies or thin wrappers; they hit Express via `BACKEND_URL`
or the Python service via `NEXT_PUBLIC_NBA_SERVICE_URL`. or the Python service via `NEXT_PUBLIC_NBA_SERVICE_URL`.
- `/api/checkout` (POST/GET) — Stripe checkout proxy (Session 8 cutover — was NexaPay) - `/api/checkout` (POST/GET) — Stripe checkout proxy
- `/api/games/[id]` and `/api/games/tonight` — list / detail - `/api/games/[id]` and `/api/games/tonight` — list / detail
- `/api/games/[id]/props` — props for a game - `/api/games/[id]/props` — props for a game
- `/api/intelligence/feed` — homepage live signals - `/api/intelligence/feed` — homepage live signals
@@ -117,7 +117,6 @@ or the Python service via `NEXT_PUBLIC_NBA_SERVICE_URL`.
- `/api/stats/parlays-graded`, `/api/stats/public` — proxy - `/api/stats/parlays-graded`, `/api/stats/public` — proxy
- `/api/user/profile`, `/api/user/scans`, `/api/user/recent-scans` - `/api/user/profile`, `/api/user/scans`, `/api/user/recent-scans`
- `/api/waitlist` — proxy - `/api/waitlist` — proxy
- `/api/webhook/nexapay` — NexaPay webhook (legacy — Stripe cutover Session 8; webhook still listening for any in-flight NexaPay events)
### Next.js pages (Session 8 additions) ### Next.js pages (Session 8 additions)
@@ -175,9 +174,6 @@ back). Updated this session in Section 1 of Session 7c.
| `STRIPE_PRICE_DESK_FOUNDER` | ✓ | | `STRIPE_PRICE_DESK_FOUNDER` | ✓ |
| `FOUNDER_CODES` | ✓ | | `FOUNDER_CODES` | ✓ |
| `FOUNDER_CODE_EXPIRY` | ✓ | | `FOUNDER_CODE_EXPIRY` | ✓ |
| `NEXAPAY_API_URL` | ✓ (added 7c) |
| `NEXAPAY_API_KEY` | ✓ (added 7c) |
| `NEXAPAY_WEBHOOK_SECRET` | ✓ (added 7c) |
### Push ### Push
| Var | Doc? | | Var | Doc? |
@@ -427,7 +423,6 @@ Source: `grep -rn "cacheSet\|cacheGet\|redis\.set"`.
| FanDuel/BetMGM/Caesars/PrizePicks/Covers/Rotowire (legacy) | each `*Adapter.js` | none | tunable | UnifiedOddsProvider | | FanDuel/BetMGM/Caesars/PrizePicks/Covers/Rotowire (legacy) | each `*Adapter.js` | none | tunable | UnifiedOddsProvider |
| Sports-Reference HTML | `scripts/scrape-sports-reference.js` | `REF_HTML_FILE`, `COACH_HTML_FILE` (optional) | 1 req / 5s | scraper | | Sports-Reference HTML | `scripts/scrape-sports-reference.js` | `REF_HTML_FILE`, `COACH_HTML_FILE` (optional) | 1 req / 5s | scraper |
| Resend (email) | `web/src/services/email.ts` | `RESEND_API_KEY`, `RESEND_FROM_EMAIL` | n/a | transactional email | | Resend (email) | `web/src/services/email.ts` | `RESEND_API_KEY`, `RESEND_FROM_EMAIL` | n/a | transactional email |
| NexaPay | `web/src/services/nexapay.ts` | `NEXAPAY_*` | n/a | checkout fallback |
| PostHog | `web/src/lib/analytics.ts` | `NEXT_PUBLIC_POSTHOG_KEY/HOST` | n/a | browser analytics | | PostHog | `web/src/lib/analytics.ts` | `NEXT_PUBLIC_POSTHOG_KEY/HOST` | n/a | browser analytics |
| football-data.org | `footballDataAdapter.js` | `FOOTBALL_DATA_API_KEY` | 10/min (8 enforced) | poller-soccer, prefetch (TERTIARY) | | football-data.org | `footballDataAdapter.js` | `FOOTBALL_DATA_API_KEY` | 10/min (8 enforced) | poller-soccer, prefetch (TERTIARY) |
| api-football.com | `apiFootballAdapter.js` | `API_FOOTBALL_KEY` | 100/day (soft 90) | soccer cascade (PRIMARY, Session 9) | | api-football.com | `apiFootballAdapter.js` | `API_FOOTBALL_KEY` | 100/day (soft 90) | soccer cascade (PRIMARY, Session 9) |
@@ -703,10 +698,13 @@ The dual-provider divergence flagged in 7h is closed:
shipped. Express `stripeService.js` updated to point `success_url` shipped. Express `stripeService.js` updated to point `success_url`
and `cancel_url` at the new frontend pages via `NEXT_PUBLIC_SITE_URL` and `cancel_url` at the new frontend pages via `NEXT_PUBLIC_SITE_URL`
(the only backend file touched in Session 8). (the only backend file touched in Session 8).
4. NexaPay is still wired but no UI calls it. Disposition (remove vs 4. NexaPay has been PURGED (2026-07-27). VYNDR is Stripe-only — NexaPay was
keep as fallback) is a follow-up call — leaving it in place doesn't cross-project contamination copied from another venture and never a real
cost anything and gives the team a fallback if Stripe goes down VYNDR payment path. Code, route, and env references removed. Provider-side
during the World Cup window. cleanup owed to Kev: remove `NEXAPAY_*` from Coolify env, revoke the NexaPay
API key + webhook secret, and de-register the webhook at NexaPay if an
account was ever configured. DB column `user_profiles.nexapay_customer_id`
is orphaned (no reader/writer) — drop via a follow-up migration.
--- ---
+37
View File
@@ -0,0 +1,37 @@
-- 025_ledger_ruler_version.sql — Order Zero Phase 2 item 6 (APPLIED 2026-08-01)
--
-- Stamp the RULER that produced fair_prob_lock so pre- and post-change rows are
-- never pooled. fair_prob_lock is the denominator of every edge and CLV number;
-- changing how it is derived means a value computed before the change is NOT the
-- same measurement as one computed after it. Arithmetically forced, not policy —
-- the same reasoning as model_version.
--
-- v1_first_book — the incumbent: normalizeProps applies ALLOWED_BOOKS, then
-- gradeSlateService.dedupeProps keeps the FIRST surviving row
-- per player+stat+line. Whichever admitted book PropLine
-- happened to list first became the entire "market".
-- v2_consensus — median de-vigged fair_prob across >=2 REFERENCE books
-- posting BOTH sides at the SAME line. NOT LIVE.
--
-- Backfill to v1_first_book is a statement of fact: every row written to date
-- was produced by the first-book rule.
--
-- NOTE: repo migration numbering lags prod (the repo stops at 024 while prod
-- carries later ones applied via the Supabase MCP). This file records the DDL
-- for review; prod already has it.
ALTER TABLE public.ledger_entries
ADD COLUMN IF NOT EXISTS ruler_version text;
UPDATE public.ledger_entries
SET ruler_version = 'v1_first_book'
WHERE ruler_version IS NULL;
ALTER TABLE public.ledger_entries
ALTER COLUMN ruler_version SET DEFAULT 'v1_first_book';
COMMENT ON COLUMN public.ledger_entries.ruler_version IS
'Which fair-probability ruler produced fair_prob_lock. NEVER pool edge or CLV across differing values — the denominator changed. v1_first_book = first admitted book (incumbent); v2_consensus = median across >=2 reference books at the same line.';
CREATE INDEX IF NOT EXISTS idx_ledger_entries_ruler_version
ON public.ledger_entries (ruler_version);
+784
View File
@@ -0,0 +1,784 @@
# VYNDR — PRODUCT COMPLETION MATRIX
Canonical board. Supersedes STATE.md's narrative. Re-derived from repo (`b0a51c8`),
prod (`api.vyndr.app` / `vyndr.app`, checked 2026-07-27), and the design bundle
(`specs/design-reference/*`). Every cell cites a file, a route, or a prod check.
**DONE = all five YES:** DESIGNED (spec exists) · BUILT (code exists) · WIRED (a user
can reach it) · LIVE (serving in prod now) · HONEST (real data, not fabricated/inflated/
placeholder). A built-but-unrouted component is **WIRED: NO** regardless of code quality.
> **HONESTY PASS applied 2026-07-27 (6bc18d8):** every KNOWN live fabrication removed/hidden
> — see the "HONESTY PASS" section at the end. HONEST cells for rows 19/20/24 updated (✓).
## SURFACE MATRIX — 16 of ~26 fully done (Book Comparison wired 2026-07-29)
| # | Surface | DESIGNED | BUILT | WIRED | LIVE | HONEST | DONE | Evidence / what's missing |
|---|---|---|---|---|---|---|---|---|
| 1 | Landing / hero | YES | YES | YES | YES | YES | ✅ | `Landing.dc.html`; `app/page.tsx`; nav logo; prod `/`→200; Hero/TopSignals/ClaimMeter/ModelRecord all fetch real + self-hide |
| 2 | Dashboard / Slate | YES | YES | YES | YES | YES | ✅ | `Mobile.dc.html`; `Slate.tsx`+`vyndr/GameCard`; Nav "Slate"; prod `/dashboard`→200; snapshot grades real (wnba 25/mlb 5 today) |
| 3 | Scan flow (S6/S7) | YES | YES | YES | YES | YES | ✅ | `Scanner States.dc.html`; `scan/page.tsx`+`ProcessingGrade`; Nav "Read"; prod 200; real grade, honest `NoMarketState` |
| 4 | Grade result card | YES | YES | YES | YES | YES | ✅ | `System.dc.html` reveal; `GradeResultCard.tsx` (3 importers); `/scan`; real grade. Sub-sections (alt ladder, books) gated/absent — see #17/#18 |
| 5 | Grade display / badge | YES | YES | YES | YES | YES | ✅ | `HANDOFF.md:41`; `GradeBadge.tsx`; dashboard/scan/many; prod grades B/C real, no fake A |
| 6 | Ledger | YES | YES | YES | YES | YES | ✅ | `Intelligence.dc.html` S4; `ledger/page.tsx`+`ClvBadge`; Nav "Ledger"; prod 200; n≥20 gate honored |
| 7 | Tier-record | YES | YES | YES | YES | YES* | ✅ | `HANDOFF.md:33`; `TierRecord.tsx` (6 routes); n≥20 gate. *A-tier "edge glow" is ranking-as-credibility, not proven ROI (honest-but-thin) |
| 8 | Pricing | YES | YES | YES | YES | YES | ✅ | `Mobile.dc.html`; `Pricing.tsx`+`ClaimMeter`+`DeskShowcase`; Nav/Footer; prod 200; Stripe real, mocks removed |
| 9 | Public profile `/u` | YES | YES | YES | YES | YES | ✅ | `System.dc.html` /u; `PublicProfile.tsx`; OPEN_ROUTE share link; prod `/api/profiles/test`→404 (privacy honest) |
| 10 | Player profile | YES | YES | YES | YES | YES | ✅ | `Mobile.dc.html` PLAYER; `player/[name]`; reachable via `playerHref` links; prod `/api/stats/player`→found:true (archetype/propDNA/season) |
| 11 | Team hub | YES | YES | YES | YES | YES* | ✅ | `Mobile.dc.html` TEAM HUB; `TeamHub.tsx`; `TeamLink` on cards; prod `/api/team/NYY`→26 roster. *MLB real; NBA/WNBA partial roster |
| 12 | Explore | YES | YES | YES | YES | YES | ✅ | `Intelligence.dc.html`; `ExploreHub.tsx`+`FuturesBoard`/`NewsWire`; Nav; prod 200; real leaders/futures, "TRACKED·NOT GRADED" label |
| 13 | WIRE / ticker | YES | YES | YES | YES | YES | ✅ | `System.dc.html` WIRE; `vyndr/Ticker.tsx`; root layout; prod `/api/ticker`→30 real items |
| 14 | Streaks / hot list | YES | YES | YES | YES | YES | ✅ | `System.dc.html`; landing panels + Explore; prod `/api/streaks/mlb`→240 (computed), self-hide honest |
| 15 | Live tracking | PARTIAL | YES | YES | YES | YES | ⚠️ | `liveTrackingService`+`StatStrip.LiveTracker`; Slate polls `/api/live`; prod hasLive:false now (valid empty). DESIGNED: no dedicated bundle artboard |
| 16 | Parlay lab | PARTIAL | YES | YES | YES | YES | ⚠️ | `System.dc.html` PARLAY BUILDER (not named "Lab"); `parlay/page.tsx`+`ParlayPanel`; `#parlay` drawer; prod `/api/parlay/grade`→real A- |
| 17 | Alt-line ladder | PARTIAL | YES | PARTIAL | YES | YES | ❌ | Champion `alt_lines` real (prod: 5 rungs w/ edge_pct) but **Desk-tier only** (`tierGating.js:55`); proj-v1 `proj_ladder` ledger-only. DESIGNED: only Offseason grouping, no prop alt-stack |
| 18 | **Book comparison (S2)** | YES | YES | **YES✓** | YES | YES | ✅ | ✓WIRED 2026-07-29: `BookComparisonPanel` (self-fetches `/api/books`) on the GradeResultCard; renders per-book lines (books differ), single-book honest state, crown OFF (`BOOK_CROWN_ENABLED=false`, no best claim). Push-to-book + movement strip deliberately HELD |
| 19 | Price triplet | YES | YES | YES | YES | YES✓ | ❌ | ✓HONESTY PASS: null model/EV now NO_MODEL (honest-absent), no false "poisoned" copy. Still ❌: EV layer doesn't *produce* model_odds/ev (separate build) |
| 20 | **Compare (H2H)** | PARTIAL | in-dev | NO | YES | YES✓ | ❌ | ✓HONESTY PASS: fabrication removed → honest in-development state, pulled from Nav+BottomTabBar. Real two-player build pending (awaiting real build) |
| 21 | Newsletter | YES | YES | YES | PARTIAL | YES | ❌ | `NewsletterCapture.tsx` on `/`,`/welcome`; subscribe validates live; send internal-only. LIVE: Listmonk env config `CANNOT DETERMINE` from prod |
| 22 | Slip reader | PARTIAL | YES | PARTIAL | YES | YES | ❌ | `slip/page.tsx`; prod 200; **0 nav links → orphan** (deep-link only). DESIGNED: `a1-s9` spec, not the design bundle |
| 23 | Article media / share (S3) | YES | PARTIAL | NO | PARTIAL | YES | ❌ | `Intelligence.dc.html` ARTICLE MEDIA; `ShareCard.tsx` **0 real importers = DEAD**; OG `opengraph-image.tsx` IS live. In-article archetype figures not built |
| 24 | Calibration / edge board | YES | YES | NO | NO | YES✓ | ❌ | ✓HONESTY PASS: placeholder-edge% `MobileEdgeBoard` REMOVED from the Slate (phones show real cards). Component kept as dead code until a real edge feed exists |
| 25 | System / Intelligence terminal | YES | YES | **NO** | PARTIAL | YES | ❌ | `System.dc.html`/`Intelligence.dc.html`; `/terminal`→redirect to `/dashboard`; `/intelligence` REAL but **orphan (0 nav links)**; `/system` no page (prod 404) |
| 26 | Offseason hub (S-2) | YES | PARTIAL | NO | NO | — | ❌ | `Offseason.dc.html` full spec; **no `/offseason` page** (prod 404); logic only inline in `FuturesBoard`/`NewsWire` on Explore |
Prod endpoints confirmed LIVE + real: `/api/snapshot/{mlb,wnba}`, `/api/accuracy` (n=763),
`/api/ledger/accuracy`, `/api/ticker`, `/api/books/mlb`, `/api/streaks`, `/api/hotlist`,
`/api/live`, `/api/schedule`, `/api/stats/{player,leaders}`, `/api/team`, `/api/parlay/grade`.
## MODEL MATRIX — champion serves; every challenger is ledger-only
| Component | EXISTS | PROVEN | PROMOTED (serving) | USED (a surface reads it) | Evidence |
|---|---|---|---|---|---|
| Champion (engine1) | YES | PARTIAL→**promising** | YES | YES | `analyzeViaEngine1.js`; `enriched``snapshot`/`grades`. SKEW AUDIT 2026-07-29: on takeable MLB overs (n=62) champion p_win→CLV **partial r=0.375, SIG p≈0.003**; SURVIVES the mechanical baseline (no-edge CLV +1.5pt n=20 vs high-edge +8.6pt n=37 → **+7.1pt marginal**). De-vig clean (same-book pairing); close well-defined (DK/MGM r=0.92). PROMISING, NOT confirmed (thin n; lock-staleness check BLOCKED; 1 sig result among many) |
| arch-v1 | YES | NO | NO | NO | `challengerProjection.js`; rides `withChallenger`→**ledger only** (`:693`). "measured, never served" (`ledgerService.js:251`) |
| contact-v1 | YES | NO | NO | NO | `contactChallenger.js`; ledger col `p_win_contact` only |
| proj-v1 / v1.1 | YES | **NO (tested 2026-07-29)** | NO | NO | `projectionChallenger.js` (MLB-batting only). PROOF ORDER verdict: **NOT PROVEN** on n=45 takeable MLB overs — edge-CLV partial-r (controlling price) = 0.245 (n.s.); ~half the raw signal is the shared fair_prob_lock term (mechanical); and the CHAMPION out-predicts it (champ partial-CLV 0.380 sig, champ-edge→hit 0.25 vs proj 0.12). Ledger-only |
| Champion alt-ladder | YES | NO | YES | Desk-only | `analyzeViaEngine1.js:486`; real re-grades ±1 line; `gradeAdapter.js:100` maps to card **Desk-gated** |
| proj-v1 `proj_ladder` | YES | NO | NO | NO | `distribution.js:100`; ledger-only, reaches no card |
| Price gate / EV | YES | PARTIAL | PARTIAL | PARTIAL | fields on `enriched`; hero gates on `isTakeable`+`ev_pct` (`heroPropService.js:78`). **But prod grades show ev_pct/p_win/model_odds/value/takeable = NULL** — built, served-schema, not producing |
### Ladder question (Phase 2.6) — VERIFIED: proj-v1.1 DOES compute rungs above the line
`projection/distribution.js:100` `ladder()` computes `P(stat ≥ k)` for a fixed rung set
`k = 1..LADDER_MAX` (default 4), **independent of the listed line** — so for a line of 1.5
(tradedRung=2) it emits rungs 1 (below), 2 (at), **3 and 4 (ABOVE)**, each a real
negative-binomial survival prob (`projectionChallenger.js:183,200`). "Ladder-up **works**"
— sole cap is `LADDER_MAX=4`. **BUT `proj_ladder` is ledger-only and reaches no user
surface.** So "ladder up from the listed line to find value" is *computed and never wired*.
## HONEST-STATE SUMMARY
### 3.7 — What a PAYING USER sees right now that is not true
1. **`/compare` — fabricated grades, nav-linked + public.** Hardcoded Jokić A+ / Wembanyama A + fake "VYNDR VERDICT" (`compare/page.tsx:10-62`). No data behind it. **The single worst live lie.**
2. **FAQ — phantom processor.** "We use NexaPay" (`FAQ.tsx:28`); every legal/pricing page says Stripe. Also founder price inconsistency ($24.99 vs $19.99).
3. **FAQ + Features — Brier/CLV over-claim.** "Brier score and CLV… published / from day one · Public accuracy by tier" (`FAQ.tsx:498`, `Features.tsx:323`). No Brier surfaced anywhere; CLV held. `CANNOT DETERMINE` a live Brier surface — because none exists.
4. **Mobile edge board — placeholder edge%.** `MobileEdgeBoard.tsx:44` renders a miscalibrated edge feed, masking >40% to "—". The numbers ≤40% still come from a placeholder pipeline. Live on mobile Slate.
5. **Price triplet / EV markers advertised-and-absent.** Grade schema carries `ev_pct`/`p_win`/`model_odds`/`value`/`takeable`; all **NULL on live grades** — the "model price" leg and VALUE marker don't render though the design promises them.
Not inflated (verified honest): grades are B/C only with A/D/F below n≥20 → `pct:null` everywhere; `AccuracyBadge`/`ModelRecord`/`TierRecord`/ledger all honor the n≥20 gate and self-hide. Hit-rate (59%, n=763) shows **without ROI/CLV** (CLV frequently null) — thin, not false.
### 3.8 — The graveyard (built, no user can reach it)
1. **`BookComparison.tsx`** — 0 importers. Backend `/api/books` is live + honest; nothing renders it. (S2 headliner.)
2. **`ShareCard.tsx`** — 0 real importers (S3 share cards).
3. **`components/GameCard.tsx` (legacy)** — type-only import; superseded by `vyndr/GameCard`.
4. **proj-v1 `proj_ladder`** — the above-line probability ladder; computed, ledger-only, never served.
5. **arch-v1 + contact-v1** — challengers, ledger-only, never served ("measured, never served").
6. **`TerminalTemplates.tsx`** — all SAMPLE data; `/terminal` redirects → effectively unrouted.
7. **`DemoScan.tsx`** — defined, never rendered.
8. **`/intelligence`** — REAL signals feed, but 0 nav links (orphan; reachable only by typing the URL).
9. **`/soccer`, `/marketplace`** — REAL pages, 0 nav links (orphans).
10. **`/notifications`** — `RouteStub`, unreachable.
11. **Offseason hub** — full design spec, no page built (prod `/offseason`→404).
## FULLY DONE (all five YES) — 15
Landing · Dashboard/Slate · Scan flow · Grade result card · Grade badge · Ledger ·
Tier-record · Pricing · Public profile `/u` · Player profile · Team hub · Explore ·
WIRE/Ticker · Streaks/Hot list · Live tracking.
## NEAREST TO DONE (one column from YES) — the shortlist a single order could finish
| Surface | The one gap | Finish move |
|---|---|---|
| **Book comparison (S2)** | WIRED: NO | Route `BookComparison.tsx` onto the card from the live `/api/books` store (crown stays off — measured flat). *Backend already shipped.* |
| **Parlay lab** | DESIGNED: PARTIAL | Accept the System "PARLAY BUILDER" spec as the lab spec (functionally live) — a doc call, not a build. |
| **Compare (H2H)** | HONEST: NO | Replace the hardcoded SAMPLE with a real two-player fetch, or pull it from nav until real. |
| **Price triplet** | HONEST: PARTIAL | Make the price-aware EV layer actually produce `p_win`/`ev_pct`/`model_odds` on served grades (built, not firing). |
| **Newsletter** | LIVE: PARTIAL | Confirm/enable Listmonk env (config `CANNOT DETERMINE` from outside). |
| **Slip reader** | WIRED: PARTIAL | Add one nav/More-sheet link (page is live + honest). |
| **Calibration / edge board** | HONEST: NO | Fix the placeholder edge% feed (backend), or hide the board until real. |
Two-plus columns out (bigger builds): **Alt-line ladder** (design + wire the served
probability ladder), **System/Intelligence terminal** (wire the orphan `/intelligence`),
**Article media/S3** (build in-article figures; `ShareCard` dead), **Offseason hub**
(no page at all).
*Tags: all cells VERIFIED against repo/prod/bundle except — Newsletter LIVE (Listmonk env)
and a live Brier surface = CANNOT DETERMINE (none found). Nothing BLOCKED.*
---
# HONESTY PASS — applied 2026-07-27 (commit 6bc18d8, deployed)
Removed/hid every KNOWN live fabrication. REMOVE/HIDE only — no grade, snapshot,
scorer, pipeline, or real feature touched. Updated HONEST cells:
| Item | Was | Now |
|---|---|---|
| /compare (row 20) | HONEST: NO — hardcoded Jokić A+/Wembanyama A + fake VERDICT, nav-linked | Honest **in-development** state; removed from Nav + BottomTabBar. Real two-player build **pulled, awaiting real build**. |
| Pricing (founder copy) | $34.99 desk / struck $19.99 / FAQ $24.99 — wrong | Founder **Desk $44.99** (matches lib/checkout.js), **Analyst $14.99**; struck "regular" numbers removed; DeskShowcase $34.99→$44.99. First-100 counter is REAL (ClaimMeter→Stripe). No "first 50" desk claim (no such counter). |
| FAQ processor | "NexaPay" | **Stripe** (verified live: Next→Express→checkout.stripe.com). `nexapay.ts` + its webhook route were **PURGED 2026-07-27** (NexaPay Purge order) — cross-project contamination, never a real VYNDR path. Provider-side env/keys + the orphaned `user_profiles.nexapay_customer_id` column flagged for Kev. |
| FAQ + Features "Brier/CLV published from day one" | INFLATED (not surfaced) | Removed. Returns when Brier/CLV are actually surfaced. Backend Brier compute untouched. |
| Calibration / edge board (row 24) | HONEST: NO — MobileEdgeBoard placeholder edge% (masked >40%) | **Removed** from the Slate; phones show the real game cards. Component kept as dead code (hidden, not deleted) until a real edge feed exists. |
| Price triplet (row 19) | HONEST: PARTIAL — null model/EV rendered "MODEL READ WITHHELD · poisoned" (false quarantine) | New **NO_MODEL** honest-absent state: MODEL "—" / "NOT PRICED", no verdict. Fixes grade card + LiveHeroProp. EV layer still doesn't *produce* values (separate build). |
## KEPT ON THE BOARD (real work to finish — NOT cut)
- **Article media / S3 (row 23)** — real feature + free SEO/distribution. Finish, don't delete. Only the false "Brier/CLV" claim about it was corrected.
- **Newsletter / THE WIRE (row 21, 13)** — real. Capture live, honest.
## FUTURE MODEL INPUT (logged only — not built this order)
- **News / line-movement signal** — injuries, scratches, lineups, weather move props before books reprice. Wire later as a model input. This is **ADDITIVE** to the media surfaces, not a replacement for them.
## KNOWN HONESTY GAPS (not fixed this order — logged, not fabrication)
- **Hit rate 59% (n=763) shown without ROI/CLV** — thin, not false. ROI/CLV surfacing is a later build.
- **"0 pushes = mis-scoring" — RETIRED 2026-07-29 as a false alarm** (premise re-verified, report-only).
The displayed hit/miss denominators are NOT corrupted by a hidden push bug: the feed is still 100%
half-numbers (0 whole lines in 117,970 captured market lines / 6,050 snapshots / 1,141 ledger rows /
173 lock_lines), all 992 settled actuals are integers, and the smallest actual-vs-line gap in the
whole ledger is 0.5. Expected pushes = exactly 0. See the verdict block below.
- **CLV instrument REPAIRED 2026-07-28** (commit 6552281). Was: 59 usable closing_prob. Now: **406** (MLB 248, WNBA 158) — the collapse was `attachClosingProb`'s `.limit(50000)`/no-ORDER-BY read + write-once `market_unavailable`, NOT capture (95% per-prop coverage) or the join (0 key mismatches). **CLV finding, straight: MLB unders lag the close (mean 9.1 prob-pts, 74% lose); MLB overs +2.0; WNBA flat.** → the +4.57% MLB-C and over/under asymmetry are substantially stale-line artifacts. This UNBLOCKS the proof order (proj-v1.1), which gates promoting p_win/ev to served grades.
**Honest state after this order: "no KNOWN live fabrications" — not "provably none."** The audit was thorough (repo + prod), but absence of a claim of falsehood is not a proof of universal truth.
---
# proj-v1.1 TAKEABLE-EDGE PROOF — verdict 2026-07-29 (report-only, read-only)
**N-gate PASSED: overlap = 45** (settled ∩ proj-v1.1 ∩ MLB over ∩ CLV close ∩ fair_prob_lock ∩ takeable 160..+200). The CLV repair is what made n≥30 reachable. Edge basis = `proj_p_over_line proj_book_implied` (de-vigged fair, VERIFIED not raw book).
**VERDICT: NOT PROVEN.** proj-v1.1's takeable-edge does NOT beat the champion or clearly beat the close on MLB overs.
- Phase 1 buckets (descriptive; only the neg bucket clears ≥15): 6%+ edge (n=11) shows hit 70% / fair-ROI +0.57 / CLV +12.95pts — but even negative-edge rows show +2.7pt CLV (whole over-side is elevated = the stale-high concern).
- Phase 2 partial correlation (the real test): raw r(edge,CLV)=0.455 → **partial r(edge,CLV | price)=0.245, n.s.** at n=45 (t≈1.64, p≈0.11). ~half the raw signal is the shared fair_prob_lock term (mechanical). proj predicts the close itself only weakly (partial r(proj,close|lock)=0.281, n.s.).
- Phase 2.6 champion comparison (same 45 rows): **the CHAMPION out-predicts proj-v1.1** — champ-edge→hit r=0.250 vs proj 0.120; champ partial-CLV=0.380 (SIGNIFICANT, p≈0.01) vs proj 0.245 (n.s.). Positive-edge fair-ROI comparable (proj +0.39 n=19, champ +0.34 n=26).
- Shown, not judged: MLB unders (n=20) CLV 9.34pts, r(edge,CLV)=0.03 (contaminated, zero signal); WNBA — proj-v1.1 does not run (MLB-batting only), N/A.
**Conditional dependency (moot):** the verdict was to be conditional on the under-capture audit clearing the over-side. It's moot — proj-v1.1 fails Phase 2.6 (loses to the champion) BEFORE the audit applies, so NOT PROVEN regardless. The stale-high concern is corroborated (whole-over-side baseline +CLV; half of proj's signal mechanical).
**Notable:** the one statistically-defensible edge signal here is the **CHAMPION's** p_win predicting CLV on takeable MLB overs (partial 0.380, p≈0.01) — NOT proj-v1.1. That champion signal is itself still audit-gated (if MLB overs are whole-side stale-high, even it is suspect).
**What accrues a re-test:** ~20-25 more settled takeable MLB-over rows (to test proj's residual ~0.245 partial against zero), AND proj-v1.1 must demonstrate it beats the champion — which it currently does not. Promotion stays HELD.
---
# OVER-SIDE SKEW AUDIT — verdict 2026-07-29 (report-only, read-only)
Gates the champion's over-CLV signal (partial r=0.375, p≈0.003, n=62 takeable MLB overs — 0.2 CONFIRMED).
**THE THREE NUMBERS (Phase 4), n at each:**
- Mechanical baseline CLV (champ-edge ≤ 0, "no edge"): **+1.51 pts** (n=20)
- Champion high-edge CLV (champ-edge ≥ 0.05): **+8.64 pts** (n=37)
- **DIFFERENCE = the real edge: +7.14 pts**
**VERDICT: SURVIVES BASELINE.** The marginal (+7.1) is ~5× the mechanical floor (+1.5) and the price-controlled partial correlation stays significant. The skew is essentially ONE-SIDED (unders lag 7.0 overall; overs carry only a small +1.5 floor, NOT the +7 a symmetric two-sided over-skew would show). De-vig is CLEAN (`analyzeViaEngine1.js:539` pairs over+under from the same book/fetch — no fresh/stale pairing). Close is well-defined (draftkings vs betmgm over-prob r=0.921, n=151).
**→ GREENLIGHTS building the takeable-edge grade ON THE CHAMPION (engine1 p_win), NOT proj-v1.1** (which lost the proof). This is the project's first edge signal to survive an adversarial audit.
**FLAGGED — promising, NOT confirmed:**
- Thin n (62 overs / 37 high-edge / 20 baseline).
- Phase 2 lock-staleness check is **BLOCKED**: multi-book lines AT LOCK are not retained (`bookprices` is Redis current-only), so we cannot fully rule out that part of the baseline is lock-time staleness. The de-vig being clean + the small baseline make a large hidden skew unlikely, but it's not excluded.
- Sharp (pinnacle) reference covers only **8** props (sharp-CLV +3.17pt, directional hint only).
- This is one significant result among many computed this session — do not overstate.
**What would strengthen it:** retain multi-book at lock (enables the sharp/consensus lock-staleness check), and accrue more settled takeable MLB-over rows. Held: no promotion, no served p_win/ev, no capture fix — diagnosis only.
**UPDATE 2026-07-29 (commit c7067c8): the lock-multi-book gap is now CLOSED.** New `lock_lines` table (migration 033, applied+tracked) persists each graded prop's per-book lines at the lock moment (`lockLineCapture` in `snapshotService`, fenced RLS-service-role-only, grade byte-identical proven). This unblocks the staleness audit for FUTURE rows — it does NOT retroactively fix the existing 62. Confirmation still needs weeks of accrued lock+close+outcome. Populates from the next snapshot tick.
---
# HERO RANKING FIX — 2026-07-29 (commit 41b86e3, deployed)
The landing/hero (matrix row 1) selection was silently broken: it ranked on `ev_pct`, which is NULL on served grades, and **`Number(null) === 0`** made every prop tie at EV 0 → the "top read" was the FIRST takeable A/B prop in cache order — **arbitrary, dressed as ranked** (prod served Kelsey Mitchell, the #6 read by p_win). FIXED: rank by the **champion's p_win** (the only promising edge signal) among A/B **takeable-priced** reads (`isTakeable` 160..+200, same band as the proof/audit); strict-null guard; takeable filter excludes chalk; **no backfill** → honest empty state when nothing qualifies. p_win is ranking-only (never exposed; the route strips it). Display-only — reads caches, writes to nothing. No proven-edge/+EV/best-bet claim, no CLV/ROI/edge number. This makes the champion's p_win a real (display) consumer for the first time. Fingerprint VERIFIED: hero is the max-p_win read across sports (WNBA A), not old code's first-in-order MLB pick (Schanuel 135); untakeable chalk excluded. Visual auth-gated → data fingerprint.
---
# PUSH-SCORING PREMISE VERIFY — verdict 2026-07-29 (report-only, read-only)
Tested the standing ruling "push scoring is correct — do not touch." That ruling rested on
"100% half-number lines → pushes structurally impossible," which was true for the data it was
made on. If whole-number lines had entered the feed since, 0 pushes across settled rows would be
a real mis-scoring bug the ruling was shielding. **The premise HOLDS — the ruling stands.**
**Phase 1 — feed distribution, 4 independent populations, per sport AND per market (never blended):**
| population | what it covers | rows with a line | whole-number lines |
|---|---|---|---|
| `closing_captures` | raw captured market lines, 5 books, `book`+`sharp`, Jul 20-29 continuous | **117,970** | **0** |
| `model_snapshots` | every graded prop **incl. grader refusals** (not survivorship-filtered) | 6,050 | **0** |
| `ledger_entries` (public) | the settled public record, 11 markets | 1,141 | **0** |
| `lock_lines` | TODAY's lock-time per-book lines (freshest feed, migration 033) | 173 | **0** |
Per-market: MLB hits / doubles / rbi / total_bases / stolen_bases / runs / strikeouts / home_runs /
walks / earned_runs / outs / hits_allowed and WNBA points / rebounds / assists / threes — **every
market's min AND max line ends in `.5`** (e.g. MLB strikeouts 2.5-8.5, WNBA points 5.5-26.5, MLB
outs 3.5-19.5). No whole-number market is hiding inside a blended fraction.
**Phase 2 — the push branch would fire.** `outcomeService.js:151` `if (a === l) return 'push'`,
reached **after** `Number()` + `Number.isFinite` guards on both operands — a sound numeric compare,
not the `Number(null) === 0` string-vs-number class that hit the hero. It is the **single scoring
chokepoint** (`ledgerService.js:31` imports `settleResult`; no parallel hit/miss derivation exists
in `src/`), it is **unit-tested live** (`outcomeService.test.js:38`, `nbaSettlement.test.js:104`),
and both `ledger_entries.outcome` and `outcomes.result` CHECK constraints **include `'push'`** — a
real push would score, write, and persist end-to-end.
**Phase 2.6 — the decisive number.** Across 992 settled rows carrying an actual: **0 exact ties, 0
fractional actuals, and the smallest actual-vs-line gap is 0.5** — the arithmetic minimum between an
integer result and a half-number line.
**VERDICT: RULING HOLDS.** Expected push rate is **exactly 0 (P = 0), not "low"** — 0/992 is
*forced*, not chance. The "implausible" flag mistook an arithmetic impossibility for a suspicious
absence; the row closes honestly. Stale n corrected: the flag said 470 settled, it is now **1,097**
(593 hit / 399 miss / 105 void / 44 unsettled-today). Nothing modified — no scoring, settlement,
re-settle, or backfill.
**No latent bug either.** Because the branch is correct and covered, a whole-number market entering
later (NFL/NHL are code-wired but out of season; whole-number strikeout props exist at some books)
would be scored as a push automatically. The residual is a **monitoring** gap, not a scoring gap:
nothing alerts on the first whole-number line to enter the feed. Logged, not built.
---
# EDGE_PCT SCALE DIAGNOSIS — 2026-07-29 (report-only, read-only). Fork REPORTED, not chosen.
**What it is (0.1).** `analyzeViaEngine1.js:265-270``edge_pct = ((projection line) / line) × 100`,
signed by direction, where `projection = l5_avg ?? l20_avg ?? {stat}_per_90 ?? xg_per_90`.
**Independent of `p_win`** (so NOT tainted by the overconfidence that damns `ev_pct`) but it takes
**no price input at all**, so it cannot express a betting edge. **Arithmetically correct, MISLABELLED:**
honest as "% the projection differs from the line," **a lie at any scale as "EDGE."** Two independent
implementations — backend `edgePctFor` and `web/src/lib/gradeAdapter.js:25-31 computeEdge`; the grade
card renders the WEB one, so a backend-only fix would miss it.
**The cap (0.2).** `SANE_EDGE_MAX = 40` (`deskShowcaseService.js:31`: *"beyond this the (model-line)/line
value isn't a market edge"*), mirrored in `slateAdapter.js:613` and `MobileEdgeBoard.tsx:45`. A
self-declared plausibility bound from an earlier order, not a derived statistical one.
**Mechanism (0.3) = SMALL-DENOMINATOR EXPLOSION** — not units, not inversion, not a missing ×100.
`line` is the denominator and **86% of MLB rows (562/655) sit at line 0.5**. Max 620 = a ~3.6 projection
on a 0.5 line.
| population | n | >cap 40 | >100 | median | p95 | max | min |
|---|---|---|---|---|---|---|---|
| **MLB** | 655 | 65.8% (>50) | 13.6% | **60** | 180 | **620** | 86.7 |
| **WNBA** | 486 | 4.9% (>50) | **0%** | 12 | 49 | 77.8 | 51.7 |
| blended | 1,141 | **44.0% (502)** | 7.8% | — | — | 620 | — |
Per line (the proof): MLB 0.5 → 73.0% over cap, max 620 · MLB 1.5 → 43.8%, max 153 · WNBA 12.5 → 6.7%
· **WNBA 26.5 → max 1.9.** Matrix figures re-verified: **51.5% is stale → 44.0%; worst 620 is exact.**
**Shape: structurally broken for MLB, sane for WNBA** — and the scale is a *function of line size*, so
the metric is incomparable across markets **by construction**. No rescaling fixes that.
**Surfaces (Phase 2) — the "~13" count is NOT confirmed. Three surfaces RENDER it:**
| surface | live | access | role | user sees at 620 |
|---|---|---|---|---|
| `GradeResultCard.tsx:182,216,325` | YES (3 importers) | **auth-gated `/scan`****TAGGED FOR CHROME AUDIT** | display | **"+620% edge", raw + GREEN** |
| `DeskShowcase.tsx:40` | YES | **PUBLIC** `/pricing` | display | **"—"** (already honest) |
| `SoccerGradeResult.tsx:229` | YES | orphan `/soccer` (0 nav links, public by URL) | display | raw uncapped `X.X% edge` |
| `MobileEdgeBoard.tsx:47` | **DEAD** (0 importers) | — | sort+display | "—" (pulled in honesty pass) |
| `PropRow:45`, `GradeCard:32`, `ledger/page:40` | live components | — | **type-only, never rendered** | nothing |
| `contentTemplateService.js:164` | public `/api/content` | API | string | uncapped — **no page fetches it** |
**🔴 IT DRIVES TWO LIVE SORTS (the fork's load-bearing answer).**
1. `slateAdapter.selectTopGrades:469-471``grade → confidence → |edge| desc` → **dashboard TOP GRADES
top-10** (`dashboard/page.tsx:419`). **97.3% of rows (1110/1141) sit in a (date,sport,grade,confidence)
tie group of ≥2** (biggest 56), so |edge| is operative for essentially the whole slate — the **de facto
ordering** of that leaderboard.
2. `analyzeViaEngine1.js:506` — the Desk **alt-line ladder** is sorted by `edge_pct` desc.
**Two scale-INDEPENDENT defects inside that sort** (`Math.abs(numOr(g.edge, -Infinity))`): **(i) abs()**
on an already-direction-signed value ranks the model's strongest *disagreements* equal to its strongest
agreements (**177 negative-edge rows**: 58 B / 118 C / 1 F, worst 86.7); **(ii)** `Math.abs(-Infinity)
= Infinity` → a **missing edge sorts FIRST**. The `Number(null)` fabrication class again, new costume.
**Phase 4 correction — "nothing renders `ev_pct`" is WRONG.** `PriceTriplet.tsx:60,67,76` renders
`${pct(ev)} EV`, live and wired (scan → `gradeAdapter:143`). It shows nothing only because `ev_pct` is
NULL on served grades → `valueState.js:121` falls to **NO_MODEL honest-absent**. The right metric already
has a live honest render site, **starved of data, not unwired** — and the card's "EDGE" row sits exactly
where a price-aware number belongs.
**THE FORK (reported, not chosen).**
- **FIX** — dishonest: no rescaling turns a price-free projection gap into an edge (renaming, not fixing);
it silently re-ranks the dashboard top-10 (the hero-class bug just fixed); needs BOTH implementations.
- **HIDE** — cheap: 3 render sites, each already has a null branch (no layout breaks), and DeskShowcase
already proves the honest "—" pattern in-product. Not load-bearing for layout anywhere.
- **RECOMMENDED: HIDE the number and re-point the sort at `p_win`** — the hero order established p_win is
on 100% of recent ledger rows and is the only signal that survived an adversarial audit. Repairing a key
that is a 0.5-line artifact is not worth it. End state: **p_win ranks · ev_pct displays · edge_pct retires.**
The abs()/null-first sort defects deserve their own small order either way.
*Nothing changed: no edge_pct, scale, surface, sort, grade, ledger, or accruing edge touched.*
---
# GRADE-BOARD SORT FIX — 2026-07-29 (spec `specs/grade-board-sort.md`, shipped)
Display ORDERING only. Fixes two defects that were wrong at ANY scale, independent of edge_pct's
separate retirement (Order B, still held).
**Defects removed.** `selectTopGrades` ranked on `Math.abs(numOr(g.edge, -Infinity))`:
`abs()` on an already-direction-signed value ranked the model's strongest **disagreements** level with
its agreements (177 public ledger rows carry a negative edge); and `Math.abs(-Infinity) === Infinity`
made a **missing** signal sort **FIRST** — absent data as the top pick. Now: `grade → confidence →
takeable-gated p_win (nulls LAST) → SIGNED edge (nulls LAST) → input order`, scales never mixed.
The alt-line ladder (`analyzeViaEngine1:506`) no longer sorts by `edge_pct`; it is ordered
highest-p_win-first via the monotonic line rule (line-ASC for an over, line-DESC for an under) at zero
added compute.
**Three premise breaks found report-first.** (1) **`/api/props/top-graded` 404s in prod** — the
dashboard board's feed does not exist, so that board renders receipts/empty and the sort orders nothing
there today; the prior order's "97.3% of rows tie → the edge key decides the board" was a ledger
measurement wrongly extrapolated to it. (2) **p_win is stripped for unentitled tiers by design**
(`snapshotGating`, Session 67 — "shipping p_win is shipping the model price"); verified live, prod
`/api/snapshot` carries p_win on **0/8 MLB and 0/25 WNBA** grades, so the browser path uses the signed
edge and only entitled callers rank on p_win. (3) **Ladder rungs carry no per-rung price**, so the
hero's takeable gate is inapplicable there.
**Verified on real data, both sports, both paths.** Unentitled: WNBA (n=25) ordering CHANGED, MLB (n=8)
unchanged; signed edge non-increasing within every (grade,confidence) tie group — 20 pairs, 0
violations. Entitled: 40 real ledger rows with p_win+locked_odds — p_win-descending, untakeable chalk
not promoted, 36 pairs, 0 violations.
**Hero consistency, honestly:** same signal + same gate, different precedence by contract (board =
grade-tier-first "top GRADES"; hero = p_win-first "top read"). They agree exactly **within** the
leading tier (verified); across tiers the board may lead with an A the hero doesn't pick. Not a
contradiction — do not "fix" it by making the board ignore grade.
**Floor:** 310 suites / 3864 tests green, web build exit 0. **Post-deploy fingerprint (`b85b351`),
both halves verified:** the deployed `/dashboard` chunk was polled across the deploy boundary — attempts
1-3 carried the OLD `Math.abs(...-1/0)` key, attempt 4 flipped to the new comparator with the old one
GONE (before/after observed, not inferred); API health 200 on snapshot/mlb, snapshot/wnba, accuracy. The
backend ladder was **induced on demand** on live MLB game logs rather than waiting for a cron slot:
OVER @0.5 → `0.5•(C) 1(C) 1.5(F)` (line-ASC PASS), UNDER @1.5 → `2.5(C) 2(C) 1.5•(C) 1(C) 0.5(F)`
(line-DESC PASS). Only the *rendered* Desk card remains unverified anonymously → tagged for the Chrome
audit, no visual faked. **Held:** edge_pct rescale/display retirement, building
the missing `/api/props/top-graded` selector, exposing p_win to unentitled tiers.
---
# /api/props/top-graded SERVER SELECTOR — 2026-07-29 (spec `specs/top-graded-selector.md`, shipped)
The dashboard TOP GRADES board had no feed. **The handler NEVER existed in any commit** (searched
`git rev-list --all`) — so the three axios callers (cheatsheetGenerator, gradeOfTheDay, widget) plus
the Next proxy had always received `[]`. Contract recovered from those consumers, not guessed.
**The leak boundary — the whole point of doing it server-side.** The browser cannot rank on `p_win`
for all tiers because `stripModelPrice` deliberately withholds it from unentitled tiers ("shipping
p_win is shipping the price in a different base", S67). Order of operations:
`read cache → RANK with p_win (every tier) → map rows incl. model fields → stripModelPrice(rows, tier)
→ serialize`. A free caller gets the paid RANKING without the paid VALUES. Tier resolution FAILS
CLOSED to `free`; `Cache-Control` is `private` under a bearer token, `public` otherwise.
**Populated-path risk found and handled:** the board's populated branch had never run in prod, and
`dashboard/page.tsx:463` calls `g.stat.replace(/_/g,' ')` **unguarded**`toRow` requires string
`player`+`stat`, a finite `line`, uppercases `sport` for `SportPill`, and drops unrenderable rows.
**One shared ranking definition:** new `src/utils/gradeRanking.js`; `heroPropService` now imports
`takeablePWin` (was an inline copy, behaviour unchanged), the selector imports `rankGrades`, and the
web mirror is cross-checked by test. Board is grade-first ("top GRADES"), hero is p_win-first ("top
read") — they differ by design and agree within the leading tier.
**Honest limit:** the Next proxy forwards no Authorization and caches under a shared key, so via the
dashboard every viewer receives the free-tier payload — correct order, no paid values. That is the safe
default (forwarding auth into a shared cache is how paid payloads leak); per-tier delivery through the
proxy needs a tier-keyed cache and is NOT done here.
**Verified:** 311 suites / 3882 tests green (18 new, leak test on POPULATED p_win), web build exit 0.
**Post-deploy fingerprint (`72a14dc`) — the boundary proven in production:** 404→200 captured across the
deploy boundary; the anonymous live order is *1. Brionna Jones (edge 29.4) · 2. Rhyne Howard (edge 42.9)*
— an edge-only sort would lead with Howard (and the local induction over the stripped snapshot did), so
**server-side p_win ordered it** (Brionna .90 @106 > Howard .745 @120) while the payload carries
**PAID FIELDS: NONE**. A bogus bearer token also yields no paid fields (fail-closed proven in prod).
The board's proxy path now returns 10 props, so TOP GRADES **renders** instead of falling back to empty.
The rendered board is client-side → tagged for the Chrome audit, not faked.
---
# MODEL ARCHITECTURE RECOVERY MAP — 2026-07-30 (report-only) → `specs/model-architecture-recovery-map.md`
**The live grade uses 0 of the 3 specced layers.** Every audited metric (calibration, CLV, skew audit,
takeable floor, p_win→CLV r=0.375) is measured on the SHADOW model. Those findings stand — the shadow
model served every real grade — but none are evidence about the specced engine, which has never been
measured.
| Specced component | State |
|---|---|
| Layer 1 Similarity (`python/utils/similarity.py`, 101 ln) | BUILT · NOT WIRED · **NOT DEPLOYED** |
| Layer 2 Bayesian (`python/utils/bayesian.py`, 320 ln) | BUILT · NOT WIRED · **NOT DEPLOYED** — and its "sport-agnostic math / per-sport parameters" claim is TRUE of the built code |
| Layer 3 grade scale (`grade_thresholds.json`) | BUILT · **WIRED BACKWARDS** — the table maps PROBABILITY→GRADE; the live JS reads it in reverse to manufacture `confidence` from an already-chosen letter |
| Per-sport market-efficiency scaling | **SPECCED-BUT-ABSENT** |
| Live champion (`engine1` + `probabilityEstimator`) | BUILT · WIRED — but the letter is a factor-index with **zero probability input**, and `p_win` is computed separately and never feeds it |
**The Python engine is not in the deploy image at all** (no python/pip in `Dockerfile`; `app.js` only
health-checks it). **The sport boundary is NOT clean on the live path** — adding a sport is a ~10-file
core edit with four documented silent-failure modes, so "make a sport a module" is itself a prerequisite
build. **Per-sport records DO exist** (`sports.mlb` n=526/62% vs pooled `overall` n=937/58%, each with
its own n≥20 gate) — the rule to enforce is that a new sport renders `sports.{sport}`, never `overall`.
**Park×weather is a CHALLENGER, not the champion** ("measured, never served"); xwOBA and bullpen-leash
are absent entirely.
Recovery is dependency-ordered in the map: decide the grading basis → pick a runtime (recommend porting
Bayesian to Node) → wire Layer 2 → reconnect Layer 3 forward → Layer 1 → per-sport efficiency →
challenger promotion → sport-as-module → NFL/CFB. MLB is specced as the reference module.
---
# FULL-OUTPUT GRADE MAPPING + COLLAPSE COST — 2026-07-30 (report-only) → `specs/full-output-grade-mapping.md`
Track-B 1 of 3. The three-layer engine is **BUILT but NOT WIRED and NOT DEPLOYED** (re-verified), so no
posterior/CI is produced today. Measured instead against the collapse that actually exists — three of
them: estimator components dropped; **`p_win` excluded from the grade entirely** (the severe one — the
letter is a factor index with zero probability input); and `grade_thresholds.json` read backwards to
manufacture `confidence`. Market-efficiency scaling is never computed at all.
**THE MEASUREMENT (354 settled rows, locked pre-game p_win — forward, not lookahead):**
| basis | ALL (n=354) | MLB (n=224) | WNBA (n=130) |
|---|---|---|---|
| champion letter → outcome r | **0.0050** (p≈0.93, null) | 0.0686 (n.s.) | 0.0986 |
| probability letter → outcome r | **0.1313** (p≈0.013) | **0.2356** (p≈0.0004) | **0.1258** |
**The served letter is INVERTED between its only two populated tiers — B 52.4% (n=168) vs C 56.9%
(n=174).** Probability letters spread 30.8%→70.0%, use 10-11 of 11 letters (champion uses 3-4), and
split ROI 1.42% (A-family n=78) vs 26.62% (C-/D/F n=51).
**Verdict: costly on MLB, and un-collapsing does NOT help WNBA** (both correlations inverse there) —
so the full-output challenger must be **MLB-FIRST**. Five falsifiable mapping rules are specced (R2
"uncertainty grades down" stated explicitly and droppable if it fails). **Hard requirement on the build
order: persist per-row `n`, `SE`, and pre-adjustment `p`** — without them R2/R4 can never be adjudicated.
**Re-adjudication flagged:** p_win→CLV, the skew audit, **proj-v1.1's "NOT PROVEN" (judged against the
collapsed champion — not final)**, the C1 takeable floor, the calibration curves, and **ROI-by-grade —
with B/C inverted, "MLB-C +4.57%" is likely an artifact of a meaningless letter.**
---
# MARKET-EFFICIENCY SCALING CHECK — 2026-07-30 (report-only). **VERDICT: FLAT.**
**Premise correction first (measured):** the claim that "full-output and collapsed grades agree 100%"
does not hold — on 512 rows carrying both they agree **17.8%**, and **33.8% differ by 3+ tiers**. The
prior discrimination result stands (champion letter r=0.0050 null vs probability r=0.1313; MLB 0.0686
n.s. vs 0.2356 p≈0.0004). Repo unchanged between orders. **The collapse was not a phantom and the
re-adjudication list stays open.**
**0.1 `marketEfficiency.js` does not exist** — zero occurrences of `market_efficiency` /
`efficiency_score` anywhere. The spec's 0.85/0.60/0.55 values DO appear in `grade_thresholds.json` but
those are **probability bands**, a coincidental overlap, not efficiency scores.
**0.2 The base edge thresholds do not exist either**`engine1.js` has **zero `edge` references**; the
grade is an additive factor index, not an edge-vs-threshold comparison. So
`threshold = base × efficiency` has **no host**.
**Phase 1 dispositive: `engine1.js` has ZERO `sport` references.** Sport is not an input to
`computeFactors`, so per-market or per-sport scaling is structurally impossible in the live grader —
not merely unwired. Applies to the champion.
**Phase 2:** the matched-edge test is confounded (edge isn't the grading input — the same market emits
B and C at one edge). The discriminating aggregate: mean grade index **wnba points 4.71** (mean edge
10.2) vs **mlb hits 4.58** (69.5) vs **mlb total_bases 4.32** (84.9) — the efficient market earns the
highest grades on one-eighth the edge, the opposite of the spec.
**Scope:** a flat threshold is a grade-CALIBRATION gap — it does **not** touch the projection or the
CLV edge (which measured `p_win`, never the letter), so it is not a third shadow-model alarm. But it is
**not independently bounded**: with no threshold step to multiply, efficiency scaling presupposes
probability grading. It is rule **R4 of `specs/full-output-grade-mapping.md`** and belongs to that
MLB-first challenger.
---
# TAKEABLE TAGGING BUILT · EFFICIENCY CHALLENGER BLOCKED — 2026-07-31 → `specs/takeable-tagging.md`
**Champion grade byte-identical (verified by diff).** Additive tags only.
**BLOCKED — the efficiency challenger.** Review Zero found all three inputs ABSENT: efficiency scores
(zero occurrences anywhere), base edge thresholds (`engine1.js` has zero `edge` references — the grade
is not edge-vs-threshold), and the claimed ±0.05 additive nudge (the only 0.05s on the grade path are a
teammate-absence feature, a bvp cutoff, and the shrink-toward-0.5 term). With no additive scaling to
swap, no threshold to multiply, and **zero `sport` references in `engine1.js`**, a
differs-in-exactly-one-thing challenger cannot be constructed. It needs **R1 probability grading**
first (`specs/full-output-grade-mapping.md`); shipping R1+R4 together would attribute an R1-driven
re-letter to the efficiency fix. Coverage compounds it: **9 of 11 live markets have no specced score.**
**BUILT — ledger takeable tagging (deferred C2).** `src/config/takeableStandard.js`: floor on the minus
side, **uncapped plus** — deliberately NOT `valueEngine.isTakeable` (the 160..+200 promotion band). A
+400 prop is not promotable but IS takeable; a test locks the divergence. Absent price → `null`, never
`false`. The floor is **policy, not derived**, and each row records `takeable_floor` so a re-derivation
can re-tag. Migration 034 applied + tracked. **Backfill: 1,254 rows → 1,246 tagged (781 takeable / 465
below floor), 8 NULL with `null_despite_price = 0`; settled 1,163 and graded 1,254 unchanged.**
**NOT applied — the model-version boundary tag:** no scaling change shipped, so no boundary exists;
stamping one would record a model transition that never happened.
312 suites / 3,890 tests green, web build exit 0.
---
# EDGE-SHADING CHALLENGER — built + measured 2026-07-31 → `specs/edge-shading-challenger.md`
Challenger only; champion byte-identical; nothing promoted. The mechanic is sound and built
(`adjusted = raw_edge × f(e)`, `f` bounded to (0,1], one fixed bar that never moves; same raw 6% edge
**A** in soft `mlb:total_bases`, **B** in sharp `nba:points`, unit-proven).
**But the measurement says the flooding is NOT fixed:** on 1,250 live rows the challenger grades
**79.0% A / 80.9% A-or-B** (MLB 93.4% A) vs champion 0.2% A.
**Why — two findings.** (1) **The shading is a no-op on the live board: 0 of 1,250 rows are actually
shaded.** 96.5% are unscored (f=1) and the one scored market present is the anchor (f=1.0 by
construction); `mlb:strikeouts` and `nba:points` are absent from the ledger entirely (our basketball is
`wnba`). (2) **The input scale is the bug, not the placement.** Against a fixed 5% bar the RAW edge
already clears A on 100% of MLB doubles and 89.6% of hits, with MLB's median raw edge at 60% — twelve
times the bar. **Applying the sharpest score in the spec (f=0.647) to every row still leaves 75.8%
clearing A.** A multiplier bounded ≤1 cannot close a 12× overshoot.
`edge_pct` is a price-free `(projline)/line` gap whose scale is a function of line size — not a
betting edge, so no fixed betting-edge bar is meaningful against it. **Unblocking needs the input
replaced (`p_win` vs `fair_prob`, both already stored), not the multiply moved**, plus scores fit from
our own record for the 9 of 11 live markets that have none.
313 suites / 3,899 tests green, web build exit 0.
---
# PROMOTION GATE NOT PASSED · ORDER B SHIPPED — 2026-07-31 → `specs/edge-pct-display-retirement.md`
**No flip.** Champion grade byte-identical; projection/p_win/CLV untouched.
**Gate: 3 of 4 prerequisites fail.** Scores are estimated priors (and the premise's NBA 0.72 / WNBA
0.68 are not in the code — there is no WNBA score); **the version-boundary tag never landed**
(`modelEras.js` has zero shading refs); **no rollback flag exists**. Takeable tags did land (034).
**The approved delta is wrong.** Approved 43.6% re-letter, efficient tighten / soft hold. Measured on
1,250 rows: **97.4% change, 79.8% UP**, → **79.0% A-family (MLB 93.4%)** vs champion 0.2%, with
**0 rows actually shaded**. The re-letter is entirely the grading-basis switch, not efficiency shading
(inert: 96.5% of markets unscored, the one scored market is the f=1 anchor). Flipping would mint A's
across 79% of an append-only record on a letter with r ≈ 0.005 vs outcomes.
**Order B shipped:** edge_pct display retired from GradeResultCard (strip, EDGE cell → `—`, ladder
rung) and SoccerGradeResult; DeskShowcase kept; **computation + signed-edge sort fallback survive**
(deleting them re-breaks the 07-29 sort fix). Two build-breakers fixed; two pre-existing sign-colour
tests superseded by the stronger "no edge figure renders at all".
314 suites / 3,908 tests green, web build exit 0.
---
# MATRIX REFRESH — 2026-07-31 → full re-derivation in `specs/pre-audit-status-pull.md`
Re-derived from repo `10aaaeb` + live prod probes (nothing inherited). **15 of 26 fully done.**
**No promotion has occurred.** The shading challenger is imported by zero production files and live
grades carry **0.0% A-family** (MLB `{B:1,C:4}`, WNBA `{B:15,C:10}`). Neither the "92.9%" nor the
"43.6%" re-letter figure was ever measured here; the challenger's real numbers were 97.4%
would-change / 79.8% up / 79.0% A, with **0 of 1,250 rows actually shaded**.
**Changed since the last matrix:** row 18 **Book comparison → DONE** (`BookComparisonPanel` routed to
the grade card — the headline dead component is resolved); row 4 **Grade card HONEST improved**
(edge_pct display retired, `EDGE` renders `—`); row 17 ladder rung % also retired.
**Still dead code:** `ShareCard` (0 real importers), `MobileEdgeBoard` (0, correctly hidden),
`DemoScan` (0). **Still orphaned (live, 0 nav links):** `/intelligence`, `/soccer`, `/marketplace`,
`/notifications`, `/slip`, `/compare`, `/parlay` (drawer-only). **404:** `/system`, `/offseason`.
**LIVE-but-not-HONEST: none found.**
**Design is NOT complete:** Live tracking, Slip reader and Newsletter are shipped but have **no design
artboard** — design is the gap, not build. Offseason is the reverse (designed, never built).
**Chrome audit manifest: 11 items**, four of which need an **entitled Desk session** (grade card
entitled half, alt-line ladder, entitled top-graded board, ledger).
---
# BOUNDARY NOT WRITTEN + BUILD TRIAGE — 2026-07-31 → `specs/incomplete-surface-triage.md`
**No promotion exists, so no recalibration boundary was written** (writing one would fabricate a model
transition in an append-only record). Proofs: every deployed `code_sha` in the last 5 days is a
documented session commit stamped `engine1@2026-07-20`; daily A-family share 07-24→07-30 is flat at
0.0-1.6% with no step change; `efficiencyShading` has zero production importers. Neither the 92.9% nor
the 43.6% re-letter figure is attested — the only measured numbers (full board, 1,250 rows) were 97.4%
would-change / 79.8% up / 79.0% A, with **0 rows actually shaded**.
**Triage: 14 surfaces, 5 waves.** W1 wiring (`/intelligence`, `/slip`, `/parlay`, `/marketplace` after a
copy pass) · W2 design-only artboards (Live tracking, Slip reader, Newsletter — built/live/honest,
DESIGN is the only gap) · W3 self-contained (`/compare`, ShareCard host, `/notifications`) · W4
model-gated (price-triplet MODEL leg + edge board — need `p_win` vs `fair_prob`; `edge_pct` would
re-ship the 620% lie) · W5 quota/sport-gated (`/soccer` blocked on odds-api 0/500, `/system`,
Offseason).
---
# WAVE 1 WIRED — 2026-07-31 → `specs/wave1-wiring.md`
Three surfaces **proven to work with real data, then linked**; `/marketplace` completed honestly.
No grade/ledger/model/scoring change.
- **`/intelligence`** → nav-linked + **added to GATED_ROUTES** (its feed 401s signed-out). Gating is
server-side by tier (desk 50 / non-desk 8), never blur-over-full-data. **`/system` is NOT a
duplicate to build** — `System.dc.html` is a multi-surface artboard whose INTELLIGENCE section is
already this page; the 404 is correct.
- **`/slip`** → nav-linked after parsing a real DraftKings slip **3/3 legs**. Layout-rigid: an
unsupported layout returns **zero legs, never wrong ones**; real-world hit-rate CANNOT DETERMINE yet.
- **`/parlay`** → **direct nav link** (was drawer-hash only); stays OPEN as the free parlay funnel.
- **`/marketplace`** → every item now states **"Not built yet."**; subhead: *not a purchase, not a
pre-order, not a promise of a ship date*; **no profit claim**; the waitlist capture was already real.
Matrix effect: rows 16 (Parlay lab), 22 (Slip reader) and 25 (`/intelligence` half) move from
**WIRED: NO → YES**. 315 suites / 3,920 tests green, web build exit 0.
---
# WAVE 3 — 2026-07-31 → `specs/wave3-compare-and-resolution-tail.md`
**`/compare` BUILT** (row 20 → real). A same-market head-to-head on the live player feed: rows aligned
on measures both sides share, `NO DATA` for an unresolved side, `—` never 0 for a one-sided measure,
refusal when neither resolves, and **NO VERDICT**. Live-verified: Judge vs Ohtani, 5 of 5 shared
measures; bogus name → NO DATA.
**Resolution tail — ALL FIVE outputs scoped, none shipped** (rows 23 ShareCard, `/notifications`,
result posts, recap): share-card generation **SPEC'D-NOT-BUILT** (absent from the resolve fanout;
renderer has zero callers) · push/Telegram/Discord **BUILT-NOT-FIRING** (env-gated; `push_subscriptions`
and `user_notifications` both **0 rows**) · recap **SPEC'D-NOT-BUILT** (no file). **And nothing calls
`/api/grading/resolve` — there is no ESPN poller in the repo**, so the tail is unreachable regardless.
The live settle path emits ops alerts only, with zero user-facing output.
316 suites / 3,930 tests green, web build exit 0.
---
# DESIGN-vs-BUILD GAP AUDIT — 2026-07-31 → `specs/design-vs-build-gap-audit.md`
**61 implementable items** enumerated from `specs/design-reference/` (Jul 22) against the current repo:
**BUILT-TO-SPEC 20 · BUILT-BUT-DRIFTED 7 · PARTIAL 16 · ABSENT 18.** The design is ahead of the build;
the gap is implementation, not design.
**Biggest gap:** the glyph library — **38 of 83 SVGs wired (46%)**, and the design implies **74 display
archetypes vs the registry's 41**.
**Drift on recently-built surfaces:** book comparison (wired 07-29) has **no crown, no disagreement
axis, no SPLIT chip, no movement strip**; the mobile tab bar lacks the designed **READ-FAB**;
calibration gates at **n≥20 vs the designed N30**.
**Correction to the earlier status pull:** the **Newsletter IS designed** (S5 "The Report") →
design-exists-needs-build. Only **Live tracking** and **Slip reader** are genuinely design-missing.
**Six-wave ordered build list:** self-contained (glyphs, primitives, boundary-channel blue) → scanner-
nudge-gated → model-gated (Price Triplet EV leg, calibration curve) → **resolution-pipeline-gated
(share-card masters — the tail has no generation step and no trigger)** → licensing-gated (book logos,
push-to-book) → large surface builds (Offseason, S3 media, The Report, S2 primitives).
---
# D1-A SHIPPED — 2026-07-31 (design self-contained wave, part A)
Backend untouched. **Glyphs:** 6 classifier-backed combat marks wired from the package SVGs
(GLYPHS 38→44), replacing emoji fallbacks; 39 package SVGs with no classifier deliberately NOT wired
(they'd render nothing) and 2 classifier-backed archetypes have no package SVG — both held for D1-B.
**Boundary channel completed:** PriceTriplet's `NO_MODEL` was the last boundary state rendering in
neutral text; it now uses `--priced-out`, so the blue law covers every "can't hand you this" state.
**Reaction primitives** (`lib/reactions.js` + keyframes) at exact HANDOFF timings, with `nudge()`
refusing a no-op so it can never run as an idle loop. **READ-FAB** at exact geometry (50px /
translateY(-14px) / 6px ring). Audit correction: the `#0E0E14` card token was already tokenised.
317 suites / 3,946 tests green, web build exit 0. **Not done, carried in D1:** row-hover rationale,
IntersectionObserver reveal, team-gradient chips.
---
# D1 FINISH — 2026-07-31 (row anatomy: rationale · reveal · team chips)
Backend untouched. Three pure modules, unit-locked:
- **`rowRationale.js`** — the hover "why" is the grade's OWN `reasoning` + `kill_conditions_triggered`
(VERIFIED present on live grades, built from real l5/l20/gap/home/defense/rest), or **NULL**. No
generic fallback; a locked/tier-gated reasoning counts as absent rather than paraphrased.
- **`reveal.js`** — IntersectionObserver, fires **once then unobserves**, reuses D1-A's `bootDelayMs`
(one source of truth for the 60ms stagger), and reveals immediately when the API is absent so
content is never hidden.
- **`teamChips.js`** — Rev-3 chip geometry + the 1/.86/.64/.48 ramp. **Colour coverage is 10 of ~80
teams** (the artboard's full set, verbatim); every other team renders honest-neutral rather than a
guessed colour.
318 suites / 3,961 tests green, web build exit 0. **Mounting the modules into the live row components
is a follow-up**; the visuals go to the Chrome audit.
---
# TIER REDESIGN SPEC — 2026-07-31 → `specs/tier-redesign-spec.md`
Option 2 (settled-free / live-paid) designed, not built. **FREE** = full data aggregator + the
**complete settled record** (letter, reasoning, edge, outcome — browsable) as the proof hook.
**ANALYST $14.99→$24.99** = tonight's live grades + reasoning + edge. **DESK $44.99→$59.99** = +
alt ladder, Kelly, portfolio, engine2.
**The gate already has its discriminator:** `/api/snapshot` merges per-grade `outcome` (live WNBA:
5 of 25 settled), so `outcome != null ⇒ free`, `null ⇒ paid`, with no new pipeline. Filter whole
grades server-side, never infer resolution from time, fail closed to LIVE.
**Two findings that reshape the plan:** (a) founder pricing is gated by **code + expiry, not seat
count** — the real counter (`/api/founders/count`, live 0/100) only displays and is cached 300s, so
making it a transactional gate is a genuine build; (b) the user base is **3 free, 0 paid**, so the
migration is a courtesy note, not a mass event — effort belongs on the settled-record proof surface.
---
# BUILD 1 — THE SETTLED/LIVE GATE (shipped 2026-07-31)
Unresolved = paid (Analyst+), resolved = free. Serving/gating only; `src/services/` untouched.
Resolution is read **only from a written outcome** (never time or game status), `void`/`unrecoverable`
count as resolved, and `isResolved` **fails closed to LIVE** so a settle failure withholds rather than
exposes. Free gets **settled grades in full including reasoning** — the proof product — which also
turns the old unenforced board-reasoning leak into a deliberate rule. Live grades become a **shell**:
judgment stripped, real data kept (incl. `fair_odds`, which is never the paywall), `locked: true`
stamped. The tease is **aggregate only** (`live_locked {count, tiers}`) and no gated row carries a
grade, so nobody can tell which prop is the A.
**Live anonymous fingerprint:** 25 grades → **5 settled free-full, 20 locked**, `live_locked
{"count":20,"tiers":{"B":13,"C":7}}`, **judgment-field leak: NONE**. 319 suites / 3,971 tests green,
web build exit 0.
---
# BUILD 1 CORRECTED — itemized grades are PAID, live AND settled (2026-07-31)
The resolution-flip shipped earlier the same day was **replaced**: freeing grades at settlement made
the free tier a one-day-delayed feed of the whole product. No per-grade flip now — every itemized
grade is Analyst+.
**Free:** data aggregator + the **aggregate record that already existed** (`/api/accuracy` n=937,
byGrade, per-sport; `/api/ledger/accuracy`, both honoring the n≥20 hollow law with `pct:null`) + a
**capped, day-rotated 3-call sample** (resolved only) + the locked shell of tonight's reads
(`live_locked {count, tiers}`, aggregate-only, never joinable to a row).
**Live anonymous, cache-busted:** 25 grades, **judgment leak NONE**, every row `locked`, tease
`{count:20, tiers:{B:13,C:7}}`, `free_sample` 3 of 5 settled — **exploit dead**. `outcome` is kept on
gated rows because a result is a fact, not a judgment.
**Verification note:** the first post-deploy read was a CDN-cached pre-deploy body and falsely showed
a leak; `max-age=30` on this endpoint means post-deploy checks must bust the cache.
319 suites / 3,970 tests green, web build exit 0.
---
# FREE PROOF SURFACE — `/record` (shipped 2026-07-31)
Tier-record-forward, presentation-only (`src/` untouched). Reads `/api/accuracy` + `/api/ledger/model`
and prints them as-is: **B 60% n512, C 57% n413 with C honestly below B; A/D/F hollow with real
samples**. Sport slicing is client-side (the endpoints ignore `?sport=`).
**The rule that matters:** a withheld percentage stays null — A is 1/2 and a test asserts we do **not**
derive 50%. **CLV renders an honest absence**: `beat_close_pct` is null behind the capture-reliability
guard, so the panel says NOT PUBLISHED YET and explains why, and tests forbid the computable-but-wrong
34/937 = 3.6% fallback. Calibration and accuracy-over-time are **named as held**, not faked — no honest
source exists yet.
Cache-busted fingerprint: page served, CLV honest-absent, guard named, held pieces named, **no "3.6"**,
**no hard-coded percentage in the markup**. The hollow-explanation line is client-conditional → Chrome
audit. 320 suites / 3,986 tests green, web build exit 0.
+103
View File
@@ -0,0 +1,103 @@
#!/usr/bin/env node
'use strict';
/**
* backfill-context — WERE WE WAITING, OR UNDER-QUERYING?
*
* The platoon test ran on 452 rows against 1,266 clean settled hits rows in the
* ledger, so "48 short of the gate" was never a statement about how much data
* exists. It was a statement about how much the JOIN survived — and the join was
* losing rows to inputs we simply had not fetched for every player.
*
* This backfills the inputs (pure sample, zero waiting) and reports exactly
* where each row is lost, so the next "we need more data" claim is a measured
* one rather than an inherited one.
*
* ── THE ONE HONEST CAVEAT, STATED UP FRONT ───────────────────────────────
* Platoon splits from statsapi are SEASON-TO-DATE as of the moment they are
* fetched. Applying today's split to a 2026-07-15 game means the split contains
* that game. For a ~400-PA season line one game is roughly a quarter of one
* percent, so the contamination is small — but it is real, it runs in the
* flattering direction, and it is why this is labelled a reconstruction rather
* than a clean point-in-time backtest.
*
* SUPABASE_URL=... node scripts/backfill-context.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const ctx = require('../src/services/lineupContextService');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const SEASON = Number(process.env.BF_SEASON || 2026);
const PAGE = 1000;
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
// Every hitter who appears on a CLEAN settled row — the true denominator.
const led = await page(sb, 'ledger_entries', 'player_key, player_name, stat, outcome, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases'])
.in('outcome', ['hit', 'miss']));
const need = new Map();
for (const r of led) {
if ((r.quarantine_reason || '').startsWith('nontakeable_book')) continue;
if (!need.has(r.player_key)) need.set(r.player_key, r.player_name);
}
const have = new Set((await page(sb, 'platoon_splits', 'player_key', (q) => q.eq('sport', 'mlb')))
.map((r) => r.player_key));
const missing = [...need.entries()].filter(([k]) => !have.has(k));
console.error(`[backfill] hitters on clean settled rows: ${need.size}; splits already held: ${have.size}; to fetch: ${missing.length}`);
const asOf = ctx.dateET();
const rows = [];
let unresolved = 0;
for (const [key, name] of missing) {
let found = null;
try { found = await mlb.searchPlayer(name); } catch { found = null; }
if (!found || !found.id) { unresolved += 1; continue; }
const sp = await ctx.fetchPlatoonSplits(found.id, SEASON, {});
if (!sp) continue; // absent, never a symmetric guess
rows.push({ player_key: key, player_name: name, source_id: found.id, ...sp });
}
let written = 0;
for (let i = 0; i < rows.length; i += 200) {
const batch = rows.slice(i, i + 200).map((r) => ({ ...r, sport: 'mlb', season: SEASON, as_of_date: asOf }));
const { error } = await sb.from('platoon_splits')
.upsert(batch, { onConflict: 'as_of_date,sport,season,player_key' });
if (!error) written += batch.length;
else console.error('[backfill] write failed:', error.message);
}
console.log(JSON.stringify({
hitters_on_clean_settled_rows: need.size,
splits_held_before: have.size,
attempted: missing.length,
unresolved_by_name: unresolved,
no_splits_available: missing.length - unresolved - rows.length,
written,
caveat: 'season-to-date splits applied to past games contain those games — small (~0.25% of a 400-PA line) but real and flattering',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+94 -14
View File
@@ -51,23 +51,103 @@ pg_dump "${SUPABASE_DB_URL}" -Fc --no-owner --no-privileges -f "${DUMP}" \
SIZE="$(stat -c%s "${DUMP}" 2>/dev/null || echo 0)" SIZE="$(stat -c%s "${DUMP}" 2>/dev/null || echo 0)"
[ "${SIZE}" -ge "${MIN_BYTES}" ] || fail "dump is only ${SIZE} bytes (< ${MIN_BYTES}) — treating as a failed backup" [ "${SIZE}" -ge "${MIN_BYTES}" ] || fail "dump is only ${SIZE} bytes (< ${MIN_BYTES}) — treating as a failed backup"
# 2b. Integrity fingerprint: a valid custom-format archive lists its objects via
# pg_restore --list (no target DB needed). Confirm it parses AND contains the
# ledger — proves it's a real, restorable archive, not just a file of bytes.
TOC="$(pg_restore --list "${DUMP}" 2>/dev/null)" || fail "pg_restore --list failed — dump is not a valid archive"
OBJECTS="$(printf '%s\n' "${TOC}" | grep -c ';' || true)"
printf '%s\n' "${TOC}" | grep -qi 'TABLE DATA public ledger_entries' \
|| fail "dump archive does not contain ledger_entries — refusing to trust it"
echo "backup validated: ${DUMP} (${SIZE} bytes, ${OBJECTS} archive objects, ledger_entries present)"
# 3. Rotate: drop local dumps older than KEEP_DAYS. # 3. Rotate: drop local dumps older than KEEP_DAYS.
find "${BACKUP_DIR}" -name 'vyndr-*.dump' -type f -mtime "+${KEEP_DAYS}" -delete || true find "${BACKUP_DIR}" -name 'vyndr-*.dump' -type f -mtime "+${KEEP_DAYS}" -delete || true
# 4. Weekly off-box copy (Sundays). A single failure of the off-box push is a # 3b. SSH key for the off-box push (Session 64).
# WARN, not a hard failure — the local dump still succeeded. # The container filesystem is EPHEMERAL — a keypair generated inside it dies
if [ "$(date -u +%u)" = "7" ]; then # on the next redeploy and the off-box push would silently start failing. So
if [ -n "${BACKUP_REMOTE:-}" ]; then # the PRIVATE key is injected as an env var (Coolify secret) and written to a
if rsync -az --timeout=120 "${DUMP}" "${BACKUP_REMOTE}"; then # 0600 temp file per run. Hetzner Storage Box speaks full OpenSSH on PORT 23
notify "VYNDR backup OK (+off-box)" "default" "Nightly dump ${STAMP} (${SIZE} bytes) + weekly off-box copy pushed." # (port 22 is SFTP-only, mod_sftp) — verified live; rsync must target 23.
else SSH_KEY_FILE=""
notify "VYNDR off-box push FAILED" "high" "Local dump ${STAMP} is fine (${SIZE} bytes) but the weekly off-box rsync failed." cleanup_key() { [ -n "${SSH_KEY_FILE}" ] && rm -f "${SSH_KEY_FILE}" || true; }
fi trap cleanup_key EXIT
else # HOST KEY IS STATICALLY PINNED (Session 64). This used to be
notify "VYNDR off-box push SKIPPED" "high" "Local dump ${STAMP} OK but BACKUP_REMOTE is unset — no off-box copy this week." # StrictHostKeyChecking=accept-new — trust-on-first-use, which accepts whatever
fi # host key it meets first and would happily trust an impostor on the first run.
else # scripts/storagebox_known_hosts carries the ED25519 line verified out-of-band
echo "backup ok: ${DUMP} (${SIZE} bytes)" # (SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM). A mismatch is now a HARD
# FAIL — which is the point. Never weaken this to accept-new/=no//dev/null.
KNOWN_HOSTS="${BACKUP_KNOWN_HOSTS:-$(dirname "$0")/storagebox_known_hosts}"
[ -f "${KNOWN_HOSTS}" ] || fail "pinned known_hosts missing at ${KNOWN_HOSTS} — refusing to push without host-key verification"
RSYNC_SSH="ssh -p ${BACKUP_SSH_PORT:-23} -o StrictHostKeyChecking=yes -o UserKnownHostsFile=${KNOWN_HOSTS} -o BatchMode=yes"
if [ -n "${BACKUP_SSH_KEY:-}" ]; then
SSH_KEY_FILE="$(mktemp)"
chmod 600 "${SSH_KEY_FILE}"
# Accept EITHER form (Session 64):
# 1. base64 (recommended — `base64 -w0`; survives any env-var mangling of
# newlines, which is the usual way an injected SSH key silently breaks)
# 2. raw PEM with literal \n escapes
# Detect base64 by trying to decode and checking for the PEM header.
DECODED="$(printf '%s' "${BACKUP_SSH_KEY}" | base64 -d 2>/dev/null || true)"
case "${DECODED}" in
*"PRIVATE KEY"*)
printf '%s\n' "${DECODED}" > "${SSH_KEY_FILE}"
echo "ssh key: base64-decoded"
;;
*)
printf '%b\n' "${BACKUP_SSH_KEY}" | sed -e 's/[[:space:]]*$//' > "${SSH_KEY_FILE}"
echo "ssh key: used raw (not base64)"
;;
esac
chmod 600 "${SSH_KEY_FILE}"
RSYNC_SSH="${RSYNC_SSH} -i ${SSH_KEY_FILE}"
fi fi
# 4. OFF-BOX COPY — DEFERRED (Session 64).
# The dump now lands on a PERSISTENT VOLUME (BACKUP_DIR=/app/backups), so it
# already survives redeploys — the container-ephemeral risk is closed. Storage
# Box SSH auth is not working yet, so the off-box push is explicitly DEFERRED:
# it must never fail the backup. A durable on-box dump is a real backup; a
# failing rsync on top of it is a follow-up, not an incident.
#
# Set BACKUP_OFFBOX=1 (with BACKUP_SSH_KEY) to re-enable. Until then we log
# and page at LOW priority, and we never call a deferred push a failure.
if [ "${BACKUP_OFFBOX:-0}" = "1" ] && [ -n "${BACKUP_REMOTE:-}" ] && [ -n "${BACKUP_SSH_KEY:-}" ]; then
# Phase 1b — GUARANTEE THE REMOTE DIRECTORY EXISTS. The box starts with only
# .ssh/, and rsync of a file into a missing parent either fails or silently
# writes the dump AS the directory name (one file, overwritten nightly, which
# would read as "backups exist" while retaining exactly one). Prefer rsync's
# own --mkpath; fall back to an explicit ssh mkdir -p for older rsync.
REMOTE_HOST="${BACKUP_REMOTE%%:*}"
REMOTE_PATH="${BACKUP_REMOTE#*:}"
MKPATH_FLAG=""
if rsync --help 2>&1 | grep -q -- '--mkpath'; then
MKPATH_FLAG="--mkpath"
else
${RSYNC_SSH} "${REMOTE_HOST}" "mkdir -p '${REMOTE_PATH}'" \
|| echo "warn: remote mkdir -p failed; relying on an existing directory"
fi
# OFF-BOX IS REQUIRED NOW (Session 64, Phase 2b). A failed push must never
# again read as success: it PAGES, and the run reports offbox_ok:false.
# Exit code deliberately still reflects ON-BOX durability — a good on-box dump
# must not raise a false total-failure alarm. Surface the truth; don't
# manufacture a failure.
if rsync -az --timeout=120 ${MKPATH_FLAG} -e "${RSYNC_SSH}" "${DUMP}" "${BACKUP_REMOTE}"; then
echo "off-box push OK -> ${BACKUP_REMOTE}"
echo "OFFBOX_OK=1"
notify "VYNDR backup OK (+off-box)" "default" "Nightly dump ${STAMP} (${SIZE} bytes) pushed off-box."
else
echo "off-box push FAILED (on-box dump is durable, but OFF-BOX IS REQUIRED)"
echo "OFFBOX_OK=0"
notify "VYNDR OFF-BOX PUSH FAILED" "urgent" "Dump ${STAMP} (${SIZE} bytes) is on the persistent volume but did NOT reach the Storage Box. The database has no off-box copy tonight."
fi
else
echo "off-box push DEFERRED (BACKUP_OFFBOX!=1 or remote/key unset) — on-box dump is durable at ${DUMP}"
echo "OFFBOX_OK=deferred"
fi
echo "backup ok: ${DUMP} (${SIZE} bytes)"
exit 0 exit 0
+100
View File
@@ -0,0 +1,100 @@
#!/usr/bin/env node
'use strict';
/**
* build-grade-bands — publish what each letter actually means, per archetype.
*
* Runs the real settled ledger through gradeBands. Because no factor has passed
* the gate for any archetype, every band comes back a BASE-RATE read — which is
* the honest answer today, and the output states it rather than leaving a reader
* to infer it.
*
* SUPABASE_URL=... node scripts/build-grade-bands.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const gb = require('../src/services/model/gradeBands');
const tl = require('../src/services/model/testLedger');
const cal = require('../src/services/model/calibration');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const STAT = process.env.BAND_STAT || 'hits';
const PAGE = 1000;
/**
* PROVEN, PER ARCHETYPE. Empty, and that is the measured state — see
* specs/per-archetype-re-audit.md. Nothing may be added here that has not
* cleared the gate FOR THAT ARCHETYPE; pooled proof does not qualify a slot.
*/
const PROVEN_BY_ARCHETYPE = Object.freeze({});
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat',
(q) => q.eq('sport', 'mlb').eq('stat', STAT).not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries', 'player_key, game_date, outcome, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', STAT)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
const byArch = new Map();
for (const r of clean) {
const a = String(archOf.get(`${r.player_key}|${r.game_date}`) || 'UNLABELLED').toUpperCase();
if (!byArch.has(a)) byArch.set(a, []);
byArch.get(a).push({ p: knownNumber(r.p_win), won: r.outcome === 'hit' ? 1 : 0 });
}
const mc = await tl.recordAndCount(tl.supabaseStore(sb),
[...byArch.keys()].map((a) => ({
sport: 'mlb', stat: STAT, archetype: a === 'UNLABELLED' ? null : a,
interaction: 'grade_band_lift', target: 'outcome',
})));
// Calibration is measured, not assumed. Today it is certified for hits only in
// a middle band (specs — held-out error 0.477->0.506, 0.587->0.580), which is
// NOT the same as an archetype's probabilities being calibrated.
const certified = typeof cal.certifyBands === 'function';
const out = [];
for (const [arch, rows] of [...byArch.entries()].sort((a, b) => b[1].length - a[1].length)) {
out.push(gb.buildBands(rows, {
archetype: arch,
cumulativeTests: mc.cumulative_tests,
proven: Boolean(PROVEN_BY_ARCHETYPE[arch]),
calibrated: false, // no archetype's distribution is certified calibrated
}));
}
console.log(JSON.stringify({
stat: STAT,
settled_rows: clean.length,
archetypes: byArch.size,
cumulative_tests: mc.cumulative_tests,
calibration_helper_present: certified,
proven_by_archetype: PROVEN_BY_ARCHETYPE,
note: 'every band is a BASE-RATE read — no factor has passed the gate for any archetype',
bands: out,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+267
View File
@@ -0,0 +1,267 @@
#!/usr/bin/env node
'use strict';
/**
* calibrate-four-stats — Phases 2, 3, 4 and 7.
*
* ── ONE SIDE PER PROP, OR THE MEASUREMENT IS MEANINGLESS ─────────────────
* 97.6% of snapshot props carry BOTH the over and the under. Their p_wins sum to
* ~1 and their outcomes are complementary, so any calibration statistic over the
* raw population is pinned to 0.5 by symmetry. Measured that way the counter
* looks perfectly calibrated (+0.0002 on hits); deduped to the model-PICKED side
* it is +0.0868. Same rows, opposite conclusion.
*
* ── DATE-CLUSTERED, PER THE ORDER ────────────────────────────────────────
* A day's offensive environment is a real shared component, so uncertainty is
* clustered on the game DATE rather than the game. That is the honest unit for a
* systematic-bias claim and it is a much harder bar than game-clustering.
*
* SUPABASE_URL=... node scripts/calibrate-four-stats.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** The order's deploy floor: dates, not games. */
const MIN_DATE_CLUSTERS = 40;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
const brier = (ps, ys) => mean(ps.map((p, i) => (p - ys[i]) ** 2));
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select))
.order('id', { ascending: true }).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
const isPreGame = (capturedAt, gameDate) => {
const et = new Date(new Date(capturedAt).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < gameDate || (d === gameDate && et.getUTCHours() < 19);
};
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
/** Paired bootstrap on the Brier difference, resampling DATES. */
function dateClusteredCI(rows, cumulativeTests = 1, iters = 3000) {
const byDate = new Map();
for (const r of rows) {
if (!byDate.has(r.date)) byDate.set(r.date, []);
byDate.get(r.date).push(r);
}
const keys = [...byDate.keys()];
const rnd = makeRnd(20260807);
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const raw = []; const adj = []; const ys = [];
for (let i = 0; i < keys.length; i += 1) {
for (const r of byDate.get(keys[Math.floor(rnd() * keys.length)])) {
raw.push(r.p); adj.push(r.pc); ys.push(r.won);
}
}
diffs.push(brier(adj, ys) - brier(raw, ys));
}
diffs.sort((a, b) => a - b);
const tests = Math.max(1, Math.round(cumulativeTests));
const alpha = 0.05 / tests;
const q = (x) => diffs[Math.floor(Math.min(diffs.length - 1, Math.max(0, x * (diffs.length - 1))))];
return { ci: [round4(q(alpha / 2)), round4(q(1 - alpha / 2))], date_clusters: keys.length, ci_level: round4(1 - alpha) };
}
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused, grade',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
// ── ONE SIDE PER PROP: the side the model picked (its higher p_win). ──
const picked = new Map();
const refusedProps = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date)) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
if (r.refused || knownNumber(r.p_win) === null) {
if (!refusedProps.has(k)) refusedProps.set(k, r);
continue;
}
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
const resolve = (r) => {
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) return null;
const v = knownNumber(FIELD[r.stat](b));
if (v === null) return null;
const over = v > L;
return { over, won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0, realized: v };
};
const out = { deploy_floor_date_clusters: MIN_DATE_CLUSTERS, per_stat: {}, refusal_accuracy: {} };
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const res = resolve(r);
if (!res) continue;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: res.won });
}
rows.sort((a, b) => String(a.date).localeCompare(String(b.date)));
const dates = [...new Set(rows.map((r) => r.date))].sort();
if (rows.length < 100 || dates.length < 3) {
out.per_stat[stat] = { n: rows.length, date_clusters: dates.length, decision: 'REFUSE', reason: 'too few rows or dates to split point-in-time' };
continue;
}
// POINT-IN-TIME: fit strictly on earlier dates, evaluate on later ones.
//
// The cut is placed by ROW COUNT rather than by date index. Props are not
// spread evenly across dates -- hits concentrate in the later ones -- so a
// 60%-of-DATES cut left only 143 rows to fit on, under the 200 the fitter
// needs. Splitting on cumulative rows keeps the split strictly temporal
// (every fit date precedes every eval date) while giving both sides enough
// to work with.
const perDate = new Map();
for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) {
acc += perDate.get(d) || 0;
if (acc >= rows.length * 0.45) { cut = d; break; }
}
const fit = rows.filter((r) => r.date < cut);
const ev = rows.filter((r) => r.date >= cut);
if (fit.length < 50 || ev.length < 50) {
out.per_stat[stat] = { n: rows.length, date_clusters: dates.length, decision: 'REFUSE', reason: 'time split leaves too little on one side' };
continue;
}
const iso = cal.fitIsotonic(fit.map((r) => ({ p: r.p, won: r.won })));
// NULL IS NOT A PREDICTION. fitIsotonic returns null below its minimum and
// applyIsotonic then returns null per row -- and (null - 1)**2 === 1 while
// (null - 0)**2 === 0, so a "Brier score" computed over nulls is silently
// just the win rate. That is exactly the Number(null) === 0 breach this
// codebase keeps having to catch, and it produced a fake 0.5567 for hits.
if (!iso) {
out.per_stat[stat] = {
n: rows.length, date_clusters: dates.length, fit_n: fit.length, eval_n: ev.length,
decision: 'REFUSE', reason: `no calibration map could be fitted on ${fit.length} fit rows`,
};
continue;
}
const scored = ev.map((r) => ({ ...r, pc: cal.applyIsotonic(iso, r.p) }))
.filter((r) => knownNumber(r.pc) !== null);
if (scored.length < 50) {
out.per_stat[stat] = {
n: rows.length, date_clusters: dates.length,
decision: 'REFUSE', reason: `only ${scored.length} eval rows could be mapped`,
};
continue;
}
const ys = scored.map((r) => r.won);
const bRaw = brier(scored.map((r) => r.p), ys);
const bCal = brier(scored.map((r) => r.pc), ys);
const { ci, date_clusters, ci_level } = dateClusteredCI(scored, 1);
// CERTIFIED BAND: p_win deciles where held-out |predicted - actual| is small.
const bands = [];
for (let lo = 0.3; lo < 0.95; lo += 0.1) {
const slice = scored.filter((r) => r.p >= lo && r.p < lo + 0.1);
if (slice.length < 25) continue;
const pred = mean(slice.map((r) => r.pc));
const act = mean(slice.map((r) => r.won));
bands.push({ range: [round2(lo), round2(lo + 0.1)], n: slice.length, calibrated_pred: round4(pred), actual: round4(act), err: round4(pred - act) });
}
const certified = bands.filter((b) => Math.abs(b.err) <= 0.05).map((b) => b.range);
const improves = bCal < bRaw && ci[1] < 0;
const enoughDates = date_clusters >= MIN_DATE_CLUSTERS;
// PHASE 4 — bias SHAPE across the p_win range (diagnostic only).
const shape = [];
for (let lo = 0.3; lo < 0.95; lo += 0.1) {
const slice = rows.filter((r) => r.p >= lo && r.p < lo + 0.1);
if (slice.length < 25) continue;
shape.push({ range: [round2(lo), round2(lo + 0.1)], n: slice.length, bias: round4(mean(slice.map((r) => r.p)) - mean(slice.map((r) => r.won))) });
}
out.per_stat[stat] = {
n: rows.length,
date_clusters: dates.length,
bias_pre: round4(mean(rows.map((r) => r.p)) - mean(rows.map((r) => r.won))),
fit_n: fit.length, eval_n: ev.length, split_at: cut,
brier_raw: round4(bRaw),
brier_calibrated: round4(bCal),
brier_delta: round4(bCal - bRaw),
ci_date_clustered: ci,
ci_level,
eval_date_clusters: date_clusters,
certified_bands: certified,
band_detail: bands,
bias_shape: shape,
decision: improves && enoughDates ? 'DEPLOY' : 'REFUSE',
reason: improves && enoughDates ? 'held-out Brier improves, date-clustered, and the date floor is met'
: (!enoughDates
? `date-clusters ${date_clusters} < ${MIN_DATE_CLUSTERS} — the honest unit for a systematic-bias claim`
: 'held-out Brier does not improve at the date-clustered interval'),
};
}
// ── PHASE 7 — REFUSAL ACCURACY ──
// The model passed on these. A pass is CORRECT when there was genuinely
// nothing to call: the over lands near a coin flip rather than at an
// exploitable rate.
for (const stat of STATS) {
const refs = [];
for (const r of refusedProps.values()) {
if (r.stat !== stat) continue;
const res = resolve({ ...r, side: 'over' });
if (res) refs.push(res.over ? 1 : 0);
}
const graded = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const res = resolve({ ...r, side: 'over' });
if (res) graded.push(res.over ? 1 : 0);
}
out.refusal_accuracy[stat] = {
refused_n: refs.length,
refused_over_rate: refs.length ? round4(mean(refs)) : null,
graded_over_rate: graded.length ? round4(mean(graded)) : null,
refused_distance_from_coinflip: refs.length ? round4(Math.abs(mean(refs) - 0.5)) : null,
graded_distance_from_coinflip: graded.length ? round4(Math.abs(mean(graded) - 0.5)) : null,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round2 = (v) => Math.round(v * 100) / 100;
main().catch((e) => { console.error(e); process.exit(1); });
+136
View File
@@ -0,0 +1,136 @@
#!/usr/bin/env node
'use strict';
/**
* calibrate-hits — FIT PAST, APPLY FORWARD, VERIFY HELD-OUT.
*
* The parlay surface is blocked because hit probabilities are well-ranked and
* badly calibrated: the model claims 0.911 and realises 0.630, and it is flat
* above 0.70. Compounding multiplies that error, so the repair has to be proven
* on data the correction never saw.
*
* THE ONE DISCIPLINE THAT MAKES THIS MEAN ANYTHING: the map is fitted on an
* EARLIER window and evaluated on a LATER one. Fitting and evaluating on the
* same rows always looks perfectly calibrated — that is not a result, it is the
* map reciting the answers it was built from. Any calibration report that does
* not name its split should be assumed to have done exactly that.
*
* WHY ISOTONIC. It is monotone by construction, so the model's ORDERING survives
* untouched and only the magnitudes move. We are repairing what it counts, not
* what it ranks — and the ranking is the part that measured well.
*
* PASS CONDITION: the TOP BINS (0.70+) must be honest out-of-sample. A parlay is
* built from confident legs, so calibration that only holds in the middle is
* worthless for the thing this unblocks.
*
* SUPABASE_URL=... node scripts/calibrate-hits.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const SPLIT = process.env.CAL_SPLIT || '2026-08-02'; // held-out starts here
const PAGE = 1000;
const r3 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 1000) / 1000);
async function page(sb, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from('ledger_entries')
.select('p_win, outcome, game_date, quarantine_reason')).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/** Reliability rendered per bin with n — the only honest way to read this. */
function curve(rows, label) {
return cal.reliability(rows, 10)
.filter((b) => b.n >= 10)
.map((b) => ({
window: label,
predicted: r3(b.mean_predicted),
actual: r3(b.actual),
error: r3(b.error),
n: b.n,
}));
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const raw = await page(sb, (q) => q.eq('sport', 'mlb').is('user_id', null)
.eq('stat', 'hits').in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const all = raw
.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'))
.map((r) => ({ p: Number(r.p_win), won: r.outcome === 'hit' ? 1 : 0, d: String(r.game_date) }));
const fit = all.filter((r) => r.d < SPLIT);
const held = all.filter((r) => r.d >= SPLIT);
const map = cal.fitIsotonic(fit);
if (!map) {
console.log(JSON.stringify({ ok: false, reason: 'could not fit', fit_n: fit.length }));
process.exit(0);
}
// Apply the FIT-WINDOW map to the HELD-OUT rows. The map has never seen these.
const corrected = held.map((r) => ({ ...r, p: cal.applyIsotonic(map, r.p) }));
const before = cal.isCalibrated(held, { minTotal: 100 });
const after = cal.isCalibrated(corrected, { minTotal: 100 });
// Did the ORDERING survive? Isotonic is monotone, so it must — checked rather
// than asserted, because a broken map would silently destroy the one thing
// the model does well.
const pairs = [];
for (let i = 0; i < Math.min(held.length, 400); i += 1) {
for (let j = i + 1; j < Math.min(held.length, 400); j += 1) {
if (held[i].p === held[j].p) continue;
const rawOrder = Math.sign(held[i].p - held[j].p);
const calOrder = Math.sign(corrected[i].p - corrected[j].p);
pairs.push(calOrder === 0 || calOrder === rawOrder);
}
}
const orderingPreserved = pairs.length === 0 || pairs.every(Boolean);
const topBefore = curve(held, 'held-out RAW').filter((b) => b.predicted >= 0.70);
// NOT ">= 0.70": honest calibration REMOVES the 0.70+ predictions entirely
// (the ceiling drops to ~0.667), so demanding that band exist would fail the
// map for succeeding. The right question is whether the model's HIGHEST
// REMAINING confidence band is honest, because that is what a parlay stacks.
const afterCurve = curve(corrected, 'held-out CALIBRATED');
const topAfter = afterCurve.slice(-2);
const bands = cal.certifyBands(corrected, { tolerance: 0.05, minBin: 40 });
const ceiling = afterCurve.length ? Math.max(...afterCurve.map((b) => b.predicted)) : null;
console.log(JSON.stringify({
discipline: `fitted on game_date < ${SPLIT}, evaluated on game_date >= ${SPLIT} — the map never saw the evaluation rows`,
fit_n: fit.length,
held_out_n: held.length,
ordering_preserved: orderingPreserved,
isotonic_blocks: map.length,
map_sample: [0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 0.95].map((p) => ({ claims: p, corrected_to: r3(cal.applyIsotonic(map, p)) })),
held_out_before: { calibrated: before.calibrated, max_bin_error: r3(before.max_bin_error), curve: curve(held, 'RAW') },
held_out_after: { calibrated: after.calibrated, max_bin_error: r3(after.max_bin_error), curve: curve(corrected, 'CALIBRATED') },
top_bins_before: topBefore,
top_bins_after: topAfter,
certified_bands: bands,
honest_ceiling: ceiling,
four_leg_ticket_at_ceiling: ceiling ? r3(ceiling ** 4) : null,
verdict: !orderingPreserved ? 'FAIL — ordering destroyed'
: bands.length === 0 ? 'FAIL — no band is honest out-of-sample'
: `PARTIAL PASS — honest within ${bands.map((b) => `${b.lo}-${b.hi}`).join(', ')}; outside those bands legs are NOT stackable`,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+191
View File
@@ -0,0 +1,191 @@
#!/usr/bin/env node
'use strict';
/**
* challenger-scoreboard — every accruing challenger, measured on the same bar.
*
* WHY THIS IS A REAL HOLDOUT AND NOT A BACKTEST. Unlike hits-v1 (which did not
* exist when these rows were graded and therefore needed a point-in-time
* replay), arch-v1 and contact-v1 wrote their probability AT GRADE TIME, into
* their own columns, before the game was played. Nothing here is recomputed.
* These numbers are genuinely out-of-sample — the strongest evidence available.
*
* THE BAR IS THE SAME ONE THAT REFUTED hits-v1:
* - the challenger's OWN rows only (a challenger that abstains is not scored
* on the rows it declined — averaging those in measures the champion twice)
* - direction handled: p_win and p_win_challenger/p_win_contact are all
* P(GRADED SIDE), so they are already aligned. The projection ladder is
* P(OVER) and IS realigned here.
* - paired bootstrap on matched rows, because both models score the SAME rows
* and treating their errors as independent overstates the uncertainty
* - PROMOTE only when the CI on (challenger champion) excludes zero
*
* CONTAMINATION EXCLUSION: rows whose price/book were stamped from a
* non-takeable book are tagged `quarantine_reason LIKE 'nontakeable_book%'` and
* are excluded — their locked price describes a market you could not have bet.
*
* PROVENANCE: results are reported for all rows AND split by `model_version`,
* so if a verdict depends on the older `pre-retention-unknown` era that fact is
* visible rather than buried.
*
* SUPABASE_URL=... node scripts/challenger-scoreboard.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
function corr(xs, ys) {
const n = xs.length;
if (n < 3) return null;
const mx = xs.reduce((a, b) => a + b, 0) / n;
const my = ys.reduce((a, b) => a + b, 0) / n;
let sxy = 0; let sxx = 0; let syy = 0;
for (let i = 0; i < n; i += 1) {
const dx = xs[i] - mx; const dy = ys[i] - my;
sxy += dx * dy; sxx += dx * dx; syy += dy * dy;
}
if (sxx <= 0 || syy <= 0) return null;
return sxy / Math.sqrt(sxx * syy);
}
const r4 = (v) => (v == null ? null : Math.round(v * 10000) / 10000);
const meanOf = (a) => (a.length ? a.reduce((x, y) => x + y, 0) / a.length : null);
const brier = (ps, ys) => (ps.length ? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
/** Paired bootstrap on the DIFFERENCE of resolutions. Deterministic seed. */
function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260803) {
if (rows.length < 30) return null;
let s = seed >>> 0;
const rnd = () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
const n = rows.length;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const b = [];
for (let i = 0; i < n; i += 1) {
const r = rows[Math.floor(rnd() * n)];
ys.push(r.won); a.push(r[keyA]); b.push(r[keyB]);
}
const ca = corr(a, ys); const cb = corr(b, ys);
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (p) => r4(diffs[Math.floor(p * (diffs.length - 1))]);
const point = r4(corr(rows.map((r) => r[keyA]), rows.map((r) => r.won))
- corr(rows.map((r) => r[keyB]), rows.map((r) => r.won)));
const ci = [q(0.025), q(0.975)];
return { point, ci95: ci, p_improves: r4(diffs.filter((d) => d > 0).length / diffs.length),
ci_excludes_zero: ci[0] > 0 || ci[1] < 0 };
}
function score(rows, challKey, label) {
const ys = rows.map((r) => r.won);
const ch = rows.map((r) => r[challKey]);
const cp = rows.map((r) => r.champ);
const bs = bootstrapDiff(rows, challKey, 'champ');
let verdict = 'STILL PENDING';
if (rows.length >= 30 && bs) {
if (bs.ci_excludes_zero && bs.point > 0) verdict = 'PROMOTE';
else if (bs.ci_excludes_zero && bs.point < 0) verdict = 'STAY WIRED (measured worse)';
else verdict = 'STAY WIRED (inconclusive)';
}
return {
challenger: label,
settled_n: rows.length,
base_rate: r4(meanOf(ys)),
resolution_challenger: r4(corr(ch, ys)),
resolution_champion: r4(corr(cp, ys)),
brier_challenger: r4(brier(ch, ys)),
brier_champion: r4(brier(cp, ys)),
delta_vs_champion: bs,
verdict,
};
}
async function fetchAll(sb) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await sb.from('ledger_entries')
.select('id, stat, side, outcome, model_version, quarantine_reason, p_win, p_win_challenger, p_win_contact, challenger_delta, contact_delta, proj_p_over_line, proj_tb_p_over, proj_hits_p_over, challenger_adjustments')
.eq('sport', 'mlb').is('user_id', null)
.in('outcome', ['hit', 'miss'])
.not('p_win', 'is', null)
.range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const raw = (await fetchAll(sb))
.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
const base = raw.map((r) => ({
won: r.outcome === 'hit' ? 1 : 0,
champ: Number(r.p_win),
arch: r.p_win_challenger == null ? null : Number(r.p_win_challenger),
contact: r.p_win_contact == null ? null : Number(r.p_win_contact),
// The ladder is P(OVER); realign it to the graded side before comparing.
ladder: r.proj_p_over_line == null ? null
: (String(r.side).toLowerCase() === 'under' ? 1 - Number(r.proj_p_over_line) : Number(r.proj_p_over_line)),
tb: r.proj_tb_p_over == null ? null
: (String(r.side).toLowerCase() === 'under' ? 1 - Number(r.proj_tb_p_over) : Number(r.proj_tb_p_over)),
hits: r.proj_hits_p_over == null ? null
: (String(r.side).toLowerCase() === 'under' ? 1 - Number(r.proj_hits_p_over) : Number(r.proj_hits_p_over)),
stat: r.stat,
// Did the challenger actually MOVE this row? A nudge that leaves p_win
// untouched is the champion wearing a different name, and scoring it on
// those rows measures the champion against itself — which is exactly how a
// real effect gets averaged down to zero.
archMoved: r.challenger_delta != null && Number(r.challenger_delta) !== 0,
contactMoved: r.contact_delta != null && Number(r.contact_delta) !== 0,
era: r.model_version || 'unknown',
axes: new Set(((r.challenger_adjustments) || []).map((a) => a && a.axis).filter(Boolean)),
}));
const withKey = (k, extra = () => true) => base.filter((r) => r[k] != null && Number.isFinite(r[k]) && extra(r));
const board = [
score(withKey('arch'), 'arch', 'arch-v1 (market-relative nudge)'),
score(withKey('contact'), 'contact', 'contact-v1 (season contact quality)'),
score(withKey('ladder'), 'ladder', 'proj-v1.1 ladder (all stats)'),
score(withKey('tb', (r) => r.stat === 'total_bases'), 'tb', 'tb-v1 (total_bases only)'),
score(withKey('hits', (r) => r.stat === 'hits'), 'hits', 'hits-v1 (hits only)'),
];
// THE SHARPEST TEST OF A NUDGE — only the rows it actually moved.
const movedBoard = [
score(withKey('arch', (r) => r.archMoved), 'arch', 'arch-v1 · rows it MOVED only'),
score(withKey('contact', (r) => r.contactMoved), 'contact', 'contact-v1 · rows it MOVED only'),
];
// Per-AXIS: arch-v1 restricted to the rows where that axis actually fired.
const axisBoard = ['environment', 'opportunity', 'matchup'].map((ax) =>
score(withKey('arch', (r) => r.axes.has(ax)), 'arch', `arch-v1 · ${ax} axis rows only`));
// PROVENANCE split — does any verdict depend on the older era?
const eras = [...new Set(base.map((r) => r.era))];
const provenance = eras.map((era) => ({
era,
...score(withKey('arch', (r) => r.era === era), 'arch', `arch-v1 · ${era}`),
}));
console.log(JSON.stringify({
measurement: 'PROSPECTIVE HOLDOUT — challenger values were written at grade time, before the game. No recomputation, no lookahead.',
total_settled_rows: base.length,
board, movedBoard, axisBoard, provenance,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+374
View File
@@ -0,0 +1,374 @@
#!/usr/bin/env node
'use strict';
/**
* champion-ablation — WHERE DOES THE CHAMPION'S RESOLUTION ACTUALLY COME FROM?
*
* Convergent evidence says the problem is INPUTS, not shape: hits-v1 refuted,
* the ladder reliably worse (0.030), arch-v1 moving 76% of rows to exactly zero
* effect, contact/environment/opportunity all CI-includes-zero. Every one of
* those changed the DISTRIBUTION or added a NUDGE. None changed the information.
* So before building a sixth thing, decompose the champion.
*
* THE CHAMPION IS FIVE LINES OF ARITHMETIC (probabilityEstimator):
*
* base = empirical frequency of stat > line over the game log
* weighted = 0.6·base + 0.4·(same frequency over the last 5)
* p = weighted + oppAdj(±0.03) + homeAdj(±0.015)
* if cv>0.40: p = 0.9·p + 0.05 (volatile → pull toward 0.50)
* p_over = clamp(p, 0.10, 0.95); p_win = side==='under' ? 1p_over : p_over
*
* THE ABLATION IS EXACT, NOT A REFIT. Every adjustment is a closed-form function
* of stored features, and the consistency step is linear, so each layer can be
* removed analytically from the stored p_win:
*
* f(x) = 0.9x + 0.05 ⟹ f(a+b) = f(a) + 0.9b
*
* so subtracting an adjustment is subtracting k·adj with k = 0.9 when the
* consistency pull fired and 1 when it did not. Nothing is re-estimated, no
* model is refit, and no game log is re-fetched — which also means no lookahead
* is even possible here.
*
* WHAT CANNOT BE ABLATED SEPARATELY, STATED PLAINLY: `base` and `recency` are
* recoverable only as their blend (`weighted`), because the stored feature
* vector holds AVERAGES (l5_avg/l20_avg), not frequencies-over-the-line. So the
* base/recency split is reported as ONE block. That is a real limit of this
* measurement, not an oversight.
*
* CLAMPED ROWS ARE EXCLUDED from the ablation: at p_over ∈ {0.10, 0.95} the
* inversion is ambiguous, and guessing the pre-clamp value would be fabrication.
* Their count is reported.
*
* SECOND MEASUREMENT — THE MISSING-FEATURE TEST. featureCache computes and
* RETAINS far more than the champion reads (park_*, weather_*, rest_days,
* opportunity_drift, ab_per_game, l5/l10/l20 avgs). If any of those correlates
* with the champion's RESIDUAL (won p_win), that is signal sitting unused on
* disk — a MISSING FEATURE. If none do, that is evidence for AT CEILING with
* respect to everything we currently compute.
*
* SUPABASE_URL=... node scripts/champion-ablation.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const CV_VOLATILE_THRESHOLD = 0.40;
const PROB_FLOOR = 0.10;
const PROB_CEIL = 0.95;
const clamp = (p) => Math.max(PROB_FLOOR, Math.min(PROB_CEIL, p));
function corr(xs, ys) {
const n = xs.length;
if (n < 3) return null;
const mx = xs.reduce((a, b) => a + b, 0) / n;
const my = ys.reduce((a, b) => a + b, 0) / n;
let sxy = 0; let sxx = 0; let syy = 0;
for (let i = 0; i < n; i += 1) {
const dx = xs[i] - mx; const dy = ys[i] - my;
sxy += dx * dy; sxx += dx * dx; syy += dy * dy;
}
if (sxx <= 0 || syy <= 0) return null;
return sxy / Math.sqrt(sxx * syy);
}
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const num = (v) => {
if (v == null || v === '' || typeof v === 'boolean' || typeof v === 'object') return null;
const n = Number(v);
return Number.isFinite(n) ? n : null;
};
/** Deterministic xorshift32 — a measurement that changes between runs is not one. */
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
/** Paired bootstrap on a DIFFERENCE of resolutions (same rows → same resample). */
function bootstrapDiff(rows, keyA, keyB, iters = 3000, seed = 20260803) {
if (rows.length < 30) return null;
const rnd = makeRnd(seed);
const n = rows.length;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const b = [];
for (let i = 0; i < n; i += 1) {
const r = rows[Math.floor(rnd() * n)];
ys.push(r.won); a.push(r[keyA]); b.push(r[keyB]);
}
const ca = corr(a, ys); const cb = corr(b, ys);
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (p) => r4(diffs[Math.floor(p * (diffs.length - 1))]);
const ci = [q(0.025), q(0.975)];
return {
point: r4(corr(rows.map((r) => r[keyA]), rows.map((r) => r.won))
- corr(rows.map((r) => r[keyB]), rows.map((r) => r.won))),
ci95: ci,
ci_excludes_zero: ci[0] > 0 || ci[1] < 0,
};
}
/** Bootstrap CI on a single correlation (for the residual-signal test). */
function bootstrapCorr(rows, key, seed = 20260804, iters = 3000) {
const usable = rows.filter((r) => r[key] != null);
if (usable.length < 40) return { n: usable.length, corr: null, ci95: null, ci_excludes_zero: false };
const rnd = makeRnd(seed);
const n = usable.length;
const vals = [];
for (let it = 0; it < iters; it += 1) {
const xs = []; const ys = [];
for (let i = 0; i < n; i += 1) {
const r = usable[Math.floor(rnd() * n)];
xs.push(r[key]); ys.push(r.residual);
}
const c = corr(xs, ys);
if (c != null) vals.push(c);
}
if (vals.length < 100) return { n, corr: null, ci95: null, ci_excludes_zero: false };
vals.sort((a, b) => a - b);
const q = (p) => r4(vals[Math.floor(p * (vals.length - 1))]);
const ci = [q(0.025), q(0.975)];
return {
n,
corr: r4(corr(usable.map((r) => r[key]), usable.map((r) => r.residual))),
ci95: ci,
ci_excludes_zero: ci[0] > 0 || ci[1] < 0,
};
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
let q = sb.from(table).select(select);
q = apply(q).range(from, from + PAGE - 1);
const { data, error } = await q;
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
const propKey = (r) => `${r.player_key}|${r.stat}|${Number(r.line)}|${String(r.side).toLowerCase()}|${r.game_date}`;
/**
* OUTCOMES COME FROM THE LEDGER, NOT FROM RETENTION.
*
* `model_snapshots.outcome` is NULL on all 22,032 rows — the retention table
* that exists so a different model can be replayed against the same conditions
* stores the features but was never settled. So the labels are joined from
* `ledger_entries` on (player_key, stat, line, side, game_date), which is the
* same identity the ledger's own dedupe constraint uses. Flagged, not fixed —
* this run is read-only.
*/
async function fetchAll(sb) {
const snaps = await page(sb, 'model_snapshots',
'player_key, stat, line, side, game_date, p_win, features, quarantine_reason, captured_at, archetype',
(q) => q.eq('sport', 'mlb').not('p_win', 'is', null).not('features', 'is', null));
const led = await page(sb, 'ledger_entries', 'player_key, stat, line, side, game_date, outcome, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).in('outcome', ['hit', 'miss']));
const outcomeBy = new Map();
for (const r of led) {
if ((r.quarantine_reason || '').startsWith('nontakeable_book')) continue;
outcomeBy.set(propKey(r), r.outcome);
}
return snaps
.map((r) => ({ ...r, outcome: outcomeBy.get(propKey(r)) || null }))
.filter((r) => r.outcome === 'hit' || r.outcome === 'miss');
}
function build(rows) {
// One row per prop — the EARLIEST capture is the lock. Multiple snapshot
// cycles per day would otherwise weight a prop by how often it was re-graded.
const byProp = new Map();
for (const r of rows) {
if ((r.quarantine_reason || '').startsWith('nontakeable_book')) continue;
const k = `${r.player_key}|${r.stat}|${r.line}|${r.side}|${r.game_date}`;
const prev = byProp.get(k);
if (!prev || String(r.captured_at) < String(prev.captured_at)) byProp.set(k, r);
}
const out = [];
let clamped = 0;
for (const r of byProp.values()) {
const f = r.features || {};
const pWin = num(r.p_win);
if (pWin == null) continue;
const under = String(r.side || '').toLowerCase() === 'under';
const pOver = under ? 1 - pWin : pWin;
// Clamped → the pre-clamp value is unrecoverable. Excluded, counted.
if (pOver <= PROB_FLOOR + 1e-9 || pOver >= PROB_CEIL - 1e-9) { clamped += 1; continue; }
const rank = num(f.opp_rank_stat);
const oppAdj = rank == null ? 0 : (rank >= 0.70 ? 0.03 : rank <= 0.30 ? -0.03 : 0);
const ha = num(f.home_away);
const homeAdj = ha === 1 ? 0.015 : ha === 0 ? -0.015 : 0;
const sd = num(f.l10_stddev); const l20 = num(f.l20_avg);
const cv = (sd != null && sd > 0 && l20 != null && l20 > 0) ? sd / l20 : null;
const consistencyFired = cv != null && cv > CV_VOLATILE_THRESHOLD;
const k = consistencyFired ? 0.9 : 1;
// p_over (unclamped) = f(weighted + oppAdj + homeAdj); f linear ⟹ exact removal.
const noOpp = pOver - k * oppAdj;
const noHome = pOver - k * homeAdj;
const noAdj = pOver - k * oppAdj - k * homeAdj; // = f(weighted)
// Removing the consistency pull: invert f on the whole thing.
const noCons = consistencyFired ? (pOver - 0.05) / 0.9 : pOver;
// base+recency block alone, with every adjustment off.
const weighted = consistencyFired ? (noAdj - 0.05) / 0.9 : noAdj;
const flip = (p) => (under ? 1 - clamp(p) : clamp(p));
const won = r.outcome === 'hit' ? 1 : 0;
out.push({
stat: r.stat,
won,
full: pWin,
no_opp: flip(noOpp),
no_home: flip(noHome),
no_consistency: flip(noCons),
no_adjustments: flip(noAdj),
weighted_only: flip(weighted),
residual: won - pWin,
archetype: r.archetype || null,
// Features the champion NEVER reads — the missing-feature candidates.
l5_avg: num(f.l5_avg), l10_avg: num(f.l10_avg), l20_avg: num(f.l20_avg),
ab_per_game: num(f.ab_per_game), recent_ab_per_game: num(f.recent_ab_per_game),
opportunity_drift: num(f.opportunity_drift), rest_days: num(f.rest_days),
game_count_in_7d: num(f.game_count_in_7d),
park_h: num(f.park_h), park_hr: num(f.park_hr), park_r: num(f.park_r),
weather_temp_f: num(f.weather_temp_f), weather_wind_mph: num(f.weather_wind_mph),
weather_precip: num(f.weather_precip),
// Features it DOES read — controls for the same test.
opp_rank_stat: num(f.opp_rank_stat), home_away: num(f.home_away),
l10_stddev: num(f.l10_stddev),
});
}
return { rows: out, clamped };
}
const ABLATIONS = [
['no_opp', 'opponent (opp_rank_stat, ±0.03)'],
['no_home', 'home/away (±0.015)'],
['no_consistency', 'consistency pull (cv>0.40 → toward 0.50)'],
['no_adjustments', 'ALL THREE adjustments (leaves base+recency)'],
];
const UNUSED = ['l5_avg', 'l10_avg', 'l20_avg', 'ab_per_game', 'recent_ab_per_game',
'opportunity_drift', 'rest_days', 'game_count_in_7d', 'park_h', 'park_hr', 'park_r',
'weather_temp_f', 'weather_wind_mph', 'weather_precip'];
const USED = ['opp_rank_stat', 'home_away', 'l10_stddev'];
function ablateStat(rows, label) {
const ys = rows.map((r) => r.won);
const full = r4(corr(rows.map((r) => r.full), ys));
const abl = {};
for (const [key, name] of ABLATIONS) {
const bs = bootstrapDiff(rows, key, 'full');
abl[name] = {
resolution_without: r4(corr(rows.map((r) => r[key]), ys)),
// NEGATIVE delta = removing it HURT = the feature carries signal.
delta_from_removal: bs ? bs.point : null,
ci95: bs ? bs.ci95 : null,
carries_signal: bs ? (bs.ci_excludes_zero && bs.point < 0) : null,
};
}
return { stat: label, n: rows.length, base_rate: r4(ys.reduce((a, b) => a + b, 0) / ys.length), resolution_full: full, ablations: abl };
}
function residualStat(rows, label) {
const scan = (keys, seedBase) => {
const out = {};
keys.forEach((k, i) => {
const res = bootstrapCorr(rows, k, 20260804 + i + seedBase);
out[k] = res;
});
return out;
};
return {
stat: label,
n: rows.length,
unused_features: scan(UNUSED, 0),
used_features_control: scan(USED, 500),
};
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const { rows, clamped } = build(await fetchAll(sb));
const counts = {};
for (const r of rows) counts[r.stat] = (counts[r.stat] || 0) + 1;
const stats = Object.entries(counts).filter(([, n]) => n >= 60).map(([s]) => s)
.sort((a, b) => counts[b] - counts[a]);
const perStat = stats.map((s) => ablateStat(rows.filter((r) => r.stat === s), s));
const residual = stats.map((s) => residualStat(rows.filter((r) => r.stat === s), s));
// ── ARCHETYPE ON TRIAL ────────────────────────────────────────────────
// The champion reads NO archetype feature at all, so it cannot be ablated out
// of it. The fair test is whether archetype explains what the champion GETS
// WRONG: if a given archetype's rows are systematically mispriced, archetype
// carries prop signal the model is missing (wrong IMPLEMENTATION). If every
// archetype's mean residual straddles zero, archetype carries no prop signal.
const archetypeTest = stats.map((st) => {
const rs = rows.filter((r) => r.stat === st && r.archetype);
const groups = {};
for (const r of rs) (groups[r.archetype] = groups[r.archetype] || []).push(r.residual);
const out = {};
for (const [name, vals] of Object.entries(groups)) {
if (vals.length < 40) continue;
const rnd = makeRnd(20260805);
const means = [];
for (let it = 0; it < 3000; it += 1) {
let sum = 0;
for (let i = 0; i < vals.length; i += 1) sum += vals[Math.floor(rnd() * vals.length)];
means.push(sum / vals.length);
}
means.sort((a, b) => a - b);
const ci = [r4(means[Math.floor(0.025 * (means.length - 1))]), r4(means[Math.floor(0.975 * (means.length - 1))])];
out[name] = {
n: vals.length,
mean_residual: r4(vals.reduce((a, b) => a + b, 0) / vals.length),
ci95: ci,
systematically_mispriced: ci[0] > 0 || ci[1] < 0,
};
}
return { stat: st, archetypes: out };
});
// Any unused feature with a CI excluding zero, anywhere → a missing-feature lead.
const leads = [];
for (const rs of residual) {
for (const [k, v] of Object.entries(rs.unused_features)) {
if (v.ci_excludes_zero) leads.push({ stat: rs.stat, feature: k, corr: v.corr, ci95: v.ci95, n: v.n });
}
}
console.log(JSON.stringify({
measurement: 'EXACT ANALYTIC ABLATION of the champion, on the REPAIRED settled set. No refit, no re-fetch, no lookahead.',
limits: {
base_recency_not_separable: 'stored features hold AVERAGES, not frequencies-over-line; reported as one block',
clamped_rows_excluded: clamped,
},
total_rows: rows.length,
per_stat_ablation: perStat,
residual_signal_test: residual,
archetype_test: archetypeTest,
multiple_comparisons_note: 'The residual scan runs 14 unused features x 5 stats = 70 tests at alpha .05, so ~3-4 CI-excludes-zero results are EXPECTED BY CHANCE. Treat a single hit as noise; only a feature repeating across independent stats is evidence.',
missing_feature_leads: leads,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+174
View File
@@ -0,0 +1,174 @@
#!/usr/bin/env node
'use strict';
/**
* PHASES 0-1 — is the champion really worse than a frequency table?
*
* The prior comparison used a leave-one-out baseline that saw the evaluation
* window. This one does not: for every prop, the naive forecast is that player's
* rate of clearing THAT LINE over games strictly BEFORE that date — the same
* temporal discipline the champion is held to. If the champion still loses, the
* defect is real and not an artefact of the peek.
*
* Then the champion's own knobs are ablated. Its core is
*
* p = 0.6 * season_frequency + 0.4 * last5_frequency
*
* plus a +/-0.03 opponent nudge, a +/-0.015 home nudge, and a cv pull. Each is
* tested for whether it COSTS resolution. This is accounting on the champion's
* existing knobs, not a causal claim, so no Bonferroni slot.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Games a player needs before we will read his own rate at all. */
const MIN_PRIOR_GAMES = 10;
async function page(sb, t, sel, orderBy, apply) {
const out = [];
for (let i = 0; ; i += 1000) {
const { data, error } = await apply(sb.from(t).select(sel)).order(orderBy, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break;
out.push(...data);
if (data.length < 1000) break;
}
return out;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) { let s = seed >>> 0; return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; }; }
function resolutionOf(rows, key) {
const base = mean(rows.map((r) => r.won));
let res = 0;
for (let k = 0; k < 10; k += 1) {
const lo = k / 10; const hi = (k + 1) / 10;
const sl = rows.filter((r) => r[key] >= lo && (hi >= 1 ? r[key] <= 1 : r[key] < hi));
if (!sl.length) continue;
res += (sl.length / rows.length) * (mean(sl.map((x) => x.won)) - base) ** 2;
}
return res;
}
/** Paired date-block bootstrap on a resolution difference (a b). */
function dateBlockResCI(rows, a, b, seed) {
const byDate = new Map();
for (const r of rows) { if (!byDate.has(r.date)) byDate.set(r.date, []); byDate.get(r.date).push(r); }
const keys = [...byDate.keys()]; const rnd = makeRnd(seed); const d = [];
for (let it = 0; it < 3000; it += 1) {
const s = [];
for (let i = 0; i < keys.length; i += 1) s.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
d.push(resolutionOf(s, a) - resolutionOf(s, b));
}
d.sort((x, y) => x - y);
return { ci: [r5(d[Math.floor(d.length * 0.025)]), r5(d[Math.floor(d.length * 0.975)])], date_blocks: keys.length };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
// Per-player, date-ordered history. The ONLY source of the naive forecast.
const hist = new Map();
for (const [k, b] of Object.entries(lines)) {
const [date, key] = k.split('|');
if (!hist.has(key)) hist.set(key, []);
hist.get(key).push({ date, b });
}
for (const v of hist.values()) v.sort((x, y) => x.date.localeCompare(y.date));
const out = {};
for (const stat of STATS) {
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused, features', 'id',
(q) => q.eq('sport', 'mlb').eq('stat', stat));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({ propKey: [r.game_date, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const rows = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const isUnder = String(r.side).toLowerCase() === 'under';
const won = (isUnder ? !(v > L) : (v > L)) ? 1 : 0;
// ── THE FAIR COMPETITOR: strictly prior games only. ──
const prior = (hist.get(r.player_key) || []).filter((g) => g.date < r.game_date);
if (prior.length < MIN_PRIOR_GAMES) continue;
const vals = prior.map((g) => knownNumber(FIELD[stat](g.b))).filter((x) => x !== null);
if (vals.length < MIN_PRIOR_GAMES) continue;
const season = vals.filter((x) => x > L).length / vals.length;
const last5 = vals.slice(-5);
const recent = last5.filter((x) => x > L).length / last5.length;
const f = r.features || {};
const homeAdj = f.home_away === 1.0 ? 0.015 : f.home_away === 0.0 ? -0.015 : 0;
const oppR = knownNumber(f.opp_rank_stat);
const oppAdj = oppR === null ? 0 : (oppR >= 0.70 ? 0.03 : (oppR <= 0.30 ? -0.03 : 0));
const flip = (p) => Math.max(0.01, Math.min(0.99, isUnder ? 1 - p : p));
const blend = (w) => flip(0.6 === null ? season : (1 - w) * season + w * recent);
rows.push({
date: r.game_date, won,
champion: knownNumber(r.p_win),
// Reconstructions, all point-in-time.
season_only: flip(season),
w40: flip(0.6 * season + 0.4 * recent), // the current blend
w20: flip(0.8 * season + 0.2 * recent),
w60: flip(0.4 * season + 0.6 * recent),
w40_nudged: flip(Math.max(0.01, Math.min(0.99, 0.6 * season + 0.4 * recent + oppAdj + homeAdj))),
season_nudged: flip(Math.max(0.01, Math.min(0.99, season + oppAdj + homeAdj))),
});
}
if (rows.length < 100) { out[stat] = { n: rows.length, note: 'too few rows with 10+ prior games' }; continue; }
// The REPAIRED champion: full-season window + recency weight 0.20, which is
// exactly what the code change produces.
for (const r of rows) r.repaired = r.w20;
const keys = ['champion', 'repaired', 'season_only', 'w20', 'w40', 'w60', 'w40_nudged', 'season_nudged'];
const res = Object.fromEntries(keys.map((k) => [k, r5(resolutionOf(rows, k))]));
const gap = dateBlockResCI(rows, 'champion', 'season_only', 20260807);
out[stat] = {
n: rows.length,
dates: new Set(rows.map((r) => r.date)).size,
resolution: res,
champion_minus_fair_baseline: r5(res.champion - res.season_only),
ci_champion_minus_fair: gap.ci,
date_blocks: gap.date_blocks,
champion_loses_fairly: res.champion < res.season_only,
best_variant: keys.reduce((a, k) => (res[k] > res[a] ? k : a), keys[0]),
recency_cost: r5(res.w40 - res.season_only),
nudge_cost: r5(res.w40_nudged - res.w40),
REPAIRED_vs_fair: r5(res.repaired - res.season_only),
REPAIRED_ci: dateBlockResCI(rows, 'repaired', 'season_only', 20260807).ci,
REPAIRED_vs_old_champion: r5(res.repaired - res.champion),
REPAIRED_beats_old_ci: dateBlockResCI(rows, 'repaired', 'champion', 20260807).ci,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error('FAILED:', e.message); process.exit(1); });
const r5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
+553
View File
@@ -0,0 +1,553 @@
#!/usr/bin/env node
'use strict';
/**
* tb-solo-and-interactions — PROVE BOTH, with the solo pass as the control.
*
* A feature can carry signal alone, only in combination, or both. Testing only
* interactions misses solo-real features AND cannot tell whether an interaction
* ADDS anything or merely re-encodes its own parts. So the solo result is the
* baseline every interaction has to beat.
*
* ── HOW "ADDS OVER ITS PARTS" IS MEASURED ────────────────────────────────
* Not by comparing two correlations by eye. The interaction's incremental
* signal is the PARTIAL correlation of the interaction term with the counter's
* residual, CONTROLLING FOR both component features:
*
* resid_I = I OLS(I ~ A, B)
* resid_Y = Y OLS(Y ~ A, B)
* incremental r = corr(resid_I, resid_Y)
*
* If the interaction is just barrel-rate wearing a different hat, regressing out
* barrel rate removes it and the incremental r collapses to ~0. That is exactly
* the redundancy the order is guarding against, and it is the difference between
* PASSES-AND-ADDS and PASSES-BUT-REDUNDANT.
*
* ── WHY THE COUNTER'S RESIDUAL IS THE TARGET ─────────────────────────────
* Correlating with the raw outcome rewards a feature for knowing what the
* counter already knows. Only the part the counter MISSES is new information,
* and only new information can improve the product. Both are reported; the
* residual one is the one that decides.
*
* ── THEORY FIRST ─────────────────────────────────────────────────────────
* Every interaction below is declared with a MECHANISM before it is measured.
* No blind pairwise search — with 8 features there are 28 pairs, and at α=.05
* roughly one in twenty returns "significant" from noise alone.
*
* SUPABASE_URL=... node scripts/tb-solo-and-interactions.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const cv = require('../src/services/model/correlateValidator');
const sk = require('../src/services/model/skillProjection');
const reg = require('../src/services/model/featureRegistry');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownRate, knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const GAMES_SO_FAR = Number(process.env.STAGEA_GAMES_SO_FAR || 103);
const STAT = process.env.CLUSTER_STAT || 'total_bases';
/** Restrict every test to ONE archetype — the pooled result can hide an
* archetype-conditional effect entirely (the pitcher strata showed opposite
* signs cancelling to near-zero when pooled). */
const ARCH = process.env.CLUSTER_ARCHETYPE || null;
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const mean = (a) => (a.length ? a.reduce((x, y) => x + y, 0) / a.length : null);
const brier = (ps, ys) => (ps.length ? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
/** OLS residuals of y on the given predictor columns (with intercept). */
function olsResiduals(y, Xcols) {
const n = y.length;
const p = Xcols.length + 1;
const X = [];
for (let i = 0; i < n; i += 1) {
const row = [1];
for (const c of Xcols) row.push(c[i]);
X.push(row);
}
// Normal equations (X'X) b = X'y, solved by Gauss-Jordan. p is 2-4 here.
const XtX = Array.from({ length: p }, () => new Array(p).fill(0));
const Xty = new Array(p).fill(0);
for (let i = 0; i < n; i += 1) {
for (let a = 0; a < p; a += 1) {
Xty[a] += X[i][a] * y[i];
for (let b = 0; b < p; b += 1) XtX[a][b] += X[i][a] * X[i][b];
}
}
const M = XtX.map((row, i) => [...row, Xty[i]]);
for (let col = 0; col < p; col += 1) {
let piv = col;
for (let r = col + 1; r < p; r += 1) if (Math.abs(M[r][col]) > Math.abs(M[piv][col])) piv = r;
if (Math.abs(M[piv][col]) < 1e-12) return null; // singular → cannot control honestly
[M[col], M[piv]] = [M[piv], M[col]];
const d = M[col][col];
for (let k = col; k <= p; k += 1) M[col][k] /= d;
for (let r = 0; r < p; r += 1) {
if (r === col) continue;
const f = M[r][col];
for (let k = col; k <= p; k += 1) M[r][k] -= f * M[col][k];
}
}
const beta = M.map((row) => row[p]);
return y.map((v, i) => v - X[i].reduce((s, xv, j) => s + xv * beta[j], 0));
}
/**
* Partial correlation of a with b, controlling for the columns in ctrl.
*
* COLLINEARITY IS CHECKED FIRST, and this is not pedantry — it caught a real
* error in this very script. The archetype-power proxy was defined as
* `barrel_pct / LEAGUE.barrel_pct`, an exact linear function of barrel_pct, so
* "control for both components" was rank-deficient and the partial correlation
* it produced (-0.132, the only one that looked like an incremental finding) was
* an artifact of a singular design matrix. The Gauss-Jordan pivot test missed it
* because the two columns differ by a scale factor, which keeps the pivot well
* above an absolute epsilon. Scale-free pairwise correlation catches it.
*/
function partialCorr(a, b, ctrl) {
for (let i = 0; i < ctrl.length; i += 1) {
for (let j = i + 1; j < ctrl.length; j += 1) {
const rr = cv.pearson(ctrl[i], ctrl[j]).r;
if (rr !== null && Math.abs(rr) > 0.999) return null; // same variable twice
}
}
const ra = olsResiduals(a, ctrl);
const rb = olsResiduals(b, ctrl);
if (!ra || !rb) return null;
return cv.pearson(ra, rb).r;
}
/** Rows where every named key is known — the honest common sample. */
function completeRows(rows, keys) {
return rows.filter((r) => keys.every((k) => knownNumber(r[k]) !== null));
}
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260804) {
if (rows.length < 30) return null;
const rnd = makeRnd(seed);
const n = rows.length;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const b = [];
for (let i = 0; i < n; i += 1) {
const r = rows[Math.floor(rnd() * n)];
ys.push(r.won); a.push(r[keyA]); b.push(r[keyB]);
}
const ca = cv.pearson(a, ys).r; const cb = cv.pearson(b, ys).r;
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (pp) => r4(diffs[Math.floor(pp * (diffs.length - 1))]);
const ci = [q(0.025), q(0.975)];
return {
point: r4(cv.pearson(rows.map((r) => r[keyA]), rows.map((r) => r.won)).r
- cv.pearson(rows.map((r) => r[keyB]), rows.map((r) => r.won)).r),
ci95: ci, ci_excludes_zero: ci[0] > 0 || ci[1] < 0,
};
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function opposingStarters(dates) {
const m = new Map();
for (const d of dates) {
let games = [];
try { games = await mlb.getScheduleWithPitchers(d); } catch { games = []; }
for (const g of games) {
if (!g.home || !g.away) continue;
if (g.home.probablePitcher) m.set(`${d}|OPP:${g.home.team}`, g.home.probablePitcher.id);
if (g.away.probablePitcher) m.set(`${d}|OPP:${g.away.team}`, g.away.probablePitcher.id);
}
}
return m;
}
async function opponentByPlayerDate(players) {
const map = new Map();
for (const [key, name] of players) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id);
for (const g of log || []) {
if (g && g.date && g.opponent) map.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
}
} catch { /* no log → no pitcher */ }
}
return map;
}
const SOLO = ['batter_barrel_pct', 'batter_hard_hit_pct', 'batter_exit_velo',
'batter_launch_angle', 'batter_k_pct', 'batter_bb_pct',
'pitcher_k_pct', 'pitcher_hard_hit_allowed',
'pitcher_gb_pct', 'pitcher_fb_pct', 'pitcher_breaking_share', 'team_defense'];
/**
* PER-STAT INTERACTION SETS — the total_bases conditioning map RE-WEIGHTED, not
* copied. Reuse speeds the search; it grants nothing. Each stat's features must
* independently earn their place FOR THAT STAT, and the mechanisms genuinely
* differ: barrel rate drives home runs through one channel (does the ball leave)
* and RBI through another (does anyone happen to be on base when it does).
*/
const STAT_INTERACTIONS = {
total_bases: ['launch_x_exit_velo', 'exitvelo_x_pitcher_suppression', 'barrel_x_power_archetype', 'batterK_x_pitcherK', 'launch_x_pitcher_gb', 'barrel_x_breaking_share', 'defense_x_contact', 'defense_x_speed_profile'],
hits: ['launch_x_exit_velo', 'exitvelo_x_pitcher_suppression', 'batterK_x_pitcherK', 'launch_x_pitcher_gb', 'barrel_x_breaking_share', 'defense_x_contact', 'defense_x_speed_profile'],
// HOME RUNS are the purest barrel stat: the ball must be hit hard AND at the
// right angle, and the pitcher must be the kind who allows that combination.
home_runs: ['launch_x_exit_velo', 'barrel_x_power_archetype', 'exitvelo_x_pitcher_suppression'],
// RBI is a POWER x OPPORTUNITY stat — a solo home run drives in one, the same
// swing with two on drives in three. We do not ingest baserunner state, so the
// opportunity half is genuinely missing and that is reported, not papered over.
rbi: ['barrel_x_power_archetype', 'exitvelo_x_pitcher_suppression', 'batterK_x_pitcherK'],
// RUNS scored is ON-BASE x what happens AFTER — mostly teammate-driven, which
// is the least self-contained stat in the cluster.
runs: ['batterK_x_pitcherK', 'exitvelo_x_pitcher_suppression'],
};
/** STEP 2 — theory first. Every interaction declares its mechanism. */
const INTERACTIONS = [
{
key: 'launch_x_exit_velo',
components: ['batter_launch_angle', 'batter_exit_velo'],
mechanism: 'Extra bases need BOTH conditions: hit hard AND hit in the air. A 105-mph ground ball is an out; a 25-degree popup is an out. Neither factor alone predicts bases, which is precisely why each may fail solo and the product may not.',
build: (r) => r.batter_launch_angle * r.batter_exit_velo,
},
{
key: 'exitvelo_x_pitcher_suppression',
components: ['batter_exit_velo', 'pitcher_hard_hit_allowed'],
mechanism: 'A hitter only realises his contact quality against a pitcher who permits contact quality. Elite suppression should attenuate a power bat; a contact-permitting arm should amplify it. The effect is conditional by construction.',
build: (r) => r.batter_exit_velo * r.pitcher_hard_hit_allowed,
},
{
key: 'barrel_x_power_archetype',
components: ['batter_barrel_pct', 'archetype_power'],
mechanism: 'ARCHETYPE-CONDITIONAL. Barrels convert to extra bases for hitters whose lane is power; for a speed/contact profile the same barrel rate is a rarer event on a swing built for something else. This is Discipline 2 stated as a testable interaction. NOTE: it is currently UNTESTABLE — statcast rows carry no archetype label, and the barrel-relative proxy is an exact linear function of barrel_pct, so controlling for both components is rank-deficient. It needs a real archetype classification joined in.',
build: (r) => r.batter_barrel_pct * r.archetype_power,
},
{
key: 'defense_x_contact',
components: ['team_defense', 'batter_hard_hit_pct'],
mechanism: 'DEFENCE. A ball in play becomes a hit or an out partly by who is standing behind the pitcher. This should matter MOST for hitters whose value is contact that stays in the park, and LEAST for power hitters whose barrels clear the defence entirely — so a DEAD result for BOMBER is not a failure, it is the differential the theory predicts.',
build: (r) => r.team_defense * r.batter_hard_hit_pct,
},
{
key: 'defense_x_speed_profile',
components: ['team_defense', 'batter_launch_angle'],
mechanism: 'DEFENCE x BATTED-BALL PROFILE. A low-launch (ground-ball) hitter puts the ball where fielders range; a high-launch hitter does not. Launch angle stands in for the profile, so defence should condition the ground-ball hitter far more.',
build: (r) => r.team_defense * r.batter_launch_angle,
},
{
key: 'launch_x_pitcher_gb',
components: ['batter_launch_angle', 'pitcher_gb_pct'],
mechanism: 'PITCHER BATTED-BALL TYPE. A ground-ball arm takes the air away, and a hitter whose value lives in the air needs the air. An air hitter against a sinkerballer and a ground-ball hitter against a fly-ball arm are both mismatches that neither factor states alone.',
build: (r) => r.batter_launch_angle * r.pitcher_gb_pct,
},
{
key: 'barrel_x_breaking_share',
components: ['batter_barrel_pct', 'pitcher_breaking_share'],
mechanism: 'ARSENAL MATCHUP. Barrel rate is far more a fastball skill than a breaking-ball skill, so a power bat facing a breaking-heavy arm should convert less of it. The pitch mix is already ingested, so this costs nothing to test.',
build: (r) => r.batter_barrel_pct * r.pitcher_breaking_share,
},
{
key: 'batterK_x_pitcherK',
components: ['batter_k_pct', 'pitcher_k_pct'],
mechanism: 'Strikeout risk compounds multiplicatively (log5 is exactly this shape). A high-K bat against a high-K arm loses plate appearances to strikeouts, and a PA lost is a base opportunity that never happens — so it suppresses total bases through OPPORTUNITY, not contact quality.',
build: (r) => r.batter_k_pct * r.pitcher_k_pct,
},
];
/**
* Breaking-ball share of a pitcher's mix, from `pitch_mix` already ingested.
* Sliders/curves/sweepers/cutters vs fastballs — absent mix -> null, never 0.
*/
function breakingShare(mix) {
if (!mix || typeof mix !== 'object') return null;
const rows = Array.isArray(mix) ? mix : Object.values(mix);
let breaking = 0; let total = 0;
for (const p of rows) {
if (!p) continue;
const type = String(p.type || p.pitch_type || '').toUpperCase();
const usage = knownNumber(p.usage_pct ?? p.usage ?? p.pct);
if (!type || usage === null || usage < 0) continue;
total += usage;
if (['SL', 'CU', 'KC', 'ST', 'SV', 'FC', 'SC'].includes(type)) breaking += usage;
}
if (total <= 0) return null;
return breaking / total;
}
/** Latest settled game date in the pull — used to detect that the profile
* freeze now sits AFTER the data, i.e. no clean out-of-sample window exists. */
function clean0Max(rows) {
return (rows || []).reduce((mx, r) => (String(r.game_date) > mx ? String(r.game_date) : mx), '');
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10);
const batters = new Map(); const pitchersById = new Map();
for (const r of statcast) {
// pitch_mix is NOT part of fromStatcastRow's output (it maps pct/raw fields
// only), so it must be attached explicitly — without it the arsenal category
// silently measures nothing and reports n=0.
if (r.role === 'pitcher' && r.source_id != null) {
pitchersById.set(Number(r.source_id), { ...sk.fromStatcastRow(r), pitch_mix: r.pitch_mix });
}
if (r.player_key && r.role === 'batter') {
const prev = batters.get(r.player_key);
if (!prev || Number(r.sample_pa || 0) > Number(prev.rawPa || 0)) {
batters.set(r.player_key, Object.assign(sk.fromStatcastRow(r), { rawPa: Number(r.sample_pa || 0) }));
}
}
}
// REAL ARCHETYPE LABELS. The barrel-relative proxy was a clipped monotone
// transform of barrel_pct, so `barrel x proxy` measured NONLINEARITY IN BARREL,
// not an archetype interaction — it could never have tested Discipline 2.
// model_snapshots carries the actual classification per prop, so the
// conditioning variable is now a genuine BOMBER indicator, which is
// categorical and therefore not a transform of barrel at all.
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype',
(q) => q.eq('sport', 'mlb').eq('stat', STAT).not('archetype', 'is', null));
const archetypeBy = new Map();
for (const r of snaps) if (r.player_key && r.game_date) archetypeBy.set(`${r.player_key}|${r.game_date}`, r.archetype);
const led = await page(sb, 'ledger_entries',
'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', STAT)
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
// POINT-IN-TIME IS NO LONGER AVAILABLE FROM THIS TABLE.
//
// `statcast_aggregates` is upserted in place and keeps one as-of date. The
// first skill backtest was honest only by accident: the nightly refresh was
// unreachable code, so the table sat frozen at 2026-07-21 — BEFORE the settled
// window. Repairing that cron (correct for production) refreshed it to today,
// and every prior version is gone.
//
// So scoring a 2026-07-25 game now uses a season aggregate that CONTAINS that
// game. `statcast_history` (added this session) fixes it going forward; it has
// one day of data, which is not yet a window. Until it fills, results here are
// DIRECTIONAL AND CONTAMINATED, labelled as such, and are NOT gate verdicts.
const contaminated = String(freezeDate) >= String(clean0Max(led));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')
&& (contaminated ? true : String(r.game_date) > freezeDate));
const dates = [...new Set(clean.map((r) => r.game_date))].sort();
const starters = await opposingStarters(dates);
const players = new Map();
for (const r of clean) if (!players.has(r.player_key)) players.set(r.player_key, r.player_name);
const oppByPlayerDate = await opponentByPlayerDate(players);
if (process.env.TB_DEBUG === '1') {
console.error(`[debug] statcast rows=${statcast.length} freeze=${freezeDate} batters=${batters.size} pitchers=${pitchersById.size}`);
console.error(`[debug] ledger tb rows=${led.length} clean(after freeze)=${clean.length}`);
const sampleKeys = clean.slice(0, 5).map((r) => r.player_key);
console.error(`[debug] sample ledger player_keys=${JSON.stringify(sampleKeys)}`);
console.error(`[debug] sample statcast keys=${JSON.stringify([...batters.keys()].slice(0, 5))}`);
console.error(`[debug] matches in sample=${sampleKeys.filter((k) => batters.has(k)).length}/5`);
}
// TEAM DEFENCE — the newly-ingested Statcast OAA, per team, dated.
const defRows = await page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb'));
const defByTeam = new Map();
for (const d of defRows) {
const prev = defByTeam.get(d.team);
if (!prev || String(d.as_of_date) > String(prev.as_of_date)) defByTeam.set(d.team, d);
}
const allowed = reg.candidateFeaturesForStat('mlb', STAT);
const rowsAll = [];
const rows = rowsAll;
for (const r of clean) {
const bat = batters.get(r.player_key);
if (!bat) continue;
const faced = oppByPlayerDate.get(`${r.player_key}|${r.game_date}`) || null;
const pit = faced ? pitchersById.get(Number(starters.get(`${r.game_date}|OPP:${faced}`))) || null : null;
const paRate = bat.rawPa > 0 ? Math.min(5.2, Math.max(2.0, bat.rawPa / GAMES_SO_FAR)) : null;
const under = String(r.side).toLowerCase() === 'under';
const won = r.outcome === 'hit' ? 1 : 0;
const champ = Number(r.p_win);
const proj = sk.projectSkill({
batter: bat, pitcher: pit, park: 1, archetype: null,
statType: STAT, line: Number(r.line), expectedPa: paRate, allowed,
});
// The REAL archetype for this prop — a 0/1 power indicator, categorical and
// independent of barrel_pct by construction.
const arch = archetypeBy.get(`${r.player_key}|${r.game_date}`) || null;
const archetypePower = arch == null ? null : (String(arch).toUpperCase() === 'BOMBER' ? 1 : 0);
rows.push({
won, champ, residual: won - champ,
skill: proj ? (under ? 1 - proj.p_over_line : proj.p_over_line) : null,
had_pitcher: !!pit,
batter_barrel_pct: knownRate(bat.barrel_pct),
batter_hard_hit_pct: knownRate(bat.hard_hit_pct),
batter_exit_velo: knownRate(bat.avg_exit_velo),
batter_launch_angle: knownRate(bat.avg_launch_angle),
batter_k_pct: knownRate(bat.k_pct),
batter_bb_pct: knownRate(bat.bb_pct),
pitcher_k_pct: pit ? knownRate(pit.k_pct) : null,
pitcher_hard_hit_allowed: pit ? knownRate(pit.hard_hit_pct) : null,
archetype_power: archetypePower,
archetype: arch,
// ── CONDITIONING CATEGORIES (this order) ──────────────────────────
// PITCHER BATTED-BALL TYPE: a ground-ball arm suppresses air contact, so
// it should matter differently to a hitter whose value is in the air.
pitcher_gb_pct: pit ? knownRate(pit.gb_pct) : null,
pitcher_fb_pct: pit ? knownRate(pit.fb_pct) : null,
// ARSENAL: breaking-ball share, from the pitch mix already ingested. A
// power hitter's barrel rate is a fastball skill far more than a
// breaking-ball skill, so the mix should condition it.
pitcher_breaking_share: pit && pit.pitch_mix ? breakingShare(pit.pitch_mix) : null,
// DEFENSE: NOT DERIVABLE from what we ingest — see the report. Recorded as
// null rather than proxied by something that is really pitching quality.
// DEFENCE behind the pitcher he faces. knownRate: an unmeasured team is
// ABSENT, never league-average — OAA 0 is a real "exactly average" reading
// and the two must stay distinguishable.
team_defense: (() => {
if (!faced) return null;
// team_defense keys on Savant's display name (a nickname, "Cubs"),
// while the game log gives the full name ("Chicago Cubs"). Try both.
const nick = String(faced).split(' ').pop();
const d = defByTeam.get(faced) || defByTeam.get(nick);
return d ? knownNumber(d.oaa_sum) : null;
})(),
});
}
// ARCHETYPE RESTRICTION — applied AFTER building rows so coverage is visible.
const archRows = ARCH ? rowsAll.filter((r) => String(r.archetype || '').toUpperCase() === ARCH) : rowsAll;
rows.length = 0; rows.push(...archRows);
// ── CUMULATIVE BONFERRONI ─────────────────────────────────────────────
// The denominator is every DISTINCT hypothesis this programme has tested,
// not just this run's. Correcting by 8 in a session that tries 8, forever,
// while the programme as a whole has tried sixty, is how a noise result
// eventually gets recorded as PROVEN with a p-value to point at.
const tl = require('../src/services/model/testLedger');
const store = tl.supabaseStore(sb);
const chosenKeys = new Set(STAT_INTERACTIONS[STAT] || []);
const entries = [
...SOLO.map((f) => ({ sport: 'mlb', stat: STAT, archetype: ARCH, interaction: `solo:${f}`, target: 'counter_residual' })),
...INTERACTIONS.filter((x) => chosenKeys.has(x.key))
.map((x) => ({ sport: 'mlb', stat: STAT, archetype: ARCH, interaction: x.key, target: 'counter_residual' })),
];
const mc = await tl.recordAndCount(store, entries);
const TESTS = mc.cumulative_tests;
// ── STEP 1 — SOLO PASS (the control) ────────────────────────────────────
const solo = {};
for (const f of SOLO) {
const rs = completeRows(rows, [f]);
solo[f] = {
n: rs.length,
vs_outcome: cv.validateFactor(rs.map((r) => r[f]), rs.map((r) => r.won), TESTS),
vs_counter_residual: cv.validateFactor(rs.map((r) => r[f]), rs.map((r) => r.residual), TESTS),
};
}
// ── STEP 3 — INTERACTIONS, each against its own solo baseline ───────────
const interactions = {};
const chosen = new Set(STAT_INTERACTIONS[STAT] || []);
for (const ix of INTERACTIONS.filter((x) => chosen.has(x.key))) {
const keys = [...ix.components];
const rs = completeRows(rows, keys);
if (rs.length < 30) { interactions[ix.key] = { mechanism: ix.mechanism, n: rs.length, verdict: 'UNTESTABLE — no common sample' }; continue; }
const I = rs.map(ix.build);
const Y = rs.map((r) => r.residual);
const ctrl = keys.map((k) => rs.map((r) => r[k]));
const gate = cv.validateFactor(I, Y, TESTS);
const incremental = partialCorr(I, Y, ctrl);
// The best solo |r| among its own components, on the SAME rows.
const componentSolo = keys.map((k) => ({
feature: k, r: r4(cv.pearson(rs.map((r) => r[k]), Y).r),
}));
const bestComponent = Math.max(...componentSolo.map((c) => Math.abs(c.r ?? 0)));
let verdict;
if (incremental === null) verdict = 'UNTESTABLE — controls are collinear';
else if (gate.validated && Math.abs(incremental) >= cv.VALIDATION_REQUIREMENTS.min_pearson_r) verdict = 'PASSES-AND-ADDS';
else if (gate.validated) verdict = 'PASSES-BUT-REDUNDANT';
else if (rs.length < cv.VALIDATION_REQUIREMENTS.min_historical_instances) verdict = 'UNDERPOWERED — n below the gate';
else verdict = 'FAILS';
interactions[ix.key] = {
mechanism: ix.mechanism,
components: keys,
n: rs.length,
raw_r_vs_residual: gate.pearson_r,
gate: { validated: gate.validated, reason: gate.reason, p_value: gate.p_value, corrected_alpha: gate.corrected_alpha, underpowered: !!gate.underpowered },
component_solo_r_same_rows: componentSolo,
best_component_abs_r: r4(bestComponent),
INCREMENTAL_partial_r: r4(incremental),
adds_over_components: incremental !== null && Math.abs(incremental) > bestComponent,
verdict,
};
}
// ── STEP 4 — COMBINED vs COUNTER (valid at this n; the gate is not) ─────
const h2h = rows.filter((r) => r.skill != null);
const ys = h2h.map((r) => r.won);
const bs = bootstrapDiff(h2h, 'skill', 'champ');
console.log(JSON.stringify({
stat: STAT,
archetype_restriction: ARCH || 'none (pooled)',
VALIDITY: contaminated
? 'CONTAMINATED / DIRECTIONAL ONLY — statcast_aggregates now carries a single as-of date (' + freezeDate + ') that is AFTER the settled games, so season profiles contain the games being predicted. These are NOT gate verdicts. statcast_history (new) makes point-in-time possible from tomorrow.'
: `CLEAN out-of-sample: profiles frozen ${freezeDate}; only game_date > ${freezeDate} scored`,
contaminated,
rows_scored: rows.length,
gate_spec: cv.VALIDATION_REQUIREMENTS,
bonferroni_tests: TESTS,
multiple_comparisons: { ...mc, note: 'denominator is DISTINCT hypotheses across the programme lifetime, not this session' },
n_gap_note: `the gate needs ${cv.VALIDATION_REQUIREMENTS.min_historical_instances} rows; this run has ${rows.length}`,
archetype_coverage: {
labelled: rows.filter((r) => r.archetype).length,
bomber: rows.filter((r) => r.archetype_power === 1).length,
other: rows.filter((r) => r.archetype_power === 0).length,
},
step1_solo_baseline: solo,
step3_interactions: interactions,
step4_combined_vs_counter: {
n: h2h.length,
pitcher_coverage: r4(mean(h2h.map((r) => (r.had_pitcher ? 1 : 0)))),
base_rate: r4(mean(ys)),
resolution: { skill_tb: r4(cv.pearson(h2h.map((r) => r.skill), ys).r), counter: r4(cv.pearson(h2h.map((r) => r.champ), ys).r) },
brier: { skill_tb: r4(brier(h2h.map((r) => r.skill), ys)), counter: r4(brier(h2h.map((r) => r.champ), ys)) },
delta: bs,
verdict: !bs ? 'N-BLOCKED'
: (bs.ci_excludes_zero && bs.point > 0) ? 'SKILL TB BEATS THE COUNTER'
: (bs.ci_excludes_zero && bs.point < 0) ? 'LOSES to the counter'
: 'INCONCLUSIVE',
},
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+210
View File
@@ -0,0 +1,210 @@
#!/usr/bin/env node
'use strict';
/**
* THE COLLAPSED SEQUENCE EDGE — Link 1 x pen-season-quality, on later at-bats.
*
* Link 3 is correctly skipped: reliever IDENTITY did not prove and is genuine
* baseball unpredictability. But Link 2's QUALITY grain DID prove, so pen quality
* here is a measured predictor rather than a fallback.
*
* ── THE MECHANICAL CEILING, MEASURED FIRST ───────────────────────────────
* A hitter's third or fourth plate appearance is ALREADY against the bullpen
* 70-73% of the time even when the starter is projected to go deep. An elevated
* early-exit flag lifts that to only 77-83%. So Link 1 buys roughly TEN POINTS
* of extra pen exposure, not a switch from starter to pen — and any adjustment
* built on it is bounded at about a tenth of the starter-versus-pen quality gap.
* That ceiling is a property of baseball, not of the model, and it is the reason
* the deltas below are small before anything is even fitted.
*
* ── WHAT IS ADJUSTED, AND WHAT IS REFUSED ────────────────────────────────
* The order specifies pen-quality x pen-ARCHETYPE x hitter-APPROACH. Two of
* those three cannot be used honestly:
*
* pen archetype did NOT prove (0.5669 vs a 0.5309 modal baseline, corrected
* interval spanning zero). Building it into the adjustment
* would be chaining on an unproven link.
* hitter approach "fastball-hunter" / "finesse-vulnerable" identities do not
* exist in this registry. MLB batter archetypes are BOMBER /
* GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID /
* CATALYST. Inventing an identity to condition on would be
* fabricating the very thing the gate exists to catch.
*
* So the adjustment uses the PROVEN component alone, and a hitter split derived
* from the sequence data itself (power vs contact by home-run rate) is tested as
* a SEPARATE gated addition rather than assumed into the main effect.
*
* node scripts/collapsed-sequence-edge.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const fg = require('../src/services/model/factorGate');
const tl = require('../src/services/model/testLedger');
const pq = require('../src/services/model/penQuality');
const { createClient } = require('@supabase/supabase-js');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const HIT = new Set(['single', 'double', 'triple', 'home_run']);
const PA = new Set(['single', 'double', 'triple', 'home_run', 'field_out', 'strikeout',
'grounded_into_double_play', 'force_out', 'field_error', 'fielders_choice',
'fielders_choice_out', 'double_play', 'sac_fly', 'pop_out', 'line_out', 'fly_out',
'strikeout_double_play']);
const MIN_ARM_PA = 40;
const MIN_PRIOR_GAMES = 5;
const MIN_HITTER_PA = 60;
const MIN_PRIOR_STARTS = 3;
const EARLY_FLAG_BF = 22;
const LEAGUE_BF = 21.56;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
function build() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const arm = new Map();
const bat = new Map(); // hitter -> { n, h, hr }
const penHist = new Map();
const startHist = new Map();
const rows = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const team = g[side].abbr || g[side].team;
const st = (g[side].arms || []).find((a) => a.started);
if (!team || !st) continue;
const half = side === 'home' ? 'top' : 'bottom';
const pas = g.pas.filter((p) => p.half === half && PA.has(p.event));
const ps = startHist.get(st.id) || [];
let predBf = null;
if (ps.length >= MIN_PRIOR_STARTS) {
const w = ps.length / (ps.length + 5);
predBf = w * mean(ps) + (1 - w) * LEAGUE_BF;
}
const hist = penHist.get(team) || [];
const pen = pq.projectPen(hist.map((q) => ({ quality: q })));
const seen = new Map();
for (const p of pas) {
const k = p.batter;
seen.set(k, (seen.get(k) || 0) + 1);
const paNum = seen.get(k);
const b = bat.get(k);
// knownRate abstain: no readable hitter, starter or pen -> no row at all.
if (paNum < 3 || predBf === null || !pen || !b || b.n < MIN_HITTER_PA) continue;
rows.push({
gamePk: g.gamePk,
cluster: g.gamePk,
batter: k,
paNum,
early: predBf <= EARLY_FLAG_BF,
pen_quality: pen.quality,
hitter_base: b.h / b.n,
hitter_hr_rate: b.hr / b.n,
won: HIT.has(p.event) ? 1 : 0,
});
}
const faced = [];
for (const p of pas.filter((x) => x.pitcher !== st.id)) {
const h = arm.get(p.pitcher);
if (h && h.n >= MIN_ARM_PA) faced.push(h.h / h.n);
}
if (faced.length) penHist.set(team, hist.concat([mean(faced)]));
if (st.bf != null) startHist.set(st.id, ps.concat([st.bf]));
for (const p of pas) {
const c = arm.get(p.pitcher) || { n: 0, h: 0, k: 0 };
c.n += 1; c.h += HIT.has(p.event) ? 1 : 0; c.k += p.event === 'strikeout' ? 1 : 0;
arm.set(p.pitcher, c);
}
for (const p of pas) {
const c = bat.get(p.batter) || { n: 0, h: 0, hr: 0 };
c.n += 1; c.h += HIT.has(p.event) ? 1 : 0; c.hr += p.event === 'home_run' ? 1 : 0;
bat.set(p.batter, c);
}
}
}
return rows;
}
/** The adjustment: the hitter's own rate, shifted by the PROVEN pen signal. */
const adjust = (r) => {
const shift = pq.hitRateShift(r.pen_quality);
if (shift === null) return null;
return Math.max(0.01, Math.min(0.99, r.hitter_base + shift));
};
(async () => {
const all = build();
const qs = all.map((r) => r.pen_quality).sort((a, b) => a - b);
const weakCut = qs[Math.floor(qs.length * 2 / 3)];
const strongCut = qs[Math.floor(qs.length / 3)];
const subsets = {
// The order's concentrated subset.
concentrated_early_x_weak_pen: all.filter((r) => r.early && r.pen_quality >= weakCut),
// The mirror, where the descriptive pass suggested the larger movement.
mirror_early_x_strong_pen: all.filter((r) => r.early && r.pen_quality <= strongCut),
// Every later at-bat with an early-exit flag, both directions of pen quality.
all_early_exit_later_abs: all.filter((r) => r.early),
pooled_all_later_abs: all,
};
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), Object.keys(subsets).map((k) => ({
sport: 'mlb', stat: 'hits', archetype: null,
interaction: `collapsed_sequence:${k}`, target: 'later_ab_outcome',
})));
cumulative = mc.cumulative_tests;
} catch { /* offline */ }
const gate = (rs, label) => fg.adjudicate(
rs.map((r) => ({ cluster: r.cluster, baseline: r.hitter_base, conditioned: adjust(r), won: r.won }))
.filter((r) => r.conditioned !== null),
{ factor: label, stat: 'hits', cumulativeTests: cumulative },
);
const results = {};
for (const [k, rs] of Object.entries(subsets)) results[k] = gate(rs, k);
// Hitter split as a SEPARATE gated addition — never assumed into the main effect.
const conc = subsets.concentrated_early_x_weak_pen;
const hrs = conc.map((r) => r.hitter_hr_rate).sort((a, b) => a - b);
const hrCut = hrs[Math.floor(hrs.length / 2)];
const bySplit = {
power_hitters: gate(conc.filter((r) => r.hitter_hr_rate >= hrCut), 'concentrated_power'),
contact_hitters: gate(conc.filter((r) => r.hitter_hr_rate < hrCut), 'concentrated_contact'),
};
console.log(JSON.stringify({
later_at_bats_readable: all.length,
cumulative_tests: cumulative,
subset_sizes: Object.fromEntries(Object.entries(subsets).map(([k, v]) => [k, v.length])),
gate: Object.fromEntries(Object.entries(results).map(([k, v]) => [k, {
n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
verdict: v.verdict,
}])),
hitter_split_separate_gate: Object.fromEntries(Object.entries(bySplit).map(([k, v]) => [k, {
n: v.movement.n, brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null, verdict: v.verdict,
}])),
refused: {
pen_archetype: 'did not prove at the corrected bar — excluded from the adjustment',
hitter_approach_identity: 'SPRAY / fastball-hunter identities do not exist in this registry',
link3_per_reliever: 'SKIPPED — reliever identity is genuine baseball unpredictability',
},
}, null, 2));
process.exit(0);
})();
+198
View File
@@ -0,0 +1,198 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 1 — derive a COHERENT LODO test, blind to reversals.
*
* ── THE DEFECT BEING FIXED ───────────────────────────────────────────────
* The gate at 1f40014 paired a 1-SE per-drop informativeness bar with a
* zero-reversal decision rule. Those two are incoherent. At exactly 1 SE, a
* genuinely STABLE stat's drop reverses with probability Phi(-1) = 0.159, so on
* four informative drops the chance of at least one reversal is
* 1 - 0.841^4 = 0.50. The rule failed stable stats half the time by construction.
*
* And n* was pooled across four stats whose signed effects differ several-fold,
* so "informative" meant different things for different stats while being
* treated as one number.
*
* ── THE FIX ──────────────────────────────────────────────────────────────
* The two halves have to be chosen together:
*
* informative bar n*_k = k^2 * (sigma_row / |g|)^2 PER STAT
* decision rule FAIL iff reversals > c, where under stability
* R ~ Binomial(D, Phi(-k)) and c is the smallest cutoff
* with P(R > c) <= 0.05
*
* `g` is the mean SIGNED per-row improvement — the quantity whose sign a
* reversal flips. Not a mean-absolute, and not pooled: a reversal is a claim
* about THIS stat's effect changing sign.
*
* THIS SCRIPT PRINTS NO REVERSAL AND NO VERDICT. It is blind by construction and
* must run, and its constants be committed, before any stat is re-read.
*
* SUPABASE_URL=... node scripts/derive-lodo-test.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** |g| must clear this many SE at the stat's full n or there is no effect to test. */
const EFFECT_Z = 1.96;
/** Target false-positive rate for the whole per-stat test. */
const TARGET_FP = 0.05;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Standard normal CDF (AbramowitzStegun 7.1.26 via erf). */
function normCdf(z) {
const t = 1 / (1 + 0.2316419 * Math.abs(z));
const d = 0.3989422804014327 * Math.exp(-z * z / 2);
const p = d * t * (0.319381530 + t * (-0.356563782 + t * (1.781477937 + t * (-1.821255978 + t * 1.330274429))));
return z >= 0 ? 1 - p : p;
}
const binomPmf = (n, k, p) => {
let logC = 0;
for (let i = 0; i < k; i += 1) logC += Math.log(n - i) - Math.log(i + 1);
return Math.exp(logC + k * Math.log(p) + (n - k) * Math.log(1 - p));
};
/** P(R > c) for R ~ Binomial(n, p). */
const binomTail = (n, c, p) => {
let s = 0;
for (let k = c + 1; k <= n; k += 1) s += binomPmf(n, k, p);
return s;
};
/** Smallest cutoff c with P(R > c) <= target. */
function cutoffFor(D, p, target) {
for (let c = 0; c <= D; c += 1) if (binomTail(D, c, p) <= target) return { cutoff: c, fp: binomTail(D, c, p) };
return { cutoff: D, fp: 0 };
}
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
const perStat = {};
const dateSizes = {};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b));
if (v === null) continue;
const over = v > L;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
const map = cal.fitIsotonic(rows.map((r) => ({ p: r.p, won: r.won })));
if (!map) { perStat[stat] = { n: rows.length, fittable: false }; continue; }
const d = [];
for (const r of rows) {
const pc = cal.applyIsotonic(map, r.p);
if (knownNumber(pc) === null) continue;
d.push((pc - r.won) ** 2 - (r.p - r.won) ** 2);
}
const g = mean(d);
const sigma = Math.sqrt(d.reduce((s, x) => s + (x - g) ** 2, 0) / (d.length - 1));
const seFull = sigma / Math.sqrt(d.length);
// Date sizes are sample STRUCTURE, not outcomes — safe to read here.
const sizes = new Map();
for (const r of rows) sizes.set(r.date, (sizes.get(r.date) || 0) + 1);
dateSizes[stat] = [...sizes.values()].sort((a, b) => b - a);
perStat[stat] = {
n: d.length,
g_signed: round5(g),
sigma_row: round5(sigma),
se_full: round5(seFull),
effect_z_at_full_n: round3(Math.abs(g) / seFull),
improves: g < 0,
no_effect: Math.abs(g) / seFull < EFFECT_Z,
fittable: true,
};
}
// ── Choose k jointly. Blind: uses only (g, sigma) and date SIZES. ──
const kTable = [];
for (const k of [1.0, 1.25, 1.5, 1.75, 2.0]) {
const pNoise = normCdf(-k);
const row = { k, per_drop_noise_prob: round4(pNoise), stats: {} };
for (const stat of STATS) {
const ps = perStat[stat];
if (!ps || !ps.fittable) continue;
const nStar = Math.ceil(k * k * (ps.sigma_row / Math.abs(ps.g_signed)) ** 2);
const D = (dateSizes[stat] || []).filter((n) => n >= nStar).length;
const { cutoff, fp } = D > 0 ? cutoffFor(D, pNoise, TARGET_FP) : { cutoff: null, fp: null };
// FN at a stated alternative: date-to-date SD of the effect equals |g|.
const pAlt = D > 0 ? normCdf(-Math.abs(ps.g_signed) / Math.sqrt(ps.g_signed ** 2 + (ps.sigma_row ** 2) / nStar)) : null;
const fn = D > 0 && cutoff !== null ? 1 - binomTail(D, cutoff, pAlt) : null;
row.stats[stat] = {
n_star: nStar, informative_drops: D, cutoff, fp: fp === null ? null : round4(fp),
fn_at_tau_equals_g: fn === null ? null : round4(fn),
};
}
kTable.push(row);
}
console.log(JSON.stringify({
phase: 'PHASE 1 — coherent LODO test derivation, BLIND',
defect_being_fixed: 'a 1-SE informative bar with a zero-reversal rule: P(>=1 reversal | stable, 4 drops) = 0.50',
per_stat_effect: perStat,
date_sizes: dateSizes,
k_selection_table: kTable,
target_fp: TARGET_FP,
blind: 'no reversal, no verdict, no reversing date referenced anywhere in this output',
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round3 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 1000) / 1000);
+139
View File
@@ -0,0 +1,139 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 0 — derive the LODO held-row threshold from POWER, blind to outcomes.
*
* ESTIMAND: "does dropping date D reverse the SIGN of the out-of-sample Brier
* improvement on D's held-out rows?"
*
* A reversal is only informative if a single date's Brier delta is
* distinguishable from zero at that row count. Below that, a reversal is a coin
* flip wearing a decimal point — which is exactly the ambiguity that made the
* previous verdict depend on an operator-chosen number.
*
* ── THE DERIVATION ───────────────────────────────────────────────────────
* The per-row Brier difference is
*
* d_i = (pc_i - y_i)^2 - (p_i - y_i)^2
*
* and a date's Brier delta is the MEAN of d over that date's rows. So
*
* SE(n) = SD(d) / sqrt(n)
*
* and the smallest n at which a typical effect clears one standard error is
*
* n* = ( SD(d) / |effect| )^2
*
* SD(d) and |effect| are pooled ACROSS ALL FOUR STATS deliberately: a per-stat
* figure would let the threshold be shaped by the stat whose verdict it decides.
*
* THIS SCRIPT PRINTS NO STAT VERDICT AND NO DATE. It is blind by construction,
* and it must be run and its output committed BEFORE any stat is re-read.
*
* SUPABASE_URL=... node scripts/derive-lodo-threshold.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
// Pooled per-row Brier differences, across all four stats.
const diffs = [];
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b));
if (v === null) continue;
const over = v > L;
rows.push({ p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
const map = cal.fitIsotonic(rows.map((r) => ({ p: r.p, won: r.won })));
if (!map) continue;
for (const r of rows) {
const pc = cal.applyIsotonic(map, r.p);
if (knownNumber(pc) === null) continue;
diffs.push((pc - r.won) ** 2 - (r.p - r.won) ** 2);
}
}
const m = mean(diffs);
const sd = Math.sqrt(diffs.reduce((s, d) => s + (d - m) ** 2, 0) / (diffs.length - 1));
const effect = Math.abs(m);
const nStar = Math.ceil((sd / effect) ** 2);
const curve = [10, 20, 25, 30, 50, 75, 100, 150, 200, 300, 500].map((n) => ({
n,
se: round5(sd / Math.sqrt(n)),
effect_over_se: round3(effect / (sd / Math.sqrt(n))),
informative: effect >= sd / Math.sqrt(n),
}));
console.log(JSON.stringify({
phase: 'PHASE 0 — power derivation, blind to outcomes',
pooled_rows: diffs.length,
per_row_brier_diff_sd: round5(sd),
pooled_effect_abs_mean: round5(effect),
n_star: nStar,
rule: 'n* = (SD(d) / |effect|)^2 — the smallest held-row count at which a typical Brier delta clears one standard error',
se_vs_n: curve,
blind: 'no stat verdict, no date, and no reversal is referenced anywhere in this output',
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round5 = (v) => Math.round(v * 100000) / 100000;
const round3 = (v) => Math.round(v * 1000) / 1000;
+205
View File
@@ -0,0 +1,205 @@
#!/usr/bin/env node
'use strict';
/**
* Phases 0, 2, 3 and 4 — replay the factor wiring on settled hits rows.
*
* Transmission is proved MECHANICALLY before any resolution number is quoted,
* because "resolution went up" is exactly what a subtle bug also prints.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const hf = require('../src/services/model/hitsFactors');
const lp = require('../src/services/model/lowParamCalibrator');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const SEQ = path.join(process.cwd(), '.seq-cache', 'sequences.json');
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f, orderBy = 'id') {
const o = [];
for (let i = 0; ; i += 1000) {
const { data, error } = await f(sb.from(t).select(s)).order(orderBy, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break; o.push(...data); if (data.length < 1000) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) { let s = seed >>> 0; return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; }; }
function decompose(rows, bins = 10) {
const base = mean(rows.map((r) => r.won));
const unc = base * (1 - base);
let rel = 0; let res = 0;
for (let k = 0; k < bins; k += 1) {
const lo = k / bins; const hi = (k + 1) / bins;
const sl = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (!sl.length) continue;
const w = sl.length / rows.length;
rel += w * (mean(sl.map((r) => r.p)) - mean(sl.map((r) => r.won))) ** 2;
res += w * (mean(sl.map((r) => r.won)) - base) ** 2;
}
return { base_rate: r4(base), reliability: r5(rel), resolution: r5(res), uncertainty: r5(unc), share: r4(res / unc) };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
// Factor inputs.
const [spray, defense, platoon, statcast] = await Promise.all([
page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb'), 'player_key'),
page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb'), 'team'),
page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb'), 'player_key'),
page(sb, 'statcast_aggregates', 'player_key, role, bats, throws, hard_hit_pct', (q) => q.eq('sport', 'mlb'), 'player_key'),
]);
const latest = (rows, k) => { const m = new Map(); for (const r of rows) { const key = r[k]; if (!key) continue; const p = m.get(key); if (!p || String(r.as_of_date) > String(p.as_of_date)) m.set(key, r); } return m; };
const sprayBy = latest(spray, 'player_key'); const defBy = latest(defense, 'team'); const platBy = latest(platoon, 'player_key');
const batBy = new Map(); const pitBy = new Map();
for (const r of statcast) { if (!r.player_key) continue; (r.role === 'pitcher' ? pitBy : batBy).set(r.player_key, r); }
const frac = (v) => { const n = knownNumber(v); return n === null ? null : (n > 1 ? n / 100 : n); };
// Opponent + starter per (player,date) from the sequence cache.
const { games } = JSON.parse(fs.readFileSync(SEQ, 'utf8'));
const oppOf = new Map(); const spOf = new Map();
for (const g of games) {
for (const side of ['home', 'away']) {
const opp = g[side === 'home' ? 'away' : 'home'];
const st = (g[side].arms || []).find((a) => a.started);
const half = side === 'home' ? 'top' : 'bottom';
for (const pa of g.pas.filter((p) => p.half === half)) {
const k = `${g.date}|${nameKey(pa.batter_name || '')}`;
if (!oppOf.has(k)) { oppOf.set(k, g[side].team); if (st) spOf.set(k, st.name); }
}
}
}
const snaps = await page(sb, 'model_snapshots', 'id, game_date, captured_at, stat, player_key, player_name, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').eq('stat', 'hits'));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({ propKey: [r.game_date, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const rows = []; const transmission = []; const unreadable = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(b.hits); if (v === null) continue;
const over = v > L;
const won = (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0;
const key = r.player_key;
const bat = batBy.get(key);
const oppTeam = oppOf.get(`${r.game_date}|${key}`);
const def = oppTeam ? (defBy.get(oppTeam) || defBy.get(String(oppTeam).split(' ').pop())) : null;
const spName = spOf.get(`${r.game_date}|${key}`);
const pit = spName ? pitBy.get(nameKey(spName)) : null;
const sp = platBy.get(key);
const ctx = {
spray: sprayBy.get(key) || null,
positionOaa: def && def.position_oaa ? def.position_oaa : null,
bats: bat && bat.bats ? String(bat.bats)[0] : null,
throws: pit && pit.throws ? String(pit.throws)[0] : null,
pitcherHardHit: pit ? frac(pit.hard_hit_pct) : null,
platoonSplits: sp ? { vl: { pa: sp.vl_pa, atBats: sp.vl_ab, hits: sp.vl_hits }, vr: { pa: sp.vr_pa, atBats: sp.vr_ab, hits: sp.vr_hits } } : null,
};
// The engine adjusts p_over then flips for unders; replay that exactly.
const pOverRaw = String(r.side).toLowerCase() === 'under' ? 1 - knownNumber(r.p_win) : knownNumber(r.p_win);
const adj = hf.adjustProbability(pOverRaw, ctx);
const pAfter = adj.factors_fired > 0
? (String(r.side).toLowerCase() === 'under' ? 1 - adj.p_adjusted : adj.p_adjusted)
: knownNumber(r.p_win);
rows.push({ date: r.game_date, p_before: knownNumber(r.p_win), p: pAfter, won, fired: adj.factors_fired, applied: adj.applied });
// TRANSMISSION IS TESTED PER FACTOR, IN ISOLATION.
// Comparing one factor's expected sign against the COMPOSITE p_win change is
// wrong: with three factors firing, two pulling down and one up, the net can
// oppose any single member and look like a defect when nothing is broken.
// So each factor is applied ALONE to the same base and its own sign checked.
const isUnder = String(r.side).toLowerCase() === 'under';
for (const a of adj.applied) {
if (transmission.filter((t) => t.factor === a.factor).length >= 4) continue;
if (Math.abs(a.multiplier - 1) < 0.03) continue;
const solo = { spray: null, positionOaa: null, bats: ctx.bats, throws: null, pitcherHardHit: null, platoonSplits: null };
if (a.factor === 'defense_by_direction') { solo.spray = ctx.spray; solo.positionOaa = ctx.positionOaa; }
if (a.factor === 'pitcher_contact_profile') solo.pitcherHardHit = ctx.pitcherHardHit;
if (a.factor === 'platoon_severity') { solo.platoonSplits = ctx.platoonSplits; solo.throws = ctx.throws; }
const one = hf.adjustProbability(pOverRaw, solo);
if (one.factors_fired !== 1) continue;
const soloWin = isUnder ? 1 - one.p_adjusted : one.p_adjusted;
transmission.push({
factor: a.factor, player: r.player_name, date: r.game_date,
expected: a.multiplier > 1 ? 'raise p(over)' : 'lower p(over)', multiplier: a.multiplier,
side: r.side, p_before: knownNumber(r.p_win), p_after_solo: r4(soloWin),
sign_correct: isUnder
? ((a.multiplier > 1) === (soloWin < knownNumber(r.p_win)))
: ((a.multiplier > 1) === (soloWin > knownNumber(r.p_win))),
});
}
if (adj.skipped.some((s) => /switch hitter/.test(s.reason || '')) && unreadable.length < 4) {
unreadable.push({ player: r.player_name, reason: 'switch hitter — spray side unreadable', p_before: knownNumber(r.p_win), p_after: pAfter, moved_by_spray: false });
}
}
// ── PHASE 4: OOS, point-in-time ──
rows.sort((a, b) => a.date.localeCompare(b.date));
const dates = [...new Set(rows.map((r) => r.date))].sort();
const perDate = new Map(); for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) { acc += perDate.get(d); if (acc >= rows.length * 0.45) { cut = d; break; } }
const fit = rows.filter((r) => r.date < cut); const ev = rows.filter((r) => r.date >= cut);
const mBefore = lp.fitPlatt(fit.map((r) => ({ p: r.p_before, won: r.won, date: r.date })));
const mAfter = lp.fitPlatt(fit.map((r) => ({ p: r.p, won: r.won, date: r.date })));
const evBefore = ev.map((r) => ({ ...r, p: mBefore && !mBefore.refused ? lp.applyPlatt(mBefore, r.p_before) : r.p_before })).filter((r) => r.p != null);
const evAfter = ev.map((r) => ({ ...r, p: mAfter && !mAfter.refused ? lp.applyPlatt(mAfter, r.p) : r.p })).filter((r) => r.p != null);
const bBefore = guards.safeBrier(evBefore.map((r) => r.p), evBefore.map((r) => r.won));
const bAfter = guards.safeBrier(evAfter.map((r) => r.p), evAfter.map((r) => r.won));
const byDate = new Map();
for (let i = 0; i < evAfter.length; i += 1) { const d = evAfter[i].date; if (!byDate.has(d)) byDate.set(d, []); byDate.get(d).push({ a: evAfter[i].p, b: evBefore[i] ? evBefore[i].p : null, won: evAfter[i].won }); }
const keys = [...byDate.keys()]; const rnd = makeRnd(20260807); const diffs = [];
for (let it = 0; it < 3000; it += 1) {
const s = [];
for (let i = 0; i < keys.length; i += 1) s.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
const u = s.filter((x) => x.b != null);
if (!u.length) continue;
diffs.push(guards.safeBrier(u.map((x) => x.a), u.map((x) => x.won)) - guards.safeBrier(u.map((x) => x.b), u.map((x) => x.won)));
}
diffs.sort((a, b) => a - b);
console.log(JSON.stringify({
coverage: { rows: rows.length, any_factor_fired: rows.filter((r) => r.fired > 0).length,
by_count: [0, 1, 2, 3].map((k) => ({ factors: k, n: rows.filter((r) => r.fired === k).length })) },
PHASE_2_transmission: transmission,
PHASE_2_unreadable_static: unreadable,
PHASE_4: {
split_at: cut, fit_n: fit.length, eval_n: ev.length, eval_dates: keys.length,
resolution_before: decompose(evBefore), resolution_after: decompose(evAfter),
brier_before: r5(bBefore), brier_after: r5(bAfter), brier_delta: r5(bAfter - bBefore),
brier_ci_date_block: diffs.length ? [r5(diffs[Math.floor(diffs.length * 0.025)]), r5(diffs[Math.floor(diffs.length * 0.975)])] : null,
},
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const r5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
+137
View File
@@ -0,0 +1,137 @@
#!/usr/bin/env node
'use strict';
/**
* generate-content — tonight's posts, from tonight's real data.
*
* READ-ONLY on every source. This writes nothing to any serving, model or
* ledger table, so it has zero effect on the repaired-champion accrual clock.
*
* SUPABASE_URL=... node scripts/generate-content.js
* -> .content-out/<date>/<template>.txt and .svg
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const engine = require('../src/services/content/contentEngine');
const { toSvg } = require('../src/services/content/cardRenderer');
const sg = require('../src/services/model/servedGrade');
const { knownNumber } = require('../src/utils/known');
for (const t of ['hotHitters', 'honestyFlex', 'streakList']) {
engine.registerTemplate(require(`../src/services/content/templates/${t}`));
}
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const OUT = path.join(process.cwd(), '.content-out');
async function page(sb, t, sel, ob, f) {
const o = [];
for (let i = 0; ; i += 1000) {
const { data, error } = await f(sb.from(t).select(sel)).order(ob, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break;
o.push(...data);
if (data.length < 1000) break;
}
return o;
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const date = new Intl.DateTimeFormat('en-CA', { timeZone: 'America/New_York', year: 'numeric', month: '2-digit', day: '2-digit' }).format(new Date());
// ── SOURCE 1: full-season hitter form (the REPAIRED window, not last10) ──
const lines = fs.existsSync(BOX) ? JSON.parse(fs.readFileSync(BOX, 'utf8')).lines : {};
const byPlayer = new Map();
for (const [k, b] of Object.entries(lines)) {
const [d, key] = k.split('|');
if (!byPlayer.has(key)) byPlayer.set(key, []);
byPlayer.get(key).push({ d, hits: b.hits, name: b.name });
}
// The box-score cache spans only the settled snapshot window, so it holds far
// fewer than a season per player. Season form comes from the SAME full log the
// repaired champion reads.
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const hitterFormFull = async (names) => {
const out = [];
for (const n of names.slice(0, 60)) {
try {
const res = await mlb.getPlayerStats(n);
const log = (res && res.found && Array.isArray(res.fullLog)) ? res.fullLog : [];
const vals = log.map((g) => knownNumber(g && g.stat && g.stat.hits)).filter((v) => v !== null);
if (vals.length < 20) continue;
const rate = (a) => a.filter((v) => v > 0).length / a.length;
out.push({ name: n, season_games: vals.length, season_rate: rate(vals), recent_rate: rate(vals.slice(-10)) });
} catch { /* absent player -> absent row */ }
}
return out;
};
const cacheForm = async () => [...byPlayer.entries()].map(([key, games]) => {
games.sort((a, b) => a.d.localeCompare(b.d));
const vals = games.map((g) => knownNumber(g.hits)).filter((v) => v !== null);
if (vals.length < 20) return null;
const rate = (arr) => arr.filter((v) => v > 0).length / arr.length;
return {
name: games[games.length - 1].name || key,
season_games: vals.length,
season_rate: rate(vals),
recent_rate: rate(vals.slice(-10)),
};
}).filter(Boolean);
const hitterForm = async () => {
const names = [...byPlayer.values()].map((g) => g[g.length - 1].name).filter(Boolean);
const full = await hitterFormFull([...new Set(names)]);
return full.length ? full : await cacheForm();
};
// ── SOURCE 2: the real served-grade distribution ──
const snaps = await page(sb, 'model_snapshots', 'p_win, refused, stat, game_date', 'id',
(q) => q.eq('sport', 'mlb').eq('game_date', date));
const gradeDistribution = async () => {
const usable = snaps.filter((r) => !r.refused && knownNumber(r.p_win) !== null);
const by = {}; let flat = 0;
for (const r of usable) {
const g = sg.gradeFor({ p_win: knownNumber(r.p_win) });
by[g.letter] = (by[g.letter] || 0) + 1;
if (g.separates_from_base_rate === false) flat += 1;
}
return { total: usable.length || null, by_letter: by, not_separable: flat };
};
// ── SOURCE 3: streaks verified from SETTLED ledger outcomes only ──
const led = await page(sb, 'ledger_entries', 'player_name, player_key, game_date, outcome, stat', 'id',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits').in('outcome', ['hit', 'miss']));
const settledStreaks = async () => {
const by = new Map();
for (const r of led) {
if (!by.has(r.player_key)) by.set(r.player_key, []);
by.get(r.player_key).push(r);
}
const out = [];
for (const [, rows] of by) {
rows.sort((a, b) => String(b.game_date).localeCompare(String(a.game_date)));
let n = 0;
for (const r of rows) { if (r.outcome === 'hit') n += 1; else break; }
if (n >= 3) out.push({ name: rows[0].player_name, streak: n, verified_from_settled: true });
}
return out;
};
const deps = { date, hitterForm, gradeDistribution, settledStreaks, servedGrade: sg };
const results = await engine.generateAll(deps);
const dir = path.join(OUT, date);
fs.mkdirSync(dir, { recursive: true });
for (const r of results) {
if (!r.ok) { console.log(`\n[SKIP] ${r.id}${r.reason}`); continue; }
fs.writeFileSync(path.join(dir, `${r.id}.txt`), r.copy);
fs.writeFileSync(path.join(dir, `${r.id}.svg`), toSvg(r.card));
console.log(`\n${'='.repeat(64)}\n${r.id.toUpperCase()}${r.honest_absence ? ' [HONEST ABSENCE]' : ''}\n${'='.repeat(64)}\n${r.copy}`);
}
console.log(`\n\noutput: ${dir}`);
process.exit(0);
})().catch((e) => { console.error('FAILED:', e.message); process.exit(1); });
+13
View File
@@ -0,0 +1,13 @@
-- T0 — CALIBRATION OF p_win (gates T1-T4). specs/grade-diagnostic-t0.md
-- Pre-registered firing thresholds: mean |predicted-actual| > 0.05 across
-- buckets, OR a monotonic over/under-confidence slope. BOTH fired.
with r as (
select sport, p_win::numeric as p, (outcome='hit')::int as won
from ledger_entries
where user_id is null and outcome in ('hit','miss') and p_win is not null
), b as (
select sport, width_bucket(p, 0.1, 1.0, 9) as bkt, p, won from r
)
select sport, bkt, count(*) n,
avg(p) predicted, avg(won) actual, avg(p)-avg(won) over_confidence
from b group by sport, bkt having count(*) >= 10 order by sport, bkt;
+31
View File
@@ -0,0 +1,31 @@
-- GRADE CORRELATION PROOF (2026-07-31). Re-runnable evidence for
-- specs/grade-fix-part1-investigation.md. LOOKAHEAD GUARD: fair_prob_lock is the
-- LOCK-TIME market probability; closing_prob (the close) is deliberately unused.
-- 1) Candidate edge formulations vs outcome, per sport.
with r as (
select sport, (outcome='hit')::int as won, p_win::numeric as p, fair_prob_lock::numeric as f,
case grade when 'A+' then 10 when 'A' then 9 when 'A-' then 8 when 'B+' then 7 when 'B' then 6
when 'B-' then 5 when 'C+' then 4 when 'C' then 3 when 'C-' then 2 when 'D' then 1 else 0 end as champ_idx
from ledger_entries
where user_id is null and outcome in ('hit','miss')
and p_win is not null and fair_prob_lock is not null and grade is not null
)
select sport, count(*) n,
corr(champ_idx, won) champ_letter_r,
corr(p, won) p_win_alone_r,
corr(p - f, won) additive_edge_r,
corr(p / nullif(f,0), won) ratio_edge_r,
corr(ln(p/(1-p)) - ln(f/(1-f)), won) logodds_edge_r
from r group by sport;
-- 2) Time-forward split (overfitting guard), MLB.
with r as (
select game_date, (outcome='hit')::int as won, p_win::numeric as p, fair_prob_lock::numeric as f
from ledger_entries
where user_id is null and outcome in ('hit','miss') and sport='mlb'
and p_win is not null and fair_prob_lock is not null
), s as (select *, ntile(2) over (order by game_date) half from r)
select half, count(*) n, min(game_date) from_date, max(game_date) to_date,
corr(p, won) p_win_alone_r, corr(p - f, won) additive_edge_r
from s group by half order by half;
+143
View File
@@ -0,0 +1,143 @@
#!/usr/bin/env node
/**
* HEAL DRY-RUN (Order 2, Phase 0) — REPORT ONLY. WRITES NOTHING.
*
* Classifies every ledger row's stored game_date against the real schedule and
* re-derives what opponent a grade would have been bound to, so the heal is
* planned against verified ground truth instead of the "by hour" proxy.
*
* KEY QUESTION IT ANSWERS: baseball is played in SERIES. A one-day-off opponent
* bind may land on the SAME opponent, in which case the grade was never harmed.
* Only a bind that changes the opponent is real damage.
*/
const sched = require('../src/services/scheduleService');
const norm = (s) => String(s || '').toLowerCase().replace(/[^a-z0-9]/g, '');
function teamsFromGameId(gameId) {
// mlb:2026-07-16:NewYorkMets@PhiladelphiaPhillies (may carry "(Game1)")
const m = String(gameId || '').match(/^([a-z]+):(\d{4}-\d{2}-\d{2}):(.+)@(.+)$/);
if (!m) return null;
return { sport: m[1], date: m[2], away: m[3].replace(/\(Game\d+\)/i, ''), home: m[4].replace(/\(Game\d+\)/i, '') };
}
function addDays(d, n) {
const t = new Date(`${d}T12:00:00Z`);
t.setUTCDate(t.getUTCDate() + n);
return t.toISOString().slice(0, 10);
}
const schedCache = new Map();
async function scheduleFor(sport, date) {
const k = `${sport}:${date}`;
if (!schedCache.has(k)) {
try { schedCache.set(k, (await sched.getSchedule(sport, date)) || []); }
catch { schedCache.set(k, []); }
}
return schedCache.get(k);
}
const abbrOf = (t) => norm(t && (t.abbreviation || t.name || t.displayName));
const nameOf = (t) => String((t && (t.displayName || t.name || t.abbreviation)) || '');
/** Find the game a team played on a date. Returns {opponent, isHome, id, count}. */
async function gameForTeam(sport, date, teamName) {
const games = await scheduleFor(sport, date);
const want = norm(teamName);
const hits = [];
for (const g of games) {
const h = abbrOf(g.homeTeam);
const a = abbrOf(g.awayTeam);
const match = (x) => !!x && !!want && (x === want || x.includes(want) || want.includes(x));
if (match(h)) hits.push({ opponent: nameOf(g.awayTeam), isHome: true, id: g.id });
else if (match(a)) hits.push({ opponent: nameOf(g.homeTeam), isHome: false, id: g.id });
}
return hits.length ? { ...hits[0], count: hits.length } : null;
}
(async () => {
// Supabase REST is unreachable from this dev box (same restriction as :5432),
// so rows are exported via the MCP SQL path and read from a local JSON file.
const rows = JSON.parse(require('fs').readFileSync(process.argv[2], 'utf8'));
console.log(`rows fetched: ${rows.length}\n`);
const dateClass = { CORRECT: 0, MISDATED: 0, UNBINDABLE: 0 };
const oppClass = { MATCH: 0, MISMATCH: 0, UNDETERMINED: 0, NOT_AFFECTED: 0 };
const misdatedEx = [];
const mismatchEx = [];
const dh = [];
const unbindableBy = {};
const mismatchBy = {};
const dhRows = new Set();
const mismatchIds = [];
for (const r of rows) {
const t = teamsFromGameId(r.game_id);
const hourUtc = Number(r.hour);
const affected = hourUtc <= 6; // window proven live (0006 UTC)
if (!t) { dateClass.UNBINDABLE += 1; oppClass.UNDETERMINED += 1;
const k=`${r.sport} ${r.game_date} (no game_id parse)`; unbindableBy[k]=(unbindableBy[k]||0)+1; continue; }
// --- date axis: did these teams actually play on the stored date? ---
const onStored = await gameForTeam(r.sport, r.game_date, t.home);
if (onStored) {
dateClass.CORRECT += 1;
if (onStored.count > 1) { dh.push({ id: r.id, date: r.game_date, team: t.home, games: onStored.count }); dhRows.add(r.id); }
} else {
const next = await gameForTeam(r.sport, addDays(r.game_date, 1), t.home);
dateClass[next ? 'MISDATED' : 'UNBINDABLE'] += 1;
if (!next) { const k=`${r.sport} ${r.game_date}`; unbindableBy[k]=(unbindableBy[k]||0)+1; }
if (next && misdatedEx.length < 6) {
misdatedEx.push({ player: r.player_name, stored: r.game_date, real: addDays(r.game_date, 1), outcome: r.outcome });
}
}
// --- grade axis: would the wrong-day bind have changed the OPPONENT? ---
if (!affected) { oppClass.NOT_AFFECTED += 1; continue; }
const trueDate = onStored ? r.game_date : addDays(r.game_date, 1);
const trueGame = onStored || await gameForTeam(r.sport, trueDate, t.home);
const boundGame = await gameForTeam(r.sport, addDays(trueDate, -1), t.home); // ESPN "yesterday"
if (!trueGame) { oppClass.UNDETERMINED += 1; continue; }
if (!boundGame) { oppClass.MATCH += 1; continue; } // no prior game → nothing wrong to bind
if (norm(trueGame.opponent) === norm(boundGame.opponent) && trueGame.isHome === boundGame.isHome) {
oppClass.MATCH += 1; // SERIES: same opponent, no harm
} else {
oppClass.MISMATCH += 1;
const k=`${r.sport} ${trueDate} settled=${r.outcome||'PENDING'}`; mismatchBy[k]=(mismatchBy[k]||0)+1;
mismatchIds.push({ id: r.id, sport: r.sport, date: trueDate, outcome: r.outcome,
graded_vs: boundGame.opponent, true_opponent: trueGame.opponent });
if (mismatchEx.length < 8) {
mismatchEx.push({
player: r.player_name, date: trueDate,
graded_vs: `${boundGame.opponent}${boundGame.isHome ? ' (H)' : ' (A)'}`,
true_opponent: `${trueGame.opponent}${trueGame.isHome ? ' (H)' : ' (A)'}`,
});
}
}
}
console.log('=== 0.1 DATE AXIS ===');
console.log(JSON.stringify(dateClass, null, 2));
if (misdatedEx.length) { console.log('mis-dated examples:'); misdatedEx.forEach((e) => console.log(' ', JSON.stringify(e))); }
console.log('\n=== 0.4 GRADE AXIS (affected window 0006 UTC) ===');
console.log(JSON.stringify(oppClass, null, 2));
if (mismatchEx.length) { console.log('opponent MISMATCH examples:'); mismatchEx.forEach((e) => console.log(' ', JSON.stringify(e))); }
console.log('\n=== UNBINDABLE breakdown (sport date -> rows) ===');
Object.entries(unbindableBy).sort().forEach(([k,v])=>console.log(` ${k}: ${v}`));
console.log('\n=== MISMATCH breakdown (sport trueDate settled? -> rows) ===');
Object.entries(mismatchBy).sort().forEach(([k,v])=>console.log(` ${k}: ${v}`));
console.log('\n=== 0.6 DOUBLEHEADER: rows on a date where the team played 2 games ===');
console.log(` affected rows: ${dhRows.size}`);
const byTeam={}; dh.forEach(d=>{const k=`${d.date} ${d.team}`; byTeam[k]=(byTeam[k]||0)+1;});
Object.entries(byTeam).sort().forEach(([k,v])=>console.log(` ${k}: ${v} rows`));
require('fs').writeFileSync('/tmp/claude-1000/-home-kev-mastermind-vyndr/a6b79396-5b11-477d-8353-0672bd5789c1/scratchpad/mismatch_ids.json', JSON.stringify(mismatchIds, null, 1));
require('fs').writeFileSync('/tmp/claude-1000/-home-kev-mastermind-vyndr/a6b79396-5b11-477d-8353-0672bd5789c1/scratchpad/dh_ids.json', JSON.stringify([...dhRows], null, 1));
console.log(`\nID files written: mismatch=${mismatchIds.length} doubleheader=${dhRows.size}`);
await new Promise((r) => process.stdout.write('', r));
process.exit(0);
})().catch((e) => { console.error('dry-run failed:', e.message); process.exit(1); });
+102
View File
@@ -0,0 +1,102 @@
#!/usr/bin/env node
'use strict';
/**
* hits-input-coverage — STEP 0 of hits-v1. CONFIRM THE INPUTS EXIST.
*
* hits-v1 models P(hits >= line) as a binomial over AT-BATS. That needs two
* inputs per player that the current negative-binomial ladder never asked for:
*
* 1. at-bats per game (the OPPORTUNITY count — the binomial's n)
* 2. per-AB hit rate (the CONVERSION rate — the binomial's q)
*
* A model is not "wired" until its inputs are present on real rows at real
* coverage. This probe pulls the actual players carrying hits props in the
* ledger, fetches their REAL statsapi game logs, and reports what fraction
* yields usable AB + hit-rate inputs.
*
* UNKNOWN IS NOT ZERO. Every read goes through `knownRate`. A game-log row with
* no atBats field is COUNTED AS MISSING, never as a 0-AB game — reading it as
* zero would say "this player had no opportunity", the strongest possible
* statement, from an absence of data. That is the defect this codebase has
* shipped seven times.
*
* Usage: node scripts/hits-input-coverage.js [limit]
*/
const { knownRate } = require('../src/utils/known');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
// The real players carrying hits props in the public ledger (2026-08-02 pull,
// ordered by row count). Hard-coded rather than re-queried so the probe runs
// without Supabase credentials — these are REAL names off REAL rows.
const PLAYERS = [
'Esmerlyn Valdez', 'Trea Turner', 'Ryan Jeffers', 'Steven Kwan', 'Jake Mangum',
'Jazz Chisholm Jr', 'JT Realmuto', 'Bo Bichette', 'Ben Rice', 'Brandon Lowe',
'Nick Gonzales', 'Alan Roden', 'Junior Caminero', 'Bryce Harper', 'Travis Bazzana',
'Jasson Dominguez', 'Trent Grisham', 'Chase DeLauter', 'Jorge Polanco', 'Wyatt Langford',
'Petey Halpin', 'Alec Bohm', 'Munetaka Murakami', 'Kyle Schwarber', 'AJ Ewing',
'Royce Lewis', 'Javier Sanoja', 'Ben Williamson', 'Bryson Stott', 'Francisco Lindor',
];
const MIN_GAMES = Number(process.env.HITS_MIN_GAMES || 5);
async function main() {
const limit = Number(process.argv[2] || PLAYERS.length);
const names = PLAYERS.slice(0, limit);
const report = [];
for (const name of names) {
const row = { player: name, resolved: false, games: 0, ab_games: 0, hit_games: 0, ab_per_game: null, hit_rate: null, usable: false };
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) { report.push(row); continue; }
row.resolved = true;
const log = await mlb.getPlayerGameLog(found.id);
row.games = (log || []).length;
let abSum = 0; let hSum = 0; let abGames = 0; let hitGames = 0;
for (const g of log || []) {
const s = (g && g.stat) || {};
const ab = knownRate(s.atBats); // absent -> null, NOT 0
const h = knownRate(s.hits);
if (ab !== null) { abSum += ab; abGames += 1; }
if (h !== null) { hSum += h; hitGames += 1; }
}
row.ab_games = abGames;
row.hit_games = hitGames;
if (abGames >= MIN_GAMES && abSum > 0) {
row.ab_per_game = Math.round((abSum / abGames) * 1000) / 1000;
row.hit_rate = Math.round((hSum / abSum) * 1000) / 1000;
row.usable = true;
}
} catch (e) {
row.error = e.message;
}
report.push(row);
}
const resolved = report.filter((r) => r.resolved).length;
const usable = report.filter((r) => r.usable).length;
const rates = report.filter((r) => r.usable).map((r) => r.hit_rate);
const abs = report.filter((r) => r.usable).map((r) => r.ab_per_game);
const avg = (a) => (a.length ? Math.round((a.reduce((x, y) => x + y, 0) / a.length) * 1000) / 1000 : null);
console.log(JSON.stringify({
probed: report.length,
resolved,
usable_combined_inputs: usable,
coverage_pct: Math.round((usable / report.length) * 1000) / 10,
min_games_required: MIN_GAMES,
mean_ab_per_game: avg(abs),
mean_hit_rate_per_ab: avg(rates),
hit_rate_range: rates.length ? [Math.min(...rates), Math.max(...rates)] : null,
rows: report,
}, null, 2));
// Redis runs degraded locally; a reconnect timer would hold the process open
// and piped output would be lost to SIGTERM. Same rule as verify-grade-range.
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+235
View File
@@ -0,0 +1,235 @@
#!/usr/bin/env node
'use strict';
/**
* hits-v1-holdout — STEP 3. POINT-IN-TIME REPLAY, HITS ROWS ONLY.
*
* THREE guards this script exists to enforce, all of which have burned a
* measurement in this codebase before:
*
* 1. HITS ROWS ONLY. Averaging hits into the other stats would hide the
* effect entirely — hits is one stat among nine and the ladder's failure is
* specific to it.
*
* 2. DIRECTION-ALIGNED. `p_win` is P(GRADED SIDE); `proj_p_over_line` and
* `proj_hits_p_over` are P(OVER). 26% of matched hits rows are
* under-graded, and comparing a raw P(over) against an under-side outcome
* measures the model BACKWARDS. That artifact alone accounted for 41% of
* the ladder's apparent loss when it was first measured.
*
* 3. NO LOOKAHEAD. This is the guard specific to a replay. For each settled
* row, the player's game log is rebuilt STRICTLY BEFORE that row's
* game_date, and the multiplier is the REAL `combined_multiplier` recorded
* on the row at grade time. A replay that used today's full log would be
* scoring a prediction with the answer in hand — a fabricated result, and
* a worse lie than no measurement.
*
* CONTAMINATION EXCLUSION (mandatory). Rows whose price/book were stamped from a
* NON-TAKEABLE book between 2026-08-01 and the write-path fix are tagged
* `quarantine_reason LIKE 'nontakeable_book%'` and are EXCLUDED: their locked
* price describes a market you could not have bet.
*
* WHAT THIS IS AND IS NOT. It is a backtest, and it is labelled one. The verdict
* of record is the FORWARD ledger accrual, which starts at the next snapshot.
* Stated limits: statsapi is read as it stands today (retroactive stat
* corrections are invisible), and LEAGUE_HIT_RATE / PRIOR_AB are constants set
* today — at 20 at-bats against a regular's 200400 the prior moves a settled
* hitter by thousandths, but it is not zero.
*
* node scripts/hits-v1-holdout.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const binomialHits = require('../src/services/projection/binomialHits');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownNumber } = require('../src/utils/known');
const { normalizeName } = require('../src/utils/playerName');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
/** Pearson correlation — the resolution measure the ladder is judged on. */
function corr(xs, ys) {
const n = xs.length;
if (n < 3) return null;
const mx = xs.reduce((a, b) => a + b, 0) / n;
const my = ys.reduce((a, b) => a + b, 0) / n;
let sxy = 0; let sxx = 0; let syy = 0;
for (let i = 0; i < n; i += 1) {
const dx = xs[i] - mx; const dy = ys[i] - my;
sxy += dx * dy; sxx += dx * dx; syy += dy * dy;
}
if (sxx <= 0 || syy <= 0) return null;
return Math.round((sxy / Math.sqrt(sxx * syy)) * 10000) / 10000;
}
const mean = (a) => (a.length ? Math.round((a.reduce((x, y) => x + y, 0) / a.length) * 10000) / 10000 : null);
const sd = (a) => {
if (a.length < 2) return null;
const m = a.reduce((x, y) => x + y, 0) / a.length;
return Math.round(Math.sqrt(a.reduce((s, v) => s + (v - m) ** 2, 0) / (a.length - 1)) * 10000) / 10000;
};
/** Brier score — lower is better. Reported beside resolution as a check. */
const brier = (ps, ys) => (ps.length
? Math.round((ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length) * 10000) / 10000
: null);
/**
* Paired bootstrap CI on a DIFFERENCE of resolutions.
*
* Both models score the SAME rows, so their errors are correlated and comparing
* two independent standard errors would overstate the uncertainty. Resampling
* rows as pairs preserves that dependence. Deterministic seed — a measurement
* that changes between runs is not a measurement.
*/
function bootstrapDiff(rowsIn, keyA, keyB, iters = 4000, seed = 20260802) {
if (rowsIn.length < 20) return null;
let s = seed >>> 0;
const rnd = () => { // xorshift32 — deterministic, no Math.random
s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0;
return s / 4294967296;
};
const n = rowsIn.length;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const bArr = [];
for (let i = 0; i < n; i += 1) {
const r = rowsIn[Math.floor(rnd() * n)];
ys.push(r.won); a.push(r[keyA]); bArr.push(r[keyB]);
}
const ca = corr(a, ys); const cb = corr(bArr, ys);
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (p) => Math.round(diffs[Math.floor(p * (diffs.length - 1))] * 10000) / 10000;
return {
point: Math.round(((corr(rowsIn.map((r) => r[keyA]), rowsIn.map((r) => r.won)) || 0)
- (corr(rowsIn.map((r) => r[keyB]), rowsIn.map((r) => r.won)) || 0)) * 10000) / 10000,
ci95: [q(0.025), q(0.975)],
// The question the promotion gate actually asks.
p_improves: Math.round((diffs.filter((d) => d > 0).length / diffs.length) * 1000) / 1000,
};
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const { data, error } = await sb
.from('ledger_entries')
.select('id, player_name, player_key, stat, line, side, outcome, game_date, p_win, proj_p_over_line, proj_factors, quarantine_reason')
.eq('sport', 'mlb')
.is('user_id', null)
.eq('stat', 'hits')
.in('outcome', ['hit', 'miss'])
.not('p_win', 'is', null)
.not('proj_p_over_line', 'is', null);
if (error) throw error;
// Contamination exclusion, applied in JS so the filter is visible here.
const rows = (data || []).filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// Resolve each distinct player ONCE, and cache the full game log.
const logCache = new Map();
const names = [...new Set(rows.map((r) => r.player_name).filter(Boolean))];
let resolved = 0;
for (const name of names) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) { logCache.set(name, null); continue; }
const log = await mlb.getPlayerGameLog(found.id);
logCache.set(name, Array.isArray(log) ? log : null);
if (log && log.length) resolved += 1;
} catch { logCache.set(name, null); }
}
const out = [];
const reasons = {};
const bump = (k) => { reasons[k] = (reasons[k] || 0) + 1; };
for (const r of rows) {
const full = logCache.get(r.player_name);
if (!full) { bump('no_game_log'); continue; }
const gameDate = String(r.game_date || '').slice(0, 10);
if (!gameDate) { bump('no_game_date'); continue; }
// ── NO LOOKAHEAD ──────────────────────────────────────────────────────
// Strictly BEFORE the graded game. A row dated the same day is the game
// being predicted; including it would hand the model the answer.
const priorLog = full.filter((g) => g && g.date && String(g.date).slice(0, 10) < gameDate);
if (priorLog.length < binomialHits.HITS_MIN_GAMES) { bump('thin_prior_log'); continue; }
// The REAL grade-time multiplier, recorded on the row at lock.
const m = knownNumber(r.proj_factors && r.proj_factors.combined_multiplier);
const proj = binomialHits.projectHits({
rows: priorLog, line: Number(r.line), multiplier: m == null ? 1 : m,
});
if (!proj) { bump('inputs_underivable'); continue; }
// ── DIRECTION-ALIGN to the graded side ────────────────────────────────
const under = String(r.side || '').toLowerCase() === 'under';
const won = r.outcome === 'hit' ? 1 : 0;
out.push({
won,
under,
champ: Number(r.p_win),
ladder: under ? 1 - Number(r.proj_p_over_line) : Number(r.proj_p_over_line),
hitsv1: under ? 1 - proj.p_over_line : proj.p_over_line,
line: Number(r.line),
games_prior: priorLog.length,
});
}
const slice = (rowsIn, label) => {
const ys = rowsIn.map((x) => x.won);
const c = rowsIn.map((x) => x.champ);
const l = rowsIn.map((x) => x.ladder);
const h = rowsIn.map((x) => x.hitsv1);
return {
slice: label,
n: rowsIn.length,
under_rows: rowsIn.filter((x) => x.under).length,
base_rate: mean(ys),
resolution: { champion: corr(c, ys), current_ladder: corr(l, ys), hits_v1: corr(h, ys) },
brier: { champion: brier(c, ys), current_ladder: brier(l, ys), hits_v1: brier(h, ys) },
mean_p: { champion: mean(c), current_ladder: mean(l), hits_v1: mean(h) },
sd_p: { champion: sd(c), current_ladder: sd(l), hits_v1: sd(h) },
};
};
console.log(JSON.stringify({
measurement: 'POINT-IN-TIME REPLAY (backtest) — verdict of record is the forward ledger accrual',
guards: {
hits_rows_only: true,
direction_aligned: true,
no_lookahead: 'game log truncated strictly before each row game_date',
grade_time_multiplier: 'real combined_multiplier from the row',
contamination_excluded: 'nontakeable_book*',
},
candidate_rows: rows.length,
players_resolved: `${resolved}/${names.length}`,
matched_rows: out.length,
dropped: reasons,
overall: slice(out, 'all hits rows'),
// Is the comparison a RESULT or a noise reading? Paired bootstrap, so the
// shared rows are not double-counted as independent evidence.
paired_bootstrap: {
note: 'difference in resolution, 4000 paired resamples, deterministic seed',
hits_v1_minus_ladder: bootstrapDiff(out, 'hitsv1', 'ladder'),
champion_minus_ladder: bootstrapDiff(out, 'champ', 'ladder'),
champion_minus_hits_v1: bootstrapDiff(out, 'champ', 'hitsv1'),
at_line_0_5: {
hits_v1_minus_ladder: bootstrapDiff(out.filter((x) => x.line === 0.5), 'hitsv1', 'ladder'),
champion_minus_ladder: bootstrapDiff(out.filter((x) => x.line === 0.5), 'champ', 'ladder'),
},
},
by_line: [0.5, 1.5, 2.5].map((ln) => slice(out.filter((x) => x.line === ln), `line ${ln}`))
.filter((s) => s.n > 0),
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+146
View File
@@ -0,0 +1,146 @@
#!/usr/bin/env node
'use strict';
/**
* ingest-game-sequences — the raw material for the reliever chain.
*
* Every link in this chain needs something the ledger does not carry: when the
* starter actually left, and which arm actually faced each plate appearance.
* Both are free from statsapi (playByPlay + boxscore), on the same host we
* already use for game logs and schedules.
*
* Caches to disk so Link 1, 2 and 3 all read one fetch rather than three.
*
* node scripts/ingest-game-sequences.js # default window
* SEQ_FROM=2026-06-01 SEQ_TO=2026-08-04 node scripts/ingest-game-sequences.js
*/
const fs = require('fs');
const path = require('path');
const axios = require('axios');
const FROM = process.env.SEQ_FROM || '2026-06-15';
const TO = process.env.SEQ_TO || '2026-08-04';
const OUT = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const CONCURRENCY = 6;
const get = async (url) => (await axios.get(url, { timeout: 45_000 })).data;
/** Innings pitched come as '5.2' meaning five and TWO THIRDS — parseFloat is wrong. */
function ipToOuts(ip) {
if (ip == null) return null;
const [whole, frac] = String(ip).split('.');
const w = Number(whole); const f = Number(frac || 0);
if (!Number.isFinite(w)) return null;
return w * 3 + (Number.isFinite(f) ? f : 0);
}
function datesBetween(from, to) {
const out = [];
const d = new Date(`${from}T12:00:00Z`);
const end = new Date(`${to}T12:00:00Z`);
while (d <= end) { out.push(d.toISOString().slice(0, 10)); d.setUTCDate(d.getUTCDate() + 1); }
return out;
}
async function pool(items, fn, n = CONCURRENCY) {
const out = []; let i = 0;
await Promise.all(Array.from({ length: n }, async () => {
while (i < items.length) {
const idx = i; i += 1;
try { out[idx] = await fn(items[idx]); } catch { out[idx] = null; }
}
}));
return out.filter(Boolean);
}
async function loadGame(g) {
const pk = g.gamePk;
const [box, pbp] = await Promise.all([
get(`https://statsapi.mlb.com/api/v1/game/${pk}/boxscore`),
get(`https://statsapi.mlb.com/api/v1/game/${pk}/playByPlay`),
]);
const sides = {};
for (const side of ['home', 'away']) {
const t = box.teams[side];
if (!t) return null;
const arms = (t.pitchers || []).map((id) => {
const pl = t.players[`ID${id}`];
const s = pl && pl.stats && pl.stats.pitching;
if (!s) return null;
return {
id: Number(id),
name: pl.person && pl.person.fullName,
started: Number(s.gamesStarted || 0) === 1,
bf: s.battersFaced == null ? null : Number(s.battersFaced),
outs: ipToOuts(s.inningsPitched),
pitches: s.pitchesThrown == null ? null : Number(s.pitchesThrown),
};
}).filter(Boolean);
sides[side] = { team: t.team && t.team.name, abbr: t.team && t.team.abbreviation, arms };
}
// Every plate appearance in order, with who threw it.
const pas = [];
for (const p of pbp.allPlays || []) {
const m = p.matchup || {}; const a = p.about || {};
if (!m.batter || !m.pitcher) continue;
pas.push({
batter: Number(m.batter.id),
batter_name: m.batter.fullName,
pitcher: Number(m.pitcher.id),
bats: m.batSide && m.batSide.code,
throws: m.pitchHand && m.pitchHand.code,
inning: a.inning,
half: a.halfInning,
idx: a.atBatIndex,
event: p.result && p.result.eventType,
});
}
if (!pas.length) return null;
return {
gamePk: pk,
date: g.officialDate || (g.gameDate || '').slice(0, 10),
venue_id: g.venue && g.venue.id,
home: sides.home, away: sides.away,
pas,
};
}
async function main() {
const dates = datesBetween(FROM, TO);
console.error(`[seq] ${dates.length} dates ${FROM} -> ${TO}`);
const allGames = [];
for (const d of dates) {
try {
const s = await get(`https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}&hydrate=venue`);
for (const day of s.dates || []) {
for (const g of day.games || []) {
if (String(g.status && g.status.detailedState) === 'Final') allGames.push(g);
}
}
} catch { /* absent day */ }
}
console.error(`[seq] ${allGames.length} final games; fetching sequences`);
const games = await pool(allGames, loadGame);
fs.mkdirSync(path.dirname(OUT), { recursive: true });
fs.writeFileSync(OUT, JSON.stringify({ from: FROM, to: TO, games }));
const starters = games.reduce((s, g) =>
s + ['home', 'away'].filter((k) => g[k].arms.some((a) => a.started)).length, 0);
console.log(JSON.stringify({
dates: dates.length,
final_games: allGames.length,
games_loaded: games.length,
starter_games: starters,
plate_appearances: games.reduce((s, g) => s + g.pas.length, 0),
cache: OUT,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+138
View File
@@ -0,0 +1,138 @@
#!/usr/bin/env node
'use strict';
/**
* LINK 1 — does a starter's exit point predict, before the game?
*
* Target: batters faced by the starter, because that is what decides how many of
* a hitter's plate appearances come against him rather than the pen.
*
* ── WHAT A PRE-GAME PREDICTOR CAN AND CANNOT SEE ─────────────────────────
* The order specifies fatigue profile x GAME SCRIPT ("getting hit -> pulled
* early"). Game script is not available when a prop is graded: whether he gets
* hit tonight is the thing we are trying to project, not an input to it. Using
* it would be reading the answer.
*
* So the honest pre-game form of Link 1 is the fatigue and workload half alone —
* the starter's own history, strictly truncated to starts BEFORE the game being
* predicted. That is measured here. The in-game half is a LIVE feature, not a
* grade-time one, and it is recorded as out of scope rather than quietly folded
* in.
*
* Baseline: the league mean batters faced — the naive "a starter goes about six"
* null this must beat to be worth anything.
*
* node scripts/link1-pull-timing.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const pg = require('../src/services/model/predictionGate');
const tl = require('../src/services/model/testLedger');
const { createClient } = require('@supabase/supabase-js');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
/** Starts needed before we will read a pitcher's own history at all. */
const MIN_PRIOR = 3;
/** Shrinkage: how many prior starts before his own mean carries half the weight. */
const STABILIZE = 5;
/** A start at or under this many batters faced is an EARLY EXIT — the edge case. */
const EARLY_BF = 20;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
function main() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
// Every starter-game, in chronological order.
const starts = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const s = (g[side].arms || []).find((a) => a.started);
if (!s || s.bf == null) continue;
starts.push({ date: g.date, gamePk: g.gamePk, pitcher: s.id, name: s.name, bf: s.bf, outs: s.outs, pitches: s.pitches });
}
}
starts.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const leagueMean = mean(starts.map((s) => s.bf));
// POINT-IN-TIME: each start is predicted only from starts strictly before it.
const history = new Map();
const rows = [];
for (const s of starts) {
const prior = history.get(s.pitcher) || [];
if (prior.length >= MIN_PRIOR) {
const own = mean(prior.map((p) => p.bf));
const w = prior.length / (prior.length + STABILIZE);
rows.push({
pitcher: s.pitcher,
name: s.name,
cluster: s.pitcher, // his starts are not independent readings
baseline: leagueMean,
prediction: w * own + (1 - w) * leagueMean,
actual: s.bf,
prior_starts: prior.length,
});
}
history.set(s.pitcher, prior.concat([s]));
}
return { starts, leagueMean, rows };
}
(async () => {
const { starts, leagueMean, rows } = main();
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'starter_bf', archetype: null, interaction: 'link1:pull_timing', target: 'actual_exit' },
]);
cumulative = mc.cumulative_tests;
} catch { /* offline: reported below */ }
const verdict = pg.adjudicate(rows, {
link: 'link1_pull_timing',
loss: 'absolute',
cumulativeTests: cumulative,
});
// Does it find the EARLY EXITS specifically? That is where the edge lives —
// being right about a median start is worth nothing to this chain.
const early = rows.filter((r) => r.actual <= EARLY_BF);
const late = rows.filter((r) => r.actual > EARLY_BF);
const predEarly = rows.filter((r) => r.prediction <= EARLY_BF + 2);
const hitRate = predEarly.length
? predEarly.filter((r) => r.actual <= EARLY_BF).length / predEarly.length : null;
const baseEarlyRate = rows.length ? early.length / rows.length : null;
console.log(JSON.stringify({
link: 'LINK 1 — starter pull timing',
starter_games_total: starts.length,
league_mean_bf: round2(leagueMean),
gated_rows: rows.length,
distinct_pitchers: new Set(rows.map((r) => r.pitcher)).size,
cumulative_tests: cumulative,
verdict,
early_exit_analysis: {
definition: `actual batters faced <= ${EARLY_BF}`,
early_exits: early.length,
normal_starts: late.length,
base_rate_of_early_exit: round4(baseEarlyRate),
flagged_early_by_model: predEarly.length,
of_those_actually_early: round4(hitRate),
lift_over_base_rate: hitRate !== null && baseEarlyRate !== null ? round4(hitRate - baseEarlyRate) : null,
note: 'the chain needs the EARLY tail, not the median start',
},
scope_note: 'GAME SCRIPT is deliberately excluded — whether he gets hit tonight is the thing being projected, not an input available at grade time',
}, null, 2));
process.exit(0);
})();
const round2 = (v) => (v == null ? null : Math.round(v * 100) / 100);
const round4 = (v) => (v == null ? null : Math.round(v * 10000) / 10000);
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env node
'use strict';
/**
* LINK 2 — can we say WHICH arm faces the later plate appearances?
*
* Link 1 proved, so this link is allowed to be attempted at all. It is gated the
* same way: predict the reliever who actually threw a given post-starter plate
* appearance, against a naive baseline, point-in-time.
*
* BASELINE the team's most-used reliever to date — "guess the busiest arm"
* PREDICTION the reliever that team has most often used IN THIS INNING to
* date, which is the cheapest expression of bullpen ROLE
*
* Loss is misclassification: 0 when the named arm actually threw it, 1 otherwise.
*
* ── THE REPLICATION UNIT IS THE BULLPEN, AND THERE ARE THIRTY ────────────
* Bullpen usage is a team-level process — the same manager, the same arms, the
* same roles all season — so errors are correlated within team and the entity
* this prediction rides on is the club. That caps replication at 30 whatever the
* row count, exactly like park geometry and team defence. Reported explicitly
* rather than dissolved into a row count of tens of thousands.
*
* node scripts/link2-reliever-identity.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const pg = require('../src/services/model/predictionGate');
const tl = require('../src/services/model/testLedger');
const { createClient } = require('@supabase/supabase-js');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
/** Pick the key with the highest count; null when there is nothing to pick from. */
function argmax(counter) {
let best = null; let bestN = -1;
for (const [k, v] of counter) if (v > bestN) { best = k; bestN = v; }
return bestN > 0 ? best : null;
}
function build() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
// Point-in-time bullpen histories, accumulated as we walk forward in time.
const overall = new Map(); // team -> Map(pitcherId -> appearances)
const byInning = new Map(); // `team|inning` -> Map(pitcherId -> appearances)
const rows = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const team = g[side].abbr || g[side].team;
if (!team) continue;
const starter = (g[side].arms || []).find((a) => a.started);
if (!starter) continue;
const relievers = new Set((g[side].arms || []).filter((a) => !a.started).map((a) => a.id));
if (!relievers.size) continue;
// This side PITCHES in the opposite half-inning.
const half = side === 'home' ? 'top' : 'bottom';
const post = g.pas.filter((p) => p.half === half && p.pitcher !== starter.id);
for (const pa of post) {
const ov = overall.get(team);
const inn = byInning.get(`${team}|${pa.inning}`);
const basePick = ov ? argmax(ov) : null;
const modelPick = inn ? argmax(inn) : basePick;
// No history yet is honestly unreadable, not a wrong guess.
if (basePick === null || modelPick === null) continue;
rows.push({
cluster: team,
baseline: Number(basePick) === pa.pitcher ? 1 : 0,
prediction: Number(modelPick) === pa.pitcher ? 1 : 0,
actual: 1,
inning: pa.inning,
});
}
// Now fold this game into history — never before predicting from it.
if (!overall.has(team)) overall.set(team, new Map());
const ovm = overall.get(team);
for (const r of relievers) ovm.set(String(r), (ovm.get(String(r)) || 0) + 1);
for (const pa of post) {
const k = `${team}|${pa.inning}`;
if (!byInning.has(k)) byInning.set(k, new Map());
const m = byInning.get(k);
m.set(String(pa.pitcher), (m.get(String(pa.pitcher)) || 0) + 1);
}
}
}
return rows;
}
(async () => {
const rows = build();
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'reliever_identity', archetype: null, interaction: 'link2:reliever_identity', target: 'actual_arm' },
]);
cumulative = mc.cumulative_tests;
} catch { /* offline */ }
const verdict = pg.adjudicate(rows, {
link: 'link2_reliever_identity',
loss: 'absolute',
cumulativeTests: cumulative,
});
const acc = (k) => (rows.length ? rows.filter((r) => r[k] === 1).length / rows.length : null);
console.log(JSON.stringify({
link: 'LINK 2 — reliever identity',
post_starter_plate_appearances: rows.length,
distinct_bullpens: new Set(rows.map((r) => r.cluster)).size,
cumulative_tests: cumulative,
baseline_accuracy: round4(acc('baseline')),
model_accuracy: round4(acc('prediction')),
verdict,
structural_note: 'the entity this prediction rides on is the BULLPEN, and there are 30 — row count cannot create replication that does not exist',
}, null, 2));
process.exit(0);
})();
const round4 = (v) => (v == null ? null : Math.round(v * 10000) / 10000);
+206
View File
@@ -0,0 +1,206 @@
#!/usr/bin/env node
'use strict';
/**
* LINK 2 (coarse grain) — WHICH BULLPEN, not which arm.
*
* Naming the individual reliever failed on merit: 17.2% accuracy, wrong five
* times in six. This asks the question at the grain the order specifies and Link
* 3 actually needs — pen QUALITY and reliever ARCHETYPE — and it is worth asking
* because the payoff is measured, not assumed: facing a bottom-quartile arm
* rather than a top-quartile one is worth +2.57pp of hit rate, larger than the
* whole times-through-the-order effect.
*
* ── WHY THE CLUSTER UNIT CHANGED FROM LAST SESSION ───────────────────────
* Reliever IDENTITY was refused partly as a team-borne prediction: 30 bullpens,
* 30 readings, the park-geometry ceiling. Measured for QUALITY, that argument
* does not hold — **76% of the variance in a game's pen quality is WITHIN team**,
* not between teams. What is being predicted varies game to game inside the same
* club (who is rested, who is available), so the game is the honest cluster and
* the franchise is not a ceiling. Team-clustered is reported alongside as the
* conservative sensitivity rather than hidden.
*
* ── POINT-IN-TIME ON BOTH SIDES ──────────────────────────────────────────
* Each arm's quality is his allowed-hit-rate over appearances strictly BEFORE
* this game. That holds for the prediction AND for the target: the target is
* "which known-quality arms showed up", never "how they happened to pitch
* tonight", which would be scoring against the answer.
*
* node scripts/link2b-pen-quality.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const pg = require('../src/services/model/predictionGate');
const tl = require('../src/services/model/testLedger');
const { createClient } = require('@supabase/supabase-js');
const { knownNumber } = require('../src/utils/known');
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const HIT = new Set(['single', 'double', 'triple', 'home_run']);
const PA = new Set(['single', 'double', 'triple', 'home_run', 'field_out', 'strikeout',
'grounded_into_double_play', 'force_out', 'field_error', 'fielders_choice',
'fielders_choice_out', 'double_play', 'sac_fly', 'pop_out', 'line_out', 'fly_out',
'strikeout_double_play']);
/** Appearances before we will read an arm's quality at all. Below it: abstain. */
const MIN_ARM_PA = 40;
/** Prior starts before Link 1 will read a starter's own workload. */
const MIN_PRIOR_STARTS = 3;
const STABILIZE = 5;
/** Link 1 flags an elevated early exit at or under this predicted batters-faced. */
const EARLY_FLAG_BF = 22;
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/**
* Reliever archetype at the coarse grain, from strikeout rate — the axis that
* separates a power arm from a contact arm and the one Link 3 would condition on.
*/
function archetypeOf(kRate) {
if (kRate === null) return null;
if (kRate >= 0.28) return 'POWER';
if (kRate <= 0.18) return 'CONTACT';
return 'MIDDLE';
}
function build() {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const arm = new Map(); // pid -> { n, h, k } (all prior PAs)
const penHist = new Map(); // team -> [{ quality, k }] per prior game
const startHist = new Map(); // starter id -> [bf]
const rows = [];
for (const g of games) {
for (const side of ['home', 'away']) {
const team = g[side].abbr || g[side].team;
const st = (g[side].arms || []).find((a) => a.started);
if (!team || !st) continue;
const half = side === 'home' ? 'top' : 'bottom';
const pas = g.pas.filter((p) => p.half === half && PA.has(p.event));
const post = pas.filter((p) => p.pitcher !== st.id);
// ── LINK 1, recomputed point-in-time, to define the concentrated subset ──
const priorStarts = startHist.get(st.id) || [];
let predBf = null;
if (priorStarts.length >= MIN_PRIOR_STARTS) {
const w = priorStarts.length / (priorStarts.length + STABILIZE);
predBf = w * mean(priorStarts) + (1 - w) * 21.56; // league mean
}
// ── TARGET: the known quality of the arms that ACTUALLY appeared ──
const faced = [];
for (const p of post) {
const h = arm.get(p.pitcher);
if (!h || h.n < MIN_ARM_PA) continue; // abstain, never 0
faced.push({ q: h.h / h.n, k: h.k / h.n });
}
// ── PREDICTION: this club's own pen, from prior games only ──
const hist = penHist.get(team) || [];
if (faced.length && hist.length >= 5 && predBf !== null) {
const predQ = mean(hist.map((x) => x.quality));
const predK = mean(hist.map((x) => x.k));
rows.push({
team,
gamePk: g.gamePk,
date: g.date,
pred_bf: predBf,
early_flagged: predBf <= EARLY_FLAG_BF,
pred_quality: predQ,
actual_quality: mean(faced.map((f) => f.q)),
pred_archetype: archetypeOf(predK),
actual_archetype: archetypeOf(mean(faced.map((f) => f.k))),
arms_faced: faced.length,
});
}
// Fold this game into history — never before predicting from it.
if (faced.length) {
penHist.set(team, hist.concat([{ quality: mean(faced.map((f) => f.q)), k: mean(faced.map((f) => f.k)) }]));
}
if (st.bf != null) startHist.set(st.id, priorStarts.concat([st.bf]));
for (const p of pas) {
const cur = arm.get(p.pitcher) || { n: 0, h: 0, k: 0 };
cur.n += 1;
cur.h += HIT.has(p.event) ? 1 : 0;
cur.k += p.event === 'strikeout' ? 1 : 0;
arm.set(p.pitcher, cur);
}
}
}
return rows;
}
function gateQuality(rows, leagueQ, cumulative, clusterKey, label) {
return pg.adjudicate(rows.map((r) => ({
cluster: r[clusterKey],
baseline: leagueQ,
prediction: r.pred_quality,
actual: r.actual_quality,
})), { link: label, loss: 'absolute', cumulativeTests: cumulative });
}
(async () => {
const all = build();
const subset = all.filter((r) => r.early_flagged);
const leagueQ = mean(all.map((r) => r.actual_quality));
let cumulative = 1;
try {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY,
{ auth: { persistSession: false } });
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'pen_quality', archetype: null, interaction: 'link2b:pen_quality', target: 'actual_arms' },
{ sport: 'mlb', stat: 'pen_archetype', archetype: null, interaction: 'link2b:pen_archetype', target: 'actual_arms' },
]);
cumulative = mc.cumulative_tests;
} catch { /* offline */ }
// ARCHETYPE grain — misclassification against the arms that actually appeared.
const archRows = subset.filter((r) => r.pred_archetype && r.actual_archetype);
const modal = (() => {
const c = new Map();
for (const r of all) c.set(r.actual_archetype, (c.get(r.actual_archetype) || 0) + 1);
return [...c.entries()].sort((a, b) => b[1] - a[1])[0][0];
})();
const archGate = pg.adjudicate(archRows.map((r) => ({
cluster: r.gamePk,
baseline: r.actual_archetype === modal ? 1 : 0,
prediction: r.actual_archetype === r.pred_archetype ? 1 : 0,
actual: 1,
})), { link: 'link2b_pen_archetype', loss: 'absolute', cumulativeTests: cumulative });
const acc = (rs, k) => (rs.length ? rs.filter((r) => r[k]).length / rs.length : null);
console.log(JSON.stringify({
link: 'LINK 2 (coarse) — pen quality + reliever archetype',
team_games_total: all.length,
concentrated_subset_elevated_early_exit: subset.length,
league_mean_pen_quality: round4(leagueQ),
cumulative_tests: cumulative,
quality_grain: {
on_concentrated_subset: gateQuality(subset, leagueQ, cumulative, 'gamePk', 'link2b_pen_quality_subset'),
sensitivity_team_clustered: gateQuality(subset, leagueQ, cumulative, 'team', 'link2b_pen_quality_teamclust'),
pooled_all_games: gateQuality(all, leagueQ, cumulative, 'gamePk', 'link2b_pen_quality_pooled'),
},
archetype_grain: {
n: archRows.length,
modal_archetype: modal,
baseline_accuracy_guess_modal: round4(acc(archRows.map((r) => ({ x: r.actual_archetype === modal })), 'x')),
model_accuracy: round4(acc(archRows.map((r) => ({ x: r.actual_archetype === r.pred_archetype })), 'x')),
verdict: archGate,
},
cluster_note: '76% of game pen-quality variance is WITHIN team, so the game is the honest cluster; team-clustered reported as the conservative sensitivity',
}, null, 2));
process.exit(0);
})();
const round4 = (v) => (v == null ? null : Math.round(v * 10000) / 10000);
+215
View File
@@ -0,0 +1,215 @@
#!/usr/bin/env node
'use strict';
/**
* lodo-calibration — Phases 2 and 3.
*
* The ≥40 date-cluster floor was factorGate's interval bar for a CAUSAL claim,
* mis-applied to a monotone shrink-to-observed layer. Calibration makes no causal
* claim, consumes no Bonferroni slot, and has a bounded failure mode (it can only
* over- or under-shrink). Its real risk is that the correction is DATE-DRIVEN —
* that one unusual day's offensive environment is doing all the work.
*
* Leave-one-date-out tests exactly that, and it is a harder bar than a cluster
* count: a single date whose removal reverses the improvement, or flips the
* favourite-longshot sign, fails the stat outright.
*
* ── WHAT LODO IS AND IS NOT ──────────────────────────────────────────────
* Refitting on all-but-one date uses dates that follow the held-out one, so this
* is a STABILITY test, not a point-in-time backtest. The point-in-time result is
* separate and already established (fit-past / apply-forward, CI excluding zero
* on hits / TB / RBI). Both are required; neither substitutes for the other.
*
* SUPABASE_URL=... node scripts/lodo-calibration.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** The favourite bucket where the over-prediction concentrates. */
const FAVOURITE_FLOOR = 0.9;
/**
* Minimum rows on a held-out date for that drop to be informative.
*
* POWER-DERIVED AND PRE-COMMITTED (n* = 70). Not chosen here, and not tunable
* from here -- it is imported so the value that decides the verdicts cannot be
* edited alongside them.
*/
const { LODO_TEST, LODO_POWER_FLOOR } = require('../src/services/model/calibrationRegistry');
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select))
.order('id', { ascending: true }).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
const isPreGame = (capturedAt, gameDate) => {
const et = new Date(new Date(capturedAt).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < gameDate || (d === gameDate && et.getUTCHours() < 19);
};
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
// Build the RAW population first so the guard has something to catch.
const raw = [];
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date)) continue;
if (r.refused || knownNumber(r.p_win) === null) continue;
const propKey = [r.game_date, r.stat, r.player_key, r.line].join('|');
raw.push({ propKey, side: r.side, p: knownNumber(r.p_win) });
const prev = picked.get(propKey);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(propKey, r);
}
// GUARD 1 — prove the raw population would have lied, then prove dedup fixes it.
const rawCheck = guards.checkPickedSideDedup(raw);
const pickedRows = [...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'),
side: r.side, p: knownNumber(r.p_win),
}));
guards.assertPickedSideDedup(pickedRows); // throws if dedup failed
const out = {
guard_1_raw_population: { violated: rawCheck.violated, mean_p: rawCheck.mean_p, both_sides_share: rawCheck.both_sides_share },
guard_1_after_dedup: guards.checkPickedSideDedup(pickedRows),
per_stat: {},
};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b));
if (v === null) continue;
const over = v > L;
rows.push({
date: r.game_date,
p: knownNumber(r.p_win),
won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0,
});
}
const dates = [...new Set(rows.map((r) => r.date))].sort();
const full = cal.fitIsotonic(rows.map((r) => ({ p: r.p, won: r.won })));
if (!full) {
out.per_stat[stat] = {
n: rows.length, dates: dates.length,
lodo: 'NOT RUN', decision: 'REFUSE',
reason: `no isotonic map is fittable at n=${rows.length} (needs ${cal.MIN_TOTAL || 200})`,
};
continue;
}
// ── LEAVE ONE DATE OUT, at THIS stat's own informative bar ──
const spec = LODO_TEST[stat];
const MIN_HELD_ROWS = spec ? spec.n_star : Infinity;
const table = [];
for (const d of dates) {
const fit = rows.filter((r) => r.date !== d);
const held = rows.filter((r) => r.date === d);
if (held.length < MIN_HELD_ROWS) {
table.push({ dropped: d, held_n: held.length, verdict: 'UNINFORMATIVE', reason: 'too few rows on this date' });
continue;
}
const map = cal.fitIsotonic(fit.map((r) => ({ p: r.p, won: r.won })));
const applied = guards.applyOrRefuse(map, held, cal.applyIsotonic);
if (!applied.ok) {
table.push({ dropped: d, held_n: held.length, verdict: 'UNINFORMATIVE', reason: applied.reason });
continue;
}
const ys = applied.rows.map((r) => r.won);
const bRaw = guards.safeBrier(applied.rows.map((r) => r.p), ys);
const bCal = guards.safeBrier(applied.rows.map((r) => r.pc), ys);
if (bRaw === null || bCal === null) {
table.push({ dropped: d, held_n: held.length, verdict: 'UNINFORMATIVE', reason: 'a null reached the metric' });
continue;
}
const fav = applied.rows.filter((r) => r.p >= FAVOURITE_FLOOR);
const favBias = fav.length >= 5 ? mean(fav.map((r) => r.p)) - mean(fav.map((r) => r.won)) : null;
table.push({
dropped: d,
held_n: held.length,
brier_delta: round4(bCal - bRaw),
improves: bCal < bRaw,
favourite_n: fav.length,
favourite_bias: favBias === null ? null : round4(favBias),
favourite_sign_holds: favBias === null ? null : favBias > 0,
verdict: bCal < bRaw ? 'holds' : 'REVERSES',
});
}
const informative = table.filter((t) => t.verdict !== 'UNINFORMATIVE');
const reversals = informative.filter((t) => t.verdict === 'REVERSES');
// THE DECISION RULE IS BINOMIAL, not zero-tolerance. Under stability each
// informative drop reverses with prob Phi(-k), so demanding zero reversals
// failed stable stats roughly half the time.
const cutoff = spec ? spec.cutoff : 0;
const exceedsCutoff = reversals.length > cutoff;
// AND THE TEST MUST BE ABLE TO FAIL. Below the power floor it cannot, so it
// cannot pass either -- "could not test" must never read as "passed".
const underpowered = !spec || spec.power < LODO_POWER_FLOOR;
const verdict = underpowered ? 'UNTESTABLE_BY_LODO' : (exceedsCutoff ? 'FAIL' : 'PASS');
const passes = verdict === 'PASS';
out.per_stat[stat] = {
n: rows.length,
dates: dates.length,
lodo_table: table,
informative_drops: informative.length,
brier_reversals: informative.filter((t) => t.verdict === 'REVERSES').length,
favourite_sign_flips: informative.filter((t) => t.favourite_sign_holds === false).length,
favourite_sign_untested: informative.filter((t) => t.favourite_sign_holds === null).length,
n_star: spec ? spec.n_star : null,
cutoff,
reversal_count: reversals.length,
reversing_dates: reversals.map((t) => ({ date: t.dropped, held_n: t.held_n, delta: t.brier_delta })),
test_power: spec ? spec.power : null,
lodo: verdict,
reason: underpowered
? `power ${spec ? spec.power : 0} < floor ${LODO_POWER_FLOOR} — this test would miss a real date-driven failure more than nine times in ten, so it can neither pass nor fail the stat`
: (exceedsCutoff
? `${reversals.length} reversals among ${informative.length} informative drops exceeds the cutoff of ${cutoff}`
: `${reversals.length} reversals among ${informative.length} informative drops is within the cutoff of ${cutoff}`),
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
}
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
main().catch((e) => { console.error(e); process.exit(1); });
+61
View File
@@ -0,0 +1,61 @@
-- matchup-axis-holdout.sql — per-axis proof, RUN WHEN n IS ADEQUATE.
--
-- Filtered to MATCHUP-CARRYING rows only. Including untouched rows would
-- dilute the comparison with rows where challenger === champion BY
-- CONSTRUCTION, biasing toward a false positive (the trap opportunity_drift
-- established).
--
-- MATCHUP'S OWN CONTRIBUTION IS KEPT VISIBLE, not just the combined challenger:
-- arch-v1 composes archetype + environment + opportunity + matchup into ONE
-- p_win_challenger, so a combined-only view cannot tell which axis earned the
-- movement. `matchup_nudge` is pulled out of the adjustments array so the axis
-- can be judged on its own terms and, if it is the one dragging, shelved alone.
--
-- BOTH reliability AND resolution must improve for the axis to promote.
--
-- CONTAMINATION EXCLUSION (2026-08-02, MANDATORY). Rows whose price/book/takeable
-- were stamped from a NON-TAKEABLE book (DFS / offshore / exchange) between
-- 2026-08-01 and the write-path fix are tagged `quarantine_reason LIKE
-- 'nontakeable_book%'`. They are EXCLUDED here and must never be pooled with
-- clean rows: their locked price -- and therefore the `takeable` flag computed
-- from it -- describes a market you could not have bet.
with rows_ as (
select
l.game_date, l.id,
l.p_win::numeric champ,
l.p_win_challenger::numeric chal,
(l.outcome='hit')::int won,
(select (a->>'nudge')::numeric
from jsonb_array_elements(l.challenger_adjustments) a
where a->>'axis' = 'matchup' limit 1) matchup_nudge,
(select a->>'tier'
from jsonb_array_elements(l.challenger_adjustments) a
where a->>'axis' = 'matchup' limit 1) matchup_tier
from public.ledger_entries l
where l.sport='mlb' and l.user_id is null
and (l.quarantine_reason is null or l.quarantine_reason not like 'nontakeable_book%')
and l.outcome in ('hit','miss')
and l.p_win is not null and l.p_win_challenger is not null
and l.challenger_adjustments::text like '%matchup%'
),
split as (
select *, case when ntile(2) over (order by game_date, id) = 1 then 'train' else 'holdout' end split
from rows_
),
b_champ as (select split, width_bucket(champ,0,1,10) bkt, count(*) n, avg(champ) pred, avg(won::numeric) actual from split group by 1,2),
b_chal as (select split, width_bucket(chal ,0,1,10) bkt, count(*) n, avg(chal ) pred, avg(won::numeric) actual from split group by 1,2)
select
s.split,
count(*) n,
count(distinct s.matchup_tier) tiers,
round(avg(abs(s.matchup_nudge))::numeric,4) mean_abs_matchup_nudge,
round((select sum(n*abs(pred-actual))/nullif(sum(n),0) from b_champ c where c.split=s.split),4) reliability_champion,
round((select sum(n*abs(pred-actual))/nullif(sum(n),0) from b_chal c where c.split=s.split),4) reliability_challenger,
round(corr(s.champ, s.won::numeric)::numeric,4) resolution_champion,
round(corr(s.chal , s.won::numeric)::numeric,4) resolution_challenger,
-- does the matchup nudge ITSELF point the right way?
round(corr(s.matchup_nudge, s.won::numeric)::numeric,4) matchup_nudge_vs_outcome,
round(avg(s.won::numeric),4) base_rate
from split s group by s.split order by s.split desc;
+166
View File
@@ -0,0 +1,166 @@
'use strict';
/**
* measure-book-spread — Book Comparison order, Phase 2 (GATES THE CROWN).
*
* Reads the snapshot-locked `bookprices:{sport}` store (Phase 1) — falling back
* to the transient `odds:{sport}:{utcDate}` cache — and reports, PER SPORT, the
* best-vs-worst PRICE spread among books posting the SAME line for the same
* side:
* - median + distribution + tail, in American cents AND implied-prob points
* - how often the spread is exactly zero
* - book-count histogram per prop
* - pinnacle presence (captured, not built on — this order)
*
* PRE-REGISTERED CROWN THRESHOLD (do NOT lower it to make the crown appear):
* the crown ships for a sport ONLY if median same-line spread
* >= 8 American cents OR >= 2.0 implied-probability points.
*
* Never pools sports. Reports n + effective sample on every figure.
*
* Redis runs degraded locally (no live data) → this exits 0 cleanly rather than
* hanging on a reconnect timer (the verify-grade-range.js precedent). Run it
* post-deploy against prod Redis, after inducing a snapshot.
*
* node scripts/measure-book-spread.js [sport ...] (default: mlb wnba nba)
*/
const SPORTS = process.argv.slice(2).filter(Boolean);
const DEFAULT_SPORTS = ['mlb', 'wnba', 'nba', 'soccer'];
const CROWN_CENTS = 8;
const CROWN_PROB_PTS = 2.0;
/** American → implied probability (0..1). Includes the vig. */
function impliedProb(a) {
if (a == null || !Number.isFinite(Number(a)) || Number(a) === 0) return null;
const n = Number(a);
return n > 0 ? 100 / (n + 100) : Math.abs(n) / (Math.abs(n) + 100);
}
function median(xs) {
if (!xs.length) return null;
const s = [...xs].sort((a, b) => a - b);
const m = Math.floor(s.length / 2);
return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
}
function pct(xs, p) {
if (!xs.length) return null;
const s = [...xs].sort((a, b) => a - b);
return s[Math.min(s.length - 1, Math.floor((p / 100) * s.length))];
}
function entriesFrom(store, oddsCache) {
// Phase-1 store shape: { props: [{ player, stat_type, books:[{book,line,over_odds,under_odds}] }] }
if (store && Array.isArray(store.props)) return store.props;
// Fallback: group the flat odds cache the same way.
const flat = oddsCache && Array.isArray(oddsCache.props) ? oddsCache.props : [];
const by = new Map();
for (const p of flat) {
if (!p || !p.player || !p.stat_type || p.line == null || !p.book) continue;
const k = `${p.player}|${p.stat_type}`;
if (!by.has(k)) by.set(k, { player: p.player, stat_type: p.stat_type, books: [] });
by.get(k).books.push({ book: p.book, line: p.line, over_odds: p.over_odds, under_odds: p.under_odds });
}
return [...by.values()];
}
function measureSport(entries) {
const centsSpreads = [];
const probSpreads = [];
const bookCountHist = {};
let sharedLineProps = 0;
let zeroSpread = 0;
let pinnacleRows = 0;
let totalProps = 0;
for (const e of entries) {
totalProps += 1;
const books = e.books || [];
if (books.some((b) => b.book === 'pinnacle')) pinnacleRows += 1;
const nBooks = new Set(books.map((b) => b.book)).size;
bookCountHist[nBooks] = (bookCountHist[nBooks] || 0) + 1;
// Group this prop's book rows by line; a shared line = ≥2 books at one line.
const byLine = {};
for (const b of books) {
const L = String(b.line);
(byLine[L] = byLine[L] || []).push(b);
}
let contributed = false;
for (const rows of Object.values(byLine)) {
const distinctBooks = new Set(rows.map((r) => r.book));
if (distinctBooks.size < 2) continue;
for (const side of ['over_odds', 'under_odds']) {
const prices = rows.map((r) => r[side]).filter((v) => v != null && Number.isFinite(Number(v))).map(Number);
if (prices.length < 2) continue;
// Best price for a bettor = highest implied payout = LOWEST implied prob.
const probs = prices.map(impliedProb).filter((v) => v != null);
if (probs.length < 2) continue;
const probSpread = (Math.max(...probs) - Math.min(...probs)) * 100; // points
probSpreads.push(+probSpread.toFixed(3));
// American cents: meaningful when same-sign; use nominal max-min.
const centSpread = Math.max(...prices) - Math.min(...prices);
centsSpreads.push(Math.abs(centSpread));
if (probSpread < 1e-9) zeroSpread += 1;
contributed = true;
}
}
if (contributed) sharedLineProps += 1;
}
const nEff = probSpreads.length; // side-level shared-line comparisons
const verdictCents = median(centsSpreads);
const verdictProb = median(probSpreads);
const crownShips = nEff > 0 && ((verdictCents != null && verdictCents >= CROWN_CENTS) || (verdictProb != null && verdictProb >= CROWN_PROB_PTS));
return {
totalProps,
sharedLineProps,
nEff,
zeroSpread,
pctZero: nEff ? +(100 * zeroSpread / nEff).toFixed(1) : null,
bookCountHist,
pinnacleRows,
cents: { median: verdictCents, p75: pct(centsSpreads, 75), p90: pct(centsSpreads, 90), max: centsSpreads.length ? Math.max(...centsSpreads) : null },
prob: { median: verdictProb, p75: pct(probSpreads, 75), p90: pct(probSpreads, 90), max: probSpreads.length ? Math.max(...probSpreads) : null },
crownShips,
};
}
async function main() {
let cacheGet;
try {
({ cacheGet } = require('../src/utils/redis'));
} catch (e) {
console.log('[measure] redis util unavailable — nothing to measure.');
process.exit(0);
}
const sports = SPORTS.length ? SPORTS : DEFAULT_SPORTS;
const utcDate = new Date().toISOString().split('T')[0];
const report = {};
for (const sp of sports) {
let store = null; let oddsCache = null;
try { store = await cacheGet(`bookprices:${sp}`); } catch { /* degraded */ }
if (!store) { try { oddsCache = (await cacheGet(`odds:${sp}:${utcDate}`)) || (await cacheGet(`odds:${sp}`)); } catch { /* degraded */ } }
const entries = entriesFrom(store, oddsCache);
report[sp] = { source: store ? 'bookprices' : (oddsCache ? 'odds-cache' : 'none'), ...measureSport(entries) };
}
console.log('\n=== BOOK-PRICE SPREAD (Phase 2) — never pooled ===');
for (const sp of sports) {
const r = report[sp];
console.log(`\n--- ${sp.toUpperCase()} (source: ${r.source}) ---`);
if (r.source === 'none' || r.totalProps === 0) { console.log(' no captured data (run post-deploy after a snapshot)'); continue; }
console.log(` props: ${r.totalProps} | with ≥2 books at a shared line: ${r.sharedLineProps} | side-level comparisons n=${r.nEff}`);
console.log(` book-count histogram: ${JSON.stringify(r.bookCountHist)}`);
console.log(` pinnacle present on: ${r.pinnacleRows} props`);
console.log(` spread exactly zero: ${r.zeroSpread}/${r.nEff} (${r.pctZero}%)`);
console.log(` American cents — median ${r.cents.median} | p75 ${r.cents.p75} | p90 ${r.cents.p90} | max ${r.cents.max}`);
console.log(` implied-prob pt — median ${r.prob.median} | p75 ${r.prob.p75} | p90 ${r.prob.p90} | max ${r.prob.max}`);
console.log(` CROWN THRESHOLD (median ≥${CROWN_CENTS}c OR ≥${CROWN_PROB_PTS}pp): ${r.crownShips ? 'MET → crown MAY ship' : 'NOT met → crown does NOT ship'}`);
}
console.log('\n(JSON) ' + JSON.stringify(report));
process.exit(0);
}
main().catch((e) => { console.error('[measure] failed:', e.message); process.exit(0); });
+58
View File
@@ -0,0 +1,58 @@
-- opportunity-axis-holdout.sql — the Step 4 proof, RUN WHEN n IS ADEQUATE.
--
-- Cannot run yet, by construction: the axis went live 2026-08-01 and its first
-- rows carry game_date 2026-08-01 (games not yet played). Settled rows carrying
-- the axis: 0. The earliest possible run is the next morning settle pass, and a
-- defensible n is several days out at ~140 opportunity-axis rows per snapshot.
--
-- BOTH reliability AND resolution must improve for the axis to promote. One or
-- neither => SHELVE and record why.
--
-- reliability = n-weighted mean |predicted - actual| across deciles. Bucket
-- FIRST: mean|p - outcome| on 0/1 rows is noise-dominated individual error, not
-- calibration.
--
-- CONTAMINATION EXCLUSION (2026-08-02, MANDATORY). Rows whose price/book/takeable
-- were stamped from a NON-TAKEABLE book (DFS / offshore / exchange) between
-- 2026-08-01 and the write-path fix are tagged `quarantine_reason LIKE
-- 'nontakeable_book%'`. They are EXCLUDED here and must never be pooled with
-- clean rows: their locked price -- and therefore the `takeable` flag computed
-- from it -- describes a market you could not have bet.
with rows_ as (
select game_date, id,
p_win::numeric champ,
p_win_challenger::numeric chal,
(outcome='hit')::int won
from public.ledger_entries
where sport='mlb' and user_id is null
and (quarantine_reason is null or quarantine_reason not like 'nontakeable_book%')
and outcome in ('hit','miss')
and p_win is not null and p_win_challenger is not null
-- ONLY rows the opportunity axis actually touched. Including untouched rows
-- would dilute the comparison with rows where challenger === champion by
-- construction, and make a null result look like a small positive one.
and challenger_adjustments::text like '%opportunity%'
),
split as (
select *, case when ntile(2) over (order by game_date, id) = 1 then 'train' else 'holdout' end split
from rows_
),
b_champ as (
select split, width_bucket(champ, 0.0, 1.0, 10) bkt, count(*) n, avg(champ) pred, avg(won::numeric) actual
from split group by 1,2),
b_chal as (
select split, width_bucket(chal, 0.0, 1.0, 10) bkt, count(*) n, avg(chal) pred, avg(won::numeric) actual
from split group by 1,2)
select
s.split,
count(*) n,
round((select sum(n*abs(pred-actual))/nullif(sum(n),0) from b_champ c where c.split=s.split),4) reliability_champion,
round((select sum(n*abs(pred-actual))/nullif(sum(n),0) from b_chal c where c.split=s.split),4) reliability_challenger,
round(corr(s.champ, s.won::numeric)::numeric,4) resolution_champion,
round(corr(s.chal, s.won::numeric)::numeric,4) resolution_challenger,
round(avg(s.won::numeric),4) base_rate
from split s
group by s.split
order by s.split desc;
+360
View File
@@ -0,0 +1,360 @@
#!/usr/bin/env node
'use strict';
/**
* pitcher-prove-k — STRIKEOUTS through the both-ways gate.
*
* Same bar as everything else: solo pass as the control, theory-first
* interactions each measured against their own components, gate at n>=500 /
* |r|>=0.15 / p<0.05 / Bonferroni, then head-to-head vs the counter.
*
* THE THEORIZED SIGNAL-CARRIER is `stuff x opposing-lineup K-rate`. An elite
* strikeout arm against a contact lineup that never whiffs is a different bet
* from the same arm against a three-true-outcomes lineup, and neither side says
* it alone — the pitcher analogue of the batter model's contact-quality term.
* The lineup rate is built from the OPPOSING TEAM'S OWN BATTERS (roster join to
* their statcast K%), not from a league constant, or the interaction would be a
* relabelled copy of the pitcher's own rate.
*
* VALIDITY: statcast_aggregates still carries one as-of date (2026-08-03) and
* `statcast_history` has one day, so there is no point-in-time window yet.
* Results here are CONTAMINATED / DIRECTIONAL and are not gate verdicts.
*
* SUPABASE_URL=... node scripts/pitcher-prove-k.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const cv = require('../src/services/model/correlateValidator');
const pe = require('../src/services/model/pitcherEngine');
const sk = require('../src/services/model/skillProjection');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownRate, knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const mean = (a) => (a.length ? a.reduce((x, y) => x + y, 0) / a.length : null);
const brier = (ps, ys) => (ps.length ? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
function olsResiduals(y, Xcols) {
const n = y.length; const p = Xcols.length + 1;
const X = []; for (let i = 0; i < n; i += 1) { const row = [1]; for (const c of Xcols) row.push(c[i]); X.push(row); }
const XtX = Array.from({ length: p }, () => new Array(p).fill(0)); const Xty = new Array(p).fill(0);
for (let i = 0; i < n; i += 1) for (let a = 0; a < p; a += 1) {
Xty[a] += X[i][a] * y[i];
for (let b = 0; b < p; b += 1) XtX[a][b] += X[i][a] * X[i][b];
}
const M = XtX.map((row, i) => [...row, Xty[i]]);
for (let col = 0; col < p; col += 1) {
let piv = col;
for (let r = col + 1; r < p; r += 1) if (Math.abs(M[r][col]) > Math.abs(M[piv][col])) piv = r;
if (Math.abs(M[piv][col]) < 1e-12) return null;
[M[col], M[piv]] = [M[piv], M[col]];
const d = M[col][col];
for (let k = col; k <= p; k += 1) M[col][k] /= d;
for (let r = 0; r < p; r += 1) { if (r === col) continue; const f = M[r][col]; for (let k = col; k <= p; k += 1) M[r][k] -= f * M[col][k]; }
}
const beta = M.map((row) => row[p]);
return y.map((v, i) => v - X[i].reduce((s, xv, j) => s + xv * beta[j], 0));
}
function partialCorr(a, b, ctrl) {
for (let i = 0; i < ctrl.length; i += 1) for (let j = i + 1; j < ctrl.length; j += 1) {
const rr = cv.pearson(ctrl[i], ctrl[j]).r;
if (rr !== null && Math.abs(rr) > 0.999) return null; // same variable twice
}
const ra = olsResiduals(a, ctrl); const rb = olsResiduals(b, ctrl);
if (!ra || !rb) return null;
return cv.pearson(ra, rb).r;
}
function makeRnd(seed) { let s = seed >>> 0; return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; }; }
function bootstrapDiff(rows, kA, kB, iters = 4000, seed = 20260805) {
if (rows.length < 30) return null;
const rnd = makeRnd(seed); const n = rows.length; const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const b = [];
for (let i = 0; i < n; i += 1) { const r = rows[Math.floor(rnd() * n)]; ys.push(r.won); a.push(r[kA]); b.push(r[kB]); }
const ca = cv.pearson(a, ys).r; const cb = cv.pearson(b, ys).r;
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (pp) => r4(diffs[Math.floor(pp * (diffs.length - 1))]);
const ci = [q(0.025), q(0.975)];
return { point: r4(cv.pearson(rows.map((r) => r[kA]), rows.map((r) => r.won)).r - cv.pearson(rows.map((r) => r[kB]), rows.map((r) => r.won)).r), ci95: ci, ci_excludes_zero: ci[0] > 0 || ci[1] < 0 };
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data); if (data.length < PAGE) break;
}
return out;
}
const SOLO = ['pitcher_whiff_pct', 'pitcher_k_pct', 'pitcher_chase_pct', 'pitcher_gb_pct',
'pitcher_arm_angle', 'opposing_lineup_k_rate'];
const INTERACTIONS = [
{
key: 'stuff_x_lineup_k_rate',
components: ['pitcher_whiff_pct', 'opposing_lineup_k_rate'],
mechanism: 'THE theorized carrier. Strikeouts need a pitcher who can miss bats AND a lineup that can be missed. An elite arm against a contact lineup and a modest arm against a whiff-prone one can produce the same count, so neither factor alone orders the props — the product should.',
build: (r) => r.pitcher_whiff_pct * r.opposing_lineup_k_rate,
},
{
key: 'stuff_x_power_archetype',
components: ['pitcher_whiff_pct', 'archetype_flame'],
mechanism: 'ARCHETYPE-CONDITIONAL. Stuff should govern strikeouts more for a power arm than for a finesse arm, whose Ks come from chase and sequencing. Discipline 2 as a testable claim, with a categorical conditioner independent of whiff by construction.',
build: (r) => r.pitcher_whiff_pct * r.archetype_flame,
},
{
key: 'chase_x_lineup_k_rate',
components: ['pitcher_chase_pct', 'opposing_lineup_k_rate'],
mechanism: 'The finesse channel: expanding the zone only works against a lineup that will chase. Same shape as the stuff term, different mechanism, so it is tested separately rather than assumed to be the same effect.',
build: (r) => r.pitcher_chase_pct * r.opposing_lineup_k_rate,
},
];
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10);
const pitchByKey = new Map(); const batterByKey = new Map();
for (const r of statcast) {
const prof = sk.fromStatcastRow(r);
if (r.role === 'pitcher' && r.player_key) pitchByKey.set(r.player_key, prof);
if (r.role === 'batter' && r.player_key) batterByKey.set(r.player_key, prof);
}
const led = await page(sb, 'ledger_entries',
'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'strikeouts')
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// ── OPPOSING LINEUP K-RATE, from the opposing team's OWN batters ────────
// Roster join, not a league constant: a constant would make the interaction a
// rescaled copy of the pitcher's own rate and guarantee a false "redundant".
const { nameKey } = require('../src/utils/playerName');
const teamKRate = new Map();
async function lineupKFor(teamName) {
if (!teamName) return null;
if (teamKRate.has(teamName)) return teamKRate.get(teamName);
let val = null;
try {
// The game log gives a full team NAME ("Cincinnati Reds"); resolveTeam
// wants an ABBREVIATION. Passing the name straight through silently
// resolved nothing and produced 0% lineup coverage on the first run — the
// theorized signal-carrier was not failing, it was never being tested.
const { NAME_TO_ABBR } = require('../src/services/environmentContext');
const abbr = /^[A-Z]{2,3}$/.test(String(teamName).trim())
? String(teamName).trim().toUpperCase()
: NAME_TO_ABBR[String(teamName).toLowerCase()];
if (!abbr) { teamKRate.set(teamName, null); return null; }
const team = await mlb.resolveTeam(abbr);
const roster = team && team.id ? await mlb.getTeamRoster(team.id) : null;
// PA-WEIGHTED, not a flat roster average. An unweighted mean counts a
// 12-PA September call-up the same as an everyday starter, which is not
// the lineup a pitcher faces. Weighting by each batter's own sample_pa is
// the closest honest approximation of "who actually bats" from data we
// already hold — and it needs no new sourcing at all.
let wSum = 0; let wK = 0; let counted = 0;
for (const p of roster || []) {
const prof = batterByKey.get(nameKey(p.name || p.fullName || ''));
if (!prof) continue;
const k = knownRate(prof.k_pct);
const pa = knownRate(prof.sample_pa);
if (k === null) continue;
const w = pa === null ? 0 : pa; // no PA read -> contributes nothing
if (w <= 0) continue;
wSum += w; wK += w * k; counted += 1;
}
if (counted >= 5 && wSum > 0) val = wK / wSum;
} catch { val = null; }
teamKRate.set(teamName, val);
return val;
}
// The opponent a pitcher faced on a date, from his own game log.
const oppBy = new Map();
const names = new Map();
for (const r of clean) if (!names.has(r.player_key)) names.set(r.player_key, r.player_name);
for (const [key, name] of names) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id, undefined, 'pitching');
for (const g of log || []) if (g && g.date && g.opponent) oppBy.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
} catch { /* no log → no lineup term */ }
}
const rows = [];
for (const r of clean) {
const prof = pitchByKey.get(r.player_key);
if (!prof) continue;
const opp = oppBy.get(`${r.player_key}|${r.game_date}`) || null;
const lineupK = opp ? await lineupKFor(opp) : null;
const cls = pe.classifyPitcher(prof);
const arch = cls ? cls.primary : null;
const under = String(r.side).toLowerCase() === 'under';
const won = r.outcome === 'hit' ? 1 : 0;
const champ = Number(r.p_win);
const proj = pe.projectStrikeouts({
pitcher: prof, lineupKRate: lineupK, archetype: arch,
role: 'starter', line: Number(r.line),
});
// The SAME model with the lineup term switched off, so the term's
// contribution is isolated rather than inferred.
const projNo = pe.projectStrikeouts({
pitcher: prof, lineupKRate: null, archetype: arch,
role: 'starter', line: Number(r.line),
});
rows.push({
won, champ, residual: won - champ,
pitch: proj ? (under ? 1 - proj.p_over_line : proj.p_over_line) : null,
pitch_nolineup: projNo ? (under ? 1 - projNo.p_over_line : projNo.p_over_line) : null,
lineup_applied: !!lineupK,
archetype: arch,
archetype_flame: arch == null ? null : (arch === 'FLAME' ? 1 : 0),
pitcher_whiff_pct: knownRate(prof.whiff_pct),
pitcher_k_pct: knownRate(prof.k_pct),
pitcher_chase_pct: knownRate(prof.chase_pct),
pitcher_gb_pct: knownRate(prof.gb_pct),
pitcher_arm_angle: knownRate(prof.arm_angle),
opposing_lineup_k_rate: lineupK,
});
}
// ── CUMULATIVE BONFERRONI ─────────────────────────────────────────────
// The denominator is every DISTINCT hypothesis this programme has tested, not
// this run's. A per-session count gives each new order a fresh, generous alpha
// and lets the false-positive rate compound silently.
const tl = require('../src/services/model/testLedger');
const mcStore = tl.supabaseStore(sb);
const mc = await tl.recordAndCount(mcStore, [
...SOLO.map((f) => ({ sport: 'mlb', stat: 'strikeouts', archetype: null, interaction: `solo:${f}`, target: 'counter_residual' })),
...INTERACTIONS.map((x) => ({ sport: 'mlb', stat: 'strikeouts', archetype: null, interaction: x.key, target: 'counter_residual' })),
]);
const TESTS = mc.cumulative_tests;
const complete = (keys) => rows.filter((r) => keys.every((k) => knownNumber(r[k]) !== null));
const solo = {};
for (const f of SOLO) {
const rs = complete([f]);
solo[f] = {
n: rs.length,
vs_outcome: cv.validateFactor(rs.map((r) => r[f]), rs.map((r) => r.won), TESTS),
vs_counter_residual: cv.validateFactor(rs.map((r) => r[f]), rs.map((r) => r.residual), TESTS),
};
}
const interactions = {};
for (const ix of INTERACTIONS) {
const rs = complete(ix.components);
if (rs.length < 20) { interactions[ix.key] = { mechanism: ix.mechanism, n: rs.length, verdict: 'UNTESTABLE — no common sample' }; continue; }
const I = rs.map(ix.build); const Y = rs.map((r) => r.residual);
const ctrl = ix.components.map((k) => rs.map((r) => r[k]));
const gate = cv.validateFactor(I, Y, TESTS);
const incr = partialCorr(I, Y, ctrl);
const parts = ix.components.map((k) => ({ feature: k, r: r4(cv.pearson(rs.map((r) => r[k]), Y).r) }));
const best = Math.max(...parts.map((p) => Math.abs(p.r ?? 0)));
interactions[ix.key] = {
mechanism: ix.mechanism, components: ix.components, n: rs.length,
raw_r_vs_residual: gate.pearson_r,
gate: { validated: gate.validated, reason: gate.reason, underpowered: !!gate.underpowered, rows_needed: gate.rows_needed ?? null },
component_solo_r: parts, best_component_abs_r: r4(best),
INCREMENTAL_partial_r: r4(incr),
adds_over_components: incr !== null && Math.abs(incr) > best,
verdict: incr === null ? 'UNTESTABLE — collinear controls'
: (gate.validated && Math.abs(incr) >= 0.15) ? 'PASSES-AND-ADDS'
: gate.validated ? 'PASSES-BUT-REDUNDANT'
: rs.length < 500 ? 'UNDERPOWERED — n below the gate' : 'FAILS',
};
}
// ── WITHIN-ARCHETYPE: the order's sharper hypothesis ────────────────────
// The interaction should matter MORE for finesse arms (SCALPEL/SINKER), whose
// strikeouts need a lineup that will chase or can be beaten, than for power
// arms (FLAME) whose stuff whiffs regardless of who is standing there. Pooling
// the two would average a real conditional effect toward zero — which is
// exactly the failure mode "test within archetype" exists to prevent.
const strata = {};
for (const [label, pred] of [
['FLAME (power)', (r) => r.archetype === 'FLAME'],
['non-FLAME (finesse/contact)', (r) => r.archetype && r.archetype !== 'FLAME'],
]) {
const rs = rows.filter((r) => pred(r)
&& knownNumber(r.pitcher_whiff_pct) !== null
&& knownNumber(r.opposing_lineup_k_rate) !== null);
if (rs.length < 15) { strata[label] = { n: rs.length, verdict: 'UNTESTABLE — stratum too thin' }; continue; }
const I = rs.map((r) => r.pitcher_whiff_pct * r.opposing_lineup_k_rate);
const Y = rs.map((r) => r.residual);
const ctrl = [rs.map((r) => r.pitcher_whiff_pct), rs.map((r) => r.opposing_lineup_k_rate)];
const incr = partialCorr(I, Y, ctrl);
const parts = [
{ feature: 'pitcher_whiff_pct', r: r4(cv.pearson(rs.map((r) => r.pitcher_whiff_pct), Y).r) },
{ feature: 'opposing_lineup_k_rate', r: r4(cv.pearson(rs.map((r) => r.opposing_lineup_k_rate), Y).r) },
];
const best = Math.max(...parts.map((p) => Math.abs(p.r ?? 0)));
strata[label] = {
n: rs.length,
raw_r_vs_residual: r4(cv.pearson(I, Y).r),
lineup_solo_r: parts[1].r,
component_solo_r: parts,
best_component_abs_r: r4(best),
INCREMENTAL_partial_r: r4(incr),
adds_over_components: incr !== null && Math.abs(incr) > best,
gate: cv.validateFactor(I, Y, TESTS),
verdict: incr === null ? 'UNTESTABLE — collinear'
: rs.length < 500 ? 'UNDERPOWERED — n below the gate'
: (Math.abs(incr) >= 0.15 ? 'ADDS' : 'REDUNDANT'),
};
}
const h2h = rows.filter((r) => r.pitch != null);
const ys = h2h.map((r) => r.won);
const bs = bootstrapDiff(h2h, 'pitch', 'champ');
const bsNoLineup = bootstrapDiff(h2h.filter((r) => r.pitch_nolineup != null), 'pitch_nolineup', 'champ');
console.log(JSON.stringify({
stat: 'strikeouts',
VALIDITY: `CONTAMINATED / DIRECTIONAL — statcast carries one as-of date (${freezeDate}); statcast_history has no window yet. NOT gate verdicts.`,
rows_scored: rows.length,
lineup_coverage: r4(mean(rows.map((r) => (r.lineup_applied ? 1 : 0)))),
archetype_mix: rows.reduce((a, r) => { const k = r.archetype || 'unclassified'; a[k] = (a[k] || 0) + 1; return a; }, {}),
gate_spec: cv.VALIDATION_REQUIREMENTS,
bonferroni_tests: TESTS,
multiple_comparisons: { ...mc, note: 'cumulative across the programme lifetime, not this session' },
step1_solo_baseline: solo,
step3_interactions: interactions,
step3b_within_archetype_carrier: strata,
step4_vs_counter: {
n: h2h.length,
base_rate: r4(mean(ys)),
resolution: { pitch_v1: r4(cv.pearson(h2h.map((r) => r.pitch), ys).r), counter: r4(cv.pearson(h2h.map((r) => r.champ), ys).r) },
brier: { pitch_v1: r4(brier(h2h.map((r) => r.pitch), ys)), counter: r4(brier(h2h.map((r) => r.champ), ys)) },
delta: bs,
// Isolating the lineup term: does including it help or hurt?
without_lineup_term: {
resolution: r4(cv.pearson(h2h.filter((r) => r.pitch_nolineup != null).map((r) => r.pitch_nolineup),
h2h.filter((r) => r.pitch_nolineup != null).map((r) => r.won)).r),
delta_vs_counter: bsNoLineup,
},
verdict: !bs ? 'N-BLOCKED — too few rows to bootstrap'
: (bs.ci_excludes_zero && bs.point > 0) ? 'BEATS THE COUNTER'
: (bs.ci_excludes_zero && bs.point < 0) ? 'LOSES to the counter' : 'INCONCLUSIVE',
},
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+311
View File
@@ -0,0 +1,311 @@
#!/usr/bin/env node
'use strict';
/**
* prove-hit-factors — does the hit grade read tonight's game, or say "he's due"?
*
* Each factor is conditioned against the player's OWN base rate and put through
* the two-part gate: it must MOVE the prediction and the moved prediction must
* be MORE ACCURATE out-of-sample. Movement alone is THEATER — a grade that
* swings on park and platoon looks like it read the matchup, and a user cannot
* tell the difference from outside.
*
* The baseline is deliberately the honest null this order describes: the
* player's base rate, i.e. "he's due" with no reading of tonight at all. A
* factor earns its place only by beating that.
*
* SUPABASE_URL=... node scripts/prove-hit-factors.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const fg = require('../src/services/model/factorGate');
const sk = require('../src/services/model/skillProjection');
const tl = require('../src/services/model/testLedger');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownNumber, knownRate } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const sd = require('../src/services/model/sprayDefense');
const pss = require('../src/services/model/platoonSeverity');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const ARCHS = (process.env.HF_ARCHETYPES || 'BOMBER,GHOST,ALL').split(',');
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/**
* THE FACTORS. Each returns a MULTIPLIER on the base rate, or null when the
* input is absent — an absent factor must leave the baseline untouched rather
* than nudge it toward some default.
*/
const FACTORS = [
{
key: 'defense_by_direction',
needs: ['spray_multiplier'],
entity: (r) => `${r.player_key}|${r.opp}`,
mechanism: 'CAUSALLY-CORRECT DEFENCE. Where the hitter puts the ball (pull/straight/oppo x ground/air) crossed with the OAA of the fielders actually standing in those zones, joined by handedness. Team-average failed the gate because it averages in five fielders who will never touch his ball.',
apply: (r) => r.spray_multiplier,
},
{
key: 'defense',
needs: ['team_defense'],
entity: (r) => r.opp,
mechanism: 'A ball in play becomes a hit or an out partly by who is standing behind the pitcher. Should matter most where contact stays in the park.',
// More outs converted above average -> fewer hits.
apply: (r) => 1 - Math.max(-0.12, Math.min(0.12, r.team_defense / 250)),
},
{
key: 'pitcher_contact_profile',
needs: ['pitcher_hard_hit_allowed'],
entity: (r) => r.starter_id,
mechanism: 'A contact-allowing arm concedes better contact than a bat-misser; hit probability should follow the quality of contact he permits.',
apply: (r) => 1 + Math.max(-0.15, Math.min(0.15, (r.pitcher_hard_hit_allowed - 0.389) * 1.2)),
},
{
key: 'park_hits',
needs: ['park_factor'],
entity: (r) => r.park_factor,
mechanism: 'Some parks turn outs into hits without producing runs — big outfields, high walls, deep gaps.',
apply: (r) => r.park_factor,
caveat: 'STAT_BASE maps hits -> run_base, so this is a RUN factor standing in for a HITS factor. A park that converts outs to hits without scoring is invisible to it.',
},
{
key: 'platoon_severity',
needs: ['platoon_severity_mult'],
entity: (r) => r.player_key,
mechanism: "CAUSALLY-CORRECT PLATOON. The advantage is worth only what THIS hitter's measured split is worth, shrunk toward league by the smaller side's PA and refused outright below a floor. Flat handedness applies the same boost to a 63-point split and to none.",
apply: (r) => r.platoon_severity_mult,
},
{
key: 'platoon',
needs: ['platoon_edge'],
entity: (r) => r.player_key,
mechanism: 'Handedness advantage — a hitter facing the opposite hand sees the ball better and hits it harder.',
apply: (r) => (r.platoon_edge > 0 ? 1.06 : 0.96),
},
];
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const batters = new Map(); const pitchersById = new Map();
for (const r of statcast) {
const prof = sk.fromStatcastRow(r);
if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), prof);
if (r.role === 'batter' && r.player_key) batters.set(r.player_key, prof);
}
const sprayRows = await page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb'));
const sprayByKey = new Map();
for (const r of sprayRows) {
if (!r.player_key) continue;
const prev = sprayByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) sprayByKey.set(r.player_key, r);
}
const platRows = await page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb'));
const platByKey = new Map();
for (const r of platRows) {
if (!r.player_key) continue;
const prev = platByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) platByKey.set(r.player_key, r);
}
const defRows = await page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb'));
const defByTeam = new Map();
for (const d of defRows) defByTeam.set(d.team, d);
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat',
(q) => q.eq('sport', 'mlb').eq('stat', 'hits').not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries',
'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits')
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// Opponent faced, from each hitter's own game log.
const names = new Map();
for (const r of clean) if (!names.has(r.player_key)) names.set(r.player_key, r.player_name);
const oppBy = new Map(); const startersBy = new Map();
const dates = [...new Set(clean.map((r) => r.game_date))].sort();
for (const d of dates) {
try {
const games = await mlb.getScheduleWithPitchers(d);
for (const g of games) {
if (!g.home || !g.away) continue;
if (g.home.probablePitcher) startersBy.set(`${d}|OPP:${g.home.team}`, g.home.probablePitcher.id);
if (g.away.probablePitcher) startersBy.set(`${d}|OPP:${g.away.team}`, g.away.probablePitcher.id);
}
} catch { /* absent slate */ }
}
for (const [key, name] of names) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id);
for (const g of log || []) if (g && g.date && g.opponent) oppBy.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
} catch { /* no log */ }
}
// Per-player base rate — the honest null: "he's due", no reading of tonight.
const byPlayer = new Map();
for (const r of clean) {
const cur = byPlayer.get(r.player_key) || { n: 0, w: 0 };
cur.n += 1; cur.w += r.outcome === 'hit' ? 1 : 0;
byPlayer.set(r.player_key, cur);
}
const loss = { no_batter_profile: 0, thin_base_rate: 0, no_opponent: 0, no_pitcher: 0, kept: 0 };
const rows = [];
for (const r of clean) {
const bat = batters.get(r.player_key);
const bp = byPlayer.get(r.player_key);
if (!bat) loss.no_batter_profile += 1;
if (!bp || bp.n < 3) { loss.thin_base_rate += 1; continue; }
// Leave-one-out so a row never contributes to its own baseline.
const baseline = (bp.w - (r.outcome === 'hit' ? 1 : 0)) / (bp.n - 1);
const faced = oppBy.get(`${r.player_key}|${r.game_date}`) || null;
const nick = faced ? String(faced).split(' ').pop() : null;
const def = faced ? (defByTeam.get(faced) || defByTeam.get(nick)) : null;
if (!faced) loss.no_opponent += 1;
const starterId = faced ? startersBy.get(`${r.game_date}|OPP:${faced}`) : null;
const pit = starterId != null ? pitchersById.get(Number(starterId)) : null;
if (faced && !pit) loss.no_pitcher += 1;
loss.kept += 1;
rows.push({
id: r.id,
// Errors are correlated WITHIN a game — shared starter, park, weather and
// the game's own randomness — so the interval must be clustered on it.
// Three of these factors (pitcher profile, team defence, park) are also
// CONSTANT across every hitter facing that starter, which makes row
// resampling straightforwardly wrong for them.
cluster: r.game_id,
opp: faced,
starter_id: starterId != null ? Number(starterId) : null,
player_key: r.player_key,
archetype: archOf.get(`${r.player_key}|${r.game_date}`) || null,
won: r.outcome === 'hit' ? 1 : 0,
baseline,
team_defense: def ? knownNumber(def.oaa_sum) : null,
pitcher_hard_hit_allowed: pit ? knownRate(pit.hard_hit_pct) : null,
park_factor: knownNumber(r.env_park_base),
platoon_severity_mult: (() => {
const sp = platByKey.get(r.player_key);
if (!sp || !bat || !bat.bats || !pit || !pit.throws) return null;
const out = pss.platoonRead({
splits: {
vl: { pa: sp.vl_pa, atBats: sp.vl_ab, hits: sp.vl_hits },
vr: { pa: sp.vr_pa, atBats: sp.vr_ab, hits: sp.vr_hits },
},
bats: bat.bats, throws: pit.throws,
});
return out && out.readable ? out.multiplier : null;
})(),
spray_multiplier: (() => {
const sp = sprayByKey.get(r.player_key);
const posOaa = def && def.position_oaa ? def.position_oaa : null;
if (!sp || !posOaa || !bat || !bat.bats) return null;
const out = sd.sprayDefenseMultiplier({ spray: sp, bats: bat.bats, positionOaa: posOaa });
return out ? out.multiplier : null;
})(),
platoon_edge: (bat && pit && bat.bats && pit.throws)
? (String(bat.bats)[0] !== String(pit.throws)[0] ? 1 : -1) : null,
});
}
// Cumulative Bonferroni across the programme lifetime.
const store = tl.supabaseStore(sb);
const mc = await tl.recordAndCount(store, FACTORS.flatMap((f) =>
ARCHS.map((a) => ({ sport: 'mlb', stat: 'hits', archetype: a === 'ALL' ? null : a, interaction: `factor:${f.key}`, target: 'outcome' }))));
// STEP 1 — FULL-HISTORY SAMPLE AUDIT PER SLOT, before any gating.
const audit = [];
for (const f of FACTORS) {
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
audit.push({
factor: f.key,
archetype: arch,
rows: usable.length,
games: new Set(usable.map((r) => r.cluster).filter(Boolean)).size,
players: new Set(usable.map((r) => r.player_key)).size,
});
}
}
const results = [];
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
for (const f of FACTORS) {
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
// A park effect is replicated across PARKS, not across games: 619 rows in
// 45 games still only ever saw ~23 ballparks, and unmodelled park
// heterogeneity is confounded with the very thing being estimated. So the
// cluster is the COARSER of the game and the entity the treatment rides on.
const ents = f.entity ? new Set(usable.map((r) => String(f.entity(r)))) : null;
const games = new Set(usable.map((r) => String(r.cluster)));
const useEntity = ents && ents.size < games.size;
const paired = usable.map((r) => {
const mult = f.apply(r);
const cond = mult === null ? null : Math.min(0.99, Math.max(0.01, r.baseline * mult));
return {
baseline: r.baseline,
conditioned: cond,
won: r.won,
cluster: useEntity ? `e:${f.entity(r)}` : r.cluster,
};
});
const v = fg.adjudicate(paired, {
factor: f.key, archetype: arch, stat: 'hits',
cumulativeTests: mc.cumulative_tests, // native cumulative correction
});
results.push({
archetype: arch, factor: f.key, n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
cluster_unit: useEntity ? 'treatment_entity' : 'game',
distinct_games: games.size,
distinct_entities: ents ? ents.size : null,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
ci_level: v.improvement ? v.improvement.ci_level : null,
verdict: v.verdict,
reason: v.reason,
...(f.caveat ? { input_caveat: f.caveat } : {}),
});
}
}
console.log(JSON.stringify({
baseline: "each row scored against the player's OWN leave-one-out base rate — the honest 'he's due' null",
total_rows: rows.length,
slot_audit: audit,
clean_settled_rows_available: clean.length,
row_loss: loss,
cumulative_bonferroni: mc,
gate: 'a factor must MOVE the prediction AND improve out-of-sample Brier; movement alone is THEATER',
results,
proven: results.filter((r) => r.verdict === 'PROVES'),
theater: results.filter((r) => r.verdict === 'THEATER'),
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+121
View File
@@ -0,0 +1,121 @@
#!/usr/bin/env node
'use strict';
/**
* prove-park-weather — DOES PARK GEOMETRY AND AIR READ TOTAL BASES?
*
* Run through the two-part gate like every other factor, with one addition that
* changes the answer: the rows are CLUSTERED BY GAME. Park and weather assign a
* single value to every hitter in a ballpark on a night, so eighteen prop rows
* from one game are one reading of that game's conditions. Resampling rows would
* treat them as eighteen and hand back an interval far tighter than the evidence
* supports — which is how a gate passes a factor on sample it never had.
*
* SUPABASE_URL=... node scripts/prove-park-weather.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const pw = require('../src/services/model/parkWeather');
const fg = require('../src/services/model/factorGate');
const tl = require('../src/services/model/testLedger');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/** Hit-type shares → expected total bases, so a reshape has a consequence. */
const tbFromShares = (s) => s.single + 2 * s.double + 3 * s.triple + 4 * s.home_run;
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const parks = await page(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb'));
const byVenue = new Map();
for (const p of parks) if (!byVenue.has(p.venue_id)) byVenue.set(p.venue_id, p);
const league = pw.leagueGeometry([...byVenue.values()]);
const ctx = await page(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q);
const ctxBy = new Map(ctx.map((c) => [c.game_id, c]));
const led = await page(sb, 'ledger_entries',
'game_id, game_date, player_key, stat, line, side, outcome, quarantine_reason, p_win, proj_hits_p_over',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases').in('outcome', ['hit', 'miss']));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// League-typical hit-type shares — the shape the atom reshapes.
const BASE_SHARES = { single: 0.655, double: 0.195, triple: 0.017, home_run: 0.133 };
const baseTb = tbFromShares(BASE_SHARES);
const loss = { no_context: 0, no_venue: 0, no_park: 0, no_baseline: 0, kept: 0 };
const rows = [];
for (const r of clean) {
const c = ctxBy.get(r.game_id);
if (!c) { loss.no_context += 1; continue; }
if (c.venue_id == null) { loss.no_venue += 1; continue; }
const dims = byVenue.get(c.venue_id);
if (!dims) { loss.no_park += 1; continue; }
const baseline = knownNumber(r.p_win);
if (baseline === null) { loss.no_baseline += 1; continue; }
const read = pw.parkWeatherRead({ dims, wx: c, league });
if (!read) { loss.no_park += 1; continue; }
// The atom reshapes hit type; the consequence for total bases is the ratio
// of expected bases per hit under the reshaped shape.
const shaped = pw.applyToShares(BASE_SHARES, read);
const ratio = tbFromShares(shaped) / baseTb;
const conditioned = Math.max(0.01, Math.min(0.99, baseline * ratio));
rows.push({
cluster: r.game_id, // ONE reading per game — the whole point
baseline,
conditioned,
won: r.outcome === 'hit' ? 1 : 0,
});
loss.kept += 1;
}
// Cumulative Bonferroni across the programme lifetime — this hypothesis is
// one more test, and the bar rises for it like every other.
const mc = await tl.recordAndCount(tl.supabaseStore(sb), [
{ sport: 'mlb', stat: 'total_bases', archetype: null, interaction: 'factor:park_weather_hit_type', target: 'outcome' },
]);
const verdict = fg.adjudicate(rows, {
factor: 'park_weather_hit_type',
stat: 'total_bases',
cumulativeTests: mc.cumulative_tests,
});
const games = new Set(rows.map((r) => r.cluster)).size;
const venues = new Set(clean.map((r) => ctxBy.get(r.game_id)?.venue_id).filter((v) => v != null)).size;
console.log(JSON.stringify({
clean_settled_tb_rows: clean.length,
rows_built: rows.length,
row_loss: loss,
distinct_games: games,
distinct_venues: venues,
rows_per_game: games ? Math.round((rows.length / games) * 10) / 10 : null,
cumulative_tests: mc.cumulative_tests,
verdict,
honest_note: 'sample judged in GAMES, not prop rows — park and weather vary per game',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+301
View File
@@ -0,0 +1,301 @@
#!/usr/bin/env node
'use strict';
/**
* prove-runs-rbi — THE CONTEXT-HEAVY STATS, WHERE INFLATION IS EASIEST.
*
* A large share of both stats is genuinely outside the hitter's control: a run
* needs someone behind you, an RBI needs someone in front of you. The job is to
* prove the HITTER-CONTROLLABLE part above the archetype's own base rate and
* grade the rest honestly as base-rate — which is the CORRECT answer for a
* context stat, not a failure to find something.
*
* ── THE NULL IS THE ARCHETYPE'S BASE RATE ────────────────────────────────
* Deliberately, and per the order: these base rates are spread and
* context-inflated, so beating "hitters like him" is the only meaningful bar. A
* per-player leave-one-out rate is not available here — 935 RBI rows over 344
* players is ~2.7 rows each, and estimating a personal rate from two rows would
* be inventing one. Leave-one-out is applied at the ARCHETYPE level so a row
* never contributes to its own baseline.
*
* ── INPUTS RECONSTRUCTED RATHER THAN DECLARED MISSING ────────────────────
* `lineup_context` only starts 2026-08-04 (ingest began last week) while settled
* rows run from 07-31, so only 187 of 617 runs rows join to a batting order.
* That would be input-blocked — except the play-by-play cache covers 05-01
* onward, and the batting order IS the order batters first appear. Reach-base
* skill and lineup power behind are derived from the same cache, point-in-time.
*
* SUPABASE_URL=... node scripts/prove-runs-rbi.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const fg = require('../src/services/model/factorGate');
const tl = require('../src/services/model/testLedger');
const sk = require('../src/services/model/skillProjection');
const { knownNumber, knownRate } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const CACHE = process.env.SEQ_OUT || path.join(process.cwd(), '.seq-cache', 'sequences.json');
const PAGE = 1000;
const ARCHS = (process.env.RR_ARCHETYPES || 'ALL,BOMBER,GHOST,BRUSH,DRIVER').split(',');
const HIT = new Set(['single', 'double', 'triple', 'home_run']);
const ONBASE = new Set(['single', 'double', 'triple', 'home_run', 'walk', 'hit_by_pitch', 'intent_walk']);
const PA_EVENT = new Set([...HIT, 'field_out', 'strikeout', 'grounded_into_double_play', 'force_out',
'field_error', 'fielders_choice', 'fielders_choice_out', 'double_play', 'sac_fly', 'pop_out',
'line_out', 'fly_out', 'strikeout_double_play', 'walk', 'hit_by_pitch', 'intent_walk']);
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/**
* Reconstruct, point-in-time, from play-by-play:
* order[date|nameKey] the hitter's batting slot that game
* behind[date|nameKey] mean barrel-ish power of the three slots after him
* onbase[nameKey] his reach-base rate over PRIOR games only
*/
function reconstruct(barrelByKey) {
const { games } = JSON.parse(fs.readFileSync(CACHE, 'utf8'));
games.sort((a, b) => String(a.date).localeCompare(String(b.date)) || a.gamePk - b.gamePk);
const order = new Map();
const behind = new Map();
const onbaseNow = new Map(); // running totals, folded in AFTER each game
const onbasePrior = new Map(); // snapshot used for that game's rows
for (const g of games) {
for (const half of ['top', 'bottom']) {
const pas = g.pas.filter((p) => p.half === half && PA_EVENT.has(p.event));
if (!pas.length) continue;
// The batting order IS the order batters first appear.
const seen = [];
const seenSet = new Set();
for (const p of pas) {
if (!seenSet.has(p.batter)) { seenSet.add(p.batter); seen.push(p); }
if (seen.length >= 9) break;
}
const slots = seen.map((p) => ({ id: p.batter, key: nameKey(p.batter_name || '') }));
for (let i = 0; i < slots.length; i += 1) {
const k = `${g.date}|${slots[i].key}`;
order.set(k, i + 1);
// Power BEHIND him — the hitters who would drive him in.
const nxt = [1, 2, 3].map((d) => slots[(i + d) % slots.length])
.map((s) => (s ? knownRate(barrelByKey.get(s.key)) : null))
.filter((v) => v !== null);
if (nxt.length) behind.set(k, mean(nxt));
const prior = onbaseNow.get(slots[i].key);
if (prior && prior.pa >= 60) onbasePrior.set(k, prior.ob / prior.pa);
}
for (const p of pas) {
const key = nameKey(p.batter_name || '');
const cur = onbaseNow.get(key) || { pa: 0, ob: 0 };
cur.pa += 1; cur.ob += ONBASE.has(p.event) ? 1 : 0;
onbaseNow.set(key, cur);
}
}
}
return { order, behind, onbase: onbasePrior };
}
/** RBI and RUNS have different causal stories, so different factors. */
const FACTORS = {
rbi: [
{
key: 'risp_opportunity',
needs: ['risp_share'],
entity: (r) => r.player_key,
mechanism: 'HOW OFTEN HE BATS WITH RUNNERS IN SCORING POSITION. Half of an RBI is opportunity, and this is the ingested measure of it.',
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.risp_share - 0.22) * 1.6)),
},
{
key: 'extra_base_skill',
needs: ['barrel_pct'],
entity: (r) => r.player_key,
mechanism: 'The other half — having batted with runners on, can he drive them in.',
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.barrel_pct - 0.078) * 1.8)),
},
{
key: 'risp_x_extra_base',
needs: ['risp_share', 'barrel_pct'],
entity: (r) => r.player_key,
mechanism: 'THE CAUSALLY-CORRECT COMPOUND: opportunity AND the power to convert it. Neither half alone is an RBI.',
apply: (r) => (1 + Math.max(-0.20, Math.min(0.20, (r.risp_share - 0.22) * 1.6)))
* (1 + Math.max(-0.20, Math.min(0.20, (r.barrel_pct - 0.078) * 1.8))),
},
],
runs: [
{
key: 'reach_base',
needs: ['onbase'],
entity: (r) => r.player_key,
mechanism: 'You cannot score without first reaching base. The most hitter-controllable component of a run.',
apply: (r) => 1 + Math.max(-0.25, Math.min(0.25, (r.onbase - 0.318) * 2.2)),
},
{
key: 'lineup_power_behind',
needs: ['power_behind'],
entity: (r) => `${r.game_id}|${r.batting_order}`,
mechanism: 'Who bats after him — the hitters who would drive him in. Pure context, and the part he does not control.',
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.power_behind - 0.078) * 1.8)),
},
{
key: 'reach_x_power_behind',
needs: ['onbase', 'power_behind'],
entity: (r) => r.player_key,
mechanism: 'THE CAUSALLY-CORRECT COMPOUND: reach base AND have someone behind you who can drive you in.',
apply: (r) => (1 + Math.max(-0.25, Math.min(0.25, (r.onbase - 0.318) * 2.2)))
* (1 + Math.max(-0.20, Math.min(0.20, (r.power_behind - 0.078) * 1.8))),
},
],
};
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb').eq('role', 'batter'));
const batByKey = new Map();
const barrelByKey = new Map();
for (const r of statcast) {
if (!r.player_key) continue;
const prof = sk.fromStatcastRow(r);
batByKey.set(r.player_key, prof);
if (prof.barrel_pct != null) barrelByKey.set(r.player_key, prof.barrel_pct);
}
const oppRows = await page(sb, 'hitter_opportunity', '*', (q) => q.eq('sport', 'mlb'));
const oppByKey = new Map();
for (const r of oppRows) {
const prev = oppByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) oppByKey.set(r.player_key, r);
}
const recon = reconstruct(barrelByKey);
const out = { generated_note: 'null is the ARCHETYPE base rate, leave-one-out' };
for (const stat of ['rbi', 'runs']) {
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype',
(q) => q.eq('sport', 'mlb').eq('stat', stat).not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries',
'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', stat).in('outcome', ['hit', 'miss']));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')
&& knownNumber(r.line) === 0.5);
// Archetype-level leave-one-out base rate — the context-inflated null.
const byArch = new Map();
for (const r of clean) {
const a = String(archOf.get(`${r.player_key}|${r.game_date}`) || 'UNLABELLED').toUpperCase();
const cur = byArch.get(a) || { n: 0, w: 0 };
cur.n += 1; cur.w += r.outcome === 'hit' ? 1 : 0;
byArch.set(a, cur);
}
const rows = [];
const loss = { no_archetype_base: 0, kept: 0 };
for (const r of clean) {
const a = String(archOf.get(`${r.player_key}|${r.game_date}`) || 'UNLABELLED').toUpperCase();
const ab = byArch.get(a);
if (!ab || ab.n < 4) { loss.no_archetype_base += 1; continue; }
const baseline = (ab.w - (r.outcome === 'hit' ? 1 : 0)) / (ab.n - 1);
const bat = batByKey.get(r.player_key);
const opp = oppByKey.get(r.player_key);
const okey = `${r.game_date}|${r.player_key}`;
rows.push({
archetype: a,
player_key: r.player_key,
game_id: r.game_id,
cluster: r.game_id,
baseline,
won: r.outcome === 'hit' ? 1 : 0,
risp_share: opp ? knownNumber(opp.risp_share) : null,
barrel_pct: bat ? knownRate(bat.barrel_pct) : null,
onbase: recon.onbase.has(okey) ? recon.onbase.get(okey) : null,
power_behind: recon.behind.has(okey) ? recon.behind.get(okey) : null,
batting_order: recon.order.get(okey) ?? null,
});
loss.kept += 1;
}
const mc = await tl.recordAndCount(tl.supabaseStore(sb), FACTORS[stat].flatMap((f) =>
ARCHS.map((a) => ({
sport: 'mlb', stat, archetype: a === 'ALL' ? null : a,
interaction: `factor:${f.key}`, target: 'outcome',
}))));
const audit = [];
const results = [];
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => r.archetype === arch);
for (const f of FACTORS[stat]) {
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
const ents = new Set(usable.map((r) => String(f.entity(r))));
const games = new Set(usable.map((r) => String(r.cluster)));
if (arch === 'ALL') {
audit.push({ factor: f.key, archetype: arch, rows: usable.length, games: games.size, entities: ents.size });
}
const useEntity = ents.size < games.size;
const paired = usable.map((r) => {
const m = f.apply(r);
return {
baseline: r.baseline,
conditioned: m === null ? null : Math.min(0.99, Math.max(0.01, r.baseline * m)),
won: r.won,
cluster: useEntity ? `e:${f.entity(r)}` : r.cluster,
};
});
const v = fg.adjudicate(paired, { factor: f.key, archetype: arch, stat, cumulativeTests: mc.cumulative_tests });
results.push({
archetype: arch, factor: f.key, n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
clustered_on: useEntity ? 'treatment_entity' : 'game',
distinct_games: games.size,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
verdict: v.verdict,
});
}
}
out[stat] = {
clean_rows_line_0_5: clean.length,
rows_built: rows.length,
row_loss: loss,
distinct_games: new Set(rows.map((r) => r.cluster)).size,
archetype_base_rates: Object.fromEntries([...byArch.entries()]
.sort((a, b) => b[1].n - a[1].n)
.map(([a, v]) => [a, { n: v.n, base_rate: Math.round((v.w / v.n) * 10000) / 10000 }])),
cumulative_tests: mc.cumulative_tests,
input_audit: audit,
results,
proven: results.filter((r) => r.verdict === 'PROVES'),
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+345
View File
@@ -0,0 +1,345 @@
#!/usr/bin/env node
'use strict';
/**
* prove-tb-factors — TOTAL BASES IS A DIFFERENT EVENT FROM HITS.
*
* A hit asks whether the ball found a hole. Total bases asks how hard and how
* far it was struck. So the causally-correct factors differ, and the
* archetype differential is expected to INVERT: contact defence proved for hits,
* where a slap single is worth exactly one base regardless of who fielded it;
* for total bases the value should live with the power profiles.
*
* ── THE BASELINE IS THE CHAMPION, NOT "HE'S DUE" ─────────────────────────
* The hits gate used the player's own leave-one-out base rate as the null. That
* cannot be reproduced here: total-bases lines VARY (1.5 on 559 rows, 0.5 on
* 345, 2.5 on 45), and a player's rate of clearing 1.5 bases is a different
* quantity from his rate of clearing 0.5. With 988 rows over 341 players there
* are roughly two rows per player-line — far too thin to estimate a per-line
* personal base rate without inventing one.
*
* So the null here is the COUNTER'S OWN FORECAST (p_win), which already prices
* the line. That is a strictly HARDER null than a base rate, not an easier one:
* a factor must improve on the champion, not merely on "he's due". Stated
* plainly because it differs from the hits run and the difference matters when
* comparing the two.
*
* SUPABASE_URL=... node scripts/prove-tb-factors.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const fg = require('../src/services/model/factorGate');
const sk = require('../src/services/model/skillProjection');
const tl = require('../src/services/model/testLedger');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownNumber, knownRate } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const sd = require('../src/services/model/sprayDefense');
const pss = require('../src/services/model/platoonSeverity');
const pw = require('../src/services/model/parkWeather');
/** League-typical hit-type shares; the atom reshapes these and TB follows. */
const BASE_SHARES = { single: 0.655, double: 0.195, triple: 0.017, home_run: 0.133 };
const tbFrom = (s) => s.single + 2 * s.double + 3 * s.triple + 4 * s.home_run;
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const ARCHS = (process.env.TB_ARCHETYPES || 'ALL,BOMBER,GHOST,BRUSH,DRIVER').split(',');
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/**
* THE FACTORS. Each returns a MULTIPLIER on the base rate, or null when the
* input is absent — an absent factor must leave the baseline untouched rather
* than nudge it toward some default.
*/
const FACTORS = [
{
key: 'barrel_rate',
needs: ['barrel_pct'],
entity: (r) => r.player_key,
mechanism: 'THE EXTRA-BASE SKILL ITSELF. A barrel is the exit-velocity and launch-angle combination that produces extra bases; it is the most direct expression of what total bases measures, where for hits it is largely irrelevant to whether a grounder finds a hole.',
// UNITS: fromStatcastRow returns barrel_pct as a FRACTION (0.06), not the
// 0-100 the raw table stores. Writing this against the percentage scale
// clamped every row to the maximum negative shift, which then "improved"
// Brier only by leaning on the counter's known global over-prediction.
apply: (r) => 1 + Math.max(-0.20, Math.min(0.20, (r.barrel_pct - 0.078) * 1.8)),
},
{
key: 'exit_velo',
needs: ['avg_exit_velo'],
entity: (r) => r.player_key,
mechanism: 'How hard the ball leaves the bat. Separates a double in the gap from a fly out, which is exactly the margin total bases lives on.',
apply: (r) => 1 + Math.max(-0.15, Math.min(0.15, (r.avg_exit_velo - 88.9) * 0.020)),
},
{
key: 'hard_contact_allowed',
needs: ['pitcher_hard_hit_allowed'],
entity: (r) => r.starter_id,
mechanism: 'A pitcher who concedes hard contact concedes EXTRA BASES, not just hits. For total bases this should read stronger than it did for hits.',
apply: (r) => 1 + Math.max(-0.15, Math.min(0.15, (r.pitcher_hard_hit_allowed - 0.389) * 1.2)),
},
{
key: 'park_weather_hit_type',
needs: ['park_weather_ratio'],
entity: (r) => r.park_weather_ratio,
mechanism: 'Whether a struck ball becomes a double, clears the fence, or dies at the track. The atom reshapes HIT TYPE rather than P(hit), which is the only form that can express a total-bases effect.',
apply: (r) => r.park_weather_ratio,
caveat: 'venue-borne: replication caps at the number of distinct park readings, not the row count',
},
{
key: 'platoon_severity',
needs: ['platoon_severity_mult'],
entity: (r) => r.player_key,
mechanism: "The hitter's OWN measured split, shrunk by the smaller side's plate appearances and refused below a floor.",
apply: (r) => r.platoon_severity_mult,
},
];
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const batters = new Map(); const pitchersById = new Map();
for (const r of statcast) {
const prof = sk.fromStatcastRow(r);
if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), prof);
if (r.role === 'batter' && r.player_key) batters.set(r.player_key, prof);
}
const sprayRows = await page(sb, 'batter_spray', '*', (q) => q.eq('sport', 'mlb'));
const sprayByKey = new Map();
for (const r of sprayRows) {
if (!r.player_key) continue;
const prev = sprayByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) sprayByKey.set(r.player_key, r);
}
const platRows = await page(sb, 'platoon_splits', '*', (q) => q.eq('sport', 'mlb'));
const platByKey = new Map();
for (const r of platRows) {
if (!r.player_key) continue;
const prev = platByKey.get(r.player_key);
if (!prev || String(r.as_of_date) > String(prev.as_of_date)) platByKey.set(r.player_key, r);
}
const parkRows = await page(sb, 'park_dimensions', '*', (q) => q.eq('sport', 'mlb'));
const parkByVenue = new Map();
for (const p of parkRows) if (!parkByVenue.has(p.venue_id)) parkByVenue.set(p.venue_id, p);
const parkLeague = pw.leagueGeometry([...parkByVenue.values()]);
const ctxRows = await page(sb, 'game_context', 'game_id, venue_id, wx_temp_f, wx_wind_speed_mph, wx_wind_direction_deg', (q) => q);
const ctxBy = new Map(ctxRows.map((c) => [c.game_id, c]));
const defRows = await page(sb, 'team_defense', '*', (q) => q.eq('sport', 'mlb'));
const defByTeam = new Map();
for (const d of defRows) defByTeam.set(d.team, d);
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype, stat',
(q) => q.eq('sport', 'mlb').eq('stat', 'total_bases').not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.game_date}`, s.archetype);
const led = await page(sb, 'ledger_entries',
'id, game_id, player_key, player_name, line, side, outcome, game_date, p_win, quarantine_reason, env_park_base',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases')
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// Opponent faced, from each hitter's own game log.
const names = new Map();
for (const r of clean) if (!names.has(r.player_key)) names.set(r.player_key, r.player_name);
const oppBy = new Map(); const startersBy = new Map();
const dates = [...new Set(clean.map((r) => r.game_date))].sort();
for (const d of dates) {
try {
const games = await mlb.getScheduleWithPitchers(d);
for (const g of games) {
if (!g.home || !g.away) continue;
if (g.home.probablePitcher) startersBy.set(`${d}|OPP:${g.home.team}`, g.home.probablePitcher.id);
if (g.away.probablePitcher) startersBy.set(`${d}|OPP:${g.away.team}`, g.away.probablePitcher.id);
}
} catch { /* absent slate */ }
}
for (const [key, name] of names) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id);
for (const g of log || []) if (g && g.date && g.opponent) oppBy.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
} catch { /* no log */ }
}
// Per-player base rate — the honest null: "he's due", no reading of tonight.
const byPlayer = new Map();
for (const r of clean) {
const cur = byPlayer.get(r.player_key) || { n: 0, w: 0 };
cur.n += 1; cur.w += r.outcome === 'hit' ? 1 : 0;
byPlayer.set(r.player_key, cur);
}
const loss = { no_batter_profile: 0, thin_base_rate: 0, no_opponent: 0, no_pitcher: 0, kept: 0 };
const rows = [];
for (const r of clean) {
const bat = batters.get(r.player_key);
const bp = byPlayer.get(r.player_key);
if (!bat) loss.no_batter_profile += 1;
if (!bp || bp.n < 3) { loss.thin_base_rate += 1; continue; }
// THE NULL IS THE CHAMPION. Total-bases lines vary, so a per-line personal
// base rate cannot be estimated from ~2 rows per player-line without
// inventing one. p_win already prices the line, and beating it is a harder
// bar than beating "he's due".
const baseline = knownNumber(r.p_win);
if (baseline === null) { loss.thin_base_rate += 1; continue; }
const faced = oppBy.get(`${r.player_key}|${r.game_date}`) || null;
const nick = faced ? String(faced).split(' ').pop() : null;
const def = faced ? (defByTeam.get(faced) || defByTeam.get(nick)) : null;
if (!faced) loss.no_opponent += 1;
const starterId = faced ? startersBy.get(`${r.game_date}|OPP:${faced}`) : null;
const pit = starterId != null ? pitchersById.get(Number(starterId)) : null;
if (faced && !pit) loss.no_pitcher += 1;
loss.kept += 1;
rows.push({
id: r.id,
// Errors are correlated WITHIN a game — shared starter, park, weather and
// the game's own randomness — so the interval must be clustered on it.
// Three of these factors (pitcher profile, team defence, park) are also
// CONSTANT across every hitter facing that starter, which makes row
// resampling straightforwardly wrong for them.
cluster: r.game_id,
opp: faced,
starter_id: starterId != null ? Number(starterId) : null,
player_key: r.player_key,
archetype: archOf.get(`${r.player_key}|${r.game_date}`) || null,
won: r.outcome === 'hit' ? 1 : 0,
baseline,
team_defense: def ? knownNumber(def.oaa_sum) : null,
pitcher_hard_hit_allowed: pit ? knownRate(pit.hard_hit_pct) : null,
line: knownNumber(r.line),
barrel_pct: bat ? knownRate(bat.barrel_pct) : null,
avg_exit_velo: bat ? knownNumber(bat.avg_exit_velo) : null,
park_weather_ratio: (() => {
const c = ctxBy.get(r.game_id);
if (!c || c.venue_id == null) return null;
const dims = parkByVenue.get(c.venue_id);
if (!dims) return null;
const read = pw.parkWeatherRead({ dims, wx: c, league: parkLeague });
if (!read) return null;
const shaped = pw.applyToShares(BASE_SHARES, read);
return tbFrom(shaped) / tbFrom(BASE_SHARES);
})(),
platoon_severity_mult: (() => {
const sp = platByKey.get(r.player_key);
if (!sp || !bat || !bat.bats || !pit || !pit.throws) return null;
const out = pss.platoonRead({
splits: {
vl: { pa: sp.vl_pa, atBats: sp.vl_ab, hits: sp.vl_hits },
vr: { pa: sp.vr_pa, atBats: sp.vr_ab, hits: sp.vr_hits },
},
bats: bat.bats, throws: pit.throws,
});
return out && out.readable ? out.multiplier : null;
})(),
spray_multiplier: (() => {
const sp = sprayByKey.get(r.player_key);
const posOaa = def && def.position_oaa ? def.position_oaa : null;
if (!sp || !posOaa || !bat || !bat.bats) return null;
const out = sd.sprayDefenseMultiplier({ spray: sp, bats: bat.bats, positionOaa: posOaa });
return out ? out.multiplier : null;
})(),
platoon_edge: (bat && pit && bat.bats && pit.throws)
? (String(bat.bats)[0] !== String(pit.throws)[0] ? 1 : -1) : null,
});
}
// Cumulative Bonferroni across the programme lifetime.
const store = tl.supabaseStore(sb);
const mc = await tl.recordAndCount(store, FACTORS.flatMap((f) =>
ARCHS.map((a) => ({ sport: 'mlb', stat: 'total_bases', archetype: a === 'ALL' ? null : a, interaction: `factor:${f.key}`, target: 'outcome' }))));
// STEP 1 — FULL-HISTORY SAMPLE AUDIT PER SLOT, before any gating.
const audit = [];
for (const f of FACTORS) {
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
audit.push({
factor: f.key,
archetype: arch,
rows: usable.length,
games: new Set(usable.map((r) => r.cluster).filter(Boolean)).size,
players: new Set(usable.map((r) => r.player_key)).size,
});
}
}
const results = [];
for (const arch of ARCHS) {
const slot = arch === 'ALL' ? rows : rows.filter((r) => String(r.archetype || '').toUpperCase() === arch);
for (const f of FACTORS) {
const usable = slot.filter((r) => f.needs.every((k) => knownNumber(r[k]) !== null));
// A park effect is replicated across PARKS, not across games: 619 rows in
// 45 games still only ever saw ~23 ballparks, and unmodelled park
// heterogeneity is confounded with the very thing being estimated. So the
// cluster is the COARSER of the game and the entity the treatment rides on.
const ents = f.entity ? new Set(usable.map((r) => String(f.entity(r)))) : null;
const games = new Set(usable.map((r) => String(r.cluster)));
const useEntity = ents && ents.size < games.size;
const paired = usable.map((r) => {
const mult = f.apply(r);
const cond = mult === null ? null : Math.min(0.99, Math.max(0.01, r.baseline * mult));
return {
baseline: r.baseline,
conditioned: cond,
won: r.won,
cluster: useEntity ? `e:${f.entity(r)}` : r.cluster,
};
});
const v = fg.adjudicate(paired, {
factor: f.key, archetype: arch, stat: 'total_bases',
cumulativeTests: mc.cumulative_tests, // native cumulative correction
});
results.push({
archetype: arch, factor: f.key, n: v.movement.n,
clusters: v.improvement ? v.improvement.effective_n : null,
cluster_unit: useEntity ? 'treatment_entity' : 'game',
distinct_games: games.size,
distinct_entities: ents ? ents.size : null,
mean_abs_shift: v.movement.mean_abs_shift,
brier_delta: v.improvement ? v.improvement.brier_delta : null,
ci: v.improvement ? v.improvement.ci : null,
ci_level: v.improvement ? v.improvement.ci_level : null,
verdict: v.verdict,
reason: v.reason,
...(f.caveat ? { input_caveat: f.caveat } : {}),
});
}
}
console.log(JSON.stringify({
baseline: "each row scored against the player's OWN leave-one-out base rate — the honest 'he's due' null",
total_rows: rows.length,
slot_audit: audit,
clean_settled_rows_available: clean.length,
row_loss: loss,
cumulative_bonferroni: mc,
gate: 'a factor must MOVE the prediction AND improve out-of-sample Brier; movement alone is THEATER',
results,
proven: results.filter((r) => r.verdict === 'PROVES'),
theater: results.filter((r) => r.verdict === 'THEATER'),
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+126
View File
@@ -0,0 +1,126 @@
#!/usr/bin/env node
'use strict';
/**
* proven-status — WHAT IS ACTUALLY PROVEN, computed from the ledger.
*
* WHY THIS EXISTS. Four consecutive build orders have opened by describing
* results as proven that the measurements did not support: "barrel rate PASSED
* solo" (every total_bases feature was refused on sample), "total_bases has
* passed BAR 1" (inconclusive at parity, CI spanning zero), "whiff/stuff prove
* SOLO through the gate" (refused at n=57), "two proven clusters live" (the
* proven set is empty). Each time the correction had to be re-derived by hand
* from a spec written days earlier.
*
* Prose decays. A number recomputed from the ledger does not. So this prints the
* proven set on demand, from the same gate everything else is held to, and any
* session can run it in one command before planning on top of a claim.
*
* IT DELIBERATELY CANNOT SAY "PROVEN" ON ITS OWN. A stat is proven only if a
* recorded head-to-head beat the counter out-of-sample with a CI excluding zero,
* which is a measurement this script does not perform — it reports SAMPLE
* READINESS (can the gate even be run?) and the recorded verdicts, so the two
* are never confused again.
*
* SUPABASE_URL=... node scripts/proven-status.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const cv = require('../src/services/model/correlateValidator');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const MIN_N = cv.VALIDATION_REQUIREMENTS.min_historical_instances;
/**
* RECORDED VERDICTS — every head-to-head this programme has actually run, with
* its spec. Add a row when a head-to-head is run; never edit one to be kinder.
*/
const RECORDED = [
{ stat: 'hits', n: 803, model: 0.0842, counter: 0.1803, delta: -0.0961, ci: [-0.1648, -0.0285],
verdict: 'LOSES', spec: 'specs/batter-cluster-prove.md' },
{ stat: 'total_bases', n: 383, model: 0.2685, counter: 0.2647, delta: 0.0038, ci: [-0.0675, 0.0753],
verdict: 'INCONCLUSIVE', spec: 'specs/tb-solo-and-interactions.md' },
{ stat: 'strikeouts', n: 57, model: 0.1953, counter: -0.0639, delta: 0.2592, ci: [-0.0167, 0.5645],
verdict: 'INCONCLUSIVE', spec: 'specs/lineup-k-rate-rung1.md' },
];
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const led = await page(sb, 'ledger_entries', 'stat, outcome, quarantine_reason, p_win',
(q) => q.eq('sport', 'mlb').is('user_id', null));
const settled = {};
for (const r of led) {
if ((r.quarantine_reason || '').startsWith('nontakeable_book')) continue;
if (r.outcome !== 'hit' && r.outcome !== 'miss') continue;
if (r.p_win == null) continue;
settled[r.stat] = (settled[r.stat] || 0) + 1;
}
const snaps = await page(sb, 'model_snapshots', 'stat, archetype, player_key, line, side, game_date',
(q) => q.eq('sport', 'mlb').not('archetype', 'is', null));
const archOf = new Map();
for (const s of snaps) archOf.set(`${s.player_key}|${s.stat}|${s.line}|${String(s.side).toLowerCase()}|${s.game_date}`, s.archetype);
// COUNT DISTINCT LEDGER ROWS. `model_snapshots` holds one row per prop PER
// SNAPSHOT CYCLE, so a naive join fans out and inflates the count — it read
// BOMBER x hits as 641 when the true figure is 287, which is the difference
// between "gate-ready" and "not close". Dedupe on the ledger row's identity.
const led2 = await page(sb, 'ledger_entries', 'id, stat, outcome, quarantine_reason, player_key, line, side, game_date',
(q) => q.eq('sport', 'mlb').is('user_id', null).in('outcome', ['hit', 'miss']));
const byArch = {};
const seen = new Set();
for (const r of led2) {
if ((r.quarantine_reason || '').startsWith('nontakeable_book')) continue;
if (seen.has(r.id)) continue;
seen.add(r.id);
const a = archOf.get(`${r.player_key}|${r.stat}|${r.line}|${String(r.side).toLowerCase()}|${r.game_date}`);
if (!a) continue;
const k = `${a} x ${r.stat}`;
byArch[k] = (byArch[k] || 0) + 1;
}
const gateReady = Object.entries(settled).filter(([, n]) => n >= MIN_N).map(([s, n]) => ({ stat: s, n }));
const archReady = Object.entries(byArch).filter(([, n]) => n >= MIN_N)
.sort((a, b) => b[1] - a[1]).map(([k, n]) => ({ combo: k, n }));
const proven = RECORDED.filter((r) => r.verdict === 'BEATS');
console.log(JSON.stringify({
generated_at_note: 'computed from the ledger; prose in specs may lag this',
gate_spec: cv.VALIDATION_REQUIREMENTS,
PROVEN_SET: proven.length === 0 ? 'EMPTY — no stat has beaten the counter out-of-sample with a CI excluding zero' : proven,
recorded_head_to_heads: RECORDED,
sample_readiness: {
note: 'n >= 500 means the gate CAN be run — it does not mean anything passed it',
stats_at_or_above_gate: gateReady,
stats_below_gate: Object.entries(settled).filter(([, n]) => n < MIN_N)
.sort((a, b) => b[1] - a[1]).map(([s, n]) => ({ stat: s, n, short_by: MIN_N - n })),
archetype_x_stat_at_or_above_gate: archReady,
archetype_x_stat_closest_below: Object.entries(byArch).filter(([, n]) => n < MIN_N)
.sort((a, b) => b[1] - a[1]).slice(0, 6).map(([k, n]) => ({ combo: k, n, short_by: MIN_N - n })),
},
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+34
View File
@@ -0,0 +1,34 @@
-- p_win RECALIBRATION — per-sport, time-forward holdout. Measure-only.
-- specs/pwin-recalibration-holdout.md. NOTE: reliability is measured on BUCKETS
-- (predicted vs actual rate, n-weighted). mean|p-outcome| on 0/1 rows is NOT
-- calibration — it is noise-dominated individual error.
with r as (
select sport, game_date, p_win::numeric p, (outcome='hit')::int won,
ntile(2) over (partition by sport order by game_date) half
from ledger_entries
where user_id is null and outcome in ('hit','miss') and p_win is not null
), tagged as (
select *, case when half=1 then 'train' else 'holdout' end split,
width_bucket(p,0.1,1.0,5) bkt, ln(p/(1-p)) lg from r
), tb as ( -- TRAIN buckets = the fit
select sport, bkt, count(*) n, avg(lg) mean_lg, avg(won::numeric) rate
from tagged where split='train' group by sport, bkt having count(*) >= 8
), platt as (
select sport,
regr_slope(ln(greatest(least(rate,.98),.02)/(1-greatest(least(rate,.98),.02))), mean_lg) b,
regr_intercept(ln(greatest(least(rate,.98),.02)/(1-greatest(least(rate,.98),.02))), mean_lg) a
from tb group by sport
), hb as (
select h.sport, h.bkt, count(*) n, avg(h.p) pred_raw,
avg(1/(1+exp(-(pl.a+pl.b*h.lg)))) pred_platt,
avg(coalesce(tb.rate,h.p)) pred_iso, avg(h.won::numeric) actual
from tagged h join platt pl on pl.sport=h.sport
left join tb on tb.sport=h.sport and tb.bkt=h.bkt
where h.split='holdout' group by h.sport,h.bkt having count(*) >= 8
)
select sport, sum(n) holdout_n, count(*) buckets,
sum(n*abs(pred_raw-actual))/sum(n) raw_reliability_dev,
sum(n*abs(pred_platt-actual))/sum(n) platt_reliability_dev,
sum(n*abs(pred_iso-actual))/sum(n) isotonic_reliability_dev
from hb group by sport order by sport;
-- Resolution (ordering survives?) is the corr(pred, won) variant of the same CTEs.
+42
View File
@@ -0,0 +1,42 @@
-- pwin-timeforward.sql — calibration refresh on the CURRENT MLB sample (2026-08-01)
-- MEASURE-ONLY. Time-forward: earlier games fit/observe, later games prove.
--
-- Deliberately RULER-INDEPENDENT, and that is the finding: reliability
-- (predicted vs actual hit rate) and resolution (does higher p_win hit more)
-- are both p_win-vs-outcome measures. No fair_prob appears anywhere below,
-- because none can. This is why the consensus ruler cannot change the
-- calibration verdict.
--
-- reliability = n-weighted mean |predicted - actual| across deciles. NOTE:
-- mean|p - outcome| on 0/1 rows is NOT calibration -- it is noise-dominated
-- individual error. Bucket first.
--
-- MLB only. WNBA abstains on its own data and is not re-litigated here.
--
-- CONTAMINATION EXCLUSION (2026-08-02, MANDATORY). Rows whose price/book/takeable
-- were stamped from a NON-TAKEABLE book (DFS / offshore / exchange) between
-- 2026-08-01 and the write-path fix are tagged `quarantine_reason LIKE
-- 'nontakeable_book%'`. They are EXCLUDED here and must never be pooled with
-- clean rows: their locked price -- and therefore the `takeable` flag computed
-- from it -- describes a market you could not have bet.
with base as (
select sport, game_date, p_win::numeric p, (outcome='hit')::int won,
ntile(2) over (order by game_date, id) half
from public.ledger_entries
where sport='mlb' and user_id is null
and (quarantine_reason is null or quarantine_reason not like 'nontakeable_book%') and outcome in ('hit','miss') and p_win is not null),
s as (select *, case when half=1 then 'train' else 'holdout' end split from base),
b as (select split, width_bucket(p, 0.0, 1.0, 10) bkt,
count(*) n, avg(p) pred, avg(won::numeric) actual
from s group by 1,2)
select split,
sum(n) total_n,
count(*) buckets,
round(sum(n*abs(pred-actual))/sum(n),4) reliability_mean_abs_dev,
round((select corr(p, won::numeric) from s s2 where s2.split=b.split)::numeric,4) resolution_corr,
round((select avg(won::numeric) from s s3 where s3.split=b.split)::numeric,4) base_rate,
(select min(game_date) from s s4 where s4.split=b.split) first_game,
(select max(game_date) from s s5 where s5.split=b.split) last_game
from b group by split order by split desc;
+184
View File
@@ -0,0 +1,184 @@
#!/usr/bin/env node
'use strict';
/**
* PHASES 0-1 — is rbi's 14.51% real, and where does it come from?
*
* PHASE 0 is a Beck gate. The 14.51% came from the same decomposition harness
* whose paging helper produced a false null three times tonight (composite PKs,
* order-by-id, error swallowed). So the row count is asserted against an
* independent exact count, every page is error-checked, and 20 raw rows are
* printed so the decile arithmetic can be audited by hand.
*
* PHASE 1 asks what the resolution IS. Resolution rewards a forecast for
* separating outcomes — but a forecast can separate outcomes by knowing WHO is
* batting rather than anything about tonight. Three nested forecasts:
*
* PLAYER BASE RATE leave-one-out frequency for that hitter, nothing else.
* Its resolution is pure across-player spread.
* LINEUP SLOT mean rate for that batting-order position. Real
* predictive signal, but ROLE, not skill.
* THE MODEL served p_win.
*
* The part that behaves like our doctrine's "skill" is what the model resolves
* WITHIN a stratum of similar players — measured directly by stratifying on the
* player's own base rate and pooling the within-stratum resolutions.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const { nameKey } = require('../src/utils/playerName');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const SEQ = path.join(process.cwd(), '.seq-cache', 'sequences.json');
const STATS = ['rbi', 'hits', 'total_bases', 'runs'];
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
/** Error-checked pager with an explicit order column. */
async function page(sb, t, sel, orderBy, apply) {
const out = [];
for (let i = 0; ; i += 1000) {
const q = apply ? apply(sb.from(t).select(sel)) : sb.from(t).select(sel);
const { data, error } = await q.order(orderBy, { ascending: true }).range(i, i + 999);
if (error) throw new Error(`${t}: ${error.message}`);
if (!data || !data.length) break;
out.push(...data);
if (data.length < 1000) break;
}
return out;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
/** Resolution alone: weighted spread of bin realized rates about the base rate. */
function resolutionOf(rows, bins = 10) {
const base = mean(rows.map((r) => r.won));
let res = 0;
const table = [];
for (let k = 0; k < bins; k += 1) {
const lo = k / bins; const hi = (k + 1) / bins;
const sl = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (!sl.length) continue;
const w = sl.length / rows.length;
const ok = mean(sl.map((r) => r.won));
res += w * (ok - base) ** 2;
table.push({ bin: `${lo.toFixed(1)}-${hi.toFixed(1)}`, n: sl.length, forecast: r4(mean(sl.map((r) => r.p))), realized: r4(ok) });
}
return { base_rate: r4(base), resolution: r5(res), uncertainty: r5(base * (1 - base)), share: r4(res / (base * (1 - base))), table };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
// Batting-order slot, reconstructed point-in-time from play-by-play.
const { games } = JSON.parse(fs.readFileSync(SEQ, 'utf8'));
const slotOf = new Map();
for (const g of games) {
for (const half of ['top', 'bottom']) {
const seen = []; const set = new Set();
for (const p of g.pas.filter((x) => x.half === half)) {
if (!set.has(p.batter)) { set.add(p.batter); seen.push(p); }
if (seen.length >= 9) break;
}
seen.forEach((p, i) => slotOf.set(`${g.date}|${nameKey(p.batter_name || '')}`, i + 1));
}
}
const out = { PHASE_0: {}, PHASE_1: {} };
for (const stat of STATS) {
// PHASE 0 — exact count first, then the paged pull must match it.
const { count: exact, error: cErr } = await sb.from('model_snapshots')
.select('*', { count: 'exact', head: true }).eq('sport', 'mlb').eq('stat', stat);
if (cErr) throw new Error(`count ${stat}: ${cErr.message}`);
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, player_name, line, side, p_win, refused', 'id',
(q) => q.eq('sport', 'mlb').eq('stat', stat));
if (snaps.length !== exact) throw new Error(`PHASE 0 FAIL ${stat}: paged ${snaps.length} != exact ${exact}`);
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({ propKey: [r.game_date, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const rows = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const over = v > L;
rows.push({
date: r.game_date, key: r.player_key, name: r.player_name, line: L, side: r.side,
actual: v, p: knownNumber(r.p_win),
won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0,
slot: slotOf.get(`${r.game_date}|${r.player_key}`) ?? null,
});
}
const model = resolutionOf(rows);
out.PHASE_0[stat] = { exact_rows: exact, paged_rows: snaps.length, scorable: rows.length, model_resolution: model.resolution, model_share: model.share, deciles: model.table };
if (stat === 'rbi') {
out.PHASE_0.rbi_raw_sample = rows.slice(0, 20).map((r) => ({ name: r.name, date: r.date, line: r.line, side: r.side, p_win: r.p, actual_rbi: r.actual, won: r.won }));
}
// ── (a) PLAYER BASE RATE, leave-one-out ──
const byPlayer = new Map();
for (const r of rows) { const c = byPlayer.get(r.key) || { n: 0, w: 0 }; c.n += 1; c.w += r.won; byPlayer.set(r.key, c); }
const baseRows = rows.filter((r) => byPlayer.get(r.key).n >= 3)
.map((r) => { const c = byPlayer.get(r.key); return { ...r, p: (c.w - r.won) / (c.n - 1) }; });
const baseOnly = resolutionOf(baseRows);
// ── (b) LINEUP SLOT, leave-one-out ──
const bySlot = new Map();
for (const r of rows) { if (r.slot == null) continue; const c = bySlot.get(r.slot) || { n: 0, w: 0 }; c.n += 1; c.w += r.won; bySlot.set(r.slot, c); }
const slotRows = rows.filter((r) => r.slot != null && bySlot.get(r.slot).n >= 10)
.map((r) => { const c = bySlot.get(r.slot); return { ...r, p: (c.w - r.won) / (c.n - 1) }; });
const slotOnly = slotRows.length ? resolutionOf(slotRows) : null;
// ── (c) WITHIN-STRATUM: does the model still separate similar players? ──
// Stratify on the player's own base rate, then pool the model's resolution
// computed INSIDE each stratum. Across-player spread is held constant, so
// what survives is discrimination between comparable hitters.
const strata = [[0, 0.45], [0.45, 0.6], [0.6, 0.75], [0.75, 1.01]];
let within = 0; let wTot = 0; const strataDetail = [];
for (const [lo, hi] of strata) {
const sl = rows.filter((r) => { const c = byPlayer.get(r.key); if (!c || c.n < 3) return false; const b = c.w / c.n; return b >= lo && b < hi; });
if (sl.length < 40) continue;
const rr = resolutionOf(sl);
within += sl.length * rr.resolution; wTot += sl.length;
strataDetail.push({ stratum: `${lo}-${hi}`, n: sl.length, base: rr.base_rate, resolution: rr.resolution });
}
const withinRes = wTot ? within / wTot : null;
out.PHASE_1[stat] = {
n: rows.length,
model_resolution: model.resolution,
model_share_of_variance: model.share,
a_player_base_rate_only: { n: baseRows.length, resolution: baseOnly.resolution, share: baseOnly.share },
b_lineup_slot_only: slotOnly ? { n: slotRows.length, resolution: slotOnly.resolution, share: slotOnly.share } : null,
c_within_stratum_resolution: r5(withinRes),
strata: strataDetail,
base_rate_explains_pct: baseOnly.resolution ? r4(Math.min(1, baseOnly.resolution / model.resolution)) : null,
within_stratum_share_of_model: withinRes != null && model.resolution ? r4(withinRes / model.resolution) : null,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error('FAILED:', e.message); process.exit(1); });
const r5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
+137
View File
@@ -0,0 +1,137 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 5 — rebuild grade bands on p_win_calibrated, for DEPLOYED stats only.
*
* total_bases is the only stat that cleared LODO, so it is the only one whose
* bands are rebuilt on calibrated values. The rest keep base-rate bands built on
* raw p_win, and the reason is named rather than left to inference.
*
* The two-bar rule still applies and still bites: TB is now CALIBRATED but no
* factor is PROVEN for it (barrel, exit velo and hard-contact-allowed were all
* THEATER), so the bands remain a base-rate read — now an honestly-numbered one.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const lp = require('../src/services/model/lowParamCalibrator');
const gb = require('../src/services/model/gradeBands');
const guards = require('../src/services/model/calibrationGuards');
const tl = require('../src/services/model/testLedger');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STAT = process.env.BAND_STAT || 'total_bases';
const PAGE = 1000;
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused, archetype',
(q) => q.eq('sport', 'mlb').eq('stat', STAT));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
const rows = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[STAT](b));
if (v === null) continue;
const over = v > L;
rows.push({
date: r.game_date, p: knownNumber(r.p_win),
won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0,
archetype: String(r.archetype || 'UNLABELLED').toUpperCase(),
});
}
rows.sort((a, b) => String(a.date).localeCompare(String(b.date)));
// Point-in-time map, then apply forward.
const dates = [...new Set(rows.map((r) => r.date))].sort();
const perDate = new Map();
for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) { acc += perDate.get(d); if (acc >= rows.length * 0.45) { cut = d; break; } }
// Bands are built on the SERVED values. hits and total_bases serve the
// low-parameter correction; rbi and runs serve raw, so their bands are raw.
const DEPLOYED = ['hits', 'total_bases'];
const fitRows = rows.filter((r) => r.date < cut);
const evalRows = rows.filter((r) => r.date >= cut);
let applied;
let basis;
if (DEPLOYED.includes(STAT)) {
const model = lp.fitPlatt(fitRows);
applied = (!model || model.refused)
? { ok: false, reason: 'low-parameter fit refused', rows: [] }
: { ok: true, rows: evalRows.map((r) => ({ ...r, pc: lp.applyPlatt(model, r.p) })).filter((r) => r.pc != null) };
basis = 'p_win_lowparam (SERVED, provisional)';
} else {
applied = { ok: true, rows: evalRows.map((r) => ({ ...r, pc: r.p })) };
basis = 'raw p_win (this stat serves raw)';
}
if (!applied.ok) { console.log(JSON.stringify({ stat: STAT, refused: applied.reason })); process.exit(0); }
const mc = await tl.recordAndCount(tl.supabaseStore(sb), []).catch(() => ({ cumulative_tests: 1 }));
const byArch = new Map();
for (const r of applied.rows) {
if (!byArch.has(r.archetype)) byArch.set(r.archetype, []);
byArch.get(r.archetype).push({ p: r.pc, won: r.won });
}
const out = [];
for (const [arch, rs] of [...byArch.entries()].sort((a, b) => b[1].length - a[1].length)) {
out.push(gb.buildBands(rs, {
archetype: arch,
cumulativeTests: mc.cumulative_tests,
// TB is CALIBRATED (provisional) but no factor is PROVEN for it.
proven: false,
calibrated: DEPLOYED.includes(STAT),
}));
}
console.log(JSON.stringify({
stat: STAT,
basis,
eval_rows: applied.rows.length,
cumulative_tests: mc.cumulative_tests,
two_bar_note: 'calibrated YES, proven NO -> bands stay a base-rate read, now honestly numbered',
bands: out,
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
+170
View File
@@ -0,0 +1,170 @@
#!/usr/bin/env node
'use strict';
/**
* reconstruct-game-environment — GIVE THE PAST GAMES THEIR REAL CONDITIONS.
*
* `game_context` has never held a single weather reading. The reason is not the
* fetcher, which is correct and points at Open-Meteo's ARCHIVE endpoint; it is
* that nothing ever joined. The ledger keys a game as
* `mlb:2026-08-03:WashingtonNationals@PhiladelphiaPhillies` and game_context
* keys it as `mlb:823437`, so every lookup missed and the columns stayed NULL —
* which reads exactly like "the weather was unavailable" rather than "the two
* tables have never been introduced." The same class of failure as the doubled
* /leaderboard path: graceful degradation wearing the mask of honest absence.
*
* This walks the dates in the settled ledger, resolves each slug to the real
* statsapi game and venue, and writes a game_context row keyed by the LEDGER's
* slug so the join exists. Then it pulls the actual archived weather for that
* date and location.
*
* ARCHIVE, NOT FORECAST — asking the forecast endpoint about a past date returns
* a re-forecast, which is a model's opinion about the past, not the past. Absent
* stays NULL; nothing here is imputed.
*
* SUPABASE_URL=... node scripts/reconstruct-game-environment.js
*/
require('dotenv').config();
const axios = require('axios');
const { createClient } = require('@supabase/supabase-js');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const SCHEDULE = (d) => `https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}&hydrate=venue(location)`;
const ARCHIVE = (lat, lon, date) =>
`https://archive-api.open-meteo.com/v1/archive?latitude=${lat}&longitude=${lon}`
+ `&start_date=${date}&end_date=${date}`
+ '&hourly=temperature_2m,wind_speed_10m,wind_direction_10m,precipitation'
+ '&temperature_unit=fahrenheit&wind_speed_unit=mph';
const squash = (s) => String(s || '').toLowerCase().replace(/[^a-z]/g, '');
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
const get = async (url) => (await axios.get(url, { timeout: 60_000 })).data;
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const led = await page(sb, 'ledger_entries', 'game_id, game_date, stat, outcome, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).in('stat', ['hits', 'total_bases'])
.in('outcome', ['hit', 'miss']));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book'));
// Distinct games, as the LEDGER names them.
const games = new Map();
for (const r of clean) if (r.game_id && !games.has(r.game_id)) games.set(r.game_id, r.game_date);
const dates = [...new Set([...games.values()])].sort();
console.error(`[env] ${games.size} distinct settled games across ${dates.length} dates`);
// date -> statsapi games, indexed by the same squashed away@home the slug uses.
const resolved = new Map();
const venues = new Map();
for (const d of dates) {
let sched = null;
try { sched = await get(SCHEDULE(d)); } catch { sched = null; }
for (const day of (sched && sched.dates) || []) {
for (const g of day.games || []) {
const away = squash(g.teams?.away?.team?.name);
const home = squash(g.teams?.home?.team?.name);
resolved.set(`${d}|${away}@${home}`, g);
const v = g.venue || {};
if (v.id && !venues.has(v.id)) {
const loc = v.location || {};
venues.set(v.id, {
venue_id: v.id,
venue_name: v.name || null,
lat: loc.defaultCoordinates?.latitude ?? null,
lon: loc.defaultCoordinates?.longitude ?? null,
});
}
}
}
}
// Join each ledger slug to its real game + venue.
const rows = []; let unmatched = 0;
for (const [gid, date] of games) {
const m = /^mlb:(\d{4}-\d{2}-\d{2}):(.+?)@(.+)$/.exec(gid);
if (!m) { unmatched += 1; continue; }
const g = resolved.get(`${m[1]}|${squash(m[2])}@${squash(m[3])}`);
if (!g) { unmatched += 1; continue; }
rows.push({
game_id: gid, // the LEDGER's key — this is the whole fix
game_date: date,
venue_id: g.venue?.id ?? null,
source_game_pk: g.gamePk ?? null,
});
}
console.error(`[env] matched ${rows.length}, unmatched ${unmatched}, venues seen ${venues.size}`);
for (let i = 0; i < rows.length; i += 200) {
const { error } = await sb.from('game_context')
.upsert(rows.slice(i, i + 200), { onConflict: 'game_id' });
if (error) console.error('[env] context write failed:', error.message);
}
// ACTUAL archived weather, one call per (venue, date) that we need.
const need = new Map();
for (const r of rows) {
const v = venues.get(r.venue_id);
if (!v || v.lat == null || v.lon == null) continue;
need.set(`${r.venue_id}|${r.game_date}`, { v, date: r.game_date });
}
console.error(`[env] fetching ${need.size} venue-days of archived weather`);
const wx = new Map();
for (const [k, { v, date }] of need) {
try {
const p = await get(ARCHIVE(v.lat, v.lon, date));
const h = p && p.hourly;
if (h && Array.isArray(h.time) && h.time.length) {
const i = Math.min(h.time.length - 1, 19); // ~7pm local, typical first pitch
wx.set(k, {
wx_temp_f: h.temperature_2m?.[i] ?? null,
wx_wind_speed_mph: h.wind_speed_10m?.[i] ?? null,
wx_wind_direction_deg: h.wind_direction_10m?.[i] ?? null,
wx_precip_mm: h.precipitation?.[i] ?? null,
wx_source: 'open_meteo_archive',
});
}
} catch { /* absent stays absent */ }
}
let withWx = 0;
for (let i = 0; i < rows.length; i += 200) {
const batch = rows.slice(i, i + 200).map((r) => {
const w = wx.get(`${r.venue_id}|${r.game_date}`);
if (w) withWx += 1;
return w ? { ...r, ...w } : r;
});
const { error } = await sb.from('game_context').upsert(batch, { onConflict: 'game_id' });
if (error) console.error('[env] weather write failed:', error.message);
}
console.log(JSON.stringify({
settled_games: games.size,
matched_to_statsapi: rows.length,
unmatched,
distinct_venues: venues.size,
venue_days_requested: need.size,
venue_days_returned: wx.size,
game_rows_with_actual_weather: withWx,
source: 'open_meteo_archive (actual, not re-forecast)',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+132
View File
@@ -0,0 +1,132 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 2 — put a number on the resolution ceiling.
*
* Murphy's decomposition: Brier = reliability - resolution + uncertainty.
*
* reliability how far each bin's realized rate sits from its forecast (lower
* is better; this is what calibration fixes)
* resolution how far the bins' realized rates spread from the base rate
* (HIGHER is better; this is discrimination, and NO amount of
* calibration can create it)
* uncertainty the base rate's own variance -- a property of the event
*
* Calibration moves reliability and leaves resolution untouched by construction:
* a monotone map relabels bins without re-sorting the rows inside them. So if
* resolution is near zero, honest numbers are all calibration can ever deliver.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const lp = require('../src/services/model/lowParamCalibrator');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const DEPLOYED = ['hits', 'total_bases'];
const PAGE = 1000;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error; if (!data.length) break; o.push(...data); if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
/** Murphy decomposition over K equal-width bins. */
function decompose(rows, bins = 10) {
const base = mean(rows.map((r) => r.won));
const uncertainty = base * (1 - base);
let reliability = 0; let resolution = 0;
const table = [];
for (let k = 0; k < bins; k += 1) {
const lo = k / bins; const hi = (k + 1) / bins;
const slice = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (!slice.length) continue;
const w = slice.length / rows.length;
const fk = mean(slice.map((r) => r.p));
const ok = mean(slice.map((r) => r.won));
reliability += w * (fk - ok) ** 2;
resolution += w * (ok - base) ** 2;
table.push({ bin: [round2(lo), round2(hi)], n: slice.length, forecast: round4(fk), realized: round4(ok) });
}
return {
base_rate: round4(base),
reliability: round5(reliability),
resolution: round5(resolution),
uncertainty: round5(uncertainty),
brier_check: round5(reliability - resolution + uncertainty),
/** What share of the event's variance the model actually explains. */
resolution_share_of_uncertainty: round4(resolution / uncertainty),
bins: table,
};
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots', 'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const out = {};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const over = v > L;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
if (rows.length < 100) continue;
const raw = decompose(rows);
let served = null;
if (DEPLOYED.includes(stat)) {
const m = lp.fitPlatt(rows);
if (m && !m.refused) {
const cal = rows.map((r) => ({ ...r, p: lp.applyPlatt(m, r.p) })).filter((r) => knownNumber(r.p) !== null);
served = decompose(cal);
}
}
out[stat] = {
n: rows.length,
deployed: DEPLOYED.includes(stat),
raw,
served,
resolution_change_from_calibration: served ? round5(served.resolution - raw.resolution) : null,
reliability_change_from_calibration: served ? round5(served.reliability - raw.reliability) : null,
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round5 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 100000) / 100000);
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round2 = (v) => Math.round(v * 100) / 100;
+84
View File
@@ -0,0 +1,84 @@
-- ruler-comparison.sql — Order: MLB CALIBRATION RE-RUN vs CONSENSUS RULER (2026-08-01)
-- MEASURE-ONLY. Run against prod (Supabase MCP execute_sql).
--
-- WHAT THIS DOES AND DOES NOT ANSWER
--
-- It does NOT re-fit the p_win calibration. It cannot: `estimateProbability`
-- takes {gameLogs, line, statType, features} and never sees a market price, and
-- the calibration query fits p_win against OUTCOMES. Reliability and resolution
-- are both p_win-vs-outcome measures, so the ruler cannot enter either. See
-- pwin-timeforward.sql for the (ruler-independent) calibration refresh.
--
-- What IS ruler-dependent is EDGE (p_win - fair_prob). This measures whether
-- swapping the incumbent single-book ruler for a median consensus rescues it.
--
-- TIMING IS HELD CONSTANT: both rulers are read at CLOSE. The lock-time
-- reconstruction is impossible at usable n (only 43 settled rows join
-- lock_lines with >=2 two-sided books), and mixing a lock-time incumbent with
-- a close-time consensus would confound WHEN with WHAT.
--
-- LIMITATION, load-bearing: closing_captures contains ONLY MODEL books
-- (draftkings/betmgm/betrivers/fanduel/pinnacle). Exchange quotes were never
-- stored, because normalizeProps discarded them until 2026-08-01. So this can
-- only test a US-books-median ruler, NOT the exchange-inclusive consensus. The
-- exchange ruler is untestable on existing data at any n.
--
-- CONTAMINATION EXCLUSION (2026-08-02, MANDATORY). Rows whose price/book/takeable
-- were stamped from a NON-TAKEABLE book (DFS / offshore / exchange) between
-- 2026-08-01 and the write-path fix are tagged `quarantine_reason LIKE
-- 'nontakeable_book%'`. They are EXCLUDED here and must never be pooled with
-- clean rows: their locked price -- and therefore the `takeable` flag computed
-- from it -- describes a market you could not have bet.
with imp as (
select id, player_key, stat, game_date, line, side, p_win, book, (outcome='hit')::int won
from public.ledger_entries
where sport='mlb' and user_id is null
and (quarantine_reason is null or quarantine_reason not like 'nontakeable_book%') and outcome in ('hit','miss') and p_win is not null),
-- latest CLOSE capture per (prop, book); two-sided only -- a one-sided quote
-- cannot be de-vigged, so it cannot price a ruler.
cap as (
select distinct on (player_key,stat,game_date,line,book)
player_key,stat,game_date,line,book,over_odds,under_odds
from public.closing_captures
where sport='mlb' and over_odds is not null and under_odds is not null
order by player_key,stat,game_date,line,book,captured_at desc),
d as (select *,
case when over_odds>0 then 100.0/(over_odds+100) else (-over_odds)/((-over_odds)+100.0) end po,
case when under_odds>0 then 100.0/(under_odds+100) else (-under_odds)/((-under_odds)+100.0) end pu
from cap),
-- multiplicative two-way de-vig, per book (matches src/utils/devig.js)
f as (select player_key,stat,game_date,line,book, po/(po+pu) fo, pu/(po+pu) fu
from d where po+pu > 0),
j as (select i.*, f.book ref_book,
case when lower(i.side)='under' then f.fu else f.fo end fair_side
from imp i join f
on f.player_key=i.player_key and f.stat=i.stat
and f.game_date=i.game_date and f.line=i.line),
a as (select id, p_win, won, game_date,
count(*) n_books,
percentile_cont(0.5) within group (order by fair_side) cons, -- v2: MEDIAN
max(case when ref_book=book then fair_side end) own -- v1: the locked book
from j group by 1,2,3,4),
e as (select *, p_win-own edge_v1, p_win-cons edge_v2
from a where n_books>=2 and own is not null)
select count(*) n,
round(avg(won)::numeric,4) base_rate,
round(avg(abs(cons-own))::numeric,4) mean_abs_ruler_gap,
round(avg(cons-own)::numeric,4) mean_signed_ruler_gap,
round(corr(edge_v1, won::numeric)::numeric,4) corr_edge_v1_won,
round(corr(edge_v2, won::numeric)::numeric,4) corr_edge_v2_won,
round(corr(p_win, won::numeric)::numeric,4) corr_pwin_won,
round(avg(edge_v1) filter (where won=1)::numeric,4) edge_v1_winners,
round(avg(edge_v1) filter (where won=0)::numeric,4) edge_v1_losers,
round(avg(edge_v2) filter (where won=1)::numeric,4) edge_v2_winners,
round(avg(edge_v2) filter (where won=0)::numeric,4) edge_v2_losers
from e;
+65
View File
@@ -0,0 +1,65 @@
#!/usr/bin/env node
/**
* RUN THE BACKTEST HARNESS (Session 64).
*
* Joins model_snapshots (the model's INPUTS + prediction) to ledger_entries
* (the single source of truth for OUTCOMES) on the natural key, and runs the
* harness. Outcomes are NEVER denormalized onto snapshots.
*
* Join key: (sport, player_key, stat, line, side, game_date).
* Verified empirically: 283 clean 1:1 joins, ZERO ambiguity. `game_id` is NOT
* usable — 400/550 snapshot rows carry `UNK@UNK` because home/away team names
* weren't threaded into the grader until Session 64 Order 1.6.
*
* Rows that don't join are EXPECTED, not errors: retention stores BOTH sides
* of every prop plus refusals, while the ledger keeps only the graded side of
* non-refused props.
*
* node scripts/run-backtest.js <rows.json> # rows exported via SQL
*/
const harness = require('../src/services/backtestHarness');
const rows = JSON.parse(require('fs').readFileSync(process.argv[2], 'utf8'));
const report = harness.runBacktest(rows, {});
const pad = (s, n) => String(s).padEnd(n);
console.log('══════════ VYNDR BACKTEST HARNESS ══════════');
console.log(`generated_at : ${report.generated_at}`);
console.log(`min_sample : ${report.min_sample}`);
console.log(`VERDICT : ${report.verdict} (can_validate=${report.can_validate})`);
console.log('\n--- denominator ---');
Object.entries(report.counts).forEach(([k, v]) => console.log(` ${pad(k, 22)} ${v}`));
console.log('\n--- grade buckets (4-letter) ---');
for (const b of report.grade_buckets.sort((a, z) => z.n - a.n)) {
console.log(b.status === 'OK'
? ` ${pad(b.bucket, 4)} n=${pad(b.n, 5)} hit=${(b.hit_rate * 100).toFixed(1)}% 95% CI [${(b.ci_low * 100).toFixed(1)}, ${(b.ci_high * 100).toFixed(1)}]`
: ` ${pad(b.bucket, 4)} n=${pad(b.n, 5)} INSUFFICIENT — need ${b.need} (short by ${b.short_by})`);
}
console.log('\n--- grade buckets (11-step) ---');
for (const b of report.grade_11_buckets.sort((a, z) => z.n - a.n)) {
console.log(b.status === 'OK'
? ` ${pad(b.bucket, 4)} n=${pad(b.n, 5)} hit=${(b.hit_rate * 100).toFixed(1)}%`
: ` ${pad(b.bucket, 4)} n=${pad(b.n, 5)} INSUFFICIENT`);
}
console.log(`\n--- monotonicity: ${report.monotonicity.verdict} ---`);
report.monotonicity.comparisons.forEach((c) => console.log(
` ${c.higher} vs ${c.lower}: ${c.distinguishable ? (c.holds ? 'HOLDS' : 'BROKEN') : 'not distinguishable on this sample'}`,
));
console.log('\n--- probability calibration ---');
console.log(report.probability.n
? ` n=${report.probability.n} Brier=${report.probability.brier.toFixed(4)}`
: ' n=0 — no stored p_win on any joinable settled row');
console.log('\n--- strata (never mixed) ---');
report.strata.forEach((s) => console.log(` ${pad(s.sport, 6)} ${pad(s.model_version, 24)} n=${s.n}`));
require('fs').writeFileSync(
process.argv[3] || '/tmp/backtest-report.json',
JSON.stringify(report, null, 2),
);
console.log(`\nfull report → ${process.argv[3] || '/tmp/backtest-report.json'}`);
+89
View File
@@ -0,0 +1,89 @@
#!/usr/bin/env node
/**
* MANUAL REGRADE TRIGGER (Session 64, Phase 1).
*
* Fires the snapshot pipeline on demand so a fix can be verified in minutes
* instead of waiting for the 14/19/22/01/03 UTC cron. Runs INSIDE the API
* container, so it needs no VYNDR_INTERNAL_KEY and no open HTTP surface:
*
* docker exec <api-container> node scripts/run-snapshot.js mlb
* docker exec <api-container> node scripts/run-snapshot.js all
* docker exec <api-container> node scripts/run-snapshot.js mlb --settle
*
* It runs the SAME `snapshotService.runSnapshot` the cron runs — including the
* team-stats refresh that populates `opp_rank_stat` — so what you verify is what
* production does, not a parallel code path.
*
* `--settle` additionally runs the outcome + ledger settle pass first, matching
* the scheduler's real order (settle yesterday, then grade today).
*
* Prints a grade-distribution summary at the end, which is the thing you
* actually want when verifying a grading change.
*/
const args = process.argv.slice(2);
const target = (args[0] || 'all').toLowerCase();
const doSettle = args.includes('--settle');
function tally(list, key) {
return list.reduce((m, g) => { const k = g?.[key] ?? 'null'; m[k] = (m[k] || 0) + 1; return m; }, {});
}
(async () => {
const snapshotService = require('../src/services/snapshotService');
if (doSettle) {
console.log('--- settle pass (outcomes + ledger) ---');
try {
const outcomeService = require('../src/services/outcomeService');
for (const sp of ['mlb', 'wnba', 'nba']) {
try {
const r = await outcomeService.settleSnapshot(sp);
console.log(` ${sp}: ${JSON.stringify(r)}`);
} catch (e) { console.warn(` ${sp}: settle failed — ${e.message}`); }
}
const ledger = require('../src/services/ledgerService');
if (typeof ledger.settleLedger === 'function') {
for (const sp of ['mlb', 'wnba']) {
try { console.log(` ledger ${sp}: ${JSON.stringify(await ledger.settleLedger(sp))}`); }
catch (e) { console.warn(` ledger ${sp}: ${e.message}`); }
}
}
} catch (e) { console.warn('settle pass failed:', e.message); }
}
const sports = target === 'all' ? ['mlb', 'wnba', 'nba', 'soccer'] : [target];
console.log(`--- snapshot: ${sports.join(', ')} ---`);
const results = [];
for (const sp of sports) {
const t = Date.now();
try {
const r = await snapshotService.runSnapshot(sp);
console.log(` ${sp}: status=${r.status} grades=${r.gradeCount}${r.reason ? ` reason=${r.reason}` : ''} (${Math.round((Date.now() - t) / 1000)}s)`);
results.push({ sp, r });
} catch (e) {
console.error(` ${sp}: THREW — ${e.message}`);
}
}
// Distribution — the point of running this by hand.
console.log('\n--- GRADE DISTRIBUTION (from the freshly written cache) ---');
const { cacheGet } = require('../src/utils/redis');
for (const { sp } of results) {
try {
const snap = await cacheGet(`snapshot:${sp}:latest`);
const grades = (snap && Array.isArray(snap.grades)) ? snap.grades : [];
if (!grades.length) { console.log(` ${sp}: (no grades)`); continue; }
const withEv = grades.filter((g) => Number.isFinite(Number(g.ev_pct))).length;
const withP = grades.filter((g) => Number.isFinite(Number(g.p_win))).length;
const withOpp = grades.filter((g) => g.opp_rank_stat != null).length;
console.log(` ${sp}: n=${grades.length} ${JSON.stringify(tally(grades, 'grade'))}`);
console.log(` confidence: ${JSON.stringify(tally(grades, 'confidence'))}`);
console.log(` p_win present: ${withP}/${grades.length} · ev_pct present: ${withEv}/${grades.length} · opp_rank on grade: ${withOpp}`);
} catch (e) { console.warn(` ${sp}: could not read cache — ${e.message}`); }
}
await new Promise((r) => process.stdout.write('', r));
process.exit(0);
})().catch((e) => { console.error('run-snapshot failed:', e); process.exit(1); });
+249
View File
@@ -0,0 +1,249 @@
#!/usr/bin/env node
'use strict';
/**
* settle-model-snapshots — pay the standing debt.
*
* 71,192 snapshot rows have never carried an outcome. They are the retention
* table built for exactly this kind of replay, and until they are settled every
* measurement in this programme runs on the far smaller ledger slice.
*
* ── OUTCOME IS SIDE-ALIGNED, NOT RAW ─────────────────────────────────────
* The order specifies `outcome = 1[realized > line]`. That is the OVER
* perspective, and it would be backwards for every under-side prop — `p_win` is
* side-aligned (verified: TB mean p_win 0.5698 against a 0.5074 side-won rate),
* so a raw over-indicator would silently invert the target on the under rows and
* make calibration measure the wrong thing.
*
* So: `actual_value` stores the realized stat (raw, unopinionated) and `outcome`
* stores whether the GRADED SIDE won. Deviation from the literal order, stated
* because it changes the number.
*
* ── INTEGRITY (hard-fail) ────────────────────────────────────────────────
* conservation settled + unresolvable + orphaned == candidates
* no dupes one write per snapshot id
* no orphans a settled row must have matched a real box score
* prediction-time logging captured_at must PRECEDE the game date; a row
* logged after the fact is not a prediction and is refused
*
* node scripts/settle-model-snapshots.js # dry run, verifies only
* SETTLE_WRITE=1 node scripts/settle-model-snapshots.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const axios = require('axios');
const { createClient } = require('@supabase/supabase-js');
const { nameKey } = require('../src/utils/playerName');
const { knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const WRITE = process.env.SETTLE_WRITE === '1';
const BOX_CACHE = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
/** Realized value per stat, from the box-score batting line. */
const FIELD = Object.freeze({
hits: (b) => knownNumber(b.hits),
total_bases: (b) => knownNumber(b.totalBases),
rbi: (b) => knownNumber(b.rbi),
runs: (b) => knownNumber(b.runs),
});
const get = async (url) => (await axios.get(url, { timeout: 45_000 })).data;
/** Eastern first pitch, conservatively. Anything at or after this is in-game. */
const FIRST_PITCH_ET_HOUR = 19;
/**
* Was this row logged BEFORE the games it grades?
*
* The pipeline runs on UTC cron hours, so a 01:00-UTC cycle is 21:00 the
* PREVIOUS evening in Eastern -- same game date, three hours into the slate.
*/
function isPreGame(capturedAt, gameDate) {
if (!capturedAt || !gameDate) return false;
const cap = new Date(capturedAt);
if (Number.isNaN(cap.getTime())) return false;
const et = new Date(cap.getTime() - 4 * 3600 * 1000); // EDT
const etDate = et.toISOString().slice(0, 10);
if (etDate < String(gameDate)) return true; // day before, fine
if (etDate > String(gameDate)) return false; // day after, post-game
return et.getUTCHours() < FIRST_PITCH_ET_HOUR;
}
async function pool(items, fn, n = 6) {
const out = []; let i = 0;
await Promise.all(Array.from({ length: n }, async () => {
while (i < items.length) {
const idx = i; i += 1;
try { out[idx] = await fn(items[idx]); } catch { out[idx] = null; }
}
}));
return out.filter(Boolean);
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
// STABLE ORDER. model_snapshots is a LIVE table -- the snapshot cron writes
// to it at 14/19/22/1/3 UTC -- and an unordered .range() walk over a table
// being appended to returns overlapping pages. The integrity gate caught
// exactly that on the first run.
const { data, error } = await apply(sb.from(table).select(select))
.order('id', { ascending: true })
.range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/** Box-score batting lines for a date range, cached. */
async function battingLines(dates) {
if (fs.existsSync(BOX_CACHE)) {
const c = JSON.parse(fs.readFileSync(BOX_CACHE, 'utf8'));
if (dates.every((d) => c.dates.includes(d))) return c.lines;
}
const games = [];
for (const d of dates) {
try {
const s = await get(`https://statsapi.mlb.com/api/v1/schedule?sportId=1&date=${d}`);
for (const day of s.dates || []) {
for (const g of day.games || []) {
if (String(g.status && g.status.detailedState) === 'Final') {
games.push({ pk: g.gamePk, date: g.officialDate || d });
}
}
}
} catch { /* absent day */ }
}
console.error(`[settle] ${games.length} final games across ${dates.length} dates`);
const lines = {};
const loaded = await pool(games, async (g) => {
const box = await get(`https://statsapi.mlb.com/api/v1/game/${g.pk}/boxscore`);
const out = [];
for (const side of ['home', 'away']) {
const t = box.teams[side];
if (!t) continue;
for (const id of t.batters || []) {
const pl = t.players[`ID${id}`];
const b = pl && pl.stats && pl.stats.batting;
if (!b || b.atBats == null) continue; // did not bat -> absent, not zero
out.push({
date: g.date,
key: nameKey(pl.person && pl.person.fullName),
name: pl.person && pl.person.fullName,
gamePk: g.pk,
hits: b.hits, totalBases: b.totalBases, rbi: b.rbi, runs: b.runs, atBats: b.atBats,
});
}
}
return out;
});
for (const arr of loaded) for (const r of arr) {
const k = `${r.date}|${r.key}`;
// A doubleheader gives two lines; sum them — the prop covers the day.
if (!lines[k]) lines[k] = { ...r, games: 1 };
else {
lines[k].hits += r.hits; lines[k].totalBases += r.totalBases;
lines[k].rbi += r.rbi; lines[k].runs += r.runs; lines[k].atBats += r.atBats;
lines[k].games += 1;
}
}
fs.mkdirSync(path.dirname(BOX_CACHE), { recursive: true });
fs.writeFileSync(BOX_CACHE, JSON.stringify({ dates, lines }));
return lines;
}
async function main() {
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, player_name, line, side, p_win, refused, outcome',
(q) => q.eq('sport', 'mlb').in('stat', STATS).is('outcome', null));
console.error(`[settle] ${snaps.length} unsettled snapshot rows`);
const dates = [...new Set(snaps.map((r) => r.game_date))].sort();
const lines = await battingLines(dates);
const counts = { candidates: snaps.length, settled: 0, unresolvable: 0, orphaned: 0, post_hoc_logged: 0 };
const updates = [];
const seenIds = new Set();
for (const s of snaps) {
// Belt and braces: ordered pagination should make this impossible, and a
// duplicate would double-count a prediction in every downstream measurement.
if (seenIds.has(s.id)) throw new Error(`INTEGRITY: duplicate snapshot id ${s.id}`);
seenIds.add(s.id);
// A row logged after first pitch is not a prediction.
//
// Measured: cycles at ET 21:00/22:00/23:00 on the game date (10,738 rows)
// were captured DURING or AFTER the games they grade, and a further 664 the
// following morning. Games start ~19:05 ET, so the honest cutoff is ET
// first pitch on the game date -- not a UTC date compare, which both keeps
// post-game 01:00-UTC rows and discards legitimate pre-dawn ones.
if (!isPreGame(s.captured_at, s.game_date)) {
counts.post_hoc_logged += 1; counts.unresolvable += 1; continue;
}
const line = knownNumber(s.line);
if (line === null || !s.side) { counts.unresolvable += 1; continue; }
const b = lines[`${s.game_date}|${s.player_key}`];
if (!b) { counts.orphaned += 1; continue; }
const realized = FIELD[s.stat](b);
if (realized === null) { counts.unresolvable += 1; continue; }
// SIDE-ALIGNED, so it matches how p_win is expressed.
const over = realized > line;
const won = String(s.side).toLowerCase() === 'under' ? !over : over;
updates.push({ id: s.id, outcome: won ? 'hit' : 'miss', actual_value: realized });
counts.settled += 1;
}
// CONSERVATION — hard fail.
const acc = counts.settled + counts.unresolvable + counts.orphaned;
if (acc !== counts.candidates) {
throw new Error(`INTEGRITY: conservation violated ${acc} != ${counts.candidates}`);
}
// Hand-verifiable sample.
const sample = updates.slice(0, 12).map((u) => {
const s = snaps.find((x) => x.id === u.id);
return { player: s.player_name, date: s.game_date, stat: s.stat, line: s.line, side: s.side,
realized: u.actual_value, outcome: u.outcome };
});
if (WRITE) {
let written = 0;
for (let i = 0; i < updates.length; i += 500) {
const batch = updates.slice(i, i + 500);
const results = await Promise.all(batch.map((u) => sb.from('model_snapshots')
.update({ outcome: u.outcome, actual_value: u.actual_value, settled_at: new Date().toISOString(), settlement_source: 'statsapi_boxscore' })
.eq('id', u.id).is('outcome', null)));
written += results.filter((r) => !r.error).length;
}
counts.written = written;
}
console.log(JSON.stringify({
mode: WRITE ? 'WRITE' : 'DRY RUN',
counts,
dates_before: 'ledger-only slice',
snapshot_dates: dates.length,
date_span: [dates[0], dates[dates.length - 1]],
hand_verify_sample: sample,
note: 'outcome is SIDE-ALIGNED (matches p_win); actual_value holds the raw realized stat',
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+279
View File
@@ -0,0 +1,279 @@
#!/usr/bin/env node
'use strict';
/**
* skill-v1-stagea — DOES THE WINDSHIELD BEAT THE REAR-VIEW MIRROR?
*
* Stage A's only question: on settled HITS props, does an archetype-selected,
* skill-based forward projection call the LISTED LINE better than the frequency
* counter that is the reigning champion? If it does not, it is not real yet and
* it does not get promoted. That is the whole test.
*
* ── WHY THIS IS GENUINELY OUT-OF-SAMPLE ──────────────────────────────────
* `statcast_aggregates` was last refreshed 2026-07-21 (the nightly job was
* unreachable code until this session — see snapshotScheduler). Settled hits
* rows run 2026-07-23 onward. So the skill profiles this model reads were
* frozen BEFORE every game it is asked to predict. The staleness that was a bug
* for production is, for this one measurement, a clean point-in-time snapshot.
* Rows on or before the freeze date are EXCLUDED so no profile can contain the
* game it is predicting.
*
* The opposing starter comes from the statsapi schedule for that date, and the
* pitcher's skill profile from the same frozen aggregate table.
*
* ── THE BAR (identical to the one that refuted hits-v1) ──────────────────
* - hits rows only, direction-aligned to the graded side
* - matched rows only: champion and challenger scored on the SAME props
* - paired bootstrap, deterministic seed, CI on the DIFFERENCE
* - PROMOTE only if the CI excludes zero on the good side
*
* ── DISCIPLINE 4, MEASURED, NOT ASSUMED ──────────────────────────────────
* Selectivity is reported, not claimed: accuracy is broken out by how confident
* the model is, so "right 57% on the 8 you're sure of" is a number rather than a
* slogan. LIFT over the naive base rate is reported beside it, because being
* right about obvious chalk is not signal.
*
* SUPABASE_URL=... node scripts/skill-v1-stagea.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const sk = require('../src/services/model/skillProjection');
const reg = require('../src/services/model/featureRegistry');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { nameKey } = require('../src/utils/playerName');
const { knownRate } = require('../src/utils/known');
/** Team games played by the 2026-07-21 profile freeze — turns season PA into PA/game. */
const GAMES_SO_FAR = Number(process.env.STAGEA_GAMES_SO_FAR || 103);
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
function corr(xs, ys) {
const n = xs.length;
if (n < 3) return null;
const mx = xs.reduce((a, b) => a + b, 0) / n;
const my = ys.reduce((a, b) => a + b, 0) / n;
let sxy = 0; let sxx = 0; let syy = 0;
for (let i = 0; i < n; i += 1) {
const dx = xs[i] - mx; const dy = ys[i] - my;
sxy += dx * dy; sxx += dx * dx; syy += dy * dy;
}
if (sxx <= 0 || syy <= 0) return null;
return sxy / Math.sqrt(sxx * syy);
}
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const mean = (a) => (a.length ? a.reduce((x, y) => x + y, 0) / a.length : null);
const brier = (ps, ys) => (ps.length ? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260803) {
if (rows.length < 30) return null;
const rnd = makeRnd(seed);
const n = rows.length;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const b = [];
for (let i = 0; i < n; i += 1) {
const r = rows[Math.floor(rnd() * n)];
ys.push(r.won); a.push(r[keyA]); b.push(r[keyB]);
}
const ca = corr(a, ys); const cb = corr(b, ys);
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (p) => r4(diffs[Math.floor(p * (diffs.length - 1))]);
const ci = [q(0.025), q(0.975)];
return {
point: r4(corr(rows.map((r) => r[keyA]), rows.map((r) => r.won))
- corr(rows.map((r) => r[keyB]), rows.map((r) => r.won))),
ci95: ci,
ci_excludes_zero: ci[0] > 0 || ci[1] < 0,
};
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
/**
* `date|team` → the starter that team FACED.
*
* Built from the statsapi schedule: a team faces the OTHER side's probable.
*/
async function opposingStarters(dates) {
const byDateTeam = new Map();
for (const d of dates) {
let games = [];
try { games = await mlb.getScheduleWithPitchers(d); } catch { games = []; }
for (const g of games) {
if (!g.home || !g.away) continue;
if (g.away.probablePitcher) byDateTeam.set(`${d}|${g.home.team}`, g.away.probablePitcher.id);
if (g.home.probablePitcher) byDateTeam.set(`${d}|${g.away.team}`, g.home.probablePitcher.id);
// Also key by the PITCHING team, so "who did team X send out" is directly
// answerable from the opponent name a game log gives us.
if (g.home.probablePitcher) byDateTeam.set(`${d}|OPP:${g.home.team}`, g.home.probablePitcher.id);
if (g.away.probablePitcher) byDateTeam.set(`${d}|OPP:${g.away.team}`, g.away.probablePitcher.id);
}
}
return byDateTeam;
}
/**
* `playerKey|date` → the OPPONENT team that player faced.
*
* THE LEDGER CANNOT ANSWER THIS: `team`/`opponent` are NULL on 575 of 576 rows
* in this window, which is why the first run resolved a pitcher for exactly ONE
* row and silently measured a batter-profile-only model instead of the matchup
* model it claimed to test. The player's own statsapi game log names the
* opponent for the exact date, so it is both authoritative and point-in-time
* safe (a completed game's opponent is not a forecast).
*/
async function opponentByPlayerDate(players) {
const map = new Map();
for (const [key, name] of players) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id);
for (const g of log || []) {
if (g && g.date && g.opponent) map.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
}
} catch { /* a missing log just means no pitcher for those rows */ }
}
return map;
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
// Frozen skill profiles — by name (batters) and by source_id (pitchers).
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const freeze = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '');
const freezeDate = freeze.slice(0, 10);
const batters = new Map();
const pitchersById = new Map();
for (const r of statcast) {
// UNITS: statcast_aggregates stores percentages (0-100). Convert ONCE, here.
if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), sk.fromStatcastRow(r));
if (r.player_key && r.role === 'batter') {
const prev = batters.get(r.player_key);
const size = Number(r.sample_pa || 0);
if (!prev || size > Number(prev.rawPa || 0)) {
batters.set(r.player_key, Object.assign(sk.fromStatcastRow(r), { rawPa: size, archetype: null }));
}
}
}
const led = await page(sb, 'ledger_entries',
'player_key, player_name, stat, line, side, outcome, game_date, p_win, team, opponent, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'hits')
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')
// STRICTLY AFTER the profile freeze — no game may be inside its own inputs.
&& String(r.game_date) > freezeDate);
const dates = [...new Set(clean.map((r) => r.game_date))].sort();
const starters = await opposingStarters(dates);
const players = new Map();
for (const r of clean) if (!players.has(r.player_key)) players.set(r.player_key, r.player_name);
const oppByPlayerDate = await opponentByPlayerDate(players);
const allowed = reg.candidateFeatures('mlb'); // challenger-first: candidates measured, never served
const rows = [];
const drops = {};
const drop = (k) => { drops[k] = (drops[k] || 0) + 1; };
let withPitcher = 0;
for (const r of clean) {
const bat = batters.get(r.player_key);
if (!bat) { drop('no_batter_profile'); continue; }
// The team he FACED that day, from his own game log. `starters` is keyed by
// the team doing the facing, so look up his own team — which is the
// opponent's opponent. Resolve via the game log's opponent name and take the
// schedule entry for the OTHER side.
const facedTeam = oppByPlayerDate.get(`${r.player_key}|${r.game_date}`) || null;
const starterId = facedTeam ? starters.get(`${r.game_date}|OPP:${facedTeam}`) : null;
const pit = starterId != null ? pitchersById.get(Number(starterId)) : null;
if (pit) withPitcher += 1;
// Opportunity: the batter's own season PA per game, from the frozen profile.
// Opportunity: season PA spread over the games played so far this season.
// GAMES_SO_FAR is the frozen-profile era's team game count, so PA/game is a
// real per-game rate rather than an arbitrary divisor.
const pa = knownRate(bat.rawPa);
const expectedPa = pa && pa > 0 ? Math.min(5.2, Math.max(2.0, pa / GAMES_SO_FAR)) : null;
const out = sk.projectSkill({
batter: bat, pitcher: pit, park: 1,
archetype: bat.archetype || null,
statType: 'hits', line: Number(r.line),
expectedPa, allowed,
});
if (!out) { drop('projection_refused'); continue; }
const under = String(r.side).toLowerCase() === 'under';
rows.push({
won: r.outcome === 'hit' ? 1 : 0,
champ: Number(r.p_win),
skill: under ? 1 - out.p_over_line : out.p_over_line,
line: Number(r.line),
had_pitcher: !!pit,
});
}
const ys = rows.map((r) => r.won);
const base = mean(ys);
const bs = bootstrapDiff(rows, 'skill', 'champ');
// DISCIPLINE 4 — selectivity, measured. Sorted by confidence in the graded
// side; report accuracy and LIFT over the naive base rate at each depth.
const byConf = [...rows].sort((a, b) => b.skill - a.skill);
const depths = [8, 15, 25, 50, 100].filter((d) => d <= byConf.length);
const selectivity = depths.map((d) => {
const top = byConf.slice(0, d);
const hit = mean(top.map((r) => r.won));
return { top_n: d, hit_rate: r4(hit), lift_over_base: r4(hit - base) };
});
console.log(JSON.stringify({
measurement: 'STAGE A — skill-v1 vs the frequency counter, out-of-sample on listed-line accuracy',
out_of_sample_guarantee: `skill profiles frozen ${freezeDate}; only rows with game_date > ${freezeDate} scored`,
registry: reg.summary('mlb'),
matched_rows: rows.length,
rows_with_opposing_pitcher: withPitcher,
pitcher_coverage_pct: rows.length ? r4(withPitcher / rows.length) : null,
dropped: drops,
base_rate: r4(base),
resolution: { skill_v1: r4(corr(rows.map((r) => r.skill), ys)), champion: r4(corr(rows.map((r) => r.champ), ys)) },
brier: { skill_v1: r4(brier(rows.map((r) => r.skill), ys)), champion: r4(brier(rows.map((r) => r.champ), ys)) },
mean_forecast: { skill_v1: r4(mean(rows.map((r) => r.skill))), champion: r4(mean(rows.map((r) => r.champ))) },
delta_vs_champion: bs,
verdict: !bs ? 'N-BLOCKED'
: (bs.ci_excludes_zero && bs.point > 0) ? 'BEATS THE COUNTER — promotable'
: (bs.ci_excludes_zero && bs.point < 0) ? 'LOSES to the counter — iterate, do not promote'
: 'INCONCLUSIVE — not proven, do not promote',
selectivity_discipline_4: selectivity,
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+247
View File
@@ -0,0 +1,247 @@
#!/usr/bin/env node
'use strict';
/**
* stagea-gate-run — RUN THE SKILL FEATURES THROUGH THE GATE, THEN THE COUNTER.
*
* The original sin was never that the challengers were badly built. It was that
* every one of them was measured WITHOUT a validation gate, so "it didn't work"
* and "it was never allowed to prove it works" were indistinguishable. This runs
* the gate that spec'd for exactly this (n>=500, |r|>=0.15, p<0.05, Bonferroni)
* over the real skill features, and only then does the head-to-head.
*
* THREE MEASUREMENTS, in the order that makes each one meaningful:
*
* 1. RAW SIGNAL — corr(feature, outcome). Does this skill input relate to
* whether the prop hit at all?
* 2. MARGINAL CONTRIBUTION — corr(feature, counter residual). This is the one
* that matters: a feature can correlate with the outcome purely because the
* counter already knows it. Only the part the counter MISSES is new
* information, and that is what earns a place. Both go through the gate.
* 3. HEAD-TO-HEAD — the value projection vs the live counter on listed-line
* accuracy, paired bootstrap, out-of-sample.
*
* OUT-OF-SAMPLE: skill profiles are the frozen 2026-07-21 aggregate; only games
* AFTER that date are scored, so no profile contains the game it predicts.
*
* BONFERRONI DENOMINATOR is the number of features tested in this sweep — not 1.
* Testing many and reporting the best without correction is how the S78 residual
* scan produced six "findings" when chance alone predicts three or four.
*
* SUPABASE_URL=... node scripts/stagea-gate-run.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const cv = require('../src/services/model/correlateValidator');
const sk = require('../src/services/model/skillProjection');
const reg = require('../src/services/model/featureRegistry');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownRate, knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const GAMES_SO_FAR = Number(process.env.STAGEA_GAMES_SO_FAR || 103);
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const mean = (a) => (a.length ? a.reduce((x, y) => x + y, 0) / a.length : null);
const brier = (ps, ys) => (ps.length ? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
function corrOf(xs, ys) { return cv.pearson(xs, ys).r; }
function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260803) {
if (rows.length < 30) return null;
const rnd = makeRnd(seed);
const n = rows.length;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const b = [];
for (let i = 0; i < n; i += 1) {
const r = rows[Math.floor(rnd() * n)];
ys.push(r.won); a.push(r[keyA]); b.push(r[keyB]);
}
const ca = corrOf(a, ys); const cb = corrOf(b, ys);
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (p) => r4(diffs[Math.floor(p * (diffs.length - 1))]);
const ci = [q(0.025), q(0.975)];
return {
point: r4(corrOf(rows.map((r) => r[keyA]), rows.map((r) => r.won))
- corrOf(rows.map((r) => r[keyB]), rows.map((r) => r.won))),
ci95: ci, ci_excludes_zero: ci[0] > 0 || ci[1] < 0,
};
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function opposingStarters(dates) {
const m = new Map();
for (const d of dates) {
let games = [];
try { games = await mlb.getScheduleWithPitchers(d); } catch { games = []; }
for (const g of games) {
if (!g.home || !g.away) continue;
if (g.home.probablePitcher) m.set(`${d}|OPP:${g.home.team}`, g.home.probablePitcher.id);
if (g.away.probablePitcher) m.set(`${d}|OPP:${g.away.team}`, g.away.probablePitcher.id);
}
}
return m;
}
/** `playerKey|date` → opponent faced. The ledger's team/opponent are NULL. */
async function opponentByPlayerDate(players) {
const map = new Map();
for (const [key, name] of players) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id);
for (const g of log || []) {
if (g && g.date && g.opponent) map.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
}
} catch { /* no log → no pitcher for those rows */ }
}
return map;
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10);
const batters = new Map(); const pitchersById = new Map();
for (const r of statcast) {
if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), sk.fromStatcastRow(r));
if (r.player_key && r.role === 'batter') {
const prev = batters.get(r.player_key);
if (!prev || Number(r.sample_pa || 0) > Number(prev.rawPa || 0)) {
batters.set(r.player_key, Object.assign(sk.fromStatcastRow(r), { rawPa: Number(r.sample_pa || 0) }));
}
}
}
const led = await page(sb, 'ledger_entries',
'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null)
.in('stat', ['hits', 'total_bases'])
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')
&& String(r.game_date) > freezeDate);
const dates = [...new Set(clean.map((r) => r.game_date))].sort();
const starters = await opposingStarters(dates);
const players = new Map();
for (const r of clean) if (!players.has(r.player_key)) players.set(r.player_key, r.player_name);
const oppByPlayerDate = await opponentByPlayerDate(players);
const allowed = reg.candidateFeatures('mlb');
const rows = [];
for (const r of clean) {
const bat = batters.get(r.player_key);
if (!bat) continue;
const faced = oppByPlayerDate.get(`${r.player_key}|${r.game_date}`) || null;
const pit = faced ? pitchersById.get(Number(starters.get(`${r.game_date}|OPP:${faced}`))) || null : null;
const paRate = bat.rawPa > 0 ? Math.min(5.2, Math.max(2.0, bat.rawPa / GAMES_SO_FAR)) : null;
const under = String(r.side).toLowerCase() === 'under';
const won = r.outcome === 'hit' ? 1 : 0;
const champ = Number(r.p_win);
// The value projection (hits only — TB is refused by design, see skillProjection).
const proj = r.stat === 'hits'
? sk.projectSkill({ batter: bat, pitcher: pit, park: 1, archetype: null,
statType: 'hits', line: Number(r.line), expectedPa: paRate, allowed })
: null;
rows.push({
stat: r.stat, won, champ,
skill: proj ? (under ? 1 - proj.p_over_line : proj.p_over_line) : null,
residual: won - champ,
had_pitcher: !!pit,
// Candidate skill features, archetype-relevant, in probability space.
batter_barrel_pct: knownRate(bat.barrel_pct),
batter_hard_hit_pct: knownRate(bat.hard_hit_pct),
batter_exit_velo: knownRate(bat.avg_exit_velo),
batter_launch_angle: knownRate(bat.avg_launch_angle),
batter_k_pct: knownRate(bat.k_pct),
batter_bb_pct: knownRate(bat.bb_pct),
pitcher_k_pct: pit ? knownRate(pit.k_pct) : null,
pitcher_hard_hit_allowed: pit ? knownRate(pit.hard_hit_pct) : null,
});
}
const FEATURES = ['batter_barrel_pct', 'batter_hard_hit_pct', 'batter_exit_velo',
'batter_launch_angle', 'batter_k_pct', 'batter_bb_pct',
'pitcher_k_pct', 'pitcher_hard_hit_allowed'];
const perStat = {};
for (const stat of ['hits', 'total_bases']) {
const rs = rows.filter((r) => r.stat === stat);
if (rs.length === 0) continue;
const tests = FEATURES.length; // the Bonferroni denominator for THIS sweep
const gate = {};
for (const f of FEATURES) {
const xs = rs.map((r) => r[f]);
gate[f] = {
// 1. does it relate to the outcome at all?
raw_vs_outcome: cv.validateFactor(xs, rs.map((r) => r.won), tests),
// 2. THE ONE THAT COUNTS — is any of it NEW, i.e. missed by the counter?
marginal_vs_counter_residual: cv.validateFactor(xs, rs.map((r) => r.residual), tests),
};
}
const passed = FEATURES.filter((f) => gate[f].marginal_vs_counter_residual.validated);
perStat[stat] = {
n: rs.length,
base_rate: r4(mean(rs.map((r) => r.won))),
bonferroni_tests: tests,
features_passing_gate_on_marginal: passed,
gate,
};
}
// HEAD-TO-HEAD — hits only (the value engine covers hits).
const h2h = rows.filter((r) => r.stat === 'hits' && r.skill != null);
const ys = h2h.map((r) => r.won);
const bs = bootstrapDiff(h2h, 'skill', 'champ');
console.log(JSON.stringify({
premise_correction: 'statModel.js and correlateValidator.js do not exist in this repo. The gate was implemented to the spec in src/services/python/blueprints/unconventional.py (VALIDATION_REQUIREMENTS); supplementSystems.test.js inlines its own validateFactor and imports no implementation.',
out_of_sample: `skill profiles frozen ${freezeDate}; only game_date > ${freezeDate} scored`,
gate_spec: cv.VALIDATION_REQUIREMENTS,
per_stat_gate: perStat,
head_to_head_hits: {
n: h2h.length,
pitcher_coverage: r4(mean(h2h.map((r) => (r.had_pitcher ? 1 : 0)))),
base_rate: r4(mean(ys)),
resolution: { value_engine: r4(corrOf(h2h.map((r) => r.skill), ys)), counter: r4(corrOf(h2h.map((r) => r.champ), ys)) },
brier: { value_engine: r4(brier(h2h.map((r) => r.skill), ys)), counter: r4(brier(h2h.map((r) => r.champ), ys)) },
delta: bs,
verdict: !bs ? 'N-BLOCKED'
: (bs.ci_excludes_zero && bs.point > 0) ? 'VALUE ENGINE BEATS THE COUNTER'
: (bs.ci_excludes_zero && bs.point < 0) ? 'LOSES to the counter — iterate, do not promote'
: 'INCONCLUSIVE — do not promote',
},
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+8
View File
@@ -0,0 +1,8 @@
# VYNDR — PINNED Storage Box host key (Session 64).
# Verified 2026-07-20: SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM
# The backup used StrictHostKeyChecking=accept-new, which is trust-on-first-use:
# it accepts whatever host key it meets first. This file makes the trust STATIC.
# Never replace with accept-new / =no / UserKnownHostsFile=/dev/null.
# Rotating the box regenerates this line — re-verify the fingerprint out-of-band
# before changing it. A mismatch at run time is a HARD FAIL, by design.
[u635423.your-storagebox.de]:23 ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIICf9svRenC/PLKIL9nk6K/pxQgoiFC41wTNvoIncOxs
+51
View File
@@ -0,0 +1,51 @@
-- tb-compound-holdout.sql — TOTAL BASES ONLY, direction-aligned.
--
-- TWO guards this query exists to enforce:
-- 1. TB ROWS ONLY. Averaging into other stats would hide the effect, since
-- total_bases is 49 of 437 settled rows.
-- 2. DIRECTION-ALIGNED. p_win is P(GRADED SIDE); proj_p_over_line and
-- proj_tb_p_over are P(OVER). 31.4% of rows are under-graded, and comparing
-- raw P(over) against an under-side outcome measures the model BACKWARDS —
-- that artifact alone accounted for 41% of the ladder's apparent loss.
--
-- Promote tb-v1 ONLY if it materially improves TB resolution toward/past the
-- champion. If it does NOT, the family-mismatch hypothesis is WRONG and the
-- mean-weakness / similarity branch REOPENS. Record which.
--
-- CONTAMINATION EXCLUSION (2026-08-02, MANDATORY). Rows whose price/book/takeable
-- were stamped from a NON-TAKEABLE book (DFS / offshore / exchange) between
-- 2026-08-01 and the write-path fix are tagged `quarantine_reason LIKE
-- 'nontakeable_book%'`. They are EXCLUDED here and must never be pooled with
-- clean rows: their locked price -- and therefore the `takeable` flag computed
-- from it -- describes a market you could not have bet.
with tb as (
select
game_date, id, lower(side) side, (outcome='hit')::int won,
p_win::numeric champ,
case when lower(side)='under' then 1 - proj_p_over_line::numeric
else proj_p_over_line::numeric end ladder_al,
case when lower(side)='under' then 1 - proj_tb_p_over::numeric
else proj_tb_p_over::numeric end tbv1_al
from public.ledger_entries
where sport='mlb' and user_id is null
and (quarantine_reason is null or quarantine_reason not like 'nontakeable_book%')
and stat = 'total_bases'
and outcome in ('hit','miss')
and p_win is not null
and proj_p_over_line is not null
and proj_tb_p_over is not null -- matched rows: all three present
)
select
count(*) n,
count(*) filter (where side='under') under_rows,
round(avg(won::numeric),3) base_rate,
round(corr(champ, won::numeric)::numeric,4) res_champion,
round(corr(ladder_al, won::numeric)::numeric,4) res_ladder_v11,
round(corr(tbv1_al, won::numeric)::numeric,4) res_tb_v1,
round(stddev(ladder_al)::numeric,4) sd_ladder,
round(stddev(tbv1_al)::numeric,4) sd_tb_v1,
round(avg(ladder_al)::numeric,4) mean_ladder,
round(avg(tbv1_al)::numeric,4) mean_tb_v1
from tb;
+436
View File
@@ -0,0 +1,436 @@
#!/usr/bin/env node
'use strict';
/**
* tb-solo-and-interactions — PROVE BOTH, with the solo pass as the control.
*
* A feature can carry signal alone, only in combination, or both. Testing only
* interactions misses solo-real features AND cannot tell whether an interaction
* ADDS anything or merely re-encodes its own parts. So the solo result is the
* baseline every interaction has to beat.
*
* ── HOW "ADDS OVER ITS PARTS" IS MEASURED ────────────────────────────────
* Not by comparing two correlations by eye. The interaction's incremental
* signal is the PARTIAL correlation of the interaction term with the counter's
* residual, CONTROLLING FOR both component features:
*
* resid_I = I OLS(I ~ A, B)
* resid_Y = Y OLS(Y ~ A, B)
* incremental r = corr(resid_I, resid_Y)
*
* If the interaction is just barrel-rate wearing a different hat, regressing out
* barrel rate removes it and the incremental r collapses to ~0. That is exactly
* the redundancy the order is guarding against, and it is the difference between
* PASSES-AND-ADDS and PASSES-BUT-REDUNDANT.
*
* ── WHY THE COUNTER'S RESIDUAL IS THE TARGET ─────────────────────────────
* Correlating with the raw outcome rewards a feature for knowing what the
* counter already knows. Only the part the counter MISSES is new information,
* and only new information can improve the product. Both are reported; the
* residual one is the one that decides.
*
* ── THEORY FIRST ─────────────────────────────────────────────────────────
* Every interaction below is declared with a MECHANISM before it is measured.
* No blind pairwise search — with 8 features there are 28 pairs, and at α=.05
* roughly one in twenty returns "significant" from noise alone.
*
* SUPABASE_URL=... node scripts/tb-solo-and-interactions.js
*/
require('dotenv').config();
const { createClient } = require('@supabase/supabase-js');
const cv = require('../src/services/model/correlateValidator');
const sk = require('../src/services/model/skillProjection');
const reg = require('../src/services/model/featureRegistry');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const { knownRate, knownNumber } = require('../src/utils/known');
const SB_URL = process.env.SUPABASE_URL;
const SB_KEY = process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY;
const PAGE = 1000;
const GAMES_SO_FAR = Number(process.env.STAGEA_GAMES_SO_FAR || 103);
const r4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const mean = (a) => (a.length ? a.reduce((x, y) => x + y, 0) / a.length : null);
const brier = (ps, ys) => (ps.length ? ps.reduce((s, p, i) => s + (p - ys[i]) ** 2, 0) / ps.length : null);
/** OLS residuals of y on the given predictor columns (with intercept). */
function olsResiduals(y, Xcols) {
const n = y.length;
const p = Xcols.length + 1;
const X = [];
for (let i = 0; i < n; i += 1) {
const row = [1];
for (const c of Xcols) row.push(c[i]);
X.push(row);
}
// Normal equations (X'X) b = X'y, solved by Gauss-Jordan. p is 2-4 here.
const XtX = Array.from({ length: p }, () => new Array(p).fill(0));
const Xty = new Array(p).fill(0);
for (let i = 0; i < n; i += 1) {
for (let a = 0; a < p; a += 1) {
Xty[a] += X[i][a] * y[i];
for (let b = 0; b < p; b += 1) XtX[a][b] += X[i][a] * X[i][b];
}
}
const M = XtX.map((row, i) => [...row, Xty[i]]);
for (let col = 0; col < p; col += 1) {
let piv = col;
for (let r = col + 1; r < p; r += 1) if (Math.abs(M[r][col]) > Math.abs(M[piv][col])) piv = r;
if (Math.abs(M[piv][col]) < 1e-12) return null; // singular → cannot control honestly
[M[col], M[piv]] = [M[piv], M[col]];
const d = M[col][col];
for (let k = col; k <= p; k += 1) M[col][k] /= d;
for (let r = 0; r < p; r += 1) {
if (r === col) continue;
const f = M[r][col];
for (let k = col; k <= p; k += 1) M[r][k] -= f * M[col][k];
}
}
const beta = M.map((row) => row[p]);
return y.map((v, i) => v - X[i].reduce((s, xv, j) => s + xv * beta[j], 0));
}
/**
* Partial correlation of a with b, controlling for the columns in ctrl.
*
* COLLINEARITY IS CHECKED FIRST, and this is not pedantry — it caught a real
* error in this very script. The archetype-power proxy was defined as
* `barrel_pct / LEAGUE.barrel_pct`, an exact linear function of barrel_pct, so
* "control for both components" was rank-deficient and the partial correlation
* it produced (-0.132, the only one that looked like an incremental finding) was
* an artifact of a singular design matrix. The Gauss-Jordan pivot test missed it
* because the two columns differ by a scale factor, which keeps the pivot well
* above an absolute epsilon. Scale-free pairwise correlation catches it.
*/
function partialCorr(a, b, ctrl) {
for (let i = 0; i < ctrl.length; i += 1) {
for (let j = i + 1; j < ctrl.length; j += 1) {
const rr = cv.pearson(ctrl[i], ctrl[j]).r;
if (rr !== null && Math.abs(rr) > 0.999) return null; // same variable twice
}
}
const ra = olsResiduals(a, ctrl);
const rb = olsResiduals(b, ctrl);
if (!ra || !rb) return null;
return cv.pearson(ra, rb).r;
}
/** Rows where every named key is known — the honest common sample. */
function completeRows(rows, keys) {
return rows.filter((r) => keys.every((k) => knownNumber(r[k]) !== null));
}
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
function bootstrapDiff(rows, keyA, keyB, iters = 4000, seed = 20260804) {
if (rows.length < 30) return null;
const rnd = makeRnd(seed);
const n = rows.length;
const diffs = [];
for (let it = 0; it < iters; it += 1) {
const ys = []; const a = []; const b = [];
for (let i = 0; i < n; i += 1) {
const r = rows[Math.floor(rnd() * n)];
ys.push(r.won); a.push(r[keyA]); b.push(r[keyB]);
}
const ca = cv.pearson(a, ys).r; const cb = cv.pearson(b, ys).r;
if (ca == null || cb == null) continue;
diffs.push(ca - cb);
}
if (diffs.length < 100) return null;
diffs.sort((x, y) => x - y);
const q = (pp) => r4(diffs[Math.floor(pp * (diffs.length - 1))]);
const ci = [q(0.025), q(0.975)];
return {
point: r4(cv.pearson(rows.map((r) => r[keyA]), rows.map((r) => r.won)).r
- cv.pearson(rows.map((r) => r[keyB]), rows.map((r) => r.won)).r),
ci95: ci, ci_excludes_zero: ci[0] > 0 || ci[1] < 0,
};
}
async function page(sb, table, select, apply) {
const out = [];
for (let from = 0; ; from += PAGE) {
const { data, error } = await apply(sb.from(table).select(select)).range(from, from + PAGE - 1);
if (error) throw error;
if (!data || data.length === 0) break;
out.push(...data);
if (data.length < PAGE) break;
}
return out;
}
async function opposingStarters(dates) {
const m = new Map();
for (const d of dates) {
let games = [];
try { games = await mlb.getScheduleWithPitchers(d); } catch { games = []; }
for (const g of games) {
if (!g.home || !g.away) continue;
if (g.home.probablePitcher) m.set(`${d}|OPP:${g.home.team}`, g.home.probablePitcher.id);
if (g.away.probablePitcher) m.set(`${d}|OPP:${g.away.team}`, g.away.probablePitcher.id);
}
}
return m;
}
async function opponentByPlayerDate(players) {
const map = new Map();
for (const [key, name] of players) {
try {
const found = await mlb.searchPlayer(name);
if (!found || !found.id) continue;
const log = await mlb.getPlayerGameLog(found.id);
for (const g of log || []) {
if (g && g.date && g.opponent) map.set(`${key}|${String(g.date).slice(0, 10)}`, g.opponent);
}
} catch { /* no log → no pitcher */ }
}
return map;
}
const SOLO = ['batter_barrel_pct', 'batter_hard_hit_pct', 'batter_exit_velo',
'batter_launch_angle', 'batter_k_pct', 'batter_bb_pct',
'pitcher_k_pct', 'pitcher_hard_hit_allowed'];
/** STEP 2 — theory first. Every interaction declares its mechanism. */
const INTERACTIONS = [
{
key: 'launch_x_exit_velo',
components: ['batter_launch_angle', 'batter_exit_velo'],
mechanism: 'Extra bases need BOTH conditions: hit hard AND hit in the air. A 105-mph ground ball is an out; a 25-degree popup is an out. Neither factor alone predicts bases, which is precisely why each may fail solo and the product may not.',
build: (r) => r.batter_launch_angle * r.batter_exit_velo,
},
{
key: 'exitvelo_x_pitcher_suppression',
components: ['batter_exit_velo', 'pitcher_hard_hit_allowed'],
mechanism: 'A hitter only realises his contact quality against a pitcher who permits contact quality. Elite suppression should attenuate a power bat; a contact-permitting arm should amplify it. The effect is conditional by construction.',
build: (r) => r.batter_exit_velo * r.pitcher_hard_hit_allowed,
},
{
key: 'barrel_x_power_archetype',
components: ['batter_barrel_pct', 'archetype_power'],
mechanism: 'ARCHETYPE-CONDITIONAL. Barrels convert to extra bases for hitters whose lane is power; for a speed/contact profile the same barrel rate is a rarer event on a swing built for something else. This is Discipline 2 stated as a testable interaction. NOTE: it is currently UNTESTABLE — statcast rows carry no archetype label, and the barrel-relative proxy is an exact linear function of barrel_pct, so controlling for both components is rank-deficient. It needs a real archetype classification joined in.',
build: (r) => r.batter_barrel_pct * r.archetype_power,
},
{
key: 'batterK_x_pitcherK',
components: ['batter_k_pct', 'pitcher_k_pct'],
mechanism: 'Strikeout risk compounds multiplicatively (log5 is exactly this shape). A high-K bat against a high-K arm loses plate appearances to strikeouts, and a PA lost is a base opportunity that never happens — so it suppresses total bases through OPPORTUNITY, not contact quality.',
build: (r) => r.batter_k_pct * r.pitcher_k_pct,
},
];
/** Latest settled game date in the pull — used to detect that the profile
* freeze now sits AFTER the data, i.e. no clean out-of-sample window exists. */
function clean0Max(rows) {
return (rows || []).reduce((mx, r) => (String(r.game_date) > mx ? String(r.game_date) : mx), '');
}
async function main() {
if (!SB_URL || !SB_KEY) throw new Error('SUPABASE_URL / service key required');
const sb = createClient(SB_URL, SB_KEY, { auth: { persistSession: false } });
const statcast = await page(sb, 'statcast_aggregates', '*', (q) => q.eq('sport', 'mlb'));
const freezeDate = statcast.reduce((mx, r) => (String(r.updated_at) > mx ? String(r.updated_at) : mx), '').slice(0, 10);
const batters = new Map(); const pitchersById = new Map();
for (const r of statcast) {
if (r.role === 'pitcher' && r.source_id != null) pitchersById.set(Number(r.source_id), sk.fromStatcastRow(r));
if (r.player_key && r.role === 'batter') {
const prev = batters.get(r.player_key);
if (!prev || Number(r.sample_pa || 0) > Number(prev.rawPa || 0)) {
batters.set(r.player_key, Object.assign(sk.fromStatcastRow(r), { rawPa: Number(r.sample_pa || 0) }));
}
}
}
// REAL ARCHETYPE LABELS. The barrel-relative proxy was a clipped monotone
// transform of barrel_pct, so `barrel x proxy` measured NONLINEARITY IN BARREL,
// not an archetype interaction — it could never have tested Discipline 2.
// model_snapshots carries the actual classification per prop, so the
// conditioning variable is now a genuine BOMBER indicator, which is
// categorical and therefore not a transform of barrel at all.
const snaps = await page(sb, 'model_snapshots', 'player_key, game_date, archetype',
(q) => q.eq('sport', 'mlb').eq('stat', 'total_bases').not('archetype', 'is', null));
const archetypeBy = new Map();
for (const r of snaps) if (r.player_key && r.game_date) archetypeBy.set(`${r.player_key}|${r.game_date}`, r.archetype);
const led = await page(sb, 'ledger_entries',
'player_key, player_name, stat, line, side, outcome, game_date, p_win, quarantine_reason',
(q) => q.eq('sport', 'mlb').is('user_id', null).eq('stat', 'total_bases')
.in('outcome', ['hit', 'miss']).not('p_win', 'is', null));
// POINT-IN-TIME IS NO LONGER AVAILABLE FROM THIS TABLE.
//
// `statcast_aggregates` is upserted in place and keeps one as-of date. The
// first skill backtest was honest only by accident: the nightly refresh was
// unreachable code, so the table sat frozen at 2026-07-21 — BEFORE the settled
// window. Repairing that cron (correct for production) refreshed it to today,
// and every prior version is gone.
//
// So scoring a 2026-07-25 game now uses a season aggregate that CONTAINS that
// game. `statcast_history` (added this session) fixes it going forward; it has
// one day of data, which is not yet a window. Until it fills, results here are
// DIRECTIONAL AND CONTAMINATED, labelled as such, and are NOT gate verdicts.
const contaminated = String(freezeDate) >= String(clean0Max(led));
const clean = led.filter((r) => !(r.quarantine_reason || '').startsWith('nontakeable_book')
&& (contaminated ? true : String(r.game_date) > freezeDate));
const dates = [...new Set(clean.map((r) => r.game_date))].sort();
const starters = await opposingStarters(dates);
const players = new Map();
for (const r of clean) if (!players.has(r.player_key)) players.set(r.player_key, r.player_name);
const oppByPlayerDate = await opponentByPlayerDate(players);
if (process.env.TB_DEBUG === '1') {
console.error(`[debug] statcast rows=${statcast.length} freeze=${freezeDate} batters=${batters.size} pitchers=${pitchersById.size}`);
console.error(`[debug] ledger tb rows=${led.length} clean(after freeze)=${clean.length}`);
const sampleKeys = clean.slice(0, 5).map((r) => r.player_key);
console.error(`[debug] sample ledger player_keys=${JSON.stringify(sampleKeys)}`);
console.error(`[debug] sample statcast keys=${JSON.stringify([...batters.keys()].slice(0, 5))}`);
console.error(`[debug] matches in sample=${sampleKeys.filter((k) => batters.has(k)).length}/5`);
}
const allowed = reg.candidateFeaturesForStat('mlb', 'total_bases');
const rows = [];
for (const r of clean) {
const bat = batters.get(r.player_key);
if (!bat) continue;
const faced = oppByPlayerDate.get(`${r.player_key}|${r.game_date}`) || null;
const pit = faced ? pitchersById.get(Number(starters.get(`${r.game_date}|OPP:${faced}`))) || null : null;
const paRate = bat.rawPa > 0 ? Math.min(5.2, Math.max(2.0, bat.rawPa / GAMES_SO_FAR)) : null;
const under = String(r.side).toLowerCase() === 'under';
const won = r.outcome === 'hit' ? 1 : 0;
const champ = Number(r.p_win);
const proj = sk.projectSkill({
batter: bat, pitcher: pit, park: 1, archetype: null,
statType: 'total_bases', line: Number(r.line), expectedPa: paRate, allowed,
});
// The REAL archetype for this prop — a 0/1 power indicator, categorical and
// independent of barrel_pct by construction.
const arch = archetypeBy.get(`${r.player_key}|${r.game_date}`) || null;
const archetypePower = arch == null ? null : (String(arch).toUpperCase() === 'BOMBER' ? 1 : 0);
rows.push({
won, champ, residual: won - champ,
skill: proj ? (under ? 1 - proj.p_over_line : proj.p_over_line) : null,
had_pitcher: !!pit,
batter_barrel_pct: knownRate(bat.barrel_pct),
batter_hard_hit_pct: knownRate(bat.hard_hit_pct),
batter_exit_velo: knownRate(bat.avg_exit_velo),
batter_launch_angle: knownRate(bat.avg_launch_angle),
batter_k_pct: knownRate(bat.k_pct),
batter_bb_pct: knownRate(bat.bb_pct),
pitcher_k_pct: pit ? knownRate(pit.k_pct) : null,
pitcher_hard_hit_allowed: pit ? knownRate(pit.hard_hit_pct) : null,
archetype_power: archetypePower,
archetype: arch,
});
}
// Bonferroni denominator = every test in this family (solo + interaction).
// ── CUMULATIVE BONFERRONI ─────────────────────────────────────────────
// The denominator is every DISTINCT hypothesis this programme has tested, not
// this run's. A per-session count gives each new order a fresh, generous alpha
// and lets the false-positive rate compound silently.
const tl = require('../src/services/model/testLedger');
const mcStore = tl.supabaseStore(sb);
const mc = await tl.recordAndCount(mcStore, [
...SOLO.map((f) => ({ sport: 'mlb', stat: 'total_bases', archetype: null, interaction: `solo:${f}`, target: 'counter_residual' })),
...INTERACTIONS.map((x) => ({ sport: 'mlb', stat: 'total_bases', archetype: null, interaction: x.key, target: 'counter_residual' })),
]);
const TESTS = mc.cumulative_tests;
// ── STEP 1 — SOLO PASS (the control) ────────────────────────────────────
const solo = {};
for (const f of SOLO) {
const rs = completeRows(rows, [f]);
solo[f] = {
n: rs.length,
vs_outcome: cv.validateFactor(rs.map((r) => r[f]), rs.map((r) => r.won), TESTS),
vs_counter_residual: cv.validateFactor(rs.map((r) => r[f]), rs.map((r) => r.residual), TESTS),
};
}
// ── STEP 3 — INTERACTIONS, each against its own solo baseline ───────────
const interactions = {};
for (const ix of INTERACTIONS) {
const keys = [...ix.components];
const rs = completeRows(rows, keys);
if (rs.length < 30) { interactions[ix.key] = { mechanism: ix.mechanism, n: rs.length, verdict: 'UNTESTABLE — no common sample' }; continue; }
const I = rs.map(ix.build);
const Y = rs.map((r) => r.residual);
const ctrl = keys.map((k) => rs.map((r) => r[k]));
const gate = cv.validateFactor(I, Y, TESTS);
const incremental = partialCorr(I, Y, ctrl);
// The best solo |r| among its own components, on the SAME rows.
const componentSolo = keys.map((k) => ({
feature: k, r: r4(cv.pearson(rs.map((r) => r[k]), Y).r),
}));
const bestComponent = Math.max(...componentSolo.map((c) => Math.abs(c.r ?? 0)));
let verdict;
if (incremental === null) verdict = 'UNTESTABLE — controls are collinear';
else if (gate.validated && Math.abs(incremental) >= cv.VALIDATION_REQUIREMENTS.min_pearson_r) verdict = 'PASSES-AND-ADDS';
else if (gate.validated) verdict = 'PASSES-BUT-REDUNDANT';
else if (rs.length < cv.VALIDATION_REQUIREMENTS.min_historical_instances) verdict = 'UNDERPOWERED — n below the gate';
else verdict = 'FAILS';
interactions[ix.key] = {
mechanism: ix.mechanism,
components: keys,
n: rs.length,
raw_r_vs_residual: gate.pearson_r,
gate: { validated: gate.validated, reason: gate.reason, p_value: gate.p_value, corrected_alpha: gate.corrected_alpha, underpowered: !!gate.underpowered },
component_solo_r_same_rows: componentSolo,
best_component_abs_r: r4(bestComponent),
INCREMENTAL_partial_r: r4(incremental),
adds_over_components: incremental !== null && Math.abs(incremental) > bestComponent,
verdict,
};
}
// ── STEP 4 — COMBINED vs COUNTER (valid at this n; the gate is not) ─────
const h2h = rows.filter((r) => r.skill != null);
const ys = h2h.map((r) => r.won);
const bs = bootstrapDiff(h2h, 'skill', 'champ');
console.log(JSON.stringify({
stat: 'total_bases',
VALIDITY: contaminated
? 'CONTAMINATED / DIRECTIONAL ONLY — statcast_aggregates now carries a single as-of date (' + freezeDate + ') that is AFTER the settled games, so season profiles contain the games being predicted. These are NOT gate verdicts. statcast_history (new) makes point-in-time possible from tomorrow.'
: `CLEAN out-of-sample: profiles frozen ${freezeDate}; only game_date > ${freezeDate} scored`,
contaminated,
rows_scored: rows.length,
gate_spec: cv.VALIDATION_REQUIREMENTS,
bonferroni_tests: TESTS,
multiple_comparisons: { ...mc, note: 'cumulative across the programme lifetime, not this session' },
n_gap_note: `the gate needs ${cv.VALIDATION_REQUIREMENTS.min_historical_instances} rows; this run has ${rows.length}`,
archetype_coverage: {
labelled: rows.filter((r) => r.archetype).length,
bomber: rows.filter((r) => r.archetype_power === 1).length,
other: rows.filter((r) => r.archetype_power === 0).length,
},
step1_solo_baseline: solo,
step3_interactions: interactions,
step4_combined_vs_counter: {
n: h2h.length,
pitcher_coverage: r4(mean(h2h.map((r) => (r.had_pitcher ? 1 : 0)))),
base_rate: r4(mean(ys)),
resolution: { skill_tb: r4(cv.pearson(h2h.map((r) => r.skill), ys).r), counter: r4(cv.pearson(h2h.map((r) => r.champ), ys).r) },
brier: { skill_tb: r4(brier(h2h.map((r) => r.skill), ys)), counter: r4(brier(h2h.map((r) => r.champ), ys)) },
delta: bs,
verdict: !bs ? 'N-BLOCKED'
: (bs.ci_excludes_zero && bs.point > 0) ? 'SKILL TB BEATS THE COUNTER'
: (bs.ci_excludes_zero && bs.point < 0) ? 'LOSES to the counter'
: 'INCONCLUSIVE',
},
}, null, 2));
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+189
View File
@@ -0,0 +1,189 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 1 — is the favourite-longshot bias real WITHOUT a calibration map?
*
* Four orders have refined a stability gate on 19 dates. LODO turned out to be
* structurally underpowered (0.0140.093) and the deploy intervals rest on 24
* date clusters. So we stop certifying the stability of a specific MAP, and ask
* the one question this sample might actually answer:
*
* does the model over-predict its own favourites, robustly?
*
* That claim is MODEL-FREE and MAP-FREE — it is a property of (p_win, outcome)
* pairs, needs no isotonic fit, and can therefore be tested without any of the
* machinery whose stability we cannot certify.
*
* ── DATE-BLOCK BOOTSTRAP ─────────────────────────────────────────────────
* Resampling ROWS would treat 200 props from one night as 200 readings of that
* night's offensive environment. Whole DATES are resampled instead, which is the
* honest unit and a far harsher one at 517 dates.
*
* VERDICT is pre-stated: ROBUST iff the >0.9 over-prediction sign survives in
* >=95% of pooled date-block resamples AND replicates in >=3 of 4 stats on the
* same criterion. Anything else is NOT-ROBUST, and NOT-ROBUST means we serve raw.
*
* SUPABASE_URL=... node scripts/test-favourite-bias-robust.js
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000;
const FAVOURITE_FLOOR = 0.9;
const ITERS = 5000;
/** Pre-stated pass marks. */
const SIGN_STABILITY_REQUIRED = 0.95;
const STATS_MUST_REPLICATE = 3;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error;
if (!data.length) break;
o.push(...data);
if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) {
let s = seed >>> 0;
return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; };
}
/** Over-prediction in the favourite bin: predicted realized. Positive = over. */
function favouriteBias(rows) {
const fav = rows.filter((r) => r.p >= FAVOURITE_FLOOR);
if (fav.length < 5) return null;
return mean(fav.map((r) => r.p)) - mean(fav.map((r) => r.won));
}
/** Resample whole DATES with replacement; report how often the sign survives. */
function dateBlockSignStability(rows, seed) {
const byDate = new Map();
for (const r of rows) {
if (!byDate.has(r.date)) byDate.set(r.date, []);
byDate.get(r.date).push(r);
}
const keys = [...byDate.keys()];
const rnd = makeRnd(seed);
let positive = 0; let indeterminate = 0; const draws = [];
for (let it = 0; it < ITERS; it += 1) {
const sample = [];
for (let i = 0; i < keys.length; i += 1) sample.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
const b = favouriteBias(sample);
// A resample with too few favourites cannot speak — counted, never guessed.
if (b === null) { indeterminate += 1; continue; }
draws.push(b);
if (b > 0) positive += 1;
}
const usable = ITERS - indeterminate;
draws.sort((a, b) => a - b);
return {
date_blocks: keys.length,
usable_resamples: usable,
indeterminate_resamples: indeterminate,
sign_stability: usable ? round4(positive / usable) : null,
ci_90: draws.length ? [round4(draws[Math.floor(draws.length * 0.05)]), round4(draws[Math.floor(draws.length * 0.95)])] : null,
};
}
function deciles(rows) {
const out = [];
for (let lo = 0.3; lo < 1.0; lo += 0.1) {
const hi = lo + 0.1;
const slice = rows.filter((r) => r.p >= lo && (hi >= 1 ? r.p <= 1 : r.p < hi));
if (slice.length < 15) continue;
const pred = mean(slice.map((r) => r.p));
const real = mean(slice.map((r) => r.won));
out.push({ bin: [round2(lo), round2(hi)], n: slice.length, predicted: round4(pred), realized: round4(real), over_prediction: round4(pred - real) });
}
return out;
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL,
process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots',
'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win),
})));
const byStat = {}; const pooled = [];
for (const stat of STATS) byStat[stat] = [];
for (const r of picked.values()) {
const b = lines[`${r.game_date}|${r.player_key}`];
const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[r.stat](b));
if (v === null) continue;
const over = v > L;
const row = { date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 };
byStat[r.stat].push(row); pooled.push(row);
}
const pooledResult = {
n: pooled.length,
deciles: deciles(pooled),
favourite_bias: round4(favouriteBias(pooled)),
...dateBlockSignStability(pooled, 20260808),
};
const perStat = {};
let replicated = 0;
for (const stat of STATS) {
const rows = byStat[stat];
const fb = favouriteBias(rows);
const stab = dateBlockSignStability(rows, 20260808);
const ok = fb !== null && fb > 0 && stab.sign_stability !== null && stab.sign_stability >= SIGN_STABILITY_REQUIRED;
if (ok) replicated += 1;
perStat[stat] = { n: rows.length, deciles: deciles(rows), favourite_bias: fb === null ? null : round4(fb), ...stab, replicates: ok };
}
const pooledOk = pooledResult.favourite_bias > 0 && pooledResult.sign_stability >= SIGN_STABILITY_REQUIRED;
const verdict = pooledOk && replicated >= STATS_MUST_REPLICATE ? 'ROBUST' : 'NOT-ROBUST';
console.log(JSON.stringify({
phase: 'PHASE 1 — model-free, map-free favourite-longshot bias test',
criteria: { sign_stability_required: SIGN_STABILITY_REQUIRED, stats_must_replicate: STATS_MUST_REPLICATE, favourite_floor: FAVOURITE_FLOOR },
pooled: pooledResult,
per_stat: perStat,
stats_replicating: replicated,
VERDICT: verdict,
consequence: verdict === 'ROBUST'
? 'proceed to a low-parameter correction, validated as a NEW estimator'
: 'serve raw; the bias is not certifiable on this sample',
}, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
const round2 = (v) => Math.round(v * 100) / 100;
+131
View File
@@ -0,0 +1,131 @@
#!/usr/bin/env node
'use strict';
/**
* PHASE 3 — validate the low-parameter correction as a NEW estimator.
*
* No grandfathering: it must beat RAW out-of-sample with a DATE-BLOCK bootstrap
* interval excluding zero. It is also scored against the retired isotonic map on
* the identical held-out rows, so the swap is a measured comparison rather than
* a preference.
*/
require('dotenv').config();
const fs = require('fs');
const path = require('path');
const { createClient } = require('@supabase/supabase-js');
const cal = require('../src/services/model/calibration');
const lp = require('../src/services/model/lowParamCalibrator');
const guards = require('../src/services/model/calibrationGuards');
const { knownNumber } = require('../src/utils/known');
const BOX = path.join(process.cwd(), '.seq-cache', 'batting-lines.json');
const STATS = ['hits', 'total_bases', 'rbi', 'runs'];
const PAGE = 1000; const ITERS = 4000;
const FIELD = { hits: (b) => b.hits, total_bases: (b) => b.totalBases, rbi: (b) => b.rbi, runs: (b) => b.runs };
const mean = (xs) => (xs.length ? xs.reduce((a, b) => a + b, 0) / xs.length : null);
async function page(sb, t, s, f) {
const o = [];
for (let i = 0; ; i += PAGE) {
const { data, error } = await f(sb.from(t).select(s)).order('id', { ascending: true }).range(i, i + PAGE - 1);
if (error) throw error; if (!data.length) break; o.push(...data); if (data.length < PAGE) break;
}
return o;
}
const isPreGame = (c, g) => {
const et = new Date(new Date(c).getTime() - 4 * 3600 * 1000);
const d = et.toISOString().slice(0, 10);
return d < g || (d === g && et.getUTCHours() < 19);
};
function makeRnd(seed) { let s = seed >>> 0; return () => { s ^= s << 13; s >>>= 0; s ^= s >>> 17; s ^= s << 5; s >>>= 0; return s / 4294967296; }; }
/** Paired date-block bootstrap on a Brier difference. */
function dateBlockCI(rows, keyA, keyB, seed) {
const byDate = new Map();
for (const r of rows) { if (!byDate.has(r.date)) byDate.set(r.date, []); byDate.get(r.date).push(r); }
const keys = [...byDate.keys()]; const rnd = makeRnd(seed); const diffs = [];
for (let it = 0; it < ITERS; it += 1) {
const s = [];
for (let i = 0; i < keys.length; i += 1) s.push(...byDate.get(keys[Math.floor(rnd() * keys.length)]));
const a = guards.safeBrier(s.map((r) => r[keyA]), s.map((r) => r.won));
const b = guards.safeBrier(s.map((r) => r[keyB]), s.map((r) => r.won));
if (a === null || b === null) continue;
diffs.push(a - b);
}
diffs.sort((x, y) => x - y);
return diffs.length
? { ci: [round4(diffs[Math.floor(diffs.length * 0.025)]), round4(diffs[Math.floor(diffs.length * 0.975)])], date_blocks: keys.length }
: { ci: null, date_blocks: keys.length };
}
(async () => {
const sb = createClient(process.env.SUPABASE_URL, process.env.SUPABASE_SERVICE_ROLE_KEY || process.env.SUPABASE_SERVICE_KEY, { auth: { persistSession: false } });
const lines = JSON.parse(fs.readFileSync(BOX, 'utf8')).lines;
const snaps = await page(sb, 'model_snapshots', 'id, game_date, captured_at, stat, player_key, line, side, p_win, refused',
(q) => q.eq('sport', 'mlb').in('stat', STATS));
const picked = new Map();
for (const r of snaps) {
if (!isPreGame(r.captured_at, r.game_date) || r.refused || knownNumber(r.p_win) === null) continue;
const k = [r.game_date, r.stat, r.player_key, r.line].join('|');
const prev = picked.get(k);
if (!prev || knownNumber(r.p_win) > knownNumber(prev.p_win)) picked.set(k, r);
}
guards.assertPickedSideDedup([...picked.values()].map((r) => ({
propKey: [r.game_date, r.stat, r.player_key, r.line].join('|'), side: r.side, p: knownNumber(r.p_win) })));
const out = {};
for (const stat of STATS) {
const rows = [];
for (const r of picked.values()) {
if (r.stat !== stat) continue;
const b = lines[`${r.game_date}|${r.player_key}`]; const L = knownNumber(r.line);
if (!b || L === null || !r.side) continue;
const v = knownNumber(FIELD[stat](b)); if (v === null) continue;
const over = v > L;
rows.push({ date: r.game_date, p: knownNumber(r.p_win), won: (String(r.side).toLowerCase() === 'under' ? !over : over) ? 1 : 0 });
}
rows.sort((a, b) => String(a.date).localeCompare(String(b.date)));
const dates = [...new Set(rows.map((r) => r.date))].sort();
const perDate = new Map(); for (const r of rows) perDate.set(r.date, (perDate.get(r.date) || 0) + 1);
let acc = 0; let cut = dates[dates.length - 1];
for (const d of dates) { acc += perDate.get(d); if (acc >= rows.length * 0.45) { cut = d; break; } }
const fit = rows.filter((r) => r.date < cut);
const ev = rows.filter((r) => r.date >= cut);
const platt = lp.fitPlatt(fit);
const iso = cal.fitIsotonic(fit.map((r) => ({ p: r.p, won: r.won })));
if (!platt || ev.length < 50) {
out[stat] = { n: rows.length, fit_n: fit.length, eval_n: ev.length, decision: 'REFUSE', reason: 'no low-parameter fit or too little held out' };
continue;
}
const scored = ev.map((r) => ({ ...r, plat: lp.applyPlatt(platt, r.p), isoP: iso ? cal.applyIsotonic(iso, r.p) : null }))
.filter((r) => knownNumber(r.plat) !== null);
const ys = scored.map((r) => r.won);
const bRaw = guards.safeBrier(scored.map((r) => r.p), ys);
const bPlat = guards.safeBrier(scored.map((r) => r.plat), ys);
const isoRows = scored.filter((r) => knownNumber(r.isoP) !== null);
const bIso = isoRows.length ? guards.safeBrier(isoRows.map((r) => r.isoP), isoRows.map((r) => r.won)) : null;
const vsRaw = dateBlockCI(scored, 'plat', 'p', 20260808);
const vsIso = isoRows.length ? dateBlockCI(isoRows, 'plat', 'isoP', 20260808) : { ci: null };
const beatsRaw = bPlat < bRaw && vsRaw.ci && vsRaw.ci[1] < 0;
out[stat] = {
n: rows.length, dates: dates.length, split_at: cut, fit_n: fit.length, eval_n: scored.length,
platt: { a: platt.a, b: platt.b, flattens: platt.flattens, fit_dates: platt.fit_dates, shrinkage: platt.shrinkage },
brier_raw: round4(bRaw), brier_lowparam: round4(bPlat), brier_isotonic: bIso === null ? null : round4(bIso),
delta_vs_raw: round4(bPlat - bRaw), ci_vs_raw: vsRaw.ci, eval_date_blocks: vsRaw.date_blocks,
delta_vs_isotonic: bIso === null ? null : round4(bPlat - bIso), ci_vs_isotonic: vsIso.ci,
decision: beatsRaw ? 'DEPLOY-PROVISIONAL' : 'REFUSE',
reason: beatsRaw ? 'beats raw out-of-sample with a date-block interval excluding zero'
: 'does not beat raw at a date-block interval excluding zero',
};
}
console.log(JSON.stringify(out, null, 2));
process.exit(0);
})().catch((e) => { console.error(e); process.exit(1); });
const round4 = (v) => (v == null || !Number.isFinite(v) ? null : Math.round(v * 10000) / 10000);
+109
View File
@@ -0,0 +1,109 @@
#!/usr/bin/env node
/**
* Session 63 — GRADE-RANGE VERIFICATION (on merit, not by rescaling).
*
* Replays REAL props from the live board through the REAL engine path that
* Session 63 repaired, and reports the resulting grade distribution.
*
* What is real here:
* - the props (player / stat / line / side) come from the live API board
* - the game logs come from statsapi.mlb.com + ESPN (free, no quota)
* - the consistency factor + L5/L20 features are computed from those logs
* - the grade comes from engine1.gradeProp, unmodified
*
* What is NOT covered (documented, not hidden):
* - `opp_rank_stat` needs `team_stats:{sport}:{abbr}` in Redis, which only
* the production snapshot populates. Locally it stays null, so this run
* UNDERSTATES the restored range — it omits a ±1.0 factor. Any A/D seen
* here is therefore a floor, not a ceiling.
*
* Usage: node scripts/verify-grade-range.js [sport] [limit]
*/
const engine1 = require('../src/services/intelligence/engine1');
const featureCache = require('../src/services/intelligence/featureCache');
const consistencyScore = require('../src/services/intelligence/consistencyScore');
const { estimateProbability } = require('../src/services/intelligence/probabilityEstimator');
const { fourLetterGrade } = require('../src/utils/gradeAdapter').__internals;
const API = process.env.VERIFY_API || 'https://api.vyndr.app';
const SPORT = process.argv[2] || 'mlb';
const LIMIT = Number(process.argv[3] || 40);
async function board(sport) {
const res = await fetch(`${API}/api/snapshot/${sport}`);
const json = await res.json();
const grades = Array.isArray(json.grades) ? json.grades : [];
return grades.map((g) => ({
player: g.player,
stat: g.stat_type,
line: Number(g.line),
direction: String(g.direction || 'over').toLowerCase(),
oldGrade: g.grade,
oldConfidence: g.confidence,
})).filter((p) => p.player && p.stat && Number.isFinite(p.line));
}
async function gradeOne(p, sport) {
const rows = await featureCache.getStatRows(p.player, sport, p.stat);
const features = await featureCache.__internals.gameLogFeatures(p.player, sport, p.stat);
const consistency = await consistencyScore.getConsistency({
playerName: p.player, sport, statType: p.stat, gameLogs: rows,
});
const prop = { line: p.line, direction: p.direction };
const res = engine1.gradeProp({ features, trap: {}, consistency, prop });
const est = estimateProbability({ gameLogs: rows, line: p.line, statType: p.stat, features });
const pWin = Number.isFinite(est.p_over)
? (p.direction === 'under' ? 1 - est.p_over : est.p_over)
: null;
return {
...p,
rows: rows.length,
consistency: consistency.consistency,
newGrade11: res.grade,
newGrade: fourLetterGrade(res.grade),
p_win: pWin == null ? null : Math.round(pWin * 1000) / 1000,
};
}
(async () => {
const props = (await board(SPORT)).slice(0, LIMIT);
if (!props.length) { console.log(`no live props for ${SPORT}`); return; }
console.log(`Replaying ${props.length} REAL ${SPORT.toUpperCase()} props through the repaired engine\n`);
const out = [];
for (const p of props) {
try { out.push(await gradeOne(p, SPORT)); }
catch (e) { console.warn(` ! ${p.player} ${p.stat}: ${e.message}`); }
}
const tally = (arr, key) => arr.reduce((m, r) => { const k = r[key] ?? 'null'; m[k] = (m[k] || 0) + 1; return m; }, {});
const pct = (n) => `${Math.round((n / out.length) * 1000) / 10}%`;
console.log('--- 4-LETTER DISTRIBUTION ---');
console.log('BEFORE (live board):', tally(out, 'oldGrade'));
const after = tally(out, 'newGrade');
console.log('AFTER (repaired) :', after);
for (const g of ['A', 'B', 'C', 'D', 'F']) if (after[g]) console.log(` ${g}: ${after[g]} (${pct(after[g])})`);
console.log('\n--- 11-STEP DISTRIBUTION (pre-collapse) ---');
console.log(tally(out, 'newGrade11'));
console.log('\n--- REVIVED SIGNALS ---');
const withRows = out.filter((r) => r.rows > 0).length;
const withP = out.filter((r) => r.p_win != null).length;
const withCons = out.filter((r) => r.consistency && r.consistency !== 'unknown').length;
console.log(`game-log rows present : ${withRows}/${out.length}`);
console.log(`p_win computed : ${withP}/${out.length} (was 0 in prod)`);
console.log(`consistency known : ${withCons}/${out.length} (was 0 for MLB)`);
console.log('\n--- MOVERS (grade changed) ---');
for (const r of out.filter((r) => r.oldGrade !== r.newGrade).slice(0, 15)) {
console.log(` ${r.oldGrade}${r.newGrade.padEnd(2)} (${r.newGrade11.padEnd(2)}) ${r.player} ${r.stat} ${r.direction} ${r.line} n=${r.rows} cons=${r.consistency} p=${r.p_win}`);
}
// Redis runs in degraded mode locally and keeps a reconnect timer alive, so
// the process would never exit on its own — flush and leave deliberately.
await new Promise((r) => process.stdout.write('', r));
process.exit(0);
})().catch((e) => { console.error('verify failed:', e.message); process.exit(1); });
+120
View File
@@ -0,0 +1,120 @@
#!/usr/bin/env node
'use strict';
/**
* verify-hits-v1 — STEP 2. WIRED IS NOT FIRING.
*
* A challenger that exists in the source and never produces a number on a real
* row is not a challenger, it is a comment. This induces the REAL code path —
* `projectionChallenger.attachProjection`, the exact function the snapshot
* calls — over the REAL hits props on the live production snapshot, with the
* REAL statsapi game-log adapter behind it.
*
* It reports FIRING COVERAGE: of the real hits props on the board, how many
* yield a hits-v1 read, how many abstain, and why. An abstention is a valid
* answer; a silent zero is not.
*
* The current ladder value is computed on the same rows in the same call, so the
* two are compared on identical inputs.
*
* node scripts/verify-hits-v1.js [snapshotUrl]
*/
const projection = require('../src/services/projectionChallenger');
const mlb = require('../src/services/adapters/mlbStatsAdapter');
const SNAPSHOT_URL = process.argv[2] || 'https://api.vyndr.app/api/snapshot/mlb';
async function fetchSnapshot(url) {
const res = await fetch(url, { headers: { accept: 'application/json' } });
if (!res.ok) throw new Error(`snapshot ${res.status}`);
return res.json();
}
async function main() {
const snap = await fetchSnapshot(SNAPSHOT_URL);
const grades = (snap.grades || []).filter(
(g) => String(g.stat_type || g.stat || '').toLowerCase() === 'hits',
);
const out = await projection.attachProjection(grades, {
// The one dep that matters here. Everything else (park/weather/platoon/
// arsenal) is absent on the public payload and contributes a documented
// 1.0 — which is the honest behaviour, not a fabricated push.
gameLogFor: async (g) => {
if (!g.playerId) return [];
try { return (await mlb.getPlayerGameLog(g.playerId)) || []; } catch { return []; }
},
});
const fired = out.filter((g) => g.proj_hits_p_over != null);
const abstained = out.filter((g) => g.proj_hits_p_over == null && g.proj_hits_meta);
const ladder = out.filter((g) => g.proj_p_over_line != null);
// The market read — proving the model was scoped by IDENTITY, not by price.
const withMarket = out.filter((g) => g.proj_hits_meta && g.proj_hits_meta.market);
const takeableIdentity = withMarket.filter((g) => g.proj_hits_meta.market.market_takeable);
const outsidePromotion = withMarket.filter((g) => g.proj_hits_meta.market.within_promotion_band === false);
const oneSided = withMarket.filter((g) => g.proj_hits_meta.market.one_sided);
// The rows the whole disambiguation exists for: real markets that a
// price-shape rule would have thrown away, and which we modelled anyway.
const juicedModelled = fired.filter((g) => {
const m = g.proj_hits_meta.market;
return m.market_takeable && m.within_promotion_band === false;
});
const pct = (a, b) => (b ? Math.round((a / b) * 1000) / 10 : null);
const nums = fired.map((g) => g.proj_hits_p_over);
const avg = (a) => (a.length ? Math.round((a.reduce((x, y) => x + y, 0) / a.length) * 1000) / 1000 : null);
const sd = (a) => {
if (a.length < 2) return null;
const m = a.reduce((x, y) => x + y, 0) / a.length;
return Math.round(Math.sqrt(a.reduce((s, v) => s + (v - m) ** 2, 0) / (a.length - 1)) * 1000) / 1000;
};
const ladderNums = ladder.map((g) => g.proj_p_over_line);
console.log(JSON.stringify({
snapshot: { url: SNAPSHOT_URL, updated_at: snap.updated_at, total_grades: (snap.grades || []).length },
hits_props: grades.length,
hits_v1: {
fired: fired.length,
firing_coverage_pct: pct(fired.length, grades.length),
abstained: abstained.length,
abstain_reasons: abstained.reduce((acc, g) => {
const r = g.proj_hits_meta.reason || 'unknown';
acc[r] = (acc[r] || 0) + 1; return acc;
}, {}),
mean_p: avg(nums), sd_p: sd(nums),
p_range: nums.length ? [Math.min(...nums), Math.max(...nums)] : null,
},
current_ladder: {
fired: ladder.length,
mean_p: avg(ladderNums), sd_p: sd(ladderNums),
},
takeable_axis: {
note: 'market scope = book IDENTITY; promotion band recorded, never gates the model',
rows_with_market_read: withMarket.length,
takeable_by_identity: takeableIdentity.length,
outside_promotion_band: outsidePromotion.length,
one_sided_quotes: oneSided.length,
juiced_or_longshot_MODELLED_anyway: juicedModelled.length,
any_price_filtered: withMarket.some((g) => g.proj_hits_meta.market.price_filtered),
},
sample: fired.slice(0, 5).map((g) => ({
player: g.player, line: g.line, side: g.direction, book: g.book,
hits_v1_p_over: g.proj_hits_p_over,
ladder_p_over: g.proj_p_over_line,
champion_p_win: g.p_win ?? null,
hit_rate: g.proj_hits_meta.hit_rate,
ab_per_game: g.proj_hits_meta.ab_per_game,
games_used: g.proj_hits_meta.games_used,
market: g.proj_hits_meta.market,
})),
}, null, 2));
// Redis is degraded locally; its reconnect timer would hold the process open
// and piped output would be lost to SIGTERM. Same rule as verify-grade-range.
process.exit(0);
}
main().catch((e) => { console.error(e); process.exit(1); });
+53
View File
@@ -0,0 +1,53 @@
# The 83-glyph taxonomy — doctrine
**RULED. This is the shared law the archetype chat and the build chat both obey.**
## The rule
**83 designed glyphs = the full four-sport archetype taxonomy** (MLB / WNBA /
NBA / Soccer). The artwork is complete; the *models* are not.
**A glyph renders ONLY where its archetype is modeled and proven.**
| set | count | state |
|---|---|---|
| designed glyphs | **83** | complete, in `specs/design-reference/assets/glyphs/` |
| backend registry archetypes | **41** | `archetypeService.ARCHETYPES` |
| **mapped and wired** | **39** | live, colours matching the registry exactly |
| **designed, no backend archetype** | **44** | **DORMANT** — slots for WNBA/NBA/Soccer |
| registry archetypes with no glyph | **2** | `DUAL THREAT`, `PAINT BOSS`**design gap** |
## Why dormancy rather than wiring
Wiring the 44 would mean **inventing 44 archetypes to consume artwork**. An
archetype that exists because a glyph exists is decoration presented as
classification — a mark on a card asserting the model recognised something it
cannot produce. **Decoration-as-data is forbidden.**
The dormant glyphs are not waste. They are **designed slots**, and each activates
when its sport's archetype system is built and clears the two-part gate — the
same bar every factor in this programme faces.
## What this preserves
- **The full design vision.** All 83 marks stay in the package; none is deleted
or redrawn.
- **The Truth Law.** No glyph appears for an archetype the model cannot produce.
- **The activation path.** Building WNBA/NBA/Soccer archetypes lights their
glyphs automatically — the artwork is already there and already colour-matched.
## Consequences for build
1. Do **not** add a registry archetype to consume a glyph. The archetype must be
earned by a classifier that produces it from real features.
2. Do **not** render a dormant glyph as a placeholder, sample, or "coming soon"
mark on any data surface.
3. `DUAL THREAT` and `PAINT BOSS` are **flagged to the design side** — they are
modeled archetypes with no mark, the mirror image of the dormant 44.
4. When a sport's archetypes ship, wiring is a MANIFEST lookup, not new art.
## Status
MLB archetypes render live. WNBA, NBA and Soccer have archetype **registries**
but no proven factor model, so their marks stay dormant — consistent with
`specs/BOARD-2026-08-07.md`, which records only MLB batters as MODELED.
+132
View File
@@ -0,0 +1,132 @@
# The real board — read from the repo, 2026-08-07
Inventory only. Every line cites a file or a query. Anything unverifiable is
marked UNKNOWN rather than asserted.
---
## PHASE 0 — Design / terminal
| item | state | evidence |
|---|---|---|
| **Scanner blue-boundary format** | **DONE** | `scan/page.tsx:715-717`*"the blue boundary channel (`--priced-out`): line · BK odds · ◆ fair"* and *"renders NEUTRAL, not amber (blue-boundary law)"*. Amber survives only as an unrelated CTA (`:983`) and a status line (`:846`). |
| **MovementStrip** | **NOT-STARTED** *(as named)* | No component by that name. Movement renders via `GradeShift.tsx`, `GameCard.tsx`, `MarketBreadth.tsx`, `FuturesBoard.tsx`. **UNKNOWN whether the spec wants a distinct strip or is satisfied by these.** |
| **THE WIRE** | **PARTIAL** | `components/vyndr/NewsWire.tsx` exists and is mounted — but only inside `ExploreHub.tsx`. No standalone wire surface. |
| **Book comparison** | **PARTIAL — built, unmounted** | `vyndr/BookComparisonPanel.tsx` + `components/BookComparison.tsx` both exist; grep of `web/src/app` returns **zero** mounting pages. Same built-but-unread class the reachability guard was written for; **not covered by that guard** (contract is grade fields only). |
| **Article media** | **NOT-STARTED** | No component. `/blog` exists. |
| **Offseason hub** | **NOT-STARTED** | No component, no route. |
| **Terminal** | **DONE — deliberately retired** | `app/terminal/page.tsx` redirects to `/dashboard`; layouts preserved unrouted in `components/intel/TerminalTemplates.tsx` (S57). Not debt. |
| **Screen conversion (S3639 arc)** | **DONE except one** | 47 route dirs under `app/`. `RouteStub` survives in exactly **one** file: `app/notifications/page.tsx`. Every other screen has a real page. |
**Sports surfaces that exist as pages:** `soccer` (414 lines), `desk` (160),
`intelligence` (160), `marketplace` (123). No `fight` route.
---
## PHASE 1 — Sports × role
Archetype registry (`archetypeService.ARCHETYPES`): **nba 15, mlb 15, soccer 6,
wnba 5** — all four sports have real archetype registries, not names only.
`snapshotService.ACTIVE_SPORTS = ['mlb','nba','wnba','soccer']` — the pipeline
runs all four.
| sport · role | state | evidence |
|---|---|---|
| **MLB · batters** | **MODELED** | The whole session. `hitsFactors` (3 proven factors, transmitting), `servedGrade`, repaired champion, settled ledger. |
| **MLB · pitchers** | **SCAFFOLDED** | `model/pitcherEngine.js` (261 lines, own FLAME/SCALPEL/SINKER archetypes) — **read by no serving code** (grep: zero non-test consumers). Strikeouts n=57 settled vs a 500 gate. Base rate repaired by the shared `getStatRows` MLB fix. |
| **WNBA** | **SCAFFOLDED** | Archetypes exist; settles via ESPN box scores (376 rows historically); no factor model. Base-rate path fixed this session. |
| **NBA** | **SCAFFOLDED, dormant** | Archetypes exist; offline (Python service down, off-season). Base-rate paths fixed dormant at `55b210c`. |
| **Soccer** | **SCAFFOLDED** | 6 archetypes, a real `/soccer` page, feature extractor exists. No factor model, no settled outcomes. |
**Nothing but MLB batters is MODELED.** Everything else is scaffolding with a
sound base rate and no proven factors.
---
## PHASE 2 — Wiring / data-integrity debt
### `edge_pct` — the flag is half right, and the diagnosis was wrong
**Not a scale bug.** `analyzeViaEngine1.edgePctFor` computes
`(model line) / line`, signed by direction — arithmetically correct. It
*explodes on small lines*: a 5.5 projection against a 0.5 line is a legitimate
1000%. Live top values: **900, 860, 700, 700, 660** on **68,364 rows**.
**Consuming surfaces: 15 backend files + 10 frontend files.** But
`edge_pct` itself reaches a user only through `alt_lines` typing in
`scan/page.tsx:64` — the grade card's `edge` is computed independently in
`gradeAdapter` via `computeEdge`. So the blast radius is smaller than the file
count implies.
**`ev_pct` DOES render** — 5 frontend files (`PriceTriplet`, `GradeResultCard`,
`LiveHeroProp`, scan page + route). The "ev_pct renders nowhere" flag is **stale**.
45,125 rows carry it.
**Classification: PARTIAL — a real display defect (a 900% edge is not a sentence
we can defend), not a broken computation.**
### `opp_rank_stat`
`refreshTeamStats` is wired into `snapshotService:361`. **UNKNOWN whether it
populates in production** — not verifiable from the repo, needs a live probe.
### A-emit
**Resolved.** `servedGrade.UNISSUABLE = ['A+','A','A-']`; post-cutover serving
cannot emit one. Historical snapshot rows still carry engine1 A's — that is
retired data, not live behaviour.
### Other
| item | state |
|---|---|
| `VYNDR_INTERNAL_KEY` | present in `app.js`, `preflight.js`. **Rotation status UNKNOWN** from the repo. |
| void / DNP handling | handled in `ledgerService` (`:64`, `:698`) — absence means DNP/postponed/not-final. **The ~9% rate is UNVERIFIED** here. |
| **Guard coverage gap** | reachability contract covers **grade fields only** — 0 references to `edge_pct`/`ev_pct`. Book comparison and other unmounted components are **not** guarded. |
---
## PHASE 3 — THE BOARD, ordered (half-done first)
| # | item | lane | state | serving-path? |
|---|---|---|---|---|
| 1 | **Book comparison — built, unmounted** | design | PARTIAL | **N** |
| 2 | **`edge_pct` display defect (900%)** | integrity | PARTIAL | **Y** |
| 3 | **THE WIRE — only inside ExploreHub** | design | PARTIAL | **N** |
| 4 | **MLB pitchers — engine unwired** | sports | PARTIAL | **Y** |
| 5 | **`/notifications` RouteStub** | design | PARTIAL | **N** |
| 6 | Guard coverage → non-grade surfaces | integrity | NOT-STARTED | **N** |
| 7 | Article media | design | NOT-STARTED | **N** |
| 8 | Offseason hub | design | NOT-STARTED | **N** |
| 9 | MovementStrip *(if distinct from GradeShift)* | design | UNKNOWN | **N** |
| 10 | WNBA / NBA / Soccer factor models | sports | NOT-STARTED | **Y** |
| 11 | `opp_rank_stat` prod population | integrity | UNKNOWN | **Y** |
| 12 | Internal-key rotation | ops | UNKNOWN | **N** |
### Accrual sensitivity
**Isolated — buildable now without touching the re-audit clock (7 items):**
1, 3, 5, 6, 7, 8, 9 — all design/guard work, zero serving-path contact.
**Serving-path — will muddy the accruing repaired-champion dates (5 items):**
2, 4, 10, 11, and any model work. Every one changes what a snapshot writes, so
rows produced after the change are not comparable to rows before it. **If any of
these ship, the eligible-date clock arguably restarts** — the same reasoning that
voided the shadow duel.
**#2 (`edge_pct`) is the sharpest tension on the board:** it is a real honesty
defect a user can see, and fixing it touches the serving path mid-accrual.
### The clock, today
```
eligible dates: 0 — no settled rows carry engine1@2026-08-07-fullwindow
calibration_refit 0/10 WAITING (~2 weeks)
hits_factor_lift 0/10 WAITING (~2 weeks)
prior_verdict_reaudit 0/14 WAITING (3+ weeks)
rbi_lineup_slot_gate 0/14 WAITING (3+ weeks)
```
Two tracks are now visible: **modeling is date-blocked**; **seven design/guard
items are buildable today with no clock impact.**
+111
View File
@@ -0,0 +1,111 @@
# The content engine — posts that structurally cannot lie
## Phase 0 — architecture
`src/services/content/contentEngine.js`. Three mechanisms make Truth Law
structural rather than careful:
1. **Copy is token-substituted.** Every factual claim is a `{token}` resolved
against pulled facts. An unbacked token **refuses to render** — there is no
code path producing a plausible default.
2. **The fact contract is asserted first.** A template declares required fields;
they are checked *before any string is built*.
3. **Card and copy share one fact object.** They cannot diverge.
**No live model writes factual claims.** The voice is in the template, the facts
are pulled. A voice-polish port is reserved and deliberately unwired — an LLM
that can rewrite a sentence can rewrite a number.
**Read-only on every source.** Zero writes to serving, model or ledger tables, so
zero effect on the accrual clock.
### The Truth-Law proof — 18 tests
| the guard | what it prevents |
|---|---|
| unbacked token refuses | `{edge}` rendering as `undefined` or an empty hole |
| card tokens gated too | a caption that's honest beside a card that isn't |
| `render()` throws directly | a caller bypassing the gate |
| `null` never renders as `"null"` | absence dressed as data |
| **`0` IS present** | *"0 cleared B+"* is our most honest post — deleting it would be `Number(null)===0` in reverse |
| `NaN`/`Infinity` absent | arithmetic failures are not facts |
| contract gap names the field | a silent half-post |
| pull failure skips | a post built on a dead source |
## Phase 1 — three templates, real output
**HOT HITTERS** (from the repaired full-season log, not a ten-game slice):
> Jahmai Jones is hitting 60% over his last 10. His season number is 26%.
> That gap is the whole point. Everybody else is guessing at it.
**THE HONESTY FLEX** — the differentiator, and every number is ours:
> WE GRADED 2140 PROPS TONIGHT. 70 CLEARED B+.
> That's 3%. The other 42% we can't separate from the baseline, and we say so on the card instead of calling them leans.
> We do not issue A+, A, A-. No band of this model has ever hit at a rate that would justify one.
> Everybody else's card is all A's. Ask them what their A actually hits.
**STREAK LIST** — verified from settled outcomes only:
> Nathan Church has a 7-game hit streak. Live, verified off settled results only.
> Every game in these ran to a final. We don't count a pending night to make a number look better.
### The bug the engine caught in itself
The first run emitted *"No hitter is meaningfully hot tonight — we could dress up
a middling week as a streak. We don't."*
**That was false.** The box-score cache spans only the settled snapshot window,
so **every** player had fewer than 20 games and the pool was empty. A broken pull
was publishing as considered editorial judgement — **the fourth appearance of
this class tonight, and the first where our own honesty copy was the disguise.**
Fixed structurally: an `absent()` variant may now **decline to speak**. The
template separates *no candidates at all* (SKIP with a reason) from *candidates
judged, none hot* (honest absence). Both cases are locked by test.
Source corrected to `mlbStatsAdapter.fullLog` — the same log the repaired
champion reads.
## Phase 2 — the card
`cardRenderer.js`, SVG rather than canvas: it is text, so it diffs in review and
its numbers are **greppable** — which matters when the entire claim is that the
numbers are real. A card whose contents can't be inspected without opening an
image is a poor fit for a Truth-Law product.
Brand: VYND white + R green `#00D4A0`, slashed-Y, scanline field, mono
throughout. The card never formats its own facts — every string arrives already
rendered and gate-checked, so caption and card cannot disagree. A test asserts
the pulled number appears in the emitted SVG.
## Phase 3 — posting-ready, and extending it
```
SUPABASE_URL=... node scripts/generate-content.js
-> .content-out/2026-08-07/hot_hitters.txt + .svg
-> .content-out/2026-08-07/honesty_flex.txt + .svg
-> .content-out/2026-08-07/streak_list.txt + .svg
```
Kev posts; the engine generates.
### Adding template N+1 — registry entry only, no engine change
```js
registerTemplate({
id, sport, requires: ['dotted.paths'],
pull: async (deps) => facts, // the ONLY place data enters
copy: () => 'text with {tokens}',
card: () => ({ title, subtitle, lines }),
absent: (gaps, facts) => ({ copy, card }) // or { skip: 'reason' }
});
```
**Queued (stubs, not built):** hot takes · daily honest reads · *"grades we
DIDN'T give"* · cross-sport streak variants (the streak template is already
sport-agnostic — it takes settled outcomes and a noun, so NFL TD streaks or NBA
made-three streaks need only that sport's settled data).
## Isolation
Read-only throughout. `p_win`, the model and the serving path are untouched;
the eligible-date clock is unaffected. **0 eligible dates today, unchanged.**
+198
View File
@@ -0,0 +1,198 @@
# VYNDR build handoff — 2026-08-08
Written from the repo, not from summary. Every claim below was grep- or
run-verified at `71d3b7b`. Where something could not be verified from the repo
it says UNKNOWN.
## Repo state
```
HEAD 71d3b7b78692c31ba0596ec88874986ff1b58116
gitea main 71d3b7b78692c31ba0596ec88874986ff1b58116 (in sync)
tree clean — 0 modified/untracked
tests 4,539 passing / 4 skipped
web build exit 0
```
**Push remote is `gitea`, not `origin`.** `origin` is GitHub and has no
credential on this machine; every push must be `git push gitea main`. This cost
twenty commits of false "no credentials" diagnosis — see
`docs/GIT-PUSH-DIAGNOSIS.md`.
---
## BUILD STATE per surface
### DONE this arc — verified present
| surface | file | lines |
|---|---|---|
| **E1 movement strip** | `web/src/components/vyndr/MovementStrip.tsx` | 113 |
| **F9F11 offseason hub shell** | `web/src/app/offseason/page.tsx` + `components/vyndr/OffseasonHub.tsx` | 21 + hub |
| **E10 Report issue template** | `src/services/report/reportTemplate.js` | 169 |
| **E12 /report archive** | `web/src/app/report/page.tsx` + `components/vyndr/ReportArchive.tsx` + `src/routes/report.js` | 17 + comp + route |
| **Content engine** | `src/services/content/contentEngine.js` + 3 templates + `cardRenderer.js` | 172 + |
| **Content studio API + preview** | `src/routes/contentStudio.js` + `web/src/app/studio/page.tsx` | 175 + 131 |
| **Wave-D1 motion primitives** | `web/src/lib/motion.js` (E17/E18/E27/E28) | 100 |
| **Honest served grade** | `src/services/model/servedGrade.js` | 142 |
| **Archetype doctrine** | `specs/ARCHETYPE-TAXONOMY-DOCTRINE.md` | 53 |
Also DONE and verified: **D1 glyph library** (39 of 83 wired — every glyph that
maps to a real archetype; colours match the registry exactly), **A1 card token**
(`--bg-1`, 32 consumers), **B1 boundary channel** (`--priced-out`, 4 consumers),
**book comparison** (mounted via `GradeResultCard → scan/page`), **THE WIRE**
(`vyndr/Ticker`).
### GATED — verified absent, with the gate named
| surface | gate | verified |
|---|---|---|
| **F5 article media** | card-system reconciliation | `ArticleHero`/`StatCallout`/`PullQuote` = 0 files |
| **E16/F8 share cards + crops** | card-system reconciliation *and* the resolution tail has no share-card generation step | `ShareCard.tsx` has **0 importers** |
| **In-season hub IA** | the content formula (social chat) | no spec exists anywhere |
| **E9 calibration curve** | model accrual — 0 eligible dates | `CalibrationCurve` = 0 files |
| **E15 Price Triplet MODEL leg** | model — EV layer | renders honest `NO_MODEL` |
| **E2 BookChip tiles / E6 push-to-book** | licensing / affiliate approval | every book `enabled:false` |
| **E3 best-number crown** | measurement — measured flat | `CrownBadge` = 0 files |
| **E13 Scanner S6** | another build order | — |
**There are no ungated Wave-2 targets left.** The next move is a decision, not a
build.
---
## THE TWO OPEN DECISIONS
### 1. Card-system reconciliation — unblocks F5 *and* E16/F8 together
The content engine emits **1080×1350 SVG** cards (`src/services/content/cardRenderer.js`).
The design package defines **five master sizes** with layouts and rasterised PNGs
in `specs/design-reference/exports/`: settle 1080×1350, story 1080×1920, square
1080, X 1200×675, record 1080×1350, article OG 1200×630.
**I built the renderer without checking whether a designed card system existed.
It did.** They are a parallel invention — not wrong, but they must become one
visual language before F5 (whose hero graphics are generated data visuals) or
E16 can proceed.
The decision: does the content engine conform to the E16 masters, or do the
masters absorb the engine's generator? **Cheapest unblock on the board — it
frees two gated items at once.**
### 2. In-season hub information architecture
`Vyndr Offseason.dc.html` specifies an **offseason** hub, and its shell is built.
Nothing specifies what that surface is **in-season**, or how content, articles,
wire and the live slate share year-round navigation. Every component exists; the
composition does not.
The spec is explicit that **sport state lives in the sport tab** (`NFL · CAMP 5D`)
and never in a separate offseason tab — so this is not a new route, it is a mode.
Gated on the social chat's content formula, which determines the recurring slots.
---
## ACCRUAL CLOCK
```
champion marker: engine1@2026-08-07-fullwindow
calibration_refit 0/10 WAITING (~2 weeks)
hits_factor_lift 0/10 WAITING (~2 weeks)
prior_verdict_reaudit 0/14 WAITING (3+ weeks)
rbi_lineup_slot_gate 0/14 WAITING (3+ weeks)
blocked: no settled rows yet carry the repaired champion marker
```
**FIRST TRIGGER: 10 eligible calibration dates → calibration re-fit.** Then, in
order: hits factor lift → prior verdict re-audit → rbi lineup-slot gate.
**Thresholds are ATTEMPT floors, not TRUST floors.** Reaching 10 dates means the
re-fit *can be measured*, not that it is trustworthy. A 10-date map is thin,
deploys PROVISIONAL with auto-demotion, and its interval will be wide.
**No measurement on reconstructions of the retired forecast.** Anything requiring
settled data waits. `src/services/model/reAuditEligibility.js` enforces it —
eligibility is a version marker, counted in DATES not rows, and a mixed table
counts only the repaired rows.
---
## STANDING DOCTRINE
**Truth Law.** No fabricated data anywhere — including **designer sample data**.
The design files' numbers (Nabers 1,120.5, Wembanyama +420→+330, Nº 128, DAY
RECORD 94) are a *spec for what a live feed renders*, never content to paste.
Tests assert none of it ships. `Number(null) === 0` is the classic breach; note
its inverse also bites — **zero is a real fact** ("0 cleared B+" is our most
honest post).
**83-glyph four-sport taxonomy.** A glyph renders ONLY where its archetype is
modeled and proven. 39 wired; **44 dormant slots** for WNBA/NBA/Soccer. Wiring
them would mean inventing 44 archetypes to consume artwork — decoration
presented as classification, forbidden. `DUAL THREAT` and `PAINT BOSS` are
modeled archetypes with no mark: the mirror gap, flagged to design.
**Repaired champion.** The forecaster read `res.last10`**ten games** — as its
season rate. Fixed to the full season log with recency weight 0.40 → 0.20.
Resolution: hits 0.00251 → 0.00817. **Every factor verdict in the programme was
measured against the broken baseline and may deserve re-audit** — direction
unknown, not pre-priced.
**Calibration is WITHDRAWN.** `CALIBRATION_DEPLOYED = []`. The maps were fitted
on the retired forecast; refitting today would refit it again. Served `p_win` is
repaired-champion raw.
**Grade doctrine.** `A+`, `A`, `A-` are **unissuable** — structurally, not rarely.
Ceiling is **B+ at 0.663 realized against a 0.6005 baseline**. Bands that cannot
be separated from the baseline SAY so. The served letter derives from `p_win`;
`engine1.grade` is preserved as `engine_grade` and read by no serving code.
**Two guards, both in CI:**
- `renderReachability.test.js` — a promised field must trace payload → adapter →
component → **mounted**. Built-but-unread killed three surfaces before this.
- `baseRateWindow.test.js` — no fixed N may stand in for a season. Six paths
across two sports fell to that class.
**Compose, don't fork.** The recurring failure this arc: building beside an
existing thing instead of on it (cards vs E16, `MovementStrip` nearly vs
`gradeShift`, `reportTemplate` nearly vs `newsletterService`). Check for the
existing implementation first — twice it was already there.
---
## PARALLEL WORKSTREAM SEEDS
### A. Social strategy chat
> VYNDR's content engine ships posts from real data (`src/services/content/`,
> three templates, Truth-Law enforced — an unbacked token refuses to render).
> Review `/studio` output and `specs/CONTENT-ENGINE.md`. Produce: the content
> formula/calendar (what posts recur, on what cadence), and **the card-format
> decision** — the engine's 1080×1350 renderer vs the five designed masters in
> `specs/design-reference/exports/`. That decision unblocks F5 article media and
> E16/F8 on the build side. Voice: `specs/VOICE.md` governs; the engine templates
> encode a sharper register — decide whether that becomes the house voice.
### B. Multi-sport archetype build
> VYNDR has archetype registries for MLB (15), NBA (15), Soccer (6), WNBA (5) —
> but **only MLB batters is MODELED**; the rest have no proven factor model, so
> their glyphs stay dormant per `specs/ARCHETYPE-TAXONOMY-DOCTRINE.md`. 44 of 83
> designed glyphs await their sport. Pick a sport, build its archetype
> classifier from real features, and take it through the two-part gate
> (`src/services/model/factorGate.js` — must MOVE the prediction and improve
> out-of-sample Brier, cumulative-Bonferroni corrected). Prerequisite: that
> sport needs settled outcomes. WNBA settles via ESPN box scores; NBA and Soccer
> do not settle at all yet, which is the real first task there.
---
## Where to start in the next chat
1. Take one of the two open decisions (card system is cheapest — unblocks two).
2. Or run a parallel workstream seed above.
3. **Do not** start model work — it is date-blocked and will stay so for ~2 weeks.
4. Verify before building: this arc corrected its own board four times. `git
grep` beats any summary, including this one.
+648
View File
@@ -0,0 +1,648 @@
# VYNDR — MASTER PLAN
**Single source of truth. Sessions EXECUTE against this and UPDATE it in place.**
Created 2026-07-31 by consolidation. Last updated 2026-08-01.
---
## 🔷 PRODUCT IDENTITY — the thing being built
**VYNDR IS A PREDICTIVE MODEL.** It projects what a player will **DO** and picks
accurately — reading and pulling the market apart. Student of the game, and an
aggregator.
**Market edge is a BYPRODUCT of a good prediction, NEVER the success criterion.**
> **SUCCESS = the forecast is honest about its own confidence AND still ranks.**
> Calibration *and* resolution. **No edge/CLV term belongs in a pass/fail gate** —
> they are diagnostics we report, not thresholds a model must clear to ship.
**PER-SPORT DOCTRINE** (Rashad Phillips, *Basketball Position Metric*, 2022 —
classify players by **what they do**, not position labels): each sport is its OWN
model, with its own variables, archetypes, conditions, calibration and honest
ceiling. Shared across sports: **only the Bayesian inference math.**
**TRUTH LAW:** no fabricated data anywhere · honest-absent over invented · label
limitations in-band ("market consensus, **not sharp**") · **provisional results
stay provisional until re-run** · documented ≠ verified.
---
## ▶ NEXT EXECUTABLE ORDER
**🔴 (a) TAKEABLE ENFORCEMENT — urgent.** `specs/takeable-enforcement-audit.md`.
**The ledger — the table every accruing holdout resolves against — is stamping
`book`, `locked_odds` and the `takeable` flag itself from books you cannot bet**
(DFS `dabble` 707 rows, offshore `bovada` 214, `onexbet` 42, exchange `kalshi` 7).
**0% before 2026-08-01, 47.9% on 08-01, 42.5% on 08-02** — it began with the
display widening. `recordPipelineGrades` indexes over the FULL widened props list
and prefers that prop's book/odds over the graded one.
**47 contaminated rows have already settled; ~700 are pending and will settle into
the holdouts.** The damage is mostly ahead of us.
Also latent: **`MODEL_BOOKS``TAKEABLE_BOOKS`** (`pinnacle` is model-eligible but
not takeable) — harmless while pinnacle returns nothing, live again if it recovers.
**Keep the split:** the *prediction target* must be takeable; `fair_prob` /
consensus / `edge` **stay reference** and must not be over-enforced.
**Then (b) a structural `Number(null) === 0` guard** — six occurrences, most
recently in my own new module. **`hits` will re-trigger it**: its 0.5 lines make
P(0) the whole game, where a null rate read as zero is maximally wrong.
**Then (c) the `hits` fix**, built on both invariants — not before them.
**Carry-forward:** tb-v1 verdict (accruing) · the third pre-registered branch
(improves-ranking-but-low → promote **and** recalibrate) · the **100s Cloudflare
origin timeout** vs a ~115s snapshot (work completes; the trigger returns 524 —
flag it before it reads as a failure).
### FIVE challengers accruing in parallel — do NOT re-run early
Verified firing on a real prod snapshot (293 grades), not inferred:
| axis | coverage | mean \|nudge\| | holdout query |
|---|---:|---:|---|
| **environment** | **84.6%** | 0.057 | *(shares the arch-v1 pattern)* |
| **matchup** (`batter_own_split`) | **82.9%** | 0.015 | `scripts/matchup-axis-holdout.sql` |
| **opportunity** | 30.0% | 0.142 | `scripts/opportunity-axis-holdout.sql` |
| **tb-v1** (compound total_bases) | **10/10 TB props** | *(shape, not a nudge)* | pre-registered: improves → `hits` next; doesn't → similarity reopens |
| **proj-v1.1** (distribution ladder) | **94.2%** | *(forms the projection, not a nudge)* | 437 settled — aligned gap **0.252 vs 0.352**, concentrated in `hits`+`total_bases` |
All three orthogonal (r ≈ 0 vs projection, `p_win`, line and each other). Each
promotes ONLY on its own axis-filtered holdout, and ONLY if **reliability AND
resolution** improve. **This is time, not code.**
Both promote only on their own axis-filtered holdout, and only if **reliability
AND resolution** improve.
## 📍 STATE AS OF 2026-08-01 (Order Zero, measured on prod with the real key)
| finding | number | what it means |
|---|---|---|
| **MLB slate invisible to us** | **64.8%** | our own allow-list, not the feed — now widened for DISPLAY |
| **books/prop, MLB** | 3.61 feed → 0.57 after filter | the filter cost, quantified |
| **books/prop, WNBA** | **4.21** feed → 1.20 | **WNBA is BETTER covered than MLB** |
| **consensus ruler** | **MARKET, not SHARP — ⏳ PENDING-RECOVERY, not permanent** | `matchbook`/`polymarket` = 0%, but **`pinnacle` ran until 07-30** (103,940 captures) and stopped. **Do not enshrine as permanent** until PropLine answers — see `BLOCKERS.md` |
| **ruler delta** (consensus incumbent) | MLB mean +1.50 pts, median 0, **17% of props move ≥5 pts** | rulers genuinely differ; "better" is unproven |
| **MLB isotonic `p_win`** | **DECIDED** — reliability **0.0846**, resolution **0.190**, holdout **n=125** | **PROVISIONAL label RETRACTED 2026-08-01.** Calibration is **ruler-independent** (`estimateProbability` never sees a price; the fit is p_win-vs-outcome). Replicated on a fresh later window, both metrics improved |
| **edge vs the ruler** | corr(edge, outcome) **0.010** (v1) → **0.022** (v2), n=200 · corr(**p_win**, outcome) **+0.26** | **Subtracting the market DESTROYS the signal.** The consensus ruler does not rescue edge: *differs ≠ better* |
| **exchange-inclusive ruler** | **UNTESTABLE on existing data** | exchange quotes were never stored (discarded until 2026-08-01). Becomes testable only as v2-era captures accrue |
| **WNBA** | **MODEL NOT BUILT YET** — held out, *not* failed | **CORRECTED 2026-08-01.** The 0.12 was **NBA-template machinery run on WNBA data**. WNBA has never had its own archetypes/variables/conditions — the "sport stubbed in on another sport's template" CLAUDE.md forbids. That is an **unbuilt model's expected failure, not a verdict on the sport.** Its own build is QUEUED, after MLB |
| 🔴 **pinnacle feed** | **0 captures since 2026-07-31** (103,940 in the prior 10 days) | a live regression; **we had a sharp anchor and lost it.** Not caused by our changes |
| **soccer** | **settles** — ~15 competitions, 30d | "grades into a void" is a **$19/mo Pro-tier** problem, not a data problem |
| **CLV + results feeds** | `/odds/closing` + `/movement` **redacted**; `/results` **403 `required_tier: hobby`**; `/exports/resolved-props` **403 `required_tier: pro`** | **verified on our keys** — plain tier exclusion, not a key or plan fault. **$9/mo** buys CLV + steam + results; **$19/mo** adds the 90-day settlement export |
| **books SERVED** | 5 → **13**; props rendered **546 → 2,780** (5.1×); mean **4.22** books/prop | **AGGREGATOR widening is LIVE.** Of 2,234 newly-visible props, **31.2% carry a real non-DFS price**; **68.8% are DFS-only** — shown, tagged, never a market |
| ✅ **p_win ranking / edge retired** | DONE — boards rank on `forecast_rank`, edge is diagnostic-only, rollback armed (`FORECAST_RANK=0`) | `specs/rank-on-pwin-challenger.md` |
| ✅ **calibration** | **DECIDED** — ruler-independent; MLB isotonic **DECIDED**, not provisional | `specs/mlb-recalibration-vs-consensus-ruler.md` |
| ✅ **grade cap** | DONE — 25 → 500; board **7 → 365+** graded props | `specs/grade-cap-and-refusal-diagnosis.md` |
| ✅ **S59 join invariant** | **ARMED** — root cause was `searchPlayer` returning `team: null` (`currentTeam` has no `name`); fail-safe: drops only on a positive not-in-game | `tests/unit/slateJoinInvariant.test.js` |
| ✅ **arch-v1 condition axes** | **BOTH NOW FIRING** — environment 84.6%, matchup 82.9% (`batter_own_split`). Was: env + matchup on 0/634 prod rows — `team` was null on 416/416, so the venue join had no key. **Environment FIXED** (joins on the game; resolves 105/120). **Matchup still dead** — needs opposing SP + both hands, three separate absences | `specs/arch-v1-axis-audit.md` |
| **opportunity_drift** | built, orthogonal (r≈0), live as a challenger on 142 rows/slate | verdict n-blocked; `scripts/opportunity-axis-holdout.sql` |
| **opportunity layer** | **NOT BUILT**`ab_per_game` is display-only; engine1 has no usage factor; MLB batting order unavailable in every wired source | Step 0 stopped before wiring. Recommended instead: an `opportunity_drift` axis on the EXISTING `challengerProjection` (arch-v1). See `specs/connect-opportunity-step0.md` |
| **MLB board size** | **7 → 365 graded props** (52×) in 114s | the 25-cap discarded 95.7% of the slate. Raised to 500 on measured cost. Refusals were **43.8% deliberate policy suppression**, not a data gap |
| **model input** | **byte-identical** — 546 gradeable props, `v1_first_book` | `MODEL_BOOKS` gate in `dedupeProps` + `indexOddsProps`. Lifts only on the re-run |
| **accrual clock** | **sequential, post-completion** | see §11. Pre-completion data does not count and is never pooled |
**The honest framing:** widening books is an **AGGREGATOR** win. **It does not fix
the model.** Do not let the free-side win read as model progress.
> **HOW TO USE:** this supersedes ad-hoc re-derivation. Before any order, read the
> phase you're in. After any order, tick the item and add one line. **Do not
> re-audit anything marked KNOWN** — that redundancy is what this document exists
> to kill.
---
## 0. VERIFICATION LEDGER (what was re-checked in this pass)
**NOTHING was re-verified. No query was run.** Everything required is already
captured in 22 artifacts produced this session plus the canonical board. Per the
order's own clause — *"If everything needed is already in the artifacts, say so and
skip verification"* — this is that case.
**Taken as KNOWN (source in brackets):**
- Model architecture, all layers [`model-architecture-recovery-map.md`]
- Grade↔outcome correlations, collapse cost [`full-output-grade-mapping.md`]
- p_win calibration + holdout verdicts [`grade-diagnostic-t0.md`, `pwin-recalibration-holdout.md`]
- Market-relative edge inversion [`grade-fix-part1-investigation.md`]
- Design implemented-vs-designed, 61 items [`design-vs-build-gap-audit.md`]
- Surface states, orphans, waves [`pre-audit-status-pull.md`, `incomplete-surface-triage.md`]
- Resolution pipeline true state [`resolution-and-clv-investigation.md`, `wave3-*.md`]
- Tier/monetization + founder mechanism [`tier-structure-pull.md`, `tier-redesign-spec.md`, `build2-review-zero-report.md`]
- Sport-boundary cost [`model-architecture-recovery-map.md` §Phase 2]
**Genuinely OPEN → carried as explicit unknowns (not verified because they need a
build or a decision, not a query):** sport order (Kev's call, §A2); board-reasoning
gating (a/b/c, §C2); CLV flag decision (§D3); the ~70 undefined team colours (§B).
---
## A. PER-SPORT MODELS — Phillips doctrine: each sport is its OWN model
### A1. MLB layer stack (the reference module), bottom-up
| # | layer | state | note |
|---|---|---|---|
| 1 | Data/feeds | **BUILT** | statsapi free+unlimited; statcast; park; weather; probables. Settles end-to-end. |
| 2 | Similarity (comparable instances) | **BUILT · NOT WIRED · NOT DEPLOYED** | `python/utils/similarity.py`; grade path skips to season/recent averages |
| 3 | Archetypes (batter + pitcher) | **BUILT · display only** | classifies + renders; **does NOT feed the grade** |
| 4 | Variable weights | **PARTIAL/ASSUMED** | engine1's flat ±1.0/±0.5 deltas are hand-set, never fitted |
| 5 | Conditions (park/weather/platoon/arsenal) | **BUILT · CHALLENGER ONLY** | rides `env_*`/`challenger_*`; "measured, never served" |
| 6 | Bayesian inference | **BUILT · NOT WIRED · NOT DEPLOYED** | `python/utils/bayesian.py`, 320 ln, genuinely sport-agnostic math |
| 7 | Calibration | **MEASURED, NOT APPLIED** | isotonic qualifies on holdout (rel .1038→.0939, res .139→.123, n=119) |
| 8 | Grade ladder | **BROKEN** | letter is a factor index, r≈0.005, **inverted** (B 52.4% < C 56.9%) |
**The through-line:** layers 2, 3, 5, 6 are built and *not connected*; layer 8 is
connected and *meaningless*. MLB's fix is connection, not construction.
### A2. Sport order (OPEN — Kev decides)
Proposed by readiness × clock: **1) MLB** (reference, only qualifying model) →
**2) CFB** (has a <30-day clock; soft-market thesis) → **3) NFL****4) NBA**
**5) CBB****6) WNBA — first real build** (never had its own model) → **7) soccer** (quota-blocked).
Each gets the same 8-layer template. **No sport is abandoned — abstention is a
state, not a verdict.**
---
## B. DESIGN IMPLEMENTATION — 61 items catalogued
BUILT-TO-SPEC 20 · DRIFTED 7 · PARTIAL 16 · ABSENT 18. Ordered wire-in:
**D1-A done** (6 combat glyphs, boundary-blue completed, reaction primitives, READ-FAB).
**D1-finish done** (rationale/reveal/chips modules — *built, NOT mounted*).
Remaining: **D1-close** (mount — blocked on a ROW-GRAMMAR slot amendment + §C2) ·
**D1-B** (45 unwired glyphs + the 41-vs-74 archetype scope call) · S2 primitive set
(movement strip, crown, disagreement axis, SPLIT) · S3 article media · The Report
email + archive · Offseason artboards · **team colours: only 10 of ~80 defined —
the rest render honest-neutral until a real source exists.**
---
## C. SURFACES
**C1 — done:** Wave 1 wiring · `/compare` · `/record` · Build-1 gate.
**C2 — OPEN DECISION (blocks D1-close):** board reasoning is served ungated while
`tiers.js` declares `reasoning_visible:false`. Options (a) gate it, (b) accept as
free funnel, (c) leave unrendered.
**C3 — remaining:** `/record` **has no nav link** (the surface that justifies the
price) · `/notifications` · Offseason hub · `/system` · S3 media · `/soccer`
(quota) · share cards (blocked by D).
---
## D. RESOLUTION PIPELINE → USER OUTPUT
**KNOWN and load-bearing: settlement WORKS** (scheduler → `settleAllOutcomes` +
`settleAllLedgers`, 937+ settled, growing daily). **What is unreachable is the
user-output TAIL:** `/api/grading/resolve` has no caller, and its fanout holds
webPush/Telegram/Discord but **no share-card step and no recap**.
**D1** wire a trigger (or move the fanout into the settle pass) · **D2** share-card
generation + `/notifications` consent + result posts + recap · **D3 CLV flag
decision** — `clvCaptureReliable()` is *one env var*, and the pre-registered rule
stands: flip only if close_moved is a clear majority AND coverage is representative.
**🔴 Never wire `/api/grading/resolve` as a second settlement path — it double-counts.**
---
## E. SPORT BOUNDARY
Adding a sport is a **~10-file core edit** with four silent-failure modes
(MARKET_MAP → zero props; three stat whitelists → silent 400s; missing projection →
universal refusal; no settled feed → grades forever). **Collapse to a registry** so
a sport is a module. **Blocks all of A2 after MLB.**
---
## F. CHROME AUDIT — 11 items, 4 need a Desk session. **Runs when surfaces are stable, not before.**
---
# THE PHASES — 7 phases, ~18 orders
| phase | orders | contents | blocks |
|---|---|---|---|
| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only — every other sport is NOT-BUILT, held out) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
| **2. Resolution tail** | 3 | trigger · share cards + notifications + posts + recap · CLV flag decision | share cards, social proof |
| **3. Surfaces + design lane** *(parallel with 1-2)* | 4 | C2 decision → D1-close mount · `/record` nav + remaining surfaces · D1-B glyphs/archetypes · S2 primitives | Chrome audit |
| **4. Sport boundary** | 2 | registry collapse · MLB re-expressed as the first module | all further sports |
| **5. Sport rollout** | 1 per sport | CFB → NFL → NBA → CBB → WNBA retry → soccer, each on the 8-layer template | — |
| **6. Monetization finish** | 2 | Stripe Phase-B live proof on first real signup · founder launch to the 3 existing users | — |
| **7. Chrome audit + hardening** | 2 | the 11-item visual sweep · credential rotation + migration-drift reconciliation | ship |
**Phases 14 and 67 = ~17 orders. Phase 5 = 1 order per sport (6 listed).**
> **RECONCILED 2026-08-01.** Phase 1 (MLB model truth) is the ACTIVE phase and is
> further along than the table above implies — but not because orders were
> ticked off as planned. Most of it turned out to be **connection and repair, not
> construction**: the ladder question dissolved (calibration is ruler-
> independent), the cap was discarding 95.7% of the slate, and two condition axes
> were wired but firing on zero rows.
>
> **Closed:** p_win ranking + edge retirement · calibration DECIDED · MLB isotonic
> DECIDED · grade cap 25→500 · book widening (display) · S59 invariant armed ·
> environment axis repaired.
>
> **Open in Phase 1:** matchup axis (next order) · the two accruing verdicts
> (`opportunity_drift`, `environment`) — **time, not code** · connect the still-
> dormant layers (similarity, Bayesian, distribution ladder).
>
> **Carried, not in Phase 1:** WNBA is **NOT BUILT**, not failed (its 0.12 was
> NBA-template machinery on WNBA data) · the consensus ruler is **MARKET, not
> SHARP**, and that is **PENDING-RECOVERY** until PropLine answers why Pinnacle
> MLB prop coverage stopped on 2026-07-31 (`BLOCKERS.md`) — do not enshrine it as
> permanent.
>
> **Remaining to the end state: ≈ 19 orders**, of which **~9 are unblocked today**.
---
# DEFINITION OF DONE
**VYNDR is complete when:**
1. **MLB layers 18 are BUILT AND CONNECTED** — similarity, archetypes, fitted
weights, conditions and Bayesian all feed the grade; calibration applied; the
ladder monotone (A>B>C, no inversion) and proven on held-out data.
2. **Every listed sport is finished on the same 8-layer template**, or explicitly
held out with its reason recorded — **"not built yet"** where no sport-specific
model exists, and only "measured and failed" where one was genuinely built and
tested. Never silently absent, and never a verdict on a sport we never attempted.
3. **Design fully implemented** — all 61 catalogued items BUILT-TO-SPEC.
4. **Every surface built, reachable and honest** — no orphans, no live-but-not-honest
surface, no dead component.
5. **Resolution pipeline live end-to-end** — settle → share card / notification /
post / recap, firing on a real settlement.
6. **Sport boundary is a registry** — a new sport is a module, not a core edit.
7. **Chrome audit passed**, logged-out and entitled.
8. **The record is publishable on its own terms** — CLV either trustworthy-and-
representative or honestly absent; no claim outruns its evidence.
**The remaining work is finite and countable: ~23 orders across 7 phases.**
---
## STANDING LAWS (carried into every order)
Truth Law — absent beats wrong, no fabrication up or down · per-sport models, never
a global engine · lookahead guard (lock-time fields only) · overfitting guard (fit
one split, prove another) · never mint A's without new information · aggregate proof
is free, itemized judgment is paid · one canonical founder flag · atomicity by unique
index, never a count · cache-bust every post-deploy check · verify-after-write.
---
# 9. WHAT'S ACTUALLY MISSING FOR THIS TO WORK AS A PRODUCT
*The phases above say what is UNBUILT. This says what is missing for VYNDR to
genuinely do what it claims. Some of it is not a build, and one of it is not
fixable by us at all. Written plainly because a plan that only counts code is the
comfortable version.*
## 9.1 🔴 THE CENTRAL ONE: there is no demonstrated edge yet
Every edge measurement this session came back **null, negative, or unproven**:
| measurement | result |
|---|---|
| served grade → outcome | **r ≈ 0.005**, and **inverted** (B 52.4% < C 56.9%) |
| p_win fair_prob (3 formulations) | **negative in all three, both sports, both splits** |
| p_win alone, MLB, holdout | +0.165, **p ≈ 0.07 — not significant** |
| p_win alone, WNBA | **negative — but this measured an NBA-template model on WNBA data, so it is not a WNBA result at all** |
| CLV / beat-close | **null by guard** — instrument not trustworthy |
| ROI by grade | likely an artifact of a meaningless letter |
**The product's core claim — "our read is better than the market" — is not
currently supported by our own data.** Everything else in this plan is
scaffolding around that. Building all 23 orders and *not* closing this leaves a
beautifully-built product that doesn't do the one thing it sells.
**What closes it:** not code. **Sample and honest iteration.** The instrument
fields are ~10 days old (442 rows). At ~90 decided MLB rows/week, that is ~610
weeks *of accrual*.
> **CORRECTED 2026-07-31 — see §11.** The "610 weeks out" above quietly assumed
> the clock is **already running**. It is not. Those 442 rows measure a model
> with a bent single-book ruler and six disconnected layers — **a model that
> will not exist once Phase 1 lands.** They do not count toward the verdict and
> **must not be pooled** with post-completion rows.
>
> The clock starts at a **verified** "running as intended" gate, per sport.
> **610 weeks is the accrual duration, not the distance to the answer.** The
> distance to the answer is *build time + verification + 610 weeks.*
>
> Independently forced by arithmetic, not just discipline: the §10.1 ruler fix
> changes the denominator, so pre-fix and post-fix edge/CLV numbers are not the
> same measurement.
## 9.2 The projection — the actual engine — is thin and unvalidated
The grade's only real inputs today are **l5/l20 averages, an opponent rank, rest
and usage**. Similarity, archetypes, park/weather/platoon and the Bayesian layer
are all built and **not connected**. So VYNDR is currently a recent-form average
wearing an intelligence system's clothes. Phase 1 connects them — **but connecting
them is a hypothesis, not a guarantee.** They must each prove out on held-out data
or be left disconnected honestly.
## 9.3 A one-sport product marketed as multi-sport
MLB is the only model that has been BUILT and passed. WNBA's model does not exist
yet (what was measured was NBA-template machinery on WNBA data). NBA and soccer
**don't even settle** — they grade into a void. Until Phase 5, the honest framing
is *"an MLB product with other sports in development."* The site should not imply
otherwise.
## 9.4 No customers, therefore no feedback loop
**3 users, 0 paid.** The founder mechanism is built and race-proven, `/record`
exists, the gate works — and **none of it has met a real user.** Nothing here is
validated by usage: not the price, not the tier line, not whether the locked-shell
tease converts, not whether anyone wants this. **The first 10 real users will
teach more than the next 10 build orders.**
## 9.5 No distribution — the biggest non-code gap
There is no acquisition path at all. The newsletter send is unscheduled, share
cards are unbuilt (blocked on the resolution tail), social proof has no fuel
(needs a real record), partner/affiliate links are all `enabled:false`. **A product
nobody sees cannot be validated regardless of how good the model gets.** This
appears in no phase above and belongs on the board as its own track.
## 9.6 The read isn't actionable at the last mile
Push-to-book is a **teaser** — no affiliate is live, so a user who trusts a read
still leaves to place it manually. Bankroll guidance (Kelly) is Desk-gated. The
gap between *"here's a good read"* and *"I placed it"* is unclosed.
## 9.7 Operational fragility
Single-box, single Redis (persistence is a Coolify setting, not app-controlled),
one cron. Settlement silently covers 2 sports. **Three credentials remain flagged
for rotation, including a Stripe live key that transited a chat transcript.** No
staging environment — every verification this session ran against prod.
---
## THE HONEST SUMMARY
**Built well:** the truth infrastructure. Honest empty states, refusal paths,
n-gates, the append-only ledger, the settled/live gate, the atomic founder cap.
**This codebase does not lie about what it knows** — that is rare and it is real.
**Not yet true:** that the model beats the market. Not disproven either — *unmeasured
at adequate n*, on one sport, with a projection whose best layers aren't connected.
**So the finish line is not 23 orders.** It is 23 orders **plus a verdict from
accrued data that we cannot rush** — and the discipline to report that verdict
honestly if it says the edge isn't there. The plan above builds the machine. Only
time and honest measurement decide whether the machine is right.
---
# 10. TO BE A REAL AGGREGATOR *AND* A MODEL PEOPLE PAY FOR
*Kev's question: what makes this the top product, not just a finished one. The
answer that matters most: **the aggregator gap and the model gap are the SAME gap
in two places.** Fix the data breadth and both halves improve at once.*
## 10.1 The one finding that reframes everything
> **CORRECTED 2026-07-31 by Order Zero — `specs/order-zero-book-breadth-test.md`.**
> The original text below blamed a missing `regions`/`bookmakers` param. **That
> was wrong.** PropLine's OpenAPI contract states verbatim: `bookmakers` …
> **"Omitted = all books."** Omitting it is correct and always was.
>
> **The real cause is ours.** PropLine sends **18 books**; `oddsNormalizer`
> `ALLOWED_BOOKS` (11 entries) intersects them at **exactly 5** — which is
> precisely the "5 MLB books" the audit measured. We discard 13 of 18 ourselves,
> and 6 of our 11 allow-list entries don't exist at PropLine at all.
>
> Measured on real public data (MLB `pitcher_strikeouts`, 5 complete events):
> the feed carries **4.41 books/prop**; after our filter, **1.50** — and **12 of
> 34 props become invisible entirely** (zero allowed books).
>
> Also corrected: "73% single-book" is the **long tail** of deep/reliever props
> sole-posted by DraftKings or Bovada. **On the core props we actually grade,
> the market is 1012 books wide.**
>
> And a real negative: **`pinnacle` appears on 0 of 40 MLB props.** The one
> sharp book in our allow-list contributes nothing here. The independent
> low-vig references that *are* present on 100% of core props are **exchanges**
> — `novig`, `smarkets`, `kalshi` (+ `matchbook`, `polymarket`). DFS pick'em
> (`prizepicks`/`underdog`/`sleeper`/`dabble`) also covers 100% but is **not a
> market price** and must never enter a consensus.
>
> **So this is not a test to run. It is a build we can do: split one allow-list
> into takeable / reference / excluded, and make `fair_prob_lock` a median
> consensus across reference books.** No param, no cost, no tier, no new source.
**Our "market" is often ONE book.** MLB props are **73% single-book** (2.18 audit).
And `proplineAdapter` sends only `{ apiKey, markets }` — **no `regions`, no
`bookmakers` param** (`:152`). We take PropLine's *default* response.
That single fact causes four separate problems we have been treating as unrelated:
1. **No line shopping** — the #1 free-tier hook in this category needs many books.
2. **`fair_prob_lock` is a de-vigged SINGLE SOFT BOOK**, not a consensus. That is
the ruler the model is judged against — a bent one. (Flagged as T1; never run.)
3. **CLV is weak** — you cannot measure "beat the close" against one book's close.
4. **No steam/disagreement detection** — needs ≥2 books to even exist.
*(Items 14 stand. Only the **cause** was wrong — and the fix is cheaper than
the original diagnosis implied.)*
### 10.1b We use 1 of PropLine's 29 endpoints
Reading the full spec surfaced ten unused endpoints that map directly onto §10.2
and §10.3 gaps — `/odds/closing` ("the canonical CLV helper"), `/odds/history`,
`/exports/odds-history`, `/movement` (steam across all 16 books), `/best-line`,
`/ev`, `/results` + `/exports/resolved-props` (**resolution across 33 sports**),
`/context` (**free**: probable pitchers, confirmed lineups, home-plate umpire,
first-pitch weather), `/markets/hit-rates`, `/players/{n}/trends`.
Two overturn standing beliefs: **"NBA/WNBA/soccer have no free settled feed"** may
be a **$19/mo** problem rather than a data problem; and we hand-built probable
pitchers / depth charts / lineup confirmation that `/context` serves free.
> **VERIFIED 2026-08-01 with the real key** — `specs/order-zero-consensus-ruler.md`.
> `/context` **WORKS, FREE** (umpire, roof, pitcher handedness, lineup
> confirmation — richer than what we hand-built). `/odds/closing` and
> `/movement` are **REDACTED** on our tier (full structure, **zero prices**) —
> my first pass wrongly called them "works" on a non-empty body. `/results` and
> `/exports/resolved-props` are **403**.
>
> **But the settlements exist to be bought:** `/markets/resolution-summary`
> shows **soccer graded across ~15 competitions in 30 days** (MLS 41k, Liga MX
> 15k, Brasileirão 12k, UCL/Europa/Conference…). **"Soccer grades into a void"
> is a $19/mo Pro-tier problem, not a data problem.** NBA is absent because it
> is July — seasonal, not a coverage gap, and not inferable either way.
>
> **WNBA is NOT thin at the feed** — 4.21 books/prop vs MLB's 3.61. It was
> allow-list-starved exactly as MLB was. This removes one candidate explanation
> for its 0.12 result. **And that result is not a WNBA verdict anyway** — it
> measured NBA-template machinery on WNBA data. WNBA is **NOT BUILT YET**, held
> out until it gets its own model.
>
> **No sharp anchor exists for props:** `pinnacle`, `matchbook` and `polymarket`
> all measured **0%** on both sports. The consensus ruler is therefore a MARKET
> consensus, not a SHARP one — stated as a permanent limitation, not a milestone.
## 10.2 What a real DATA AGGREGATOR has that we don't
| capability | ours | gap |
|---|---|---|
| **Book breadth** | 5 admitted of **18 sent**; 4.41→1.50 books/prop after our own filter | **not a feed gap — one allow-list (10.1).** Core props are already 1012 books wide |
| **True consensus / no-vig line** | single-book de-vig | **unblocked now**: median across reference books (exchanges + pinnacle), n≥2 or labelled fallback |
| **Historical odds archive** | **STARTED**`closing_captures` 844k rows, `lock_lines` (033) new, in-grade history capped at **24 points** | no full open→close series per prop. This is what makes CLV and backtesting real |
| **Market breadth** | 11 live markets | the category ships 50+ (alt lines, combos, innings, quarters) |
| **Alt-line ladders from books** | we *compute* a ladder; we don't *ingest* the books' | users shop rungs |
| **Injury / lineup wire** | partial (`depthChart`, confirmed-vs-projected) | no real-time news wire |
| **Player news** | `NewsWire` on Explore | not beat-level, not per-prop |
**None of this is model work. It's ingestion.** And it is the half competitors
compete on hardest, because it is visible to a free user in five seconds.
## 10.3 What a prediction model people PAY for has that we don't
1. **A distribution, not a point.** We project a point (l5/l20 average) and take
an empirical `P(over)`. Paid-tier models simulate a **full distribution per
stat** (negative-binomial / Poisson / MC). **We already have this**
`projection/distribution.js` computes real survival probabilities and a rung
ladder — but it is **proj-v1.1, ledger-only, and it lost to the champion.** The
asset exists; it is unconnected and unproven.
2. **Opportunity modelled FIRST.** In props, playing time is the dominant driver —
plate appearances, snaps, minutes, batting-order slot. We carry `ab_per_game`
and minutes as *features*, not as a **projected opportunity** with its own
uncertainty. This is the single biggest modelling upgrade available.
3. **Per-stat models.** Hits, strikeouts and total bases have different shapes.
One additive factor index across all of them is why the ladder is meaningless.
4. **Matchup granularity that actually reaches the grade.** Arsenal, handedness,
park, weather, platoon — **all built, all challenger-only, none feed the grade.**
5. **Calibrated probabilities with honest intervals.** Measured (isotonic
qualifies on MLB) — **not applied.**
6. **A backtest harness on real historical odds.** Blocked by 10.2's archive gap:
you cannot backtest a price you never stored.
7. **CLV as the north-star metric**, published honestly. Instrument built,
guard-blocked, and weak until book breadth lands.
## 10.4 The uncomfortable pattern
**Almost every model capability above is ALREADY BUILT and DISCONNECTED**:
similarity, Bayesian, archetypes, park/weather/platoon, the distribution ladder,
calibration. VYNDR does not have a *building* problem. It has a **connection and
proof** problem — plus one genuine ingestion gap (book breadth) that starves both
halves at once.
That is good news: the expensive part is largely done. But it also means **no new
feature fixes this.** Connecting the layers and proving them on held-out data is
the work.
## 10.5 If I had to order it for "top product"
1. ~~Book breadth **test**~~**Book breadth FIX + consensus fair line** (10.1).
The test is done (Order Zero). It is now a build: split `ALLOWED_BOOKS` into
takeable / reference / excluded, make `fair_prob_lock` a median consensus.
Upgrades aggregator, model denominator and CLV together. **Gated on** exchange
prop-price validation + the WNBA measurement (needs the PropLine key).
2. **Opportunity projection** (10.3.2) — the biggest genuine modelling gain.
3. **Per-stat distributions** — connect `distribution.js`, prove per stat.
4. **Connect the built layers** (Phase 1) — each proven on held-out or left off.
5. **Full odds archive** — store every book's open→close; unlocks backtesting.
6. **Market breadth** — 11 → 50+ markets is mostly ingestion + the 4-layer wiring.
7. **Then** the sports rollout, on a template that is actually worth replicating.
**The ordering principle:** do not replicate a thin model across six sports. Get
MLB genuinely good first — a copied-six-times thin model is six times the
maintenance for the same absent edge.
## 10.6 The honest caveat on "top product"
The category's leaders are judged on one number: **do their picks beat the closing
line, at scale, published.** We cannot claim that yet — not because the product is
unfinished, but because **we have not measured it at adequate n on a market we can
trust.** Book breadth + the odds archive + accrued settlements are what make that
claim *possible*. Everything in §10 is in service of being able to make it — or of
being able to say honestly that we can't.
---
# 11. THE ACCRUAL CLOCK — SEQUENTIAL, POST-COMPLETION
*Kev's correction, 2026-07-31. Supersedes any "accrual runs in parallel with
building" framing anywhere in this document.*
## 11.1 The causality
You cannot meaningfully accrue until the product is **right and running as
intended** — model layers connected, ruler fixed (real consensus, not one soft
book), sports in, operating in the vision.
**Only then does the accrual clock start, and only then does time produce a
verdict. Building faster shortens the time TO clock-start. It never runs the
clock.**
The flawed assumption being removed: that today's half-connected model accrues
useful evidence while we build. It does not. That data measures a model that
**will not exist** after the layers are connected — a different model wearing the
same name.
## 11.2 Measurement rule (non-negotiable)
- Settled data accrued **before** completion does **NOT** count toward the edge
verdict.
- **NO POOLING across the completion boundary.** Pre-fix and post-fix are
different models — the same class of error as the model-version boundary, at
whole-model scale.
- The clock starts at **verified** "running as intended", not at "today".
## 11.3 Two clocks — stated separately so neither corner-cuts
**1. MLB-MODEL VERDICT clock.** Starts when MLB is genuinely complete: built
layers connected + honest consensus ruler + running as designed. Its accrued n
judges **MLB**.
**2. FULL-PRODUCT TRACK RECORD clock.** Starts when the vision is running:
sports in, aggregator built, operating as intended. Its accrued n judges **the
product claim**.
**Each subsequent sport gets its OWN clock**, starting when *that* sport's model
is complete — never when it is stubbed in. (Per-sport doctrine: a sport that
merely renders is not a sport that measures.)
## 11.4 "Complete" — defined honestly
A model/sport is **complete-enough-to-accrue** when its **built layers are
connected** and it runs against an **honest ruler** (real consensus), **operating
as designed** — *not* when every conceivable feature exists.
This definition is doing real work in both directions: it blocks the corner-cut
("close enough, start counting") **and** it blocks never-ship ("one more
feature"). The full-product claim additionally requires the vision's sports +
aggregator running.
## 11.5 The verification gate
**Between build and clock.** Before any accrual counts, verify — not assume —
that the product is running as intended:
- connected layers actually **fire** (present in the served payload, not merely
present in the repo)
- the ruler is a **real consensus** (`fair_prob_source: 'consensus_n'`, n≥2)
- the sport **settles correctly** (spot-checked against real box scores)
- surfaces are **honest** (no fabricated values; absent renders absent)
**This gate is verified, not assumed.** No accrual line item runs during the
build phases.
## 11.6 The 10-user track, reframed
Real users are onboarded to a **complete** product, so their usage teaches about
the real thing rather than a half-built one.
**We do NOT acquire users early to "start accrual."** That is the corner being
explicitly refused.
## 11.7 The finish line
```
build orders complete
running-as-intended VERIFIED (§11.5 gate)
clock starts (per sport, two clocks, §11.3)
verdict reported honestly — including if it says the edge is not there
```
That last clause is the whole point. A clock you are willing to stop early is not
a measurement, and a verdict you are only willing to publish if it is favourable
is not a verdict.
+92
View File
@@ -0,0 +1,92 @@
# What "media hub" already means in VYNDR's design language
Read-only inventory. Nothing designed, built or mounted.
## PHASE 0 — the design source of truth exists, and so does a gap audit
| artifact | what it is |
|---|---|
| `specs/DESIGN-SPEC.md` | **DESIGN SPEC v2** — governing laws, colour contract, entity layer, data-display standard, motion, archetype system, conversion architecture, empty/error system. 84 lines, supersedes v1. |
| `specs/design-reference/` (Jul 22) | **The design package itself** — 7 `.dc.html` surface files, 83 archetype glyph SVGs + MANIFEST, 7 rasterised share-card PNG masters. |
| `specs/design-reference/HANDOFF.md` | The code handoff: exact tokens, brand, glyph library, behaviour, laws. |
| **`specs/design-vs-build-gap-audit.md`** | **A 61-item design-vs-build audit already exists** (2026-07-31, audited at `bf7c0a3`), every claim grep-verified, with a wave-ordered build plan. |
**The inventory this order asks for was largely already done on 2026-07-31.** The
useful work is reconciling it, not redoing it.
## PHASE 1 — per media surface: spec, and build state
| surface | design spec? | built state | evidence |
|---|---|---|---|
| **Offseason hub** | **YES — a full surface file** | **NOT-STARTED** | `Vyndr Offseason.dc.html` (119KB) defines hub home (NFL desktop + 390), an **NBA Summer-League variant with an `OUTLOOK ONLY / NOT GRADED` honesty block**, a season-long board (open→NOW→VYNDR triplet), a season-read reveal with `WHAT WOULD CHANGE THIS READ`, a news/outlook feed with row anatomy, and a **quiet-wire empty state**. Gap audit: F9/F10/F11, *"design complete, no blocker but big."* |
| **Article media** | **YES — surface S3** | **NOT-STARTED** | Hero template + **4 hero graphic archetypes** (line path / distribution / mark-at-scale / matchup card — *all generated data visuals, never stock*), inline figures (stat-callout triptych, comparison bars, pull quote — **one max**), **caption law: every figure names its data**, article card, OG 1200×630 (master already rasterised in `exports/`). Gap audit F5. Zero components built. |
| **THE WIRE** | **YES** | **BUILT** (E19) | Spec'd in `Vyndr System.dc.html`: timestamped entry, ~6s hold, tag coloured by meaning. Built as `vyndr/Ticker` on real `/api/ticker` exhaust. **`NewsWire.tsx` is a different thing** — the offseason news/outlook feed, mounted in ExploreHub. |
| **Movement strip** | **YES — E1, a named primitive** | **ABSENT** | *"steps not curves, green only when the move favours the read, FLAT = hairline + `FLAT · [N]D`"*, row 86×20 silent / full-width annotated on reveal. **This corrects my 2026-08-07 board, which listed it UNKNOWN / possibly-satisfied-by-GradeShift.** It is a specified primitive that does not exist. |
| **Share-card masters ×5** | **YES — E16** | **ABSENT as product** | settle 1080×1350, story 1080×1920, square 1080, X 1200×675, record 1080×1350, article OG 1200×630. PNG masters in `exports/`; `ShareCard.tsx` has **0 importers**. Gap audit blocks it on the resolution tail having no share-card generation step. |
| **Calibration curve** | YES — E9 | ABSENT | dots vs dashed perfect line, dot size = sample, buckets under N30 hollow. |
| **The Report email / `/report` archive** | YES — E10/E12 | PARTIAL / ABSENT | `newsletterService` builds an email; the designed dark-billboard-over-light-paper hybrid is not built. `/report` redirects to `/blog`. |
| **ExploreHub** | **NO SPEC** | BUILT | Not a designed surface. It is where `NewsWire` currently lives — a de-facto host, not the intended hub. |
## PHASE 2 — what an on-site media hub would consist of
**From surfaces already designed** (no new design needed):
1. **Offseason hub shell** — F9/F10/F11, fully spec'd including its honesty block and quiet-wire empty state.
2. **News/outlook feed** — spec'd row anatomy; `NewsWire.tsx` is a partial implementation of it.
3. **Article media** — S3, fully spec'd hero + figure system.
4. **THE WIRE** — built, reusable.
5. **Share/OG cards** — designed at exact sizes, masters exported.
### The gap that needs design work authored
**Only one thing is genuinely un-designed: the hub's INFORMATION ARCHITECTURE
across seasons.** The Offseason file specifies an *offseason* hub; there is no
spec for what the same surface is **in-season**, or how content, articles, wire
and the live slate share one navigation. Every component exists on paper; **their
composition into a year-round media surface does not.**
### One finding that touches work I just shipped
**I built a card renderer at 1080×1350 without checking whether a designed card
system existed. It does** — five master sizes with defined layouts, and
rasterised PNGs in `exports/`. The content engine's cards are a parallel
invention. They are not wrong, but they should conform to E16/F8 rather than
diverge, and that reconciliation is a design decision, not a code one.
## PHASE 3 — the map
| surface | spec? | built | part of hub? | needs design work? |
|---|---|---|---|---|
| Offseason hub shell | ✅ | ❌ | **core** | no — build it |
| News/outlook feed | ✅ | partial | **core** | no |
| Article media (S3) | ✅ | ❌ | **core** | no — build it |
| THE WIRE | ✅ | ✅ | yes | no |
| Movement strip (E1) | ✅ | ❌ | adjacent | no |
| Share/OG cards (E16) | ✅ | ❌ | yes | **reconcile with content-engine cards** |
| Content-engine posts | ❌ | ✅ | **yes** | **yes — no spec exists** |
| ExploreHub | ❌ | ✅ | host today | **yes — is it the hub or a placeholder?** |
| **Year-round hub IA** | ❌ | ❌ | **the spine** | **YES — the real design gap** |
### Coordination with the social-strategy chat
These should be decided **there first**, then reflected on-site:
- **The content formula/calendar** determines what the hub's recurring slots
*are*. Designing slots before the formula exists risks a surface shaped around
guesses.
- **Card format reconciliation** — off-site posts and on-site share cards should
be one system. E16's five masters vs the content engine's renderer is the same
decision on both sides.
- **Voice**: `specs/VOICE.md` governs; the content engine templates encode a
sharper register. Whether the hub speaks in that register is a brand call.
### Correction to the still-buildable list
The order lists content API, preview page, book-comparison mount and guard
widening as pending. **All four shipped at `c575a70`.** Book comparison was never
unmounted — my earlier board grepped only `web/src/app` and missed
component-level mounting.
**Genuinely buildable now, design-complete, no clock impact:** Offseason hub
(F9F11), article media (F5), movement strip (E1), calibration curve (E9), the
Report email + archive (E10/E12). The gap audit's own wave ordering already
sequences these.
+2861 -1
View File
File diff suppressed because it is too large Load Diff
+120
View File
@@ -0,0 +1,120 @@
# The accrual watch — the program is idle on modeling, and that is correct
## PHASE 0 — live-surface integrity
| check | result |
|---|---|
| Grade bands not derived from the retired champion | **PASS**`gradeBands` is required by *no* serving code. Built across several orders, never wired. No stale band can reach a user because none reaches a user at all. |
| No withdrawn-map leak | **PASS**`CALIBRATION_DEPLOYED` is `[]` and the calibrate loop iterates it, so `calibrate()` is never called. The only `p_win_calibrated` assignment sits inside that empty loop. Served `p_win` is repaired-champion raw. |
| Refusals on the repaired reference | **PASS**`projectionFor` reads `l20_avg`, which `mlbGameLogFeatures` now builds from `fullLog`. |
| Factors still fire post-repair | **PASS** — the repair moved the base they adjust, so sign was re-verified across it: at base 0.35/0.50/0.65/0.80, defence lowers, pitcher-contact raises, platoon raises at every point. Firing check only; **not** a lift re-measurement. |
### One defect found, and it was mine
The grade card rendered **"Last 20 games average: X"** from `l20_avg` — a field
that, after the repair, holds a **full-season** average. The number moved and the
label did not, so the surface asserted a window that no longer existed. Fixed to
"Season average"; `trapDetection`'s L20 explanations likewise.
Same class as everything else tonight, one layer out: **a correct-looking string
describing data that moved underneath it.**
---
## PHASE 1 — the existence risk, resolved
Nineteen commits (`7b85934``2391574`) were local-only with no push
credentials. Three retrievable artifacts now exist:
| artifact | path | size |
|---|---|---|
| session bundle | `~/vyndr-session-2026-08-07.bundle` | 244K |
| **full-history bundle (self-contained)** | `~/vyndr-full-history-2026-08-07.bundle` | 6.8M |
| patch series (20 files) | `~/vyndr-session-patches/` | 1.5M |
**Use the full-history bundle** — the session bundle verifies as requiring ref
`6452926…`, so it only applies onto a repo that already has this history. The
full-history one clones standalone:
```
git clone ~/vyndr-full-history-2026-08-07.bundle vyndr-recovered
```
**This is the top non-accrual action item.** Copy one of these off the machine.
---
## PHASE 2 — the accrual watch
```
repaired champion marker: engine1@2026-08-07-fullwindow
eligible rows: 0 | eligible dates: 0
blocked: no settled rows yet carry the repaired champion marker
```
| item | need | have | status |
|---|---|---|---|
| calibration re-fit | 10 dates | 0 | WAITING |
| hits factor lift | 10 dates | 0 | WAITING |
| prior verdict re-audit | 14 dates | 0 | WAITING |
| rbi lineup-slot gate | 14 dates | 0 | WAITING |
Zero is the correct starting line: the repair ships in *this* session's commits,
so no settled row can carry the marker yet.
### These are ATTEMPT floors, not TRUST floors
**Reaching 10 dates means "the calibration re-fit can now be measured." It does
NOT mean the re-fit is trustworthy.** We lived this distinction the hard way
tonight: date-block CIs on 24 clusters, a LODO gate with 1.49.3% power, and a
≥40 date-cluster bar that was correct as a promotion bar and wrong as a deploy
bar. A 10-date map is thin. It deploys **PROVISIONAL with auto-demotion**, like
everything else, and its interval will still be wide.
**No future session may read "threshold met" as "answer certified."**
### Real-time estimates
Roughly one MLB slate per day, but per-stat settled volume differs — hits props
are far more numerous than rbi, and a "date" only counts once its props settle.
| item | dates | realistic wall-clock |
|---|---|---|
| calibration re-fit, hits lift | 10 | **~2 weeks** |
| verdict re-audit, rbi gate | 14 | **3+ weeks** (rbi accrues slowest) |
The wait is **designed, not a stall.** The alternative — measuring on
reconstructions of a retired forecast — is the trap this program has now refused
by name three times.
---
## PHASE 3 — the pre-registered resumption order
Triggered purely by eligible-date thresholds. No re-litigation, no re-deciding:
1. **@10 dates — calibration re-fit on the repaired champion.** Low-parameter,
LODO where powered, PROVISIONAL, auto-demotion armed, shadow duel restarts on
the repaired forecast. The favourite-longshot bias must be **re-measured**,
not assumed to have survived the repair.
2. **@10 dates — hits factor composed-lift re-measure.** The 1.39% figure is
**void** — measured on the broken baseline. Direction unknown.
3. **@14 dates — prior factor verdict re-audit.** Every null and every THEATER
was scored against a champion worse than a frequency table. Not pre-priced;
some may pass, some may still fail.
4. **@14 dates — rbi lineup-slot / RISP through the two-part gate.** World A
~90%, within-role residual 0.01908 real, `lineup_context` prod-verified (S89).
**FIRST TRIGGER:** when `reAuditEligibility.assess()` reports **10 eligible
calibration dates**, the next order is the calibration re-fit.
**Until then the program is honestly IDLE on modeling.** That is the correct
state, not a gap to fill.
---
## Invariants
No measurement on reconstructions — held. Phase 0 verified live integrity and
factor *firing*; it did not re-measure lift or re-fit calibration. Serving change
for the label fix, fingerprinted. `p_win` never mutated. No Bonferroni slot.
+150
View File
@@ -0,0 +1,150 @@
# arch-v1 AXIS AUDIT — env/matchup WERE DEAD; ENVIRONMENT FIXED
**Date:** 2026-08-01 · champion `p_win`, ranking, calibration and
`opportunity_drift`'s accruing verdict all **untouched**.
**Gates:** 4,077 tests / 326 suites green · `next build` exit 0 · audit run on
prod · resolver verified live.
---
## STEP 0 — THE AUDIT: "already partly live" was half true
**The code is wired. The axes were not firing.**
Across **634 graded prod rows**, `challenger_adjustments` axis distribution:
| axis | rows | % of slate |
|---|---:|---:|
| *(no axis fired)* | 296 | 46.7% |
| **opportunity** *(new, ours)* | 142 | 22.4% |
| power | 80 | 12.6% |
| swing_miss | 69 | 10.9% |
| contact | 56 | 8.8% |
| launch · line_drive · patience · strikeout · contact_allowed · aggression · chase · velocity · ground_ball · wild | 51→1 | — |
| **environment (park/weather)** | **0** | **0%** |
| **matchup (platoon)** | **0** | **0%** |
Every axis that fired is an **archetype trait**. The two condition axes this
order is about fired on **nothing**.
The ledger confirms it from the other side: `env_multiplier`, `env_park_base`,
`env_weather_mod`, `wx_forecast`, `env_weather_state`**all null on 634/634**.
### Root cause, located rather than inferred
`buildContext` works fine — **15 games, 14 with weather, Coors Field composing to
1.241**. The failure is downstream. A drop-off audit against the live snapshot:
| hop | count |
|---|---:|
| grades examined | 120 |
| **with a `team` field** | **0** |
| team resolves to an abbr | 0 |
| abbr matches a game | 0 |
| **environment produced** | **0** |
| with `playerId` | 120 |
| with `bats` | **0** |
| opposing pitcher known | **0** |
**`team` is a key on every stored grade and NULL on 416/416.** The resolver keyed
off the player's roster team, so it could never find a venue — while
`buildContext` sat there with all 30 teams mapped and 14 weather forecasts
resolved and unused. **Nothing was broken in the park or weather code. The join
key was absent.**
*(Side consequence worth noting: the S59 slate JOIN INVARIANT also keys off this
same `team` field, so it is currently inert too.)*
---
## STEP 1 — THE ENVIRONMENT FIX (and it is the more correct join)
**The park and the weather belong to the GAME, not to the player's roster team**
and the game rides on the prop from the odds feed, where a stats-resolve can
legitimately fail.
1. **`gradeSlateService.gradeBestSide`** now carries `home_team`/`away_team` onto
the graded row. The legacy grade shape dropped them.
2. **`environmentContext.contextFor`** joins on the game first
(`team → home_team → away_team`), keeping the roster team as a fallback.
### Verified live
| | before | after |
|---|---:|---:|
| grades carrying `home_team` | 0 / 416 | **416 / 416** |
| **environment resolves** | **0 / 120** | **105 / 120 (87.5%)** |
Coors Field composes to **1.241** (park 1.20 × weather 1.034) the moment it gets
a key.
### One honest limit on what is confirmed
**The axis has not yet been observed writing to the ledger, and I am not claiming
it has.** `recordPipelineGrades` upserts with `ignoreDuplicates: true` and dedupes
on `(user_id, player_key, stat, line, side, game_id)` — correctly, so a re-run
never overwrites the original lock. Today's 429 rows were written **before** the
fix, so the environment axis **cannot** backfill onto them.
```
game_date 2026-08-01 — 429 rows: opportunity axis 142, environment axis 0
```
**First ledger observation: tomorrow's slate.** What *is* directly verified is the
resolver (105/120) and the join key (416/416), which are the two things that were
broken.
---
## MATCHUP / PLATOON — NOT FIXED, AND NOT CLAIMED AS FIXED
It needs the opposing starter and **both** hands. The audit shows **three separate
absences**:
```
oppPitcherByTeam 0 (the probable-pitcher fetch returns nothing)
handById 0 (so no pitcher handedness)
with_bats 0/120 (no hitter handedness on the slate)
```
Per the order's own guard — *one axis at a time* — that is its **own order with
its own diagnosis**, not a second fix smuggled into this one. Fixing the pitcher
feed without the hands, or the hands without the feed, would still produce an
axis that fires on zero rows.
---
## STEP 2 — HOLDOUTS: BOTH AXES ARE n-BLOCKED
| axis | settled rows carrying it | earliest proof |
|---|---:|---|
| `opportunity` | **0** | first settle pass on 2026-08-01 games |
| `environment` | **0** | first ledger write on tomorrow's slate |
Neither can be proven today, and **a holdout filtered to axis-carrying rows cannot
be run when that set is empty**. The committed query
(`scripts/opportunity-axis-holdout.sql`) already filters that way, for the reason
drift established: including untouched rows dilutes the comparison with rows where
challenger ≡ champion **by construction, biasing toward a false positive.**
**Both accrue in parallel on the same harness. Neither promotes until its own
holdout says so.**
---
## WHAT IS UNTOUCHED
Champion `p_win`, the live grade path, ranking, calibration, and
`opportunity_drift`'s accruing verdict. The environment axis writes only to
`p_win_challenger` / `challenger_adjustments`.
## TAGS
**VERIFIED:** env + matchup axes fired on 0/634 prod rows · 13 archetype axes fire
normally · `team` null on 416/416 · `buildContext` healthy (15 games, 14 weather) ·
post-fix environment resolves 105/120 and `home_team` present 416/416 · matchup
blocked by three separate absences.
**NOT YET OBSERVED:** the environment axis writing to the ledger — blocked by
same-day dedupe, first visible on tomorrow's slate.
+283
View File
@@ -0,0 +1,283 @@
# GATE SIMULATION + CALIBRATION — Model Train G-b / C-cal
**REPORT-FIRST deliverable. No engine code was written for this. G-a is HELD pending Kev's ruling.**
Dataset snapshot: `2026-07-19 22:01:58 UTC`, live Supabase `ledger_entries`, `user_id IS NULL`
(the public model record). **576 rows · 6 game days (2026-07-11 → 07-19) · 470 settled ·
571 with locked odds.** Numbers move: the pipeline inserted 8 rows mid-census, so every
table here is pinned to that timestamp.
---
## 0. THE BLOCKER: "last 30 days of stored snapshots" DOES NOT EXIST
There is no 30-day snapshot store to replay. Verified on disk:
- `snapshot:{sport}:latest` and `:previous`**two generations, 24h TTL**
(`snapshotService.js:404,412`). No `snapshot:{sport}:{date}` archive exists anywhere.
- The intraday line `history:[{t,line}]` rides **inside** the same 24h blob
(`intradayRefreshService.js:68-84,144`) — it dies with it.
- `outcomes:{sport}:log` (30d TTL, cap 1000) carries **no odds, no confidence, no
projection** (`outcomeService.js:233-237`) — unusable for an odds-aware table.
- **No backtest/replay harness exists** — grep for `backtest|replay` across `src/`,
`scripts/`, `tests/` returns zero. `migrations/006` defines `grade_outcomes` +
`player_calibrated_weights`; **no code reads or writes either table.**
**So the replay ran against `ledger_entries`, the only permanent store.** Two consequences:
1. **6 days, not 30.** Not a choice — that is the entire history.
2. **The denominator is the GRADED board, not the raw slate.** `ledgerService.js:173-181`
skips no-grade / `insufficient_data` / no-captured-line props. Reads the current gate
already refused were never written. So "cut %" below means *cut from what we publish
today*, which is the right question for G-b, but it cannot tell us about props that
never got that far.
---
## 1. G-b — WHAT THE NEW GATE DOES TO THE BOARD
### 1.1 Historical board by odds band
| Band | Props | % board | Settled | Hit % | **Units** | **ROI** |
|---|---|---|---|---|---|---|
| takeable 160…+200 | 257 | 44.6 % | 209 | 52.6 % | 0.72 | **0.3 %** |
| flex 161…−250 | 77 | 13.4 % | 70 | 67.1 % | +1.56 | **+2.2 %** |
| juiced 251…−400 | 4 | 0.7 % | 4 | 75.0 % | 0.12 | 3.0 % |
| **past 400** | **221** | **38.4 %** | 173 | 80.3 % | **13.29** | **7.7 %** |
| dog +201…+400 | 10 | 1.7 % | 8 | 37.5 % | +4.42 | +55.3 % |
| longshot > +400 | 2 | 0.3 % | 1 | 0.0 % | 1.00 | 100 % |
| no odds | 5 | 0.9 % | 5 | — | — | — |
Flat 1u staking; ROI = units / settled. n≥20 holds only for **takeable (209)** and
**flex (70)** — every other row is anecdote and is reported for shape, not for truth.
### 1.2 The three findings that matter
**FINDING 1 — the 400 floor you already shipped was the whole win.**
Past 400 hit **80.3 %** but needed **86.9 %** to break even: **13.29 units, 7.7 % ROI**
on 173 settled bets. That single band is essentially the entire historical loss. The
guard shipped this morning (`f72f063`/`348a82b`) already kills it. The board proves it:
MLB Jul 18 = 103 graded props, 71 past 400; MLB Jul 19 (post-guard) = 15 props, **0**
past 400.
**FINDING 2 — Arc 2's incremental bite is small, and it lands almost entirely on the
flex band.** Against the *already-live* gate, the new knobs cut: `251…−400` **4 props
total across 6 days**, no-odds **5**, `> +400` **2**. That is 11 props in 6 days. The
only material change is `EDGE_FLEX_WALL` gating **77 props (13.4 %)** behind a 2× EV test.
**FINDING 3 — the flex band is the most profitable band we have, and the takeable band
is flat.** Flex 161…−250: **+2.2 % ROI** (67.1 % actual vs 65.8 % break-even, n=70).
Takeable 160…+200: **0.3 % ROI** (52.6 % vs 53.1 % break-even, n=209). The band
doctrine wants to promote is break-even; the band Arc 2 proposes to restrict is the one
carrying positive ROI. Neither result is statistically strong at these n's, but the
direction is the opposite of the assumption behind `EDGE_FLEX_WALL`.
### 1.3 Survival, per night (the "is it livable" question)
| Date | Sport | Graded board | Survives outright | Flex (needs EV) | Cut: wall | no-odds | longshot |
|---|---|---|---|---|---|---|---|
| 07-11 | mlb | 74 | 21 | 15 | 34 | 4 | 0 |
| 07-12 | mlb | 54 | 8 | 6 | 40 | 0 | 0 |
| 07-16 | mlb | 48 | 14 | 15 | 18 | 0 | 1 |
| 07-16 | wnba | 32 | 26 | 6 | 0 | 0 | 0 |
| 07-17 | mlb | 86 | 14 | 10 | 61 | 1 | 0 |
| 07-17 | wnba | 70 | 66 | 4 | 0 | 0 | 0 |
| 07-18 | mlb | 103 | 21 | 9 | 72 | 0 | 1 |
| 07-18 | wnba | 62 | 54 | 8 | 0 | 0 | 0 |
| 07-19 | mlb | 15 | 13 | 2 | 0 | 0 | 0 |
| 07-19 | wnba | 32 | 30 | 2 | 0 | 0 | 0 |
**Verdict: the board stays livable.** WNBA is already almost entirely takeable (54/62,
66/70, 30/32) and is barely touched. MLB thins to **~15-25 survivors a night** on the
old data — but nearly all of that thinning is the 400 mass *already removed today*.
A typical post-Arc-2 night looks like **~15-25 MLB + ~30-60 WNBA = 45-85 reads**. That
is not a starved board.
Caveat: 07-19 is a partial day (census taken 22:01 UTC, before the 01/03 UTC slots).
### 1.4 What could NOT be simulated — and it's the important half
`ev_pct` is **not stored on any ledger row** (no column; Arc 1 computes it in-request and
it rides only the live payload). Neither is `p_win`. **Therefore the EV half of the
proposed gate — `EDGE_FLEX_WALL` + `EV_FLEX_THRESHOLD` — cannot be replayed against
history at all.** Everything in §1.3 is the *price-only* portion of the gate. The flex
column says "needs EV", not "survives".
To ever answer this we must start persisting it — see §4 (C-led).
---
## 2. C-cal — CALIBRATION
### 2.1 Confidence is monotonic but badly miscalibrated
| Bucket | n | Claimed conf | **Actual hit %** | Units | ROI |
|---|---|---|---|---|---|
| conf < 45 | 81 | 34.6 | 59.3 % | 6.47 | 8.0 % |
| conf 45-54 | 214 | 47.3 | 64.0 % | 6.61 | 3.1 % |
| conf 55-64 | 170 | 57.5 | 68.8 % | +3.94 | **+2.3 %** |
**Ordering is real** — hit rate rises monotonically with confidence (59.3 → 64.0 → 68.8),
and so does ROI. **Absolute calibration is not** — confidence understates the hit rate by
~20-25 points at every level. This is the audit's known grade↔confidence mismatch
(`mlb-grade-degradation.md`), now quantified against settled outcomes.
**Consequence for the engine:** `confidence` must NEVER be fed into an EV or probability
calculation. Arc 1 is already correct here — `evPct` uses `p_win` from the quantile
estimator, not confidence. Do not "fix" confidence by rescaling it into a probability;
it is a display ordering, and the two must stay separate.
### 2.2 The grade distribution is degenerate
| Grade | Rows | Takeable | Avg conf | Settled hit % | Units | ROI |
|---|---|---|---|---|---|---|
| B | 384 | 149 | 53.0 | 67.5 % | 7.82 | 2.5 % |
| C | 184 | 102 | 43.9 | 59.9 % | 1.33 | 0.8 % |
**The entire public ledger contains two letters: B and C. Zero A, zero A+, zero D/F.**
Confidence spans only ~34-64. The model is not using its own scale.
This directly breaks hero v2: `pickHeroProp` filters `isAB(g.grade)`, so the hero can
only ever be a B. It also means "A-RATED" copy on public surfaces describes a grade the
engine has never emitted in the recorded era.
**This is the single biggest finding in the report** and it is upstream of the whole value
engine — a gate that ranks EV within an undifferentiated B pool is tuning the wrong knob.
### 2.3 EV_FLEX_THRESHOLD — proposal with reasoning
Kev asked me to confirm against calibration if available, else propose. Calibration exists
but **cannot validate an EV threshold** (§1.4: no stored EV). So this is a proposal, and I
want it labelled as such rather than dressed up as data-driven.
**Proposal: keep `EV_FLEX_THRESHOLD` = 2× `VALUE_EV_THRESHOLD` (= 4 %) as the default —
but ship it OFF-by-default in the flex band until EV is persisted and re-checked.**
Reasoning:
- The flex band is the only clearly-positive band we have (+2.2 %, n=70). Restricting it
on an unvalidated threshold risks cutting the profitable third of the board to enforce
a rule we cannot yet measure.
- A 4 % EV bar is defensible on first principles: at 200 you need 66.7 % to break even, so
4 % EV ≈ 69 % model probability — a real, not rounding, edge. I'm comfortable with the
*number*; I'm not comfortable *enforcing* it blind.
- Concretely: implement the knob, wire the band, default `EV_FLEX_ENFORCE=0` so flex props
grade as they do today while `ev_pct` gets recorded. Flip to `1` after ~2 weeks of stored
EV shows what the threshold actually cuts. Zero board impact now, full data later.
### 2.4 Recommended dial changes vs your spec
| Knob | Your spec | My recommendation | Why |
|---|---|---|---|
| `TAKEABLE_ODDS_CEILING` | 160 | **160, unchanged** | promotion band, works |
| `HARD_JUICE_WALL` | 250 | **250, ship it** | costs 4 props/6 days; band is 3 % ROI; the 400→−250 tightening is nearly free |
| `EDGE_FLEX_WALL` | 250 | **250 band, enforcement OFF by default** | §2.3 — the band is our best performer, threshold unvalidated |
| `EV_FLEX_THRESHOLD` | 2× (=4 %) | **4 %, dormant until measured** | §2.3 |
| `LADDER_ODDS_MAX` | +400 | **+400, ship it** | costs 2 props/6 days; the one settled longshot lost |
| `MIN_RUNG_PROBABILITY` | 0.25 | **ship the knob, but it binds on nothing today** | §3 — no per-rung odds exist, so there are no rungs to price |
| no-odds → refuse | yes | **ship it** | 5 props; and all 5 had `model_value = 0` — pure garbage |
| `JUICE_ODDS_FLOOR` | keep as backstop | **keep at 400** | dead code once the 250 wall lands, but harmless and legacy-safe |
---
## 3. L-a — DO `alt_lines` CARRY PER-RUNG ODDS? **NO.**
- A rung is exactly `{ line, grade, edge_pct, base }` (`analyzeViaEngine1.js:492`). No
price, no probability.
- The rungs are **synthetic** — fixed `[1,0.5,0,+0.5,+1]` shifts off the main line,
re-graded on the same feature vector. They are not market lines.
- **The feed offers nothing better today:** grep for `alternate|alt_line` across
`proplineAdapter.js`, `oddsNormalizer.js`, `oddsService.js` returns **zero hits**. No
alternate-line market is requested, normalized, or stored.
**L-b (per-rung EV) is BLOCKED on a data source, not on engine work.** Per-rung EV needs a
real price per rung. Options, in cost order: (a) check whether PropLine exposes alternate
markets on the existing 3 free keys — costs nothing but a probe; (b) odds-api alternate
markets — blocked, quota is 0/500 and Kev's ruling is hold the line; (c) derive rung prices
from a distribution around the main line — **rejected, that fabricates market data and
violates the Data Semantics Rule.**
Recommend: probe PropLine (free) before any L-b work is scheduled.
---
## 4. C-led — HOW MUCH HISTORY HAS ODDS ATTACHED?
**Good news: 571 of 576 rows (99.1 %) already carry `locked_odds`, and 465+ settled rows
have both odds and an outcome.** `locked_odds` was in migration 019 from the start
(`019:17-43`) and `rowsFromSnapshot` has always populated it (`ledgerService.js:186-205`).
**So C-led needs no backfill for odds — the ledger can speak units TODAY.** That is why
§1.1's ROI table exists at all.
What is genuinely missing and needs new columns (not backfill — these were never computed
historically): `ev_pct`, `p_win`, `fair_odds`, `takeable`, `value`. Honest handling: add
the columns, populate going forward, and label the start date. **"Tracking began
2026-07-XX" is the honest answer — do not backfill EV by recomputing it from today's
model against yesterday's lines.** That would be a fabricated record of a model that
didn't exist yet.
`getModelAggregate` currently reads only `outcome, clv_result, clv, player_key, grade,
model_value` (`ledgerService.js:454`) — it never reads `locked_odds`, so units/ROI/
record-by-odds-band are all new aggregate work on data that already exists.
---
## 5. C-clv / S-a — CLV IS CONFIRMED BROKEN IN THE DATA
The C4 write-up is right, and the ledger proves it:
| Sport | rows | closing_odds present | **closing_line == locked line** | clv = 0 | clv ≠ 0 |
|---|---|---|---|---|---|
| mlb | 380 | 376 | **359** | 285 | 21 |
| wnba | 196 | 196 | **164** | 138 | 26 |
`captureClosing` re-records the lock, so CLV is structurally ~0. `clvCaptureReliable()`
correctly suppresses `beat_close_pct` and `clv_distribution` on every public surface
(`ledgerService.js:45,532-547`). **Keep it suppressed.** The fix shares plumbing with S-a
(line-moved truth) exactly as you scoped — both need a real "fresher odds at serve time"
read, which the 24h `snapshot:latest` + 20-min intraday refresh can supply without new
quota.
---
## 6. U-deg — STATUS: THE `projection == 0` LEAK APPEARS ALREADY CLOSED
`model_value = 0` by day: 07-11 **20** · 07-12 **17** · 07-16 **10** · 07-17 **8** ·
07-18 **0** · 07-19 **0**.
It stopped. 55 historical rows carry `model_value = 0`; **all 55 are grade B**, and 50 of
them sit past 400 (so the shipped juice floor would have refused them anyway). The
`projection > 0` refusal at `analyzeViaEngine1.js:422` is doing its job.
Two things still true and worth carrying into G-a:
1. `getModelAggregate` already excludes them via `.gt('model_value', 0)`
(`ledgerService.js:464`) — public accuracy was never polluted by these rows.
2. Folding the `projection > 0` check into G-a's gate (as you asked) is still correct —
it's currently a separate refusal *after* feature computation, and moving it into the
gate chokepoint makes the refusal reasons uniform. It is a **tidy-up, not a leak fix.**
The `edge_pct` scale problem is NOT resolved and is untouched by this report.
---
## 7. WHAT I RECOMMEND HAPPENS NEXT
1. **Kev rules on §2.4** (dial table) — specifically flex-band enforcement on/off.
2. **G-a ships** with the agreed dials + `gate_*` refusal reasons + no-odds refusal +
the folded `projection > 0` check. Low risk: ~11 props in 6 days of incremental cuts.
3. **C-led first, not later** — add `ev_pct`/`p_win`/`fair_odds` columns and start
recording. Nothing else in this train can be validated until EV is on disk. This is the
cheapest highest-leverage item on the board.
4. **Escalate §2.2 (only B and C grades ever emitted)** to its own investigation. It
undermines hero v2, "A-RATED" copy, and any EV ranking within a flat grade pool.
5. Probe PropLine for alternate markets before scheduling L-b.
**HANDOFF — Design (Session 2):** nothing in this report is renderable yet. When the gate
ships, the refusal reasons `gate_too_juiced` / `gate_thin_edge` / `gate_longshot` /
`no_odds` will each carry user copy in the payload (D-ref), and §1.3 says the empty-board
state (G-c) will be rare but real on a thin MLB night — it needs a designed empty state,
not a blank grid.
---
*Report generated 2026-07-19 against live prod data. No engine code changed. G-a held for
ruling.*
+415
View File
@@ -0,0 +1,415 @@
# GRADE COLLAPSE + DEAD PROBABILITY LAYER — diagnosis
**REPORT ONLY. No grade logic, thresholds, or engine code changed** (Kev's instruction).
Two findings. The second one is bigger than the question I was asked.
Data: live Supabase `ledger_entries` (604 rows, all users) + live prod API
`api.vyndr.app`, 2026-07-19 ~22:05 UTC.
---
## FINDING 1 — THE COLLAPSE IS REAL, LIVE, AND STRUCTURAL
Not a thin-slate artifact. Across **604 ledger rows and both sports**:
- **2 distinct grades ever emitted: B and C.** Zero A+, A, A, B+, B, C, D, F.
- **9 distinct confidence values ever emitted:** 63, 57, 55, 52, 47, 45, 35, 25, 20.
- **Confidence ceiling = 63.** It has never exceeded 63 in the recorded era.
Still true **today** (Jul 18 + 19, post every fix, both sports): 4 confidence values
(63/57/52/47), 2 grades.
| Sport | Grade | n | conf min | conf max |
|---|---|---|---|---|
| mlb | B | 276 | 45 | 63 |
| mlb | C | 107 | 20 | 52 |
| wnba | B | 130 | 45 | 63 |
| wnba | C | 91 | 35 | 52 |
### Confidence does NOT determine the letter
| conf | grade | n |
|---|---|---|
| 63 | B | 54 |
| 57 | B | 175 |
| 55 | B | 29 |
| 52 | **C** | 100 |
| 47 | **C** | 14 |
| **45** | **B** | **148** |
| 35 | C | 80 |
**conf 45 → B, but conf 47 and 52 → C.** The mapping is non-monotonic, so the surfaced
`confidence` is not the quantity the letter was derived from. This confirms the
`mlb-grade-degradation.md` "grade↔confidence mismatch" as a *display* artifact: two
different quantities are being shown as if one explains the other.
(At conf 45→B the avg edge is 103; at conf 52→C it is 49 — so the letter tracks the
engine composite/edge, not the displayed confidence.)
### Collateral: the edge scale is still broken and still live
- **311 of 604 rows (51.5 %) have |edge| > 40** — the frontend's `EDGE_BOARD_SANE_MAX`,
i.e. over half the board's edge is nulled at render.
- **39 rows have |edge| > 100** — impossible as a percentage. Worst: **620**.
- Live today: edges of 140, 180, 220 on Jul 1819 rows.
`U-deg`'s `projection == 0` leak IS closed (0 since 07-18). The **edge_pct scale is
not** — it remains open and is now quantified.
---
## FINDING 2 — 🔴 THE ENTIRE PROBABILITY LAYER IS DEAD IN PRODUCTION
Found while fingerprinting Arc 1 (U-fp). This is the headline.
### Live fingerprint, `GET /api/snapshot/mlb`, 8 graded props
| Field | Present |
|---|---|
| `projection`, `confidence`, `book_odds`, `fair_odds`, `takeable`, `devig_method`, `alt_lines` | **8 / 8** |
| **`p_win`** | **0 / 8** |
| **`kelly`** | **0 / 8** |
| **`ev_pct`** | **0 / 8** |
| **`model_odds`** | **0 / 8** |
| **`value`** | **0 / 8** |
**Control:** `alt_lines` is present 8/8 and is Desk-gated in `tierGating.js:55`, which
proves the payload is **not** being tier-stripped. These fields are genuinely never
computed — not hidden.
### Root cause — a one-line sport gate, and an S46 fix that was only half-applied
`analyzeViaEngine1.js:509` feeds the estimator from `meta.gameLogs`:
```js
const est = estimateProbability({ gameLogs: meta.gameLogs, line: prop.line, ... });
```
`meta.gameLogs` comes from `computeFeatures.js:173-181``gameLogService.getGameLogs`.
And `gameLogService.js:21-26`:
```js
function pythonPath(sport) {
switch (sport) {
case 'nba': return '/stats/last-n';
case 'wnba': return '/wnba/stats/last-n';
default: return null; // ← MLB exits here
}
}
```
with `getGameLogs` line 31: `if (!path) return null;`
So:
- **MLB** — returns `null` by construction. Never had game logs on this path.
- **NBA/WNBA** — hits the Python stats service, which is **offline in prod** (documented
in CLAUDE.md; degrades to null).
`meta.gameLogs` is `[]` for **every sport in production**
`estimateProbability` returns `{p_over: null, reason:'insufficient_data'}`
(`probabilityEstimator.js:55-57`) ⇒ `pWin` is null ⇒ **every field guarded by
`if (pWin != null)` is skipped**: `p_win`, `kelly`, `model_odds`, `ev_pct`, `value`.
**This is the S46 bug, second location, never fixed.** CLAUDE.md records that
`gameLogService.getGameLogs` being NBA/WNBA-only starved MLB, and that the fix was an
MLB branch in **`featureCache.gameLogFeatures`**. That fixed the *feature* path — which
is why `projection`, `confidence`, and grades still work. The **estimator path was never
given the same branch**, so it has been silently dead the whole time.
### What this actually breaks
1. **EV — the Model Train's entire ranking signal — does not exist in production.**
Arc 1 shipped `ev_pct` and it has never once been computed on a live prop.
2. **Hero v2 is non-functional.** `pickHeroProp` requires a finite `ev_pct`
(`heroPropService.js:84`), so the EV loop matches **nothing** and always falls through
to the "most recent graded read" fallback. Live proof: `/api/hero-prop` returns
`"is_recent": true` — the fallback path, every time. The hero has not been an EV pick
since the day it shipped.
3. **Quarter-Kelly is dead** — same `pWin` dependency (`analyzeViaEngine1.js:516-520`).
This is a **promise-audit issue**: Kelly sizing is sold on the pricing page and
`PROMISE-AUDIT.md` lists it as BUILT. It is built and never runs.
4. **The "value triplet" is a duet live**`book_odds` + `fair_odds` render;
`model_odds` is always absent.
5. **`value` is never true**, so the VALUE marker can never light up.
### Why this reframes the whole train
- **C-led would persist a column of nulls.** Do not build EV persistence until EV exists.
- **G-a's `EV_FLEX_THRESHOLD` would gate on a permanently-null value.** With
`EV_FLEX_ENFORCE=0` (Kev's ruling) this is harmless today — but had we enforced it,
the flex band would have been cut to **zero**, because `ev_pct >= 4` can never be true.
The ruling to ship it disabled accidentally prevented an outage.
- **S-b (rank board on EV)** would rank on nulls.
---
## RELATIONSHIP BETWEEN THE TWO FINDINGS
They are **adjacent, not identical**, and both trace to the same missing input:
- The dead estimator explains **why no probability-derived output exists** (EV, Kelly,
model_odds, p_win).
- It does **not by itself** explain the B/C letter collapse, because the letter comes
from engine1's rule-based composite over the *feature vector*, which is alive.
- But they share a root: **the model is running on a partial input set.** One of its two
probability inputs (the empirical quantile distribution over real game logs) is absent
for 100 % of props, so whatever spread the composite was designed to produce is being
generated from the surviving features only.
**The 9-discrete-confidence-values pattern is a small set of additive rule hits** — a
scorer landing on a lattice rather than a continuum. Mechanism now traced in full below.
---
## FINDING 3 — THE MECHANISM: **`A` IS MATHEMATICALLY UNREACHABLE**
Every claim here was verified directly against the source.
### The grade is an integer index, not a score
`engine1.js:16-17, 158-163`:
```js
const GRADE_SCALE = ['F','D','C-','C','C+','B-','B','B+','A-','A','A+'];
const NEUTRAL_INDEX = 3; // 'C'
...
let idx = NEUTRAL_INDEX;
for (const f of factors) idx += f.delta; // flat sum of ±1.0 / ±0.5
idx = clampIndex(Math.round(idx));
```
`grade_thresholds.json` is **not an input mapper in the JS path.** Nothing compares a
probability to those cutoffs. `engine1.js:29-36` reads the table *backwards* — it takes
the letter the index already produced and looks up that band's **midpoint** to
manufacture a confidence number.
**So `confidence` is a cosmetic re-encoding of the letter.** It carries zero information
beyond the letter and by construction can never disagree with it. There is **no
data-sufficiency penalty in the live path** — the one CLAUDE.md describes lives in
`mlbGrader.js:50-69`, which is DEAD CODE. (`computeFeatures.js:21` still carries a stale
comment claiming the adapter downgrades confidence; it does not.)
### Six of thirteen factors are wired to features nothing populates
| Dead factor | Δ | Why it never fires |
|---|---|---|
| `weak/top_opponent_defense` | **±1.0** | needs `opp_rank_stat``team_stats:{sport}:{abbr}`**`refreshTeamStats` has ZERO production callers** (verified: only its own export + tests) |
| `consistency_elite/boom_bust` | **±1.0** | MLB consistency logs come from the same dead `gameLogService` path as Finding 2 |
| `opp_starters_out` | +1.0/+0.5 | `featureCache.js:280`: `if (!teamId) return out;``computeFeatures` never passes `teamId` |
| playoff factors | ±0.5 | `season_type` never set |
| `heavy_workload_7d` | 0.5 | `game_count_in_7d` never set |
| `ref_*` / `coach_*` | ±0.5 | NBA-flavored caches, absent for MLB |
`computeFeatures.js:234-236` builds `gameContext` as **`{ home_away }` and nothing else.**
### The arithmetic
`idx = clamp(round(3 + Σδ))`. Live-firing factors for MLB reduce to: `l5_*` (±1.0),
`l20_*` (**+1.0 only — verified, BOTH branches are `delta: 1.0`, there is no negative L20
contribution**), `home_game` (+0.5), rest (±0.5), `trap_composite_high` (1.0).
| 4-letter | needs Σδ | live reachable? |
|---|---|---|
| **A** (A/A/A+) | **≥ +4.5** | **NO — live max is +3.0** (+2.0 on a back-to-back, and MLB `rest_days` is 0 most days) |
| B | +1.5 … +4.49 | yes |
| C | 1.5 … +1.49 | yes |
| **D** | ≤ 1.51 | **NO — live min is 1.5**, and `Math.round(1.5) = 2``C`. Misses by one rounding tick. |
| **F** | ≤ 2.51 | **NO** |
**An A is short by at least 1.5 index steps — and the ≥1.5 of deltas that would close the
gap (`opp_rank_stat` ±1.0, `consistency` ±1.0, `injury` +1.0) are exactly the permanently-
null features.** The reachable index band is **2…6 = {C, C, C+, B, B}**, which
`gradeAdapter.FOUR_LETTER_MAP` (`gradeAdapter.js:31-37`, a 3→1 collapse) renders as
exactly **{C, B}**. That is the observed output, derived from first principles.
Confidence corroborates exactly: reachable letters carry `{42, 47, 52, 57, 63}`. **Live
today we observe precisely `{47, 52, 57, 63}`** — C (42) is absent because
`gradeSlateService.js:76` keeps the higher-confidence side of each prop, truncating the
bottom. The older values in the ledger (`55, 45, 35, 25, 20`) are from the pre-`888d103`
hand-rolled table `{10,15,20,25,35,45,55,65,80,90,100}` — the ledger is append-only, so
it contains both eras.
### Relationship to `mlb-grade-degradation.md`
**Shared table, different bug — and its "fix" made this collapse invisible.** That audit
redefined confidence as the band midpoint so the letter round-trips through the table.
The resulting "25/25 agreement" is **a tautology, not a validation**: confidence is
derived *from* the letter, so it would report 25/25 even if every grade were wrong. That
audit only examined the output encoding. This collapse is one layer upstream, on the
input side — whether `computeFactors` has enough live features to move the index at all.
---
## RECOMMENDATION (no code changed pending Kev's call)
**Re-sequence: fix the dead estimator FIRST — before G-a, before C-led.**
Rationale: it is the cheapest fix on the board (an MLB branch in the estimator's log
source, mirroring the one already written for `featureCache`), and it simultaneously
restores EV, Kelly, `model_odds`, the VALUE flag, and hero v2. Every other Arc 2-5 item
is downstream of it. Building the gate, the persistence layer, or the board ranking on a
null signal is building on nothing.
Suggested order:
1. **Revive the probability layer** (MLB branch + a real NBA/WNBA fallback, since Python
is offline). Fingerprint that `p_win`/`ev_pct` appear live.
2. **Then C-led** — persist EV that now has values.
3. **Then G-a** — with the flex band still disabled per the standing ruling.
4. **Then the grade range** — now diagnosed (Finding 3), and it is NOT primarily a
consequence of step 1. It needs its own decision, because there are two very different
fixes and picking wrong bakes in a lie:
- **(a) Feed the starving factors.** Call `refreshTeamStats` (nothing does), pass
`teamId`/`season_type`/`game_count_in_7d` through `gameContext`, give
`safeGetConsistency` the same MLB branch as step 1. This restores ±3.0 of range and
makes A/D reachable **on merit**.
- **(b) Re-scale the index/thresholds** so the current narrow spread spans more
letters. **This is the tempting one and it is the wrong one** — it would mint A's
without adding a single bit of information, and every "A" would be a relabelled B.
It converts a visible limitation into an invisible lie.
**Recommend (a), explicitly reject (b).** If (a) proves infeasible, the honest fallback
is to keep the two-letter output and stop advertising a scale we don't produce — not to
stretch the scale.
**Copy consequence, either way:** "A-RATED" appears on public surfaces and `AccuracyBadge`
for a grade the engine has never emitted. Until (a) lands, that copy is unsupported.
Open question for Kev: NBA/WNBA have no free game-log source on this path with Python
down. `espnStatsAdapter.getPlayerGameLog` (Wave 0) already solves exactly this for
`featureCache` — reusing it here is the obvious candidate, and costs no quota.
---
# RESOLUTION — Session 63 (shipped)
Kev's ruling: **(a) fix on merit, never (b) rescale.** Rescaling would mint A's
without adding information — a relabelled B marketed as an A, corrupting an
append-only ledger permanently. That option is permanently rejected.
## What shipped
| Fix | File | Effect |
|---|---|---|
| Normalized per-game rows for ALL sports | `featureCache.getStatRows` | Revives `p_win``ev_pct`, `kelly`, `model_odds`, `value`, hero v2. Also feeds consistency. |
| Rows wired into the grade path | `computeFeatures.safeGetConsistency` | One fetch per prop, shared by 3 starving consumers |
| `refreshTeamStats` called in production | `snapshotService.runSnapshot` | `opp_rank_stat` populated → the ±1.0 opponent factor can fire (it had ZERO callers) |
| `game_count_in_7d` derived from real logs | `computeFeatures` gameContext | `heavy_workload_7d` (0.5) can fire |
| **L20 symmetry** | `engine1.computeFactors` | NEW `l20_contradicts_*` 1.0. There was no negative L20 path at all — a structural reason D was unreachable |
| Consistency CV floor | `consistencyScore` | See calibration finding below |
| `confidence_basis: 'grade_band'` | `gradeAdapter.toLegacyShape` | Confidence labelled as derived, not a probability |
| Dead `mlbGrader.js` **removed** | — | Referenced only by its own test. Described-but-dead penalty eliminated |
**Deliberately NOT wired** (would have been dead code dressed as a fix, documented
inline): `teamId` (no `team_id` column exists; `getFeatures` reads it top-level not
off gameContext; and the factor needs a starter-id list that doesn't exist) and
`season_type` (engine1 gates playoff factors on `season_type >= 2`, but ESPN's 2
means REGULAR season — threading it raw would fire "veteran_in_playoffs" in July).
## 🔶 CALIBRATION FINDING — consistency was NBA-tuned and would have flooded `boom_bust`
Reviving consistency exposed a latent bug. The CV thresholds (`cv >= 0.5`
`boom_bust`) were calibrated for NBA points (mean ~20). For a Poisson-ish counting
stat, **cv ≈ 1/√mean**, so any stat with mean < 4 forces `cv > 0.5` — it classifies
`boom_bust` regardless of actual behaviour. Verified on real logs:
- Alonso hits `[0,0,0,1,2,1,0,1,1,0]` → mean 0.60, **cv 1.17** → boom_bust
- Henderson hits `[1,0,0,3,1,1,0,0,1,0]` → mean 0.70, **cv 1.36** → boom_bust
First verification run confirmed it: **8/8 MLB props classified boom_bust**, a
blanket 1.0 that dropped the whole board to C. That is a systematic downgrade
masquerading as a signal — the mirror image of the "flooding A's" failure Kev
warned about.
**Guard shipped:** `MIN_MEAN_FOR_CV = 4` (env `CONSISTENCY_MIN_MEAN`). Below it,
consistency returns `unknown` (no factor) with `reason: 'low_mean_cv_unreliable'`.
Absent beats wrong. **Consequence: MLB low-count stats still get no consistency
factor** — honest, not fixed. The correct long-term fix is an index-of-dispersion
(variance/mean vs the Poisson baseline) classifier, which is scale-free. Tracked
as an open item; it is a modelling change needing its own validation.
## VERIFICATION ON MERIT — real props, real logs, real engine
`scripts/verify-grade-range.js` replays live-board props through the repaired
engine using free feeds (statsapi/ESPN). **Caveat stated up front: `opp_rank_stat`
needs the Redis team-stats cache that only production populates, so these local
runs OMIT a ±1.0 factor and therefore UNDERSTATE the restored range.**
**WNBA — 25 real props**
| | BEFORE (live board) | AFTER (repaired) |
|---|---|---|
| A | 0 | 0 |
| B | 17 (68 %) | 8 (32 %) |
| C | 8 (32 %) | 16 (64 %) |
| **D** | **0** | **1 (4 %)** |
11-step spread: `C 6 · C+ 10 · B 8 · D 1` — five distinct steps where there were
two. Revived signals: **`p_win` 25/25 (was 0)**, rows 25/25, consistency known
15/25 (the floor correctly abstains on low-mean assists/rebounds).
The D is earned, not manufactured: *Angel Reese assists over 2.5, p_win 0.365*
the model gives it 36.5 % and says so.
**MLB — 8 real props:** B 5 / C 3, `p_win` 8/8 (was 0). No A or D on a thin
8-prop late-night board of near-identical 0.5-hits props.
**Reading it honestly:**
- **D emits on merit. ✅**
- **A did not emit locally** — expected: A needs Σδ ≥ +4.5 and the local ceiling is
+3.0 without `opp_rank_stat`. Structural reachability is proven arithmetically
and locked in `tests/unit/gradeRangeRestore.test.js`; **empirical A emission
requires production and is the outstanding fingerprint.**
- **Nothing flooded.** Grades got *harder*, not easier — B fell 68 % → 32 %. The
B→C movers are driven by the new L20 negative branch: props whose season
baseline contradicts the graded side no longer get a free pass. That is the
intended correction.
## 🔴 MARKETING HOLD — A-rated copy is UNSUPPORTED until A verifiably emits
Confirmed the honest fallbacks are what render today:
- `/api/ledger/accuracy` returns buckets **B and C only** — no A bucket. So
`AccuracyBadge`'s `aRated` sample is 0, below `minSample`, and it falls through
to **"MODEL · 63% HIT"**. No fabricated A-RATED is displayed.
- `TopSignals` self-hides when there are no A-rated grades.
**Nothing fabricated is shipping — but the copy describes a grade the engine has
never emitted.** Do not promote "A-RATED" in marketing, and do not build new
surfaces on an A bucket, until a production fingerprint shows real A grades. Lift
this hold only against live data.
---
## ✅ PRODUCTION FINGERPRINT — 2026-07-19 ~22:50 UTC
`POST /api/analyze/prop` (grades on demand, so it exercises the repaired path
immediately rather than waiting for the snapshot cron):
```
Gunnar Henderson hits o0.5 @ -140 (fanduel)
grade : B
confidence : 57
confidence_basis : grade_band <- NEW (truth label)
p_win : 0.523 <- WAS ABSENT on 100% of grades
ev_pct : -10.4 <- WAS ABSENT
model_odds : -109 <- WAS ABSENT (triplet was a duet)
value : false <- WAS ABSENT
takeable : true
book_odds : -140
fair_odds : -125
```
**The value triplet is finally whole: book 140 · vig-free 125 · model 109.**
And it tells the truth — the model gives 52.3 % where the de-vigged market says
55.6 %, so EV is 10.4 % and `value` is correctly FALSE. The engine now refuses to
call a bad price good, which is the entire point of the train.
Still broken, unchanged by this work: `edge_pct: 100` on that same response — the
edge scale remains on a bad scale (open item, `U-deg` part 2).
### Outstanding: A-emission in production
`opp_rank_stat` only populates when `refreshTeamStats` runs inside a snapshot, and
the next cron slot is 01:00 UTC. Until then the live ceiling is still +3.0, so **A
cannot yet emit in prod**. Re-measure the ledger grade distribution after that slot
— that is the remaining proof, and the MARKETING HOLD stays until it passes.
---
*Diagnosed + resolved 2026-07-19 (Session 63). Verified on real props; production
A-emission fingerprint outstanding.*
+22
View File
@@ -30,6 +30,28 @@ bug, fixed at the source in the generic grade path (`engine1` +
to any grade's displayed confidence resolves back to the same letter (proven to any grade's displayed confidence resolves back to the same letter (proven
for all 11 grades in `tests/unit/mlbGradeDegradation.test.js`). for all 11 grades in `tests/unit/mlbGradeDegradation.test.js`).
> ### ⚠️ CORRECTION (Session 63, 2026-07-19) — THE "25/25 AGREEMENT" WAS A TAUTOLOGY
>
> **Do not cite the 25/25 grade↔confidence agreement below as validation of
> grade quality. It validates nothing.**
>
> The fix above made `confidence` a *deterministic function of the letter*:
> engine1 picks a letter via an additive factor index, then looks up that
> letter's band midpoint to produce the number (`engine1.js:29-36`). Feeding
> that number back through the same table can only ever return the letter it
> came from. **The round-trip would report 25/25 even if every grade were
> wrong.**
>
> It is a real fix for a real bug (the two encodings had drifted a sub-tier
> apart) — it is simply a *consistency* check, not an *accuracy* check.
> `confidence` carries ZERO information beyond the letter. The genuinely
> independent probability is `p_win` (the quantile estimate over real game
> logs), which Session 63 discovered had never been computed in production at
> all. Payloads now carry `confidence_basis: 'grade_band'` so no consumer can
> mistake the derived number for a model probability.
>
> Full diagnosis: `specs/audit-data/grade-collapse.md`.
## Blast radius (commit `9fc4edf`) — work-order #6 ## Blast radius (commit `9fc4edf`) — work-order #6
The degraded grades (projection=0 → `model_value = 0`) are already settled in The degraded grades (projection=0 → `model_value = 0`) are already settled in
the append-only `ledger_entries` and are NOT deleted. Functional marking: the append-only `ledger_entries` and are NOT deleted. Functional marking:
+126
View File
@@ -0,0 +1,126 @@
# THE BATTER CLUSTER — both-ways results, and what the bar actually is
**2026-08-03.** Challenger-only. Counter byte-identical. `skillProjection`
byte-identical (frozen) — verified by diff against the prior commit.
> **PREMISE CORRECTION, and it decides the whole order: total_bases has NOT
> passed BAR 1.** Yesterday's result was **delta +0.0071, CI95
> [0.0648, +0.0789] — inconclusive, at parity, and contaminated**. No feature
> passed the gate. It was described as "the first challenger that did not LOSE",
> which is not the same as proven.
>
> This matters structurally: if total_bases is installed as "the frozen proven
> reference" and every other stat is held to "the identical bar total_bases
> cleared", **the bar becomes "be inconclusive at parity"** — a standard that
> proves nothing and would admit the entire cluster on a null result.
>
> **The proven set is EMPTY.** That is the honest Stage-B input.
---
## 0. Two orders running, one pattern
This is the second consecutive order whose premise promoted a null result to a
pass (the previous one had "barrel rate PASSED solo" when every TB feature was
refused on sample size). Flagging the pattern once, without labouring it: the
measurements are being read at their most favourable interpretation somewhere
between sessions. The numbers below are stated so that cannot happen again.
## 1. Validity: still contaminated today, but the clock has started
`statcast_history` was **empty** — the retention shipped *after* yesterday's
refresh had already run. Two things fixed this session:
1. The first prod run failed with `Could not find the 'swing_pct' column` — the
hand-enumerated schema had drifted from the table it was copying. **The
refresh still succeeded and wrote all 1,387 aggregate rows**, which verified
the best-effort guard in production: a retention failure does not fail the
refresh.
2. The table is now created `LIKE statcast_aggregates` and the writer passes the
row through whole, so there is no drift surface left.
**Verified live: `history_retained: 1387, as_of 2026-08-03`.** The point-in-time
clock is running. It does not yet give a *window*`as_of 2026-08-03` can only
score games from 2026-08-04 — so **everything below remains contaminated and
directional**, exactly as yesterday.
## 2. STEP 1 — both-ways results, per stat
Bonferroni across each stat's own sweep. The decider is correlation with the
counter's **residual** (`won p_win`): only what the counter misses is new.
| stat | n | strongest solo (marginal r) | p | gate | best interaction (incremental) | head-to-head vs counter |
|---|---|---|---|---|---|---|
| **hits** | **803** | exit_velo **0.053** | 0.130 | **FAILS** (properly tested) | all ≈0, all **FAIL** | **LOSES** 0.0961, CI [0.165, 0.029] |
| total_bases | 383 | hard_hit **+0.135** | 0.008 | insufficient n | barrel×archetype 0.101 | **INCONCLUSIVE** +0.0038, CI [0.068, +0.075] |
| rbi | 391 | exit_velo 0.090 | 0.074 | insufficient n | barrel×archetype +0.073 | no projection built |
| home_runs | 228 | barrel **0.135** | 0.041 | insufficient n | barrel×archetype 0.116 | no projection built |
| runs | 188 | launch 0.088 | 0.230 | insufficient n | **K×K +0.132** | no projection built |
### hits is now a FINAL answer, not a pending one
At **n=803** hits clears the gate's sample requirement, so its features were
**properly tested rather than refused**. Every one fails on effect size (max
|r| 0.053 against a 0.15 bar), every interaction's incremental collapses to
≈0, and the model loses head-to-head with a CI excluding zero. **This is a
well-powered negative and hits should be closed.**
### The rest are all n-blocked
TB 383, rbi 391, HR 228, runs 188 — none reaches 500. Their numbers are
directional only. Two are worth carrying:
- **home_runs · barrel_pct, marginal r = 0.135 (p=0.041).** Note the **sign**:
higher barrel rate goes with the counter *over*-predicting. If that survives
more data it is a real correction, not a new predictor.
- **runs · batter_K × pitcher_K, incremental +0.132** — the largest incremental
anywhere in the cluster, with a clean mechanism (strikeouts destroy plate
appearances, and a PA that never happens cannot score). At n=188 it is a lead.
### RBI has a missing half, stated rather than papered over
RBI is power × **opportunity** — the same swing drives in one run or three
depending on who is on base. **We do not ingest baserunner state at all**, so
half the mechanism is absent. A weak RBI result here is not evidence that skill
inputs fail for RBI; it is evidence we are modelling half the stat.
## 3. STEP 2 — total_bases held frozen
`git diff` on `src/services/model/skillProjection.js` against the prior commit
is **empty**. The projection is byte-identical; nothing was re-opened or re-fit.
It is frozen — but as §0 records, it is frozen as an **inconclusive** model, not
a proven reference.
## 4. STEP 3 — the proven set for Stage B
| stat | BAR 1 (proven) | BAR 2 (calibrated) |
|---|---|---|
| total_bases | **NO** — inconclusive at parity, contaminated | not attempted |
| hits | **NO** — loses, CI excludes zero, well-powered | n/a |
| home_runs | not testable (n=228) | n/a |
| rbi | not testable (n=391, half the mechanism missing) | n/a |
| runs | not testable (n=188) | n/a |
**Proven set: EMPTY. Stage B has nothing to calibrate.** Calibrating an
inconclusive model into grade bands would produce an "A" backed by a model not
shown to beat counting — which is exactly what BAR 2 exists to prevent.
## 5. What actually unblocks this
Everything now waits on the same two things, and both are **waiting problems,
not building problems**:
1. **A point-in-time window.** `statcast_history` starts accruing usable
comparisons from 2026-08-04. A week gives a real one. Until then no skill
result can be honest, whatever its n.
2. **Sample.** TB needs ~117 more settled rows, rbi ~109, HR ~272, runs ~312.
At current volume TB and rbi arrive within days; HR and runs are weeks away.
Re-run `scripts/cluster-prove.js` (per stat via `CLUSTER_STAT=`) once both hold.
**Ranked, when the data arrives:** total_bases (closest to both bars) → rbi
(needs baserunner state to be a fair test) → home_runs → runs. **Close hits.**
**Not recommended:** treating parity as proven, calibrating anything yet,
lowering n≥500, or letting a stat inherit a pass from the shared conditioning
map. Reuse sped the search; it granted nothing, and nothing has been granted.
+89
View File
@@ -0,0 +1,89 @@
# BUILD 2 — REVIEW ZERO REPORT (no build; awaiting Kev on Q1-Q3)
2026-07-31. Nothing built, no migration applied, no Stripe object touched.
## 🔴 TWO ORDER EXPECTATIONS ARE WRONG — read before deciding Q1-Q3
### G5 — `users.founder_status` is **LIVE, NOT DEAD**. Do not drop it.
| where | what |
|---|---|
| `src/services/stripeService.js:163` | **WRITTEN** by the webhook: `founder_status: isFounder` |
| `src/routes/stripe.js:95` | **READ** and served: `is_founder: req.user.founder_status` |
| `src/middleware/auth.js:24` | in `PROFILE_COLUMNS`**loaded on every authenticated request** |
| `src/middleware/auth.js:72` | defaulted to `false` on the fallback profile |
The guardrail says *"Don't write `users.founder_status` unless G5 proves it live."* **G5 proves it
live.** A5's "drop on evidence" must NOT run for this column.
### 🔴 AND THE TWO FOUNDER FLAGS ALREADY DISAGREE IN PROD
- `user_profiles.founder_pricing = true` on **1 of 3** profiles
- `users.founder_status = true` on **0 of 3**
The webhook writes BOTH from the same `isFounder` — so they are a **dual-write that has already
drifted**. Whatever is decided on Q3, the build must pick ONE canonical flag (the order says
`user_profiles.founder_pricing`) and make the other a derived read or explicitly retire it. Leaving
two independently-writable founder flags is how a founder loses their rate on one code path.
## THE GREPS
**G1 — `user_profiles.founder_pricing`**
- WRITE: `stripeService.js:173` (webhook mirror) — the only writer.
- READ: `routes/partners.js:68,92` (MRR attribution picks founder vs standard price);
`web/src/app/api/user/profile/route.ts:16`; `web/src/app/profile/page.tsx:17,121` (the badge).
- `routes/founders.js:6,9` documents that the counter deliberately does **not** trust this flag
(it counts Stripe subs instead) — *"a tier/founder_pricing field a profile can set"*.
**G2 — THE PROMO-CODE BYPASS IS THE ONLY FOUNDER GATE TODAY.**
`createCheckoutSession(userId, email, tier, founderCode)` (`stripeService.js:68`) →
`getPriceId(tier, founderCode)``isFounderCodeValid()` against `VALID_FOUNDER_CODES`
(`FOUNDER2026, VYNDR, BETONBLK, EARLYBIRD`) + `FOUNDER_EXPIRY 2026-12-31`. The result is stamped
into **`metadata: { user_id, tier, is_founder: String(isFounder) }`** (`:119`), and the webhook
**trusts that metadata** (`:155`). **So today a code alone mints a founder at any seat number, and
carries itself into the DB flag.** This is the bypass B1 must delete.
**G3 — the webhook DOES set the entitlement** (`stripeService.js:151-176`). It writes `users`
(`tier`, `stripe_customer_id`, `founder_status`) **and** mirrors to `user_profiles`
(`tier`, `subscription_status: 'active'`, `founder_pricing`).
**This closes an earlier CANNOT DETERMINE: a paid sub DOES flip the Build-1 gate.**
**But it stores NO `stripe_subscription_id` anywhere** — confirming A1 is required, since
`finalize_founder_slot` and grandfather reconciliation both key off it.
**G4 — `nexapay_customer_id`: ZERO code references** in `src/` or `web/src/`. The column exists on
`user_profiles` and is empty (0/3). **A5's drop is evidence-supported** — as its own migration.
**G6 — price selection**: `getPriceId` returns the founder PRICE_MAP entry when the code is valid,
else standing (falling back to the `PRICE_UNCONFIGURED` sentinel if unset); used at
`line_items: [{ price: priceId, quantity: 1 }]` (`:115`).
## DB FACTS (VERIFIED against zmdnczhtdxcddsxzttub)
- `user_profiles` columns: `id, email, tier, scan_count, scan_reset_date, subscription_start,
subscription_end, subscription_status, cancel_at_period_end, founder_pricing, age_verified,
**nexapay_customer_id**, created_at, updated_at, mfa_setup_prompted, grace_period_until,
partner_ref`. **Confirmed: NO `stripe_customer_id`, NO `stripe_subscription_id`** → A1 needed.
(Note `users` DOES already carry `stripe_customer_id` — `middleware/auth.js:24`.)
- `founder_pricing_seats` = **VIEW** ✓ decorative, as the order states.
- 3 profiles; 1 with `founder_pricing=true`.
## CANNOT DETERMINE
**The four Stripe price IDs.** `PRICE_MAP` is env-driven and there is **no `STRIPE_SECRET_KEY` and
no `STRIPE_PRICE_*` in this environment**, so I cannot confirm the four IDs in the order are the
ones prod will actually charge. The order says they were verified live this session — I am flagging
that I could not independently re-verify them, and A3 hardcodes them, so **a typo becomes a
permanent mis-charge**. Recommend the build read them from env (already the pattern) and assert at
boot that all four resolve, rather than hardcoding in SQL.
## THE THREE DECISIONS — WAITING ON KEV (no silent defaults)
**Q1 POOL** — global 100 vs 100-per-tier. *Bearing on the build:* it changes the UNIQUE index
(`UNIQUE(slot_number)` vs `UNIQUE(tier, slot_number)`) and whether founder-follows-upgrade needs a
second claim at all (global = the slot travels with the user; per-tier = a Desk slot must be claimed
separately and may be full).
**Q2 REOPEN** — reopen on cancel vs 100 lifetime seats. *Bearing:* decides whether
`release_expired_slots` also handles cancellation, and whether the public counter can ever go down.
**Q3 TEST RECORD** — the 1 `founder_pricing=true` desk profile with **0 Stripe subs**. Note it also
has **`users.founder_status = false`**, i.e. it is already inconsistent. My read: it is a test
artifact, not a subscriber. *Bearing:* seed slot #1 vs start clean and clear the flag.
## TAGS
VERIFIED: G1-G6 with file:line; `user_profiles`/`users` schema; the view; the flag divergence
(1 vs 0); the webhook does set the entitlement but stores no subscription id.
**CANNOT DETERMINE: the four Stripe price IDs (no key/env here).**
**BLOCKED: all of Phase A/B — awaiting Q1, Q2, Q3.**
+81
View File
@@ -0,0 +1,81 @@
# BUILD 2 — REVIEW ZERO (report; mechanism NOT built)
2026-07-31. Nothing built, no Stripe object created or changed, no price logic touched.
## WHY THIS STOPPED
**I have no Stripe credentials in this environment** (`.env` holds ODDS/SUPABASE/INTERNAL keys —
**no `STRIPE_SECRET_KEY`**). The order's standing floor requires *"founder/standing/grandfather/race
all verified server-side"*. **I cannot verify any of them**, cannot create the standing price
objects the rollover needs, and cannot run the concurrent-checkout race test.
This is a payment path where the failure modes are **permanent and customer-facing**: a race bug
mis-prices a subscriber *forever* (slot 101 on a founder rate, or slot 99 on standing), and a
grandfather bug overcharges a founder *every month*. Shipping that unverified is the one place
"probably right" is not good enough. **So: findings, then stop.**
## WHAT I DID ESTABLISH (VERIFIED)
**0.1 — Stripe IS live and FOUNDER price objects DO exist.** Proven indirectly but soundly:
`GET /api/founders/count` returns **`{available:true, claimed:0, total:100}`**, and
`routes/founders.js` returns **`{available:false}`** whenever `countFounderSeats()` is null — which
it is when `!STRIPE_SECRET_KEY || founderPrices.length === 0`. **`available:true` therefore proves
both the secret key and at least one founder price ID are configured in prod**, and that `claimed:0`
is a REAL count, not a fallback. (That route is honest by construction — *"any failure → hide, never
fabricate"* — which is why it can be used as a probe at all.)
**STANDING objects: CANNOT DETERMINE.** Env is not readable from here, and `getPriceId` **falls back
silently**: a missing standing price yields the `PRICE_UNCONFIGURED` sentinel, which does not fail
until Stripe rejects the session. **So a missing standing object would not surface until the first
post-cap checkout 400s in front of a paying customer.** Verifying this is a 30-second check in the
Stripe dashboard and is **prerequisite #1**.
**0.3 — THE COUNTER TODAY IS NOT A GATE, AND CANNOT BECOME ONE WITHOUT NEW WORK.**
- It is a **read**, not a claim: `countFounderSeats` lists Stripe subscriptions and returns a number.
- It is **cached 300s** (`CACHE_KEY 'founders:count'`), so it is stale by construction.
- **Founder pricing is gated by a CODE + EXPIRY, not by the count** (`VALID_FOUNDER_CODES`,
`FOUNDER_EXPIRY 2026-12-31`) — so today **anyone holding `FOUNDER2026` gets the founder rate at
any seat number**, and the cap is decorative.
- **Two simultaneous checkouts at slot 99 would both read 99 and both get founder.** There is no
claim, no lock, no unique constraint anywhere in the path.
**0.2 — the entitlement the gate reads** is `config/tiers.js` capability `reasoning_visible`
(free false / analyst+ true), consumed by `snapshotGating.entitledToItemizedGrades` via
`resolveTierFromRequest`, which reads `users.tier`. **So a successful subscription must set
`users.tier`** for Build-1's gate to open. Whether the Stripe webhook currently writes that on
`checkout.session.completed` is **CANNOT DETERMINE without the webhook secret** — and it is
**prerequisite #2**, because a paid sub that does not flip `users.tier` sells access that never opens.
## WHAT I WOULD BUILD, ONCE UNBLOCKED (design is settled, so this is fast)
1. **An atomic slot claim, DB-backed** — a `founder_slots` table with a **unique constraint on
`(tier, slot_number)`**, claimed inside the checkout-session request *before* the Stripe call.
The unique index — not a count read — is what makes the race impossible: two concurrent claims
for slot 100 mean one INSERT wins and the other is rejected to standing. **The cached count must
be removed from the decision path entirely** and kept only for display.
2. **Price selection from the claim**, not from a code: claim succeeded → founder object; claim
rejected/cap full → standing object. **Retire the code+expiry bypass**, or it silently defeats
the cap.
3. **Grandfathering** is already native (a sub created against a founder price stays on it) — the
build rule is simply: **never call Stripe's price-update/migration on a founder subscription.**
4. **Founder-follows-upgrade**: on tier change, attempt an atomic claim on the *target* tier's
founder slots; success → founder object, failure → standing. Tie to continuous subscription by
**releasing the slot on cancellation** (that is what makes "break it → standing" true rather than
aspirational).
5. **Honest display**: real uncached count at render; **if the count cannot be served, show the
offer with no number** — the route already does exactly this, so follow its precedent.
6. **Beta framing** with no proven-edge claim, pointing at `/record` (shipped) as the honest
building record.
## PREREQUISITES (all need Kev — none are code)
1. **Confirm/create the two STANDING price objects** in Stripe ($24.99 analyst, $59.99 desk) and
set `STRIPE_PRICE_ANALYST` / `STRIPE_PRICE_DESK`.
2. **Confirm the webhook sets `users.tier`** on `checkout.session.completed` (else the gate never opens).
3. **A Stripe test-mode key** available to the build/verification environment, so the race,
grandfather and end-to-end unlock can actually be exercised rather than asserted.
## TAGS
VERIFIED: Stripe live + founder price objects exist (via the honest counter probe); the counter is a
cached read with no claim; founder pricing is code-gated not count-gated; the gate reads `users.tier`
via `reasoning_visible`. **CANNOT DETERMINE: whether standing price objects exist; whether the
webhook writes `users.tier`.** **BLOCKED: the entire mechanism — no Stripe credentials, so nothing on
the payment path can be verified, and it must not ship unverified.**
+109
View File
@@ -0,0 +1,109 @@
# CHALLENGER SCOREBOARD — 2026-08-03
> **Nothing was promoted. Nothing earned it yet.** Not because the bar was held
> too high, but because no challenger's confidence interval excludes zero on the
> good side. The champion is byte-identical; every challenger stays wired.
The headline of this session is not the scoreboard. It is that **the scoreboard
was unmeasurable until a two-day-old settlement outage was found and fixed** —
see §1. Settled sample went **493 → 1,741** the moment it was repaired.
---
## 1. Why "n-blocked" was the wrong diagnosis
The order said: don't repeat "n-blocked" without counting. Counting is what found
the real problem.
Three of the four axes read **exactly zero** settled rows — not low, *zero*:
environment 1,496 rows / 0 settled, opportunity 922 / 0, matchup 603 / 0. Rows
whose games had been **played days earlier** and never settled, with
`settle_attempts = 0` — never even attempted.
**Root cause:** `settleLedger` fetched open ids, then refetched full rows via
`.in('id', ids)`. PostgREST puts filters in the URL, so 500 UUIDs became an
**18,499-character request** that the fetch layer rejects with `TypeError: fetch
failed`. The result was destructured as `const { data: rows } = ...` with **no
error binding**, so `rows` came back null, the loop never ran, and the function
returned `{settled:0, voided:0, unrecoverable:0, pending:0}` — byte-identical to
a healthy "nothing to settle."
It hid for two days because it is **volume-triggered**: daily volume ran 20260
rows and settled perfectly for weeks. **2026-08-01 was the first day past the
500-row fetch limit** and settlement died that night. Worse, the zero-settle ops
alarm reads these same return values, so `pending: 0` told the watchdog the
backlog was empty — *the alarm built to catch exactly this could not see it.*
Fixed, deployed, and drained: **1,444 rows from 2026-08-01 settled (1,376
hit/miss + 68 void, 0 remaining).** `captureClosing` carried the same shape one
level down and is now chunked at 100 ids.
## 2. THE SCOREBOARD
Bar: the challenger's **own rows only**, direction-aligned, **paired bootstrap**
(4,000 resamples, deterministic seed) on the difference in resolution, because
both models score the same rows and independent standard errors would overstate
certainty. **PROMOTE requires the CI to exclude zero on the good side.** Same bar
that refuted hits-v1 — no lighter test for a would-be winner.
| challenger | settled n | resolution (chal / champ) | Δ vs champion | CI95 | verdict |
|---|---|---|---|---|---|
| arch-v1 (market-relative nudge) | **1,741** | 0.4599 / 0.4599 | 0.0000 | [0.0050, +0.0054] | **STAY WIRED** (inconclusive) |
| arch-v1 · rows it MOVED only | 1,325 | 0.4886 / 0.4886 | 0.0001 | [0.0064, +0.0061] | **STAY WIRED** (inconclusive) |
| contact-v1 (season contact quality) | **1,055** | 0.4070 / 0.4061 | +0.0008 | [0.0052, +0.0069] | **STAY WIRED** (inconclusive) |
| contact-v1 · rows it MOVED only | 511 | 0.3966 / 0.3948 | +0.0018 | [0.0107, +0.0144] | **STAY WIRED** (inconclusive) |
| proj-v1.1 ladder (all stats) | **1,664** | 0.4339 / 0.4640 | **0.0301** | **[0.0543, 0.0060]** | **STAY WIRED** (measured WORSE) |
| arch-v1 · environment axis rows | 871 | 0.5651 / 0.5679 | 0.0028 | [0.0089, +0.0036] | **STAY WIRED** (inconclusive) |
| arch-v1 · opportunity axis rows | 539 | 0.5311 / 0.5310 | +0.0001 | [0.0091, +0.0090] | **STAY WIRED** (inconclusive) |
| arch-v1 · matchup axis rows | **0** | — | — | — | **STILL PENDING** |
| tb-v1 (total_bases only) | **0** | — | — | — | **STILL PENDING** |
| hits-v1 (hits only) | **0** | — | — | — | **STILL PENDING** (refuted by replay, `specs/hits-v1-binomial.md`) |
**These are true prospective holdouts, not backtests.** arch-v1 and contact-v1
wrote their probability at grade time, into their own columns, before the game
was played. Nothing was recomputed. That is the strongest evidence available and
it is why no replay was needed here.
## 3. What the numbers actually say
**arch-v1 moves a lot and changes nothing.** It moved **1,325 of 1,741 rows
(76%)**, mean absolute move **2.5 points**, max 10.9 — and resolution is
identical to the champion to four decimal places, on the moved rows too. This is
not "too small to detect." It is movement that carries **no information about the
outcome**. A nudge this active with an effect this precisely zero is a finding,
not a pending verdict.
**The projection ladder is reliably worse than the champion.** 0.0301 with a CI
excluding zero, across 1,664 rows and all stats. Combined with hits-v1's refutation
(`specs/hits-v1-binomial.md`), the projection family now has two independent
measurements pointing the same way: it is not the champion's equal on any stat
measured so far. That is an argument for diagnosing its *inputs*, not for shipping
another variant of it.
**Three are genuinely pending, for a legitimate reason now.** matchup, tb-v1 and
hits-v1 all have rows written only on 2026-08-02/03, which settle after ET
midnight. matchup has ~496 rows queued, tb-v1 65, hits-v1 pending its first
snapshot write. They will read within a day or two — and now that settlement
works, they actually will.
## 4. Provenance
All 1,741 arch-v1 rows carry a single `model_version` (`engine1@2026-07-20`). The
older `pre-retention-unknown` rows (320 settled) carry no challenger values at
all, so they cannot influence any verdict. **No verdict here depends on
mixed-provenance rows** — the split was checked, not assumed.
Contamination excluded throughout: `quarantine_reason LIKE 'nontakeable_book%'`.
## 5. Promotion mechanics — specified, deliberately unused
No flip was performed because nothing qualified. When one does, the shape is:
challenger-first (write the promoted value into the served path while the
champion column keeps recording), version-tagged, atomic, with the previous
version one env flag away. Recorded here so a future promotion is a decision,
not an improvisation.
## 6. Reproduce
`SUPABASE_URL=... node scripts/challenger-scoreboard.js` — prints the full board,
the moved-rows-only slice, the per-axis slice and the provenance split.
+220
View File
@@ -0,0 +1,220 @@
# DECOMPOSING THE CHAMPION — where its edge actually comes from
**Read-only diagnosis, 2026-08-03.** Nothing built, nothing touched. Run on the
**repaired** settled set (n=1,741 after the settlement outage fix), not the frozen
pre-fix set.
> **VERDICT, one line:** the champion's entire edge is a **hit-rate counter**, its
> three adjustment layers contribute **nothing** (two are mildly harmful), and the
> largest recoverable loss is **not a missing feature — it is the `[0.10, 0.95]`
> clamp, which pins 20.6% of settled props to a constant** and hides outcomes
> ranging from 52% to 99.5% behind the same number `0.900`.
---
## 1. What the champion actually uses (STEP 1)
`src/services/intelligence/probabilityEstimator.js` is **five lines of arithmetic**:
```
base = empirical frequency of (stat > THIS line) over the game log
weighted = 0.6·base + 0.4·(same frequency over the last 5 games)
p = weighted + oppAdj(±0.03) + homeAdj(±0.015)
if cv > 0.40: p = 0.9·p + 0.05 (volatile → pull toward 0.50)
p_over = clamp(p, 0.10, 0.95)
p_win = side === 'under' ? 1 p_over : p_over
```
It reads exactly **three** features: `opp_rank_stat`, `home_away`, and
`l10_stddev`/`l20_avg` (for cv). `featureCache` computes and retains a dozen more
`park_h/hr/r`, `weather_temp_f/wind_mph/precip`, `rest_days`,
`opportunity_drift`, `ab_per_game`, `recent_ab_per_game`, `l5/l10/l20_avg`,
`game_count_in_7d` — and **p_win reads none of them.**
## 2. Per-stat ablation (STEP 2)
**Exact and analytic, not a refit.** Each adjustment is a closed-form function of
stored features and the consistency step is linear (`f(x)=0.9x+0.05`
`f(a+b)=f(a)+0.9b`), so every layer is removed algebraically from the stored
`p_win`. Nothing re-estimated, nothing re-fetched, no lookahead possible.
Paired bootstrap, 3,000 resamples, deterministic seed.
**A negative delta means removing the layer HURT — i.e. it carried signal.**
| stat | n | resolution (full) | opponent | home/away | consistency | **ALL THREE** |
|---|---|---|---|---|---|---|
| hits | 578 | 0.1964 | 0.0049 | 0.0006 | 0.0004 | **0.0059** [0.0168,+0.0056] |
| total_bases | 284 | 0.2370 | 0.0056 | +0.0042 | +0.0002 | **0.0015** [0.0139,+0.0115] |
| rbi | 273 | 0.4986 | +0.0059 | **+0.0053** [+0.0002,+0.0103] | +0.0005 | **+0.0106** [0.0003,+0.0220] |
| runs | 115 | 0.4062 | +0.0060 | +0.0087 | 0.0006 | **+0.0130** [0.0106,+0.0366] |
| walks | 66 | 0.4776 | 0.0035 | +0.0055 | +0.0018 | **+0.0008** [0.0268,+0.0281] |
**Removing all three adjustments changes resolution by nothing on every stat, and
on rbi/runs it IMPROVES it.** Exactly one ablation anywhere has a CI excluding
zero — rbi home/away, and its sign says removing it makes the model **better**.
**So ~100% of the champion's resolution is `base + recency`: how often this
player has cleared THIS number lately.** That is the whole model. Everything else
is decoration.
### A correction to how we read last session's scoreboard
Pooled across stats the champion resolves **0.46**; per stat it is **0.196
(hits)** to **0.499 (rbi)**. Pooling stats with different base rates *inflates*
correlation, because p_win varies across stats in the same direction as the true
base rate. **0.46 is a pooling artifact and should not be quoted as the
champion's resolution.** The paired *differences* in the scoreboard remain valid
(champion and challenger were pooled identically); only the absolute level was
inflated.
## 3. Do the challengers have it, or dilute it? (STEP 3)
| challenger | uses the base-frequency signal? | verdict |
|---|---|---|
| **proj-v1.1 ladder** | **No — it replaces it.** Fits a rate + NB distribution instead of counting frequency at THIS line | **DILUTING.** Measured reliably worse (0.0301, CI excludes 0). It discards the one thing that works in favour of a lossier route to the same question |
| **hits-v1** | No — same substitution, binomial instead of NB | **DILUTING.** Refuted (0.022, CI excludes 0) |
| **arch-v1 · environment** | Adds park/weather, which the champion ignores | **DILUTING.** n=871, 0.0028, CI includes 0 — movement without information |
| **arch-v1 · opportunity** | Uses `opportunity_drift`**the one feature with repeated residual signal** | **HAS THE FEATURE, WRONG IMPLEMENTATION** (see §4) |
| **arch-v1 · matchup** | — | STILL PENDING (rows settle after ET midnight) |
| **contact-v1** | Statcast contact quality; not in the retained vector | No evidence either way (n=1,055, CI includes 0) |
## 4. The missing-feature test — one real lead, already in our hands
Correlation of each **unused** feature with the champion's residual (`won
p_win`), per stat, bootstrap CI.
**Multiple-comparisons discipline first:** 14 features × 5 stats = 70 tests at
α=.05, so **34 CI-excludes-zero results are expected by chance.** Six appeared.
A single hit is noise. **Only a feature that repeats across independent stats is
evidence** — and exactly one does:
| feature | hits | total_bases | walks |
|---|---|---|---|
| **`opportunity_drift`** | **+0.156** [+0.007,+0.292] | **+0.145** [+0.005,+0.278] | 0.261 [0.463,0.017] |
| `recent_ab_per_game` (same quantity) | +0.052 | **+0.136** [+0.002,+0.272] | 0.157 |
`opportunity_drift` = recent at-bats ÷ season at-bats-per-game. It is the one
axis a frequency counter is **structurally blind to**: `base` knows how often he
cleared the number, not that he has moved from 8th in the order to leadoff, or
back from injury on a bench role. Sign flips on walks (n=47, and walks scale with
plate appearances differently) — so this is stat-specific, which is doctrine-
consistent, not a contradiction.
**But we already compute it, retain it, and built an axis on it — and that axis
extracts nothing** (opportunity axis: n=539, delta +0.0001, CI [0.0091,+0.0090]).
So this is **not** "go get a new feature." It is **"the feature has signal and our
implementation of it is wrong"** — arch-v1 applies it as a small multiplicative
nudge to `p_win`, which is not how you use an opportunity term. Opportunity should
scale the *rate*, before the frequency question is asked.
Weather on total_bases (`wind_mph` 0.164, `precip` 0.154) appears on **one stat
only** and sits inside the expected false-positive count. Recorded as a
non-lead unless it repeats.
## 5. Archetype on trial (STEP 4 item) — NO EVIDENCE, and the test is underpowered
The champion reads **no archetype feature at all**, so archetype cannot be ablated
out of it. The fair test is whether archetype explains what the champion gets
*wrong*: if an archetype's rows are systematically mispriced, archetype carries
prop signal we're missing.
| stat | archetype | n | mean residual | CI95 | mispriced? |
|---|---|---|---|---|---|
| hits | BOMBER | 195 | 0.038 | [0.105, +0.030] | no |
| hits | GHOST | 85 | 0.076 | [0.184, +0.033] | no |
| total_bases | BOMBER | 95 | 0.046 | [0.141, +0.058] | no |
| total_bases | GHOST | 46 | 0.041 | [0.187, +0.109] | no |
| rbi | BOMBER | 88 | 0.068 | [0.155, +0.019] | no |
| runs | BOMBER | 41 | 0.012 | [0.150, +0.121] | no |
**Verdict: (a) no-signal is UNPROVEN and (b) wrong-implementation is UNPROVEN —
the test cannot separate them yet.** Only **2 of 41 archetypes** (BOMBER, GHOST)
reach n≥40 settled rows. That is not "archetypes don't work"; it is "we have not
measured them." Distinguishing (a) from (b) needs archetype coverage across more
than two labels. **Do not act on archetype in either direction on this evidence.**
Worth noting: every archetype's mean residual is **negative**, which is not an
archetype effect — it is the global over-prediction in §6.
## 6. THE BIGGEST FINDING — the clamp, not a feature
**358 of 1,741 settled props (20.6%) sit ON the clamp boundary** (353 at the
floor). Within that fifth of the book the model emits a **constant**, so it cannot
rank those props at all — resolution there is zero by construction.
And the constant is hiding two *opposite* failures at once:
| stat · side | n | model says | actually wins | miscalibration |
|---|---|---|---|---|
| home_runs · under | 222 | 0.900 | **0.995** | **9.5 pts** (badly UNDER-confident) |
| hits · under | 27 | 0.900 | **0.519** | **+38.1 pts** (a coin flip sold as 90%) |
| total_bases · under | 12 | 0.900 | 0.583 | +31.7 pts |
| rbi · over | 35 | 0.100 | 0.229 | 12.9 pts |
**The same output `0.900` covers true probabilities from 52% to 99.5%.** A
near-lock and a coin flip are indistinguishable in the product. `PROB_CEIL = 0.95`
also makes it *impossible* to express the 99.5% case honestly.
Overall calibration, all settled MLB rows:
| stat | n | mean p_win | actual | over-prediction |
|---|---|---|---|---|
| **ALL POOLED** | 1,741 | 0.585 | 0.550 | **+3.5 pts** |
| total_bases | 297 | 0.534 | 0.458 | **+7.6** |
| hits | 589 | 0.606 | 0.562 | +4.4 |
| rbi | 320 | 0.399 | 0.356 | +4.3 |
| walks | 83 | 0.546 | 0.506 | +4.0 |
| runs | 124 | 0.596 | 0.605 | 0.9 (well calibrated) |
| home_runs | 228 | 0.899 | 0.996 | **9.6** |
Per the product doctrine, calibration is **half** the success criterion — "does
60% mean 60%?" Right now 58.5% means 55.0%, and on total_bases 53.4% means 45.8%.
**This is fixable with no new data at all.**
## 7. VERDICT PER STAT (STEP 4)
| stat | n | resolution | verdict |
|---|---|---|---|
| **hits** | 578 | 0.196 | **DILUTION + CALIBRATION.** Adjustments contribute nothing; one real lead (`opportunity_drift`) that we already compute and implement wrongly; +4.4pt over-prediction |
| **total_bases** | 284 | 0.237 | **DILUTION + CALIBRATION (worst).** Same lead; +7.6pt over-prediction |
| **rbi** | 273 | 0.499 | **DILUTION.** Removing home/away *improves* it (CI excludes zero). Prune |
| **runs** | 115 | 0.406 | **AT CEILING.** Adjustments neutral-to-harmful, no residual leads. The model is good here |
| **walks** | 66 | 0.478 | **AT CEILING** (underpowered, n=66) |
| **home_runs** | 228 | n/a — entirely clamped | **CALIBRATION.** 9.6pts, structurally uncorrectable while `PROB_CEIL=0.95` |
**AT CEILING is a real result here, not a shrug**: on runs and walks the champion
already resolves ~0.410.48 and nothing we compute explains its residual.
## 8. What this says about the next order
The convergent evidence was read correctly — the problem *is* inputs, not shape.
But the decomposition sharpens it, and the ranking is not what we assumed:
1. **The clamp + calibration (biggest, cheapest, no new data).** 20.6% of the
book pinned to a constant, a documented +3.5pt global over-prediction, and one
number covering 52%99.5%. This costs both halves of the success criterion.
2. **Opportunity, implemented properly** — the one feature with repeated
cross-stat residual signal. Not a new feature: a correct use of one we have.
Scale the rate, don't nudge the probability.
3. **Stop diluting** — the ladder and the environment axis add movement with no
information, and are measurably worse or flat.
4. **Archetype: measure before judging.** 2 of 41 labels have testable n.
**Explicitly NOT recommended:** another projection variant. That is the sixth
thing, and this diagnosis is why it would fail — the champion's edge is asking
the frequency question at the traded line, and every challenger so far has
replaced that question rather than improved its inputs.
## 9. Incidental finding, flagged not fixed
**`model_snapshots.outcome` is NULL on all 22,032 rows.** The retention table
built expressly so "a different model can be replayed against the same
conditions" stores features but was **never settled**, so it cannot answer the
question it exists for. This diagnosis worked around it by joining outcomes from
`ledger_entries` on (player_key, stat, line, side, game_date). Settling retention
would make every future ablation a single-table query — and would make refusals
(which the ledger drops) measurable for the first time.
## 10. Reproduce
`SUPABASE_URL=... node scripts/champion-ablation.js`
+124
View File
@@ -0,0 +1,124 @@
# The champion was reading ten games — repaired
## PHASE 0 — the defect is real past the peek
The prior baseline peeked at the evaluation window. This one does not: for every
prop the naive forecast is **that player's rate of clearing that line over games
strictly before that date**, from box scores back to 2026-05-01, requiring ≥10
prior games. Same temporal discipline the champion is held to.
| stat | n | champion | **fair PIT baseline** | gap | CI | loses |
|---|---|---|---|---|---|---|
| hits | 799 | 0.00251 | **0.00774** | 0.00523 | [0.0074, 0.0011] | **yes** |
| total_bases | 832 | 0.00393 | **0.00619** | 0.00226 | [0.0055, 0.0003] | **yes** |
| rbi | 501 | 0.02481 | **0.03133** | 0.00652 | [0.0153, 0.0005] | **yes** |
| runs | 473 | 0.00181 | **0.00683** | 0.00502 | [0.0114, +0.0008] | yes (CI touches) |
**Confirmed, not an artefact of the peek.** Three of four CIs exclude zero. The
served forecast was reliably worse than a frequency table.
---
## PHASE 1 — the cause: the window, not the weights
`estimateProbability` computes its base rate as the frequency over **every row it
is handed**. It was handed ten:
```js
// featureCache.getStatRows, MLB branch
const logs = res.last10; // <- the "season rate" was a TEN-GAME rate
```
So the forecast was `0.6 × (ten-game frequency) + 0.4 × (last five OF THOSE TEN)`
— a five-game read carrying 40% of the weight, on top of a ten-game base.
Resolution by variant, all point-in-time:
| stat | champion | season only | w=0.20 | w=0.40 | w=0.60 | best |
|---|---|---|---|---|---|---|
| hits | 0.00251 | 0.00774 | **0.00817** | 0.00688 | 0.00647 | w=0.20 |
| total_bases | 0.00393 | 0.00619 | **0.00734** | 0.00512 | 0.00485 | w=0.20 |
| rbi | 0.02481 | **0.03133** | 0.02727 | 0.02571 | 0.02559 | season only |
| runs | 0.00181 | **0.00683** | 0.00436 | 0.00318 | 0.00180 | season+nudge |
**The 0.40 recency weight costs resolution on all four stats** (0.00086,
0.00107, 0.00562, 0.00365). The nudges are mixed and small: harmful on hits
(0.00157) and rbi (0.00284), marginally helpful on TB (+0.00091) and runs
(+0.00056) — left alone, since the evidence does not support removing them.
---
## PHASE 2 — the repair
Two lines, no new data, no extra API call — **`fullLog` was already being fetched
by the same adapter call that produced `last10`**:
1. `featureCache.getStatRows` MLB branch reads `fullLog`, falling back to
`last10`.
2. `RECENCY_WEIGHT` 0.40 → **0.20**, set at the value the measurement supports.
| stat | OLD | **REPAIRED** | fair baseline | vs baseline | CI | vs old champion |
|---|---|---|---|---|---|---|
| hits | 0.00251 | **0.00817** | 0.00774 | **+0.00043** | [0.0030, +0.0025] | +0.00566 |
| total_bases | 0.00393 | **0.00734** | 0.00619 | **+0.00115** | [0.0014, +0.0046] | +0.00341, **CI [0.0020, 0.0067]** |
| rbi | 0.02481 | 0.02727 | 0.03133 | 0.00406 | [0.0091, +0.0020] | +0.00246 |
| runs | 0.00181 | 0.00436 | 0.00683 | 0.00247 | [0.0088, +0.0014] | +0.00255 |
**Hits resolution tripled; total_bases and runs roughly doubled.**
**Gate assessment, stated exactly:** hits and total_bases now exceed the fair
baseline on the point estimate; rbi and runs remain below it but **every CI now
includes zero.** So no stat *reliably loses* to a frequency table any more, which
satisfies "beat or tie, never lose" in the only sense this sample can support. It
is a tie on rbi/runs, not a win, and it is reported as one. Only total_bases'
improvement over the old champion is CI-confirmed; the rest are directional.
### The stale-fit gate — calibration is OFF
The low-parameter maps were fitted on the retired forecast, and `fromLedger`
cannot rescue them: settled ledger rows still carry OLD `p_win` values, so
refitting today would fit the retired forecast again.
**`CALIBRATION_DEPLOYED` is now empty.** Nothing is served calibrated until
enough dates settle under the repaired champion, and the favourite-longshot bias
must be **re-measured** on the new forecast rather than assumed to have survived.
The shadow duel is likewise void. Serving the raw repaired number is the honest
state, not a regression.
---
## PHASE 3 — the hits factor lift, NOT re-measured
Honest answer: it **cannot** be measured yet. The three proven hits factors were
measured against the old baseline, and re-measuring their lift on the repaired
champion requires settled rows produced *by* the repaired champion. Those do not
exist — the repair ships in this commit. Replaying it would score the factors
against a reconstruction rather than the served forecast.
**Deferred to the first order after the repaired champion has settled dates.**
The factors remain wired and transmitting (43f65d3, sign-verified, 75% coverage);
only their *lift* is unquantified on the new baseline.
---
## PHASE 4 — log and re-queue
**Standing flag, and it is a large one:** every factor verdict in this
programme — every null, every THEATER — was measured against a champion that was
worse than a frequency table. Signal added to noise reads as noise. **Prior
verdicts may deserve re-audit on the repaired champion.** Not re-run here; logged
as standing.
**Re-queued, not built — rbi lineup-slot / RISP opportunity** through the
two-part gate, now landing on a repaired champion. World A ~90%, within-role
residual 0.01908 real, `lineup_context` ingested and prod-verified (S89). That is
the next factor order.
---
## Invariants
Serving-path change by design — the byte-identical invariant inverted again, and
all four stats' numbers move. Nine frozen model modules verified unchanged.
`p_win` is the forecast itself, not mutated post-hoc. No Bonferroni slot: this is
resolution accounting on the champion's own knobs, not a causal claim.
+111
View File
@@ -0,0 +1,111 @@
# The collapsed sequence edge — two proven links whose product is too small to use
**Both components are real. Their product is 0.37pp, and detecting it would take
52 seasons.** This is the most instructive negative in the programme so far,
because nothing in it failed: link-by-link proof did not produce a usable edge.
---
## One correction to the framing
The order states Link 2 proved you *can't* predict the reliever. Half true, and
the other half matters: the **individual** grain failed (17.2% accuracy), but the
**quality** grain PROVED — predicted pen quality separates 2.70pp of realized hit
rate. Pen-season-quality here is a measured predictor, not a fallback after a
failure. That strengthened the plan going in.
## Two of the three specified inputs could not be used honestly
| specified | status |
|---|---|
| pen **quality** | PROVEN (Link 2 coarse grain) — used |
| pen **archetype** | did NOT prove (0.5669 vs 0.5309 modal baseline, corrected interval spanning zero) — **excluded**, building it in would chain on an unproven link |
| hitter **approach identity** ("fastball-hunter", "finesse-vulnerable") | **does not exist** in this registry. MLB batter archetypes are BOMBER / GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID / CATALYST. Inventing an identity to condition on is the fabrication the gate exists to catch |
A hitter power/contact split derived from the sequence data itself was tested as a
**separate gated addition** rather than assumed into the main effect. Neither half
proved (power 0.0002, contact 0.0001, both intervals spanning zero).
---
## The gate — two-part, on the concentrated subset, 114 cumulative tests
| subset | n | games | mean shift | Brier Δ | CI | verdict |
|---|---|---|---|---|---|---|
| concentrated (early-exit × WEAK pen) | 1,931 | 141 | 0.0193 | 0.0001 | [0.0014, +0.0010] | NOT_PROVEN |
| mirror (early-exit × STRONG pen) | 2,574 | 189 | 0.0186 | 0.0000 | [0.0011, +0.0010] | **THEATER** |
| all early-exit later ABs | 6,869 | 451 | 0.0140 | 0.0001 | [0.0007, +0.0005] | NOT_PROVEN |
| pooled all later ABs | 17,891 | 803 | 0.0141 | 0.0000 | [0.0004, +0.0003] | **THEATER** |
Not pooled-diluted: the concentrated subset was gated on its own and is no
better. Two subsets are THEATER by the gate's own definition — the adjustment
moves the number ~1.9pp and improves accuracy by essentially nothing.
---
## Why: the mechanical ceiling
The descriptive pass found the direction the order predicted (early-exit + weak
pen +0.74pp, early-exit + strong pen 0.79pp vs a deep-starter baseline). The
signs are right. The magnitude is the problem, and it is structural:
```
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925 <- NOT a switch to the pen
hit-rate swing across pen quality 0.0394 (weak 0.2491 vs strong 0.2038)
MAX JUSTIFIABLE ADJUSTMENT = 0.0925 x 0.0394 = 0.00365 (0.37pp)
adjustment actually applied (mean |shift|) = 0.01930 (1.93pp)
OVER-MOVEMENT FACTOR = 5.3x
```
**A hitter's third or fourth plate appearance is ALREADY against the bullpen 71%
of the time even when the starter is projected to go deep.** Link 1 lifts that to
81%. It buys nine points of extra pen exposure, not a change of opponent — so any
adjustment riding on it is capped at about a tenth of the pen-quality swing.
The 5.3× over-movement is exactly why the mirror subset reads as THEATER rather
than as a small true effect: the adjustment asserts five times more than the
mechanism can support.
### And a correctly-scaled version is not detectable either
```
concentrated subset n = 1,931 SE of hit rate = 0.00985
max justifiable effect 0.00365 = 0.37 SE
to detect at the corrected bar (z~3.46 for 114 tests): n = 168,488
shortfall 87x -> ~52 seasons of concentrated-subset accrual
```
**This line is structurally closed, not sample-blocked.** Waiting does not fix it.
---
## Not wired, and the self-check deliberately not wired either
The adjustment does not prove, so it feeds nothing. The order also asks for a
self-check flagging where our sequence read diverges from the line's
starter-script, as an opportunity signal. **That is not wired**, because flagging
divergence on an adjustment measured as absent would advertise an edge we have
just shown does not exist — the same failure as fabricated reasoning, one layer up.
## The lesson worth keeping
Link 1 proved (MAE 3.22 → 2.80 batters faced). Link 2's quality grain proved
(2.70pp of realized separation). Both are real, both are point-in-time, both
survived cumulative correction. **Their product is still too small to use.**
Link-by-link validation guarantees each link is real. It does not guarantee the
chain transmits anything. The multiplicative structure has to be sized BEFORE
building — one exposure term of 0.09 is enough to reduce a genuine 3.94pp signal
to noise, and no amount of downstream care recovers it.
## Parallel track — total_bases per-archetype (logged, not run)
Unchanged: `total_bases` settled n=948 pooled, BOMBER × TB **340**, short by 160.
Sample-readiness only, not a verdict. The `specs/per-archetype-grade-bands.md`
blocker still stands — the grade does not yet separate within any archetype.
Link 3 confirmed SKIPPED. Counter and frozen clusters byte-identical.
+147
View File
@@ -0,0 +1,147 @@
# THE CONDITIONING REGISTRY — built, and what it currently holds
**2026-08-03.** Challenger-only. Counter, batter model and pitcher engine
byte-identical (verified by diff).
> **No archetype × stat combination reaches the gate. The best is BOMBER × hits
> at n=287, short by 213.** So the conditioning categories were tested and are
> all UNDERPOWERED — none proved, none died. Nothing was recalibrated and nothing
> shipped, because nothing earned it.
>
> **The durable deliverables are the registry itself and a status probe** that
> makes "what is proven" a query instead of a memory.
---
## 0. The recurring premise problem, and a structural fix
This order opens with "two proven clusters live (batter contact stats + pitcher
strikeouts)". They are not proven. **The proven set is empty**, and this is the
fourth consecutive order to start from a stronger claim than the measurements
support:
| order said | measurement said |
|---|---|
| "barrel rate PASSED solo" | every total_bases feature refused on sample |
| "total_bases has passed BAR 1" | inconclusive at parity, CI spanning zero |
| "whiff/stuff prove SOLO through the gate" | all refused at n=57 |
| "two proven clusters live" | **proven set EMPTY** |
Correcting it in prose four times has not worked, so this session added
**`scripts/proven-status.js`** — it recomputes the answer from the ledger:
```
PROVEN_SET: EMPTY — no stat has beaten the counter out-of-sample
with a CI excluding zero
hits n=803 delta 0.0961 [0.165, 0.029] LOSES
total_bases n=383 delta +0.0038 [0.068, +0.075] INCONCLUSIVE
strikeouts n=57 delta +0.2592 [0.017, +0.564] INCONCLUSIVE
stats at/above the gate: hits (806) — and hits is a closed negative
archetype × stat at/above the gate: NONE
closest: BOMBER×hits 287 (short 213) · BOMBER×TB 142 · BOMBER×rbi 128
```
**Run it before planning on top of a claim.** It deliberately cannot say
"proven" on its own — it reports sample readiness and *recorded* verdicts, so the
two can never be conflated again.
## 1. The structured registry (STEP 1) — built
`featureRegistry.recordConditioning()` keys **archetype × underlying-skill ×
conditioning-interaction × status**, with measured lift.
**The skill tag is mandatory and enforced.** An untagged entry is refused
(`untagged_or_unknown_skill`), and a `PROVEN` entry without sufficient evidence is
refused (`insufficient_evidence_for_proven`). Skills: POWER, CONTACT, SPEED,
WHIFF, COMMAND, OPPORTUNITY.
Why the tag matters: a proven interaction is not merely "this helps this stat" —
it is evidence that **one underlying skill is real and measurable for this
archetype**. `validatedSkills(sport, archetype)` returns the coherent profile as
it currently stands. **It returns `{}` for every archetype**, because nothing has
been proven, and seeding it with hopeful rows would defeat its purpose exactly as
seeding PROVEN features would.
## 2. Top-volume selection (STEP 2)
**Batter: BOMBER** — the highest-volume archetype by a distance (287 settled hits
rows; next is GHOST×hits at 124).
**Pitcher: none testable.** Strikeouts total 58 settled rows across *all*
archetypes, so no pitcher archetype has a sample. STEP 4 could not be run.
### A counting error caught, worth recording
The first read said BOMBER × hits was **641** — gate-ready. It is **287**. The
join to `model_snapshots` fans out, because that table holds one row per prop
**per snapshot cycle**, so each ledger row was counted once per cycle it appeared
in. Deduping on the ledger row's identity gives the true figure. **That is the
difference between "run the gate" and "not close", and my own status script had
the same bug until it was fixed.**
## 3. BOMBER × hits conditioning (STEP 3) — all underpowered
n=282, Bonferroni across 17 tests. Every result refused on sample.
| conditioning | category | n | raw r | best part | **incremental** |
|---|---|---|---|---|---|
| barrel × breaking share | **ARSENAL** | 282 | 0.063 | 0.105 | +0.043 |
| launch × pitcher GB% | **BATTED-BALL** | 282 | 0.054 | 0.052 | +0.001 |
| launch × exit velo | contact quality | 282 | 0.060 | 0.087 | 0.020 |
| exit velo × pitcher suppression | contact quality | 282 | +0.008 | 0.087 | 0.015 |
| batter K × pitcher K | opportunity | 282 | 0.080 | 0.077 | 0.063 |
Solo, within BOMBER, the strongest is `batter_barrel_pct` at 0.105 (p=0.077);
`pitcher_breaking_share` is +0.025. **Head-to-head within BOMBER: 0.1599 vs the
counter's 0.2176 — the counter still leads on hits even inside its best
archetype**, consistent with the closed pooled negative.
**A bug fixed mid-run:** `pitcher_breaking_share` first reported **n=0** for every
row. `fromStatcastRow` maps percentage and raw fields only — it does not carry
`pitch_mix` — so the arsenal category was silently measuring nothing rather than
failing. Attaching the mix explicitly gave full coverage. Had it not been caught,
"arsenal doesn't matter" would have been recorded from a column that was never
populated.
### DEFENSE — the honest answer after looking
**We ingest no fielding data at all.** `statcast_aggregates` holds batter
offensive skill and pitcher stuff; there is no OAA, DRS, range, or positional
metric anywhere in it. I checked for a derivable proxy before declaring it
unsourceable, and the candidates all fail on construction:
- opposing pitchers' hits-allowed conflates *pitching* with *defense*, so it
would validate the wrong skill and could quietly "prove" defense using pitching;
- there is no team-level balls-in-play or expected-vs-actual column to difference.
**So defense is genuinely not derivable from what we hold** — it needs Baseball
Savant's fielding endpoint (free, same host as the five feeds already ingested,
so it is cheap). **Not sourced this order**, because sourcing it to test at n=282
would answer nothing.
## 4. Ship + recalibrate (STEP 5)
**Nothing proved, so nothing was recalibrated and nothing shipped.** The counter
continues to grade everything. The registry records the tested interactions as
CANDIDATE with their measured lift, so re-running at n≥500 compares against a
recorded baseline rather than starting over.
## 5. What actually unblocks this
Everything is one constraint: **sample per archetype**. Two things move it:
1. **The cap fix is already compounding** — 907 grades/snapshot vs 334, so
archetype cells fill ~2.7× faster than the rates that produced today's counts.
BOMBER × hits needs 213 more rows.
2. **A point-in-time window** from `statcast_history`, which starts producing
usable comparisons 2026-08-04.
**Ranked next:** BOMBER × hits (closest by far) → BOMBER × total_bases → pitcher
archetypes once strikeouts clear. **Add the Savant fielding feed before the
defense category is tested**, not before it can be.
**Not recommended:** recording anything as proven, sourcing defense to test at
n=282, or reading the arsenal incremental (+0.043) as encouraging — it is inside
noise at this sample, and the category only became measurable at all because a
silent n=0 was caught.
+173
View File
@@ -0,0 +1,173 @@
# CONNECT PROJECTION LAYERS — STEP 0: OPPORTUNITY/USAGE INPUT CHECK
**Date:** 2026-08-01 · **READ-ONLY** · live grade path byte-identical ·
probe committed (`src/services/featureCoverage.js`,
`GET /api/internal/feature-coverage`).
---
## VERDICT: STOPPED AT STEP 0 — AND THE REASON IS THE DELIVERABLE
**Inputs are 100% populated. The layer still should not be wired as specified**,
for four reasons the input check surfaced. Three of them would have made the
work either unmeasurable or wrong.
---
## 1. INPUT COVERAGE (MLB, n=80 real props, through the grader's own path)
| feature | populated | rate |
|---|---:|---:|
| `ab_per_game` | 80/80 | **100%** |
| `rest_days` | 80/80 | 100% |
| `l5_avg` | 80/80 | 100% |
| `l20_avg` | 80/80 | 100% |
| `l10_stddev` | 80/80 | 100% |
| `game_count_in_7d` | 80/80 | 100% |
| `opp_rank_stat` | 52/80 | 65% |
| `minutes_per_game` / `usage_rate` | 0/80 | 0% *(NBA-only — correct for MLB)* |
**No honest-degradation problem exists** for the opportunity input: there is
nothing sparse to fall back from. The one real hole is `opp_rank_stat`, which is
**0% for `stolen_bases`** (no opponent SB-defence rank) and ~84% elsewhere.
---
## 2. 🔴 THE PREMISE: there is no built opportunity layer to connect
The order says the expensive layers are *"BUILT and DISCONNECTED"* and that this
is *"CONNECTION, not construction."* **For opportunity, that is not the case.**
`ab_per_game` is computed in `featureCache` and consumed in exactly one place —
`analyzeViaEngine1:379`, which renders **"4.3 AB/G"** on the grade card. It is a
**display field**. `engine1` has **no opportunity or usage factor at all**.
**A projected opportunity — expected plate appearances with its own uncertainty —
was never built.** §10.3 of MASTER-PLAN said exactly this and called it *"the
single biggest modelling upgrade available."* Building it is construction.
## 3. 🔴 The available input is the wrong shape for the job
`ab_per_game = season atBats ÷ games`. Two consequences:
- **It is a per-player constant.** Measured: it varies across a player's own
props for **3 of 20 players** (and those three are almost certainly a cache /
refresh artifact, since the value cannot legitimately depend on `stat_type`).
**A constant can only move all of a player's props together — it cannot
separate them**, which is what a per-prop opportunity signal has to do.
- **It is collinear with the projection that already exists.** `l20_avg` is
`seasonTotal ÷ games` — the *same denominator*. For a batter,
`hits/game ≈ (hits/AB) × (AB/game)`, so **`l20_avg` already embeds
opportunity multiplicatively.** Adding `ab_per_game` as an independent additive
factor double-counts it rather than adding information.
**What is actually missing is tonight's deviation from the season baseline**
batting-order slot, a platoon sit, a role change.
## 4. 🔴 That input does not exist in any wired source
- `depthChartService.getLineup` returns, for MLB, **only the probable pitcher**,
with `battingOrder: null`. Its own comment: *"the one lineup slot the free
schedule feed exposes."*
- PropLine `/context` (free, verified this session) carries `lineup_confirmed`
**a boolean**, not the order.
**Tonight's batting order, the actual driver of MLB plate appearances, is not
available from anything we have wired.**
## 5. 🔴 ARCHITECTURE: wiring it into `engine1` would be unmeasurable by this order's own test
Step 2 requires proving **reliability and resolution** improve. Both are measured
on **`p_win`**.
**`engine1` factors move the grade LETTER. They do not touch `p_win`.**
`p_win` comes from `probabilityEstimator.estimateProbability`, which builds from
`frequencyOver(gameLogs, line)` plus feature adjustments.
So an opportunity factor added to `engine1` would produce a change that **Step 2
literally cannot measure**. The layer belongs in `probabilityEstimator` (or the
challenger below), which *does* consume features — it already adjusts on
`opp_rank_stat`, `home_away`, and a consistency pull off `l10_stddev/l20_avg`.
---
## 6. THE SEQUENCING ASSUMPTION IS STALE — matchup is already connected
The order sequences *opportunity → matchup granularity (archetype × opponent,
park/weather/platoon) → distribution ladder.*
**`challengerProjection` (`arch-v1`) is already live in production** and already
carries three axes: **archetype**, **matchup (platoon)** and **environment
(park)**. It writes `p_win_challenger`, `challenger_delta`,
`challenger_adjustments` and `challenger_version` to the ledger on every graded
prop.
So step 2 of the sequence is **partly done**, and — more usefully — **the harness
this order needed already exists.** Any new axis should be added there, not
invented.
---
## WHAT I RECOMMEND INSTEAD (its own order, per "one layer at a time")
**An `opportunity` axis on `challengerProjection`, driven by opportunity DRIFT
rather than by the season level:**
```
opportunity_drift = recent AB/G (last 5) ÷ season AB/G
```
- **>1** — batting higher / playing more than his baseline → lean over
- **<1** — reduced role, platoon, lower slot → lean under
- **null** — no at-bat data → **no adjustment**, challenger ≡ champion on that row
This is a **deviation**, so it is not collinear with `l20_avg` the way the raw
level is, and it *does* vary per player over time.
**The input exists but is not extracted.** Per-game `atBats` is present in the
statsapi game-log rows (the Session-56 box-score audit lists `atBats` among the
batting fields), but `MLB_LOG_FIELD` has no entry for it and nothing computes a
recent AB/G. That is a small, contained build — **and it is a build**, which is
why it belongs in its own order rather than being smuggled into a connection
order.
**Honest caveat to carry into it:** this is still a *proxy* for tonight's
opportunity, not tonight's opportunity. The real input is the confirmed batting
order, and that needs a lineup source we do not have.
---
## WHAT DID NOT HAPPEN, DELIBERATELY
No layer was wired. **The live grade path is byte-identical.** No threshold moved,
no challenger was added, no holdout was run — running one would have measured a
change that could not have occurred.
## A PROBE BUG WORTH RECORDING
The first coverage run reported **0% for every feature, including `l5_avg`** — on
a pipeline that had just graded 365 props, which `projectionFor` cannot do
without a positive `l5_avg` or `l20_avg`. **Impossible, therefore the probe was
wrong.**
Two bugs, both mine: `featureCache.getFeatures` takes **camelCase**
(`playerName`/`statType`) and I passed the prop's snake_case shape; and it
returns **`{ features: {...} }`** while I read the top level. Either alone yields
all zeros. Fixed by calling `computeFeaturesForProp` — the grader's own entry
point.
Same class as the earlier harness that returned a silent `false`: **a measurement
that makes working code look broken is more dangerous than no measurement**, because
it invites you to "fix" something that was never broken.
## TAGS
**VERIFIED:** `ab_per_game` 100% populated · per-player constant (varies for
3/20) · `opp_rank_stat` 0% on stolen_bases · engine1 has no opportunity factor ·
`ab_per_game` is display-only · MLB batting order unavailable in depthChart and
`/context` · `probabilityEstimator` consumes features, `engine1` does not affect
`p_win` · `challengerProjection` arch-v1 already carries archetype/matchup/
environment.
**CORRECTED:** "the opportunity layer is built and disconnected" — it is not
built. "Connect opportunity, then matchup" — matchup is already connected.
+68
View File
@@ -0,0 +1,68 @@
# D1 CLOSE — REVIEW ZERO FINDINGS (report; mount NOT performed)
2026-07-31. Nothing changed: no mount, no row edit, no data threading.
`git diff` = docs only.
## Why this stopped at Review Zero
The order scopes this as "mount only — NO row refactor." Reading the real components shows the
mount is **not** a mount, and one finding needs a decision before anything is surfaced.
### 0.1/0.2 — THE RATIONALE DOES NOT REACH THE ROW (VERIFIED)
`StripProp` (`components/vyndr/StatStrip.tsx`) carries stat · line · side · grade · gradedAt ·
delta · awaiting · outcome · movement · revisedFrom · book · bestBook · dead · history —
**no `reasoning`, no `kill_conditions_triggered`.** `buildPlayerStripsFromProps` never threads
them either. So mounting the hover requires **adding a field to the strip contract and threading
it through the slate adapter** — additive, but a data-path change, not a mount.
### 0.3 — THE PREMISE INVERTS: THE RATIONALE IS **ALREADY PUBLIC** (VERIFIED LIVE)
The order asks me to ensure "a free-tier user's hover shows nothing, not the paid reasoning."
Measured against prod, anonymously:
GET https://api.vyndr.app/api/snapshot/wnba (no auth)
reasoning.summary present : True
reasoning.locked : None
kill_conditions_triggered : 1
**`stripModelPrice` strips `model_odds`/`p_win`/`ev_pct`/`value`/`takeable` — it does NOT strip
`reasoning` or `kill_conditions_triggered`.** So the full model rationale is already shipped to
every anonymous browser on the main board endpoint.
**Two consequences, and they pull in opposite directions:**
1. Mounting the hover **leaks nothing new** — it renders bytes the browser already has.
2. But it **surfaces** something currently shipped-but-unrendered, and the same content **IS**
tier-gated on the scan path (`utils/tierGating.js` redacts `reasoning` for free tier). So the
product currently gates the reasoning in one place and serves it openly in another.
**This is a monetization/consistency decision, not an implementation detail** — which is why it
is reported rather than resolved unilaterally. Three coherent options:
- **(a)** Gate `reasoning` on `/api/snapshot` the way scan does, then mount the hover for entitled
viewers only. Consistent, but removes content free users already receive.
- **(b)** Accept it as intentionally free (board reasoning is the funnel) and mount for everyone.
Consistent the other way; makes the scan-path gating the odd one out.
- **(c)** Leave the payload alone and don't surface it. Status quo.
### THE THIRD BLOCKER — ROW-GRAMMAR IS LAW AND LOCKS StatStrip's SOURCE ORDER (VERIFIED)
`tests/unit/rowGrammar.test.js` asserts StatStrip's element order by **`src.indexOf(...)` on the
component source** (slots: identity → viability → stat+line → market context → model output →
outcome → actions → provenance). Adding a rationale affordance or a team chip **moves those
offsets**, so `specs/ROW-GRAMMAR.md` and the test must be amended **in the same commit** — per
CLAUDE.md, that spec is LAW. That makes this a spec-amending change, not an additive mount, and
the amendment needs its own slot decision (where does a rationale affordance sit in the grammar?
where does a team chip sit relative to identity?).
## WHAT IS SAFELY MOUNTABLE WITHOUT ANY OF THE ABOVE
- **`reveal.js`** — wraps the row LIST, touches no StatStrip internals, no new data, no grammar
slot. This one is genuinely a mount.
- **`teamChips.js`** — needs a grammar slot (it sits before the team abbr, inside identity), so it
is small but spec-touching.
- **`rowRationale.js`** — needs the data threading AND the gating decision AND a grammar slot.
## RECOMMENDATION
Split D1-close into: **(1)** mount `reveal` now (no blockers), **(2)** a ROW-GRAMMAR amendment
order that adds the team-chip slot + the rationale-affordance slot to the spec and test, and
**(3)** the rationale mount, after the gating decision above is made.
## TAGS
VERIFIED: strip contract lacks reasoning; anon snapshot serves reasoning + kills; ROW-GRAMMAR
locks source order. CANNOT DETERMINE: none. **BLOCKED: the rationale mount — on a
gating decision (a/b/c) and a ROW-GRAMMAR amendment.**
@@ -0,0 +1,138 @@
# DEFENCE INGESTED + CUMULATIVE CORRECTION LOCKED
**2026-08-03.** Challenger-only. Counter, batter model and pitcher engine
byte-identical (verified by diff).
> **Two durable things shipped, and they stand regardless of sample:** the free
> Statcast fielding feed is ingested and persisted (514 fielders, 31 teams,
> verified in production), and Bonferroni is now corrected against the
> **programme's lifetime test count**, not the session's.
>
> **The differential the theory predicted actually appears:** defence correlates
> with the counter's residual for GHOST (contact/speed, **+0.130**) and is flat
> for BOMBER (power, **0.018**). That is "defence matters, and for whom" showing
> up in the data — at n=104 and n=245, so it is a signal shape, not a result.
>
> **Nothing proved. Nothing recalibrated. Nothing shipped.**
---
## 1. Defence ingested (STEP 1)
Statcast Outs Above Average, free, same host as the six feeds already pulled.
```
feed rows 514 fielders
team defence 31 teams
best Cubs oaa_sum +56
worst Mariners oaa_sum 29
prod verified fielding rows 514 · team_defense_written 31
```
Stored per fielder (`oaa`, `runs_prevented`, `success_diff`, position, team) and
aggregated to **team level**, which is the unit a batter's prop needs: the
defence behind the pitcher he faces. Summed OAA is the team's outs converted
above average; the mean rides along because a team with more measured fielders
would otherwise look better merely for being measured more. Under three measured
fielders → **absent**, not thin.
**Unknown is not zero, and it bites unusually hard here.** An OAA of 0 is a REAL
reading meaning *exactly average*. Coercing absence to 0 would assert that every
unmeasured fielder is league-average — the most common defensive profile there
is — which is a fabricated fact wearing the costume of a neutral default. Every
read goes through `knownRate`/`knownNumber`.
**`team_defense` carries `as_of_date` in its primary key from the first row.**
`statcast_aggregates` was built upsert-in-place with a single as-of date, which
silently made every backtest leak the games it was predicting and cost a full
session to discover. Point-in-time is available here *before* it is needed.
### A bug worth recording as a class
The first prod run reported **`fielding_oaa: 0 rows`**. `BASE` already ends in
`/leaderboard`, so the new feed built `.../leaderboard/leaderboard/...` and 404'd.
Because a failing feed **degrades to an empty index by design** — correct, so one
broken source cannot fail the whole mechanism pull — it surfaced as *zero
fielders*, which reads exactly like "Statcast has no fielding data."
**Graceful degradation makes a wiring bug look like an honest absence.** Any feed
reporting 0 should be treated as suspect until the URL is fetched by hand.
## 2. Defence conditioning per archetype (STEP 2)
Cumulative Bonferroni (see §3): denominator **38**, corrected α = **0.0013**.
| archetype | n | **defence solo r** | p | defence × contact (incr) | defence × launch profile (incr) |
|---|---|---|---|---|---|
| **GHOST** (contact/speed) | 104 | **+0.130** | 0.188 | 0.148 | +0.029 |
| **BOMBER** (power) | 245 | **0.018** | 0.782 | +0.030 | +0.025 |
**The differential is the point, and it is present.** Defence carries a signal
against the counter's residual for the contact/speed archetype and essentially
nothing for the power archetype — which is the causal story: a GHOST's hits
depend on whether anyone can range to the ball, while a BOMBER's barrels clear
the defence entirely. **A flat BOMBER result is the theory working, not the test
failing.**
**But neither is a result.** GHOST is n=104 against a 500 bar with p=0.188 against
a corrected α of 0.0013 — three orders of magnitude short. The direction matches
the theory, which is worth carrying forward; it is not worth acting on.
Both recorded in the registry as CANDIDATE with measured lift, tagged
`CONTACT`-skill, so re-running at n≥500 compares against a recorded baseline.
## 3. Cumulative multiple-comparisons correction (STEP 4) — locked
`src/services/model/testLedger.js` + `mc_test_ledger`.
Bonferroni had been applied **per session** throughout: a run testing 8 features
corrected by 8. Across a programme's lifetime that is wrong in the dangerous
direction — every order gets a fresh, generous alpha, so the false-positive rate
compounds quietly. **Correcting by 8 when sixty have been tried is exactly how a
noise result eventually gets recorded as PROVEN, with a p-value to point at.**
The denominator is now the count of **distinct hypotheses ever tested**,
persisted. Demonstrated live this session:
```
GHOST run → cumulative 19 (19 new)
BOMBER run → cumulative 38 (19 new) corrected α: 0.0026 → 0.0013
```
**Re-tests do not inflate it.** Re-running the same hypothesis on more data is
the same question asked again, not a new shot on goal — counting it again would
punish the discipline of waiting for sample, which is the behaviour this
programme depends on. `times_tested` increments; the denominator does not.
**The alpha only ever shrinks**, which is the correct ordering: an interaction
proved late has cleared a genuinely higher bar than one proved on day one,
because by then we have had far more chances to get lucky. Seven tests lock this.
## 4. Ship + registry (STEP 5)
**Nothing proved → nothing recalibrated, nothing shipped.** The counter continues
to grade everything. `validatedSkills()` returns `{}` for every archetype.
## 5. Programme state
| | status |
|---|---|
| proven set | **EMPTY** (`node scripts/proven-status.js`) |
| gate-ready archetype × stat | **none** — BOMBER×hits 287 is closest, short by 213 |
| defence | **ingested**, testable, underpowered |
| cumulative correction | **locked**, α now 0.0013 and falling |
| point-in-time window | `statcast_history` 1 day; `team_defense` dated from row one |
## 6. Next
1. **Sample is still the only constraint.** The cap fix (907 grades/snapshot vs
334) is compounding it; GHOST × hits needs ~396 more rows, BOMBER × hits ~213.
2. **Re-run `scripts/cluster-prove.js` per archetype at n≥500.** The GHOST
defence differential is the single most theory-consistent signal the programme
has produced — it deserves a fair test, and it will get one.
3. **Note the moving bar:** every new hypothesis tightens α for everything that
follows. Prefer re-testing the standing candidates over inventing new ones —
that is now mathematically, not just methodologically, the disciplined choice.
**Not recommended:** reading GHOST's +0.130 as evidence, recalibrating anything,
or adding new hypotheses while the standing ones are unresolved.
+22 -5
View File
@@ -1,10 +1,29 @@
# VYNDR VISUAL SYSTEM — CODE HANDOFF # VYNDR VISUAL SYSTEM — CODE HANDOFF
## SESSION 3 (Jul 22) — Scanner surfaces S6 + S7
`Vyndr Scanner States.dc.html` — spec for the two scanner states Code flagged missing (converges at the scanner-nudge build order).
- **S6 scan empty (marketless scan)**: read-card frame, ZERO market numbers (no empty triplet frame — it doesn't render), chip `NO MARKET` in the blue boundary channel (`rgba(106,147,200,.12)` bg / `.34` border / text `#8FB2DE` — same tokens as PRICED OUT), dashed blue void box (refusal frame shape, blue not red: nothing failed), copy "no live market for this line" (never "coming soon"). Path forward built in: green primary CTA to tonight's board with live priced-prop count; player-specific path (`HERBERT · 3 PRICED PROPS ▸`) leads when the player is on the slate. Mobile 390: both paths 44px.
- **S7 priced-line display (the nudge)**: strip under the scan field showing ONLY genuinely-priced lines per exact player+stat. Case A `NONE PRICED` is the DEFAULT/common case — quiet dashed blue box + "we price these for [player]" stat chips (with amber fair previews). Case B priced: 44px rows `o 27.5 · BOOK 114 · ◆ FAIR 105 · OPEN READ ▸`, one tap to the real triplet. Case C off-slate: text-tokens-only fact line, no blue, no chips, stops. Typed-line mismatch: query text preserved, `o 30.5 ISN'T PRICED` blue fact line (not an error), nearest priced line = same player+stat only.
- New color law: **blue #8FB2DE = the boundary channel** — the system being honest about what it can't hand you (edge priced out, no market, line not priced). Distinct from red REFUSAL (model decision) and amber QUARANTINE (model leg suppressed).
## SESSION 2 (Jul 17) — Offseason & Intelligence expansion
Two new design files, same package format. System law unchanged (tokens, grade colors, glow-A-only, THE VERB IS "READ").
- `Vyndr Offseason.dc.html` — Surface 1. Hub home (NFL desktop + 390, NBA Summer-League variant with OUTLOOK ONLY / NOT GRADED honesty block), season-long board (open→NOW→VYNDR triplet — open `#707080`, now `#F0F0F0` 700, model `#00D4A0` 800; ladder grouping: player header row + nested market lines at `padding-left:52px`), season read reveal (news-annotated movement, NEWS-DRIVEN FACTORS, WHAT WOULD CHANGE THIS READ kill conditions), news/outlook feed + row anatomy + quiet-wire empty state. Sport state lives IN the sport tab (`NFL · CAMP 5D`), never a separate offseason tab. Kickoff countdown is computed live from the real date (`Sep 10 2026`) — ambient, top-right, never the hero. New mobile tab bar: SLATE · EXPLORE · READ-FAB (50px green circle, void-colored V, `translateY(-14px)`, 6px void ring) · LEDGER · MORE.
- `Vyndr Intelligence.dc.html` — Surfaces 25 + M2.
- S2 line shopping: MOVEMENT STRIP primitive defined once (row 86×20 silent / reveal full-width annotated / laws: steps not curves, green only when move favors the read, FLAT = hairline + `FLAT · [N]D`). Book comparison: 7 BookChips (24px tile, brand-color mono wordmark, `image-slot` overlay for licensed logos — DK `#53D337`, FD `#1493FF`, MGM `#C4A45E`, CZR `#AD9660`, B365 `#FFE100`, ESPN BET `#D50A0A`, Fanatics `#2E6BFF`), best number crowned (green inset bar + BEST NUMBER chip + row tint `rgba(0,212,160,.04)`), disagreement mark (dots on an axis around the VYNDR reference line; SPLIT chip amber when spread wide). Push-to-book: primary green deep-link + secondary book buttons, geo-gated (buttons never render disabled — they don't render), affiliate microcopy verbatim. Geo-empty honest state included.
- S3 article media: hero template + 4 hero graphic archetypes (A line path · B distribution · C mark-at-scale · D matchup card — all generated data visuals, never stock), inline figures (stat callout triptych, comparison bars, pull quote — one max, caption law: every figure names its data), article card (thumb = hero graphic re-cropped), article OG 1200×630.
- S4 CLV: per-read chip `BEAT CLOSE +2.0` / `MISSED CLOSE 1.0` — color follows CLV sign, never the result; aggregate joins record line (`BEAT CLOSE 61%` + `AVG CLV`); no-data state names the tracking start date + N≥100 threshold, never a fake 0%. Calibration curve: headline claim ("When we say 70%, it hits 68.2%"), dots vs dashed perfect line, dot size = sample, buckets under N 30 render hollow/unlabeled.
- S5 The Report: hybrid email — dark billboard header + dark record band survive every client; light paper body (`#F7F6F2` / ink `#1A1A22` / green shifts to `#00A57D` on paper). 600px single column, system-font fallbacks, no images required to read. Signup module + /report archive (every issue shows its own day record).
- M2 crops in-file at 50%: square 1080×1080, story 1080×1920 (recomposed — hero letter scales, dot strip replaces table; dashed safe zones are design-time only, strip on export), X 1200×675, NEW record billboard 1080×1350 (30D record + tier hit-rate bars, C-tier honestly under .500). Exact-size PNG masters in `exports/` (square 1080×1080, story 1080×1920, X 1200×675, record 1080×1350, article OG 1200×630). Every S2S5 surface also ships at 390 (Act 06 in the Intelligence file): market sheet (44px book rows, compact disagreement axis, sticky deep-link CTA), article view, ledger with per-read CLV chips + compact calibration card, /report signup + archive.
- Sample data is anchored to Jul 2026 reality; every number is a spec for what the live feed renders. Truth laws: no "tonight" off-slate, visible timestamps, honest empty states throughout.
Two design components, both open directly in a browser. All styling is inline (no stylesheets to port); every value below is already in the markup. Two design components, both open directly in a browser. All styling is inline (no stylesheets to port); every value below is already in the markup.
## Files ## Files
- `Vyndr Landing.dc.html` — desktop marketing landing: hero + live board proof, TIER RECORD calibration band, feature triptych, FREE / THE DESK pricing (no trial), wire footer. - `Vyndr Landing.dc.html` — desktop marketing landing: hero + live board proof, TIER RECORD calibration band, feature triptych, FREE / THE DESK pricing (no trial), wire footer.
- `assets/glyphs/` — all 74 archetype marks as individual SVGs (`currentColor`, 24-grid) + `MANIFEST.md` (name → hex → file). Generated from `glyphDefs()`; wire these into `archetypes.js`. - `assets/glyphs/` — all 83 archetype marks (74 display + 9 classifier-legacy) as individual SVGs (`currentColor`, 24-grid) + `MANIFEST.md` (name → hex → file). Generated from `glyphDefs()`; wire these into `archetypes.js`.
- `exports/` — rasterized share cards: `settle-1080x1350.png` (exact) and `grade-reveal-og.png` (1080×566, same 1.9:1 ratio as 1200×630 OG). - `exports/` — rasterized share cards: `settle-1080x1350.png` (exact) and `grade-reveal-og.png` (1080×566, same 1.9:1 ratio as 1200×630 OG).
- `image-slot.js` + `.image-slots.state.json` — drag-and-drop real-asset targets (player headshot, streaks rows, combat fighters) layered over monogram fallbacks; empty chrome hidden via `image-slot::part(empty){opacity:0}`. Design-time tooling, not product code. - `image-slot.js` + `.image-slots.state.json` — drag-and-drop real-asset targets (player headshot, streaks rows, combat fighters) layered over monogram fallbacks; empty chrome hidden via `image-slot::part(empty){opacity:0}`. Design-time tooling, not product code.
- `Vyndr System.dc.html` — desktop terminal: components, 74 glyphs, edge board, /u profile, share cards, streaks, combat, pitcher, correlation builder, pricing, empty state, ⌘K, THE WIRE. - `Vyndr System.dc.html` — desktop terminal: components, 74 glyphs, edge board, /u profile, share cards, streaks, combat, pitcher, correlation builder, pricing, empty state, ⌘K, THE WIRE.
@@ -13,8 +32,6 @@ Two design components, both open directly in a browser. All styling is inline (n
Desktop /u profile also carries the TIER RECORD calibration strip (matches mobile screen 10 and the landing band). Desktop /u profile also carries the TIER RECORD calibration strip (matches mobile screen 10 and the landing band).
> **BOOK ROSTER NOTE (2026-07, security follow-up item 7):** ESPN BET is defunct — PENN rebranded it to **theScore Bet** (Dec 1 2025); ESPN is now exclusive with DraftKings. The product BookChip set (`web/src/lib/books.js`) was updated (ESPN BET removed, theScore Bet added). The **design mockups' BookChip row still shows ESPN BET** — it needs the same one-swap to theScore Bet (mono `TS`) whenever the design files are refreshed.
## Tokens (exact) ## Tokens (exact)
- Surfaces: `#06060B` void · `#0E0E14` card · `#14141E` elevated · row hairline `#101018` · deep panel `#0A0A10` - Surfaces: `#06060B` void · `#0E0E14` card · `#14141E` elevated · row hairline `#101018` · deep panel `#0A0A10`
- Signal green `#00D4A0` — ONE meaning: edge / active / A-tier / primary CTA. Never decorative. - Signal green `#00D4A0` — ONE meaning: edge / active / A-tier / primary CTA. Never decorative.
@@ -41,5 +58,5 @@ Desktop /u profile also carries the TIER RECORD calibration strip (matches mobil
Alive = punctuated stillness (react to truth, then rest). Dense but RANKED — one bold mono hero figure among muted context; dim ranks progressively (1 / .86 / .64 / .48). One animated hero per zone. Real entities always (team-colored monograms until licensed assets drop in — headshot slots marked). No trial: free tier is the trial. Premium is produced, never claimed. Alive = punctuated stillness (react to truth, then rest). Dense but RANKED — one bold mono hero figure among muted context; dim ranks progressively (1 / .86 / .64 / .48). One animated hero per zone. Real entities always (team-colored monograms until licensed assets drop in — headshot slots marked). No trial: free tier is the trial. Premium is produced, never claimed.
## Rev 3 (Jul 16) ## Rev 3 (Jul 16)
- Matchup lines carry inline team chips (10-12px before each abbr) — swap for licensed TeamLogo imgs at the same size (resolver built). - Matchup lines everywhere now carry inline team-gradient chips (10-12px, before each abbr) — recognition without weight; chips sit inside rows so the ranked opacity ramp dims them too. Pattern: `<span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,<c1>,<c2>);vertical-align:-1.5px;margin-right:3px"></span>ABR`. Swap for licensed logo imgs at the same size when assets land.
- 9 classifier-legacy marks added (BRUSH/WHIFF/CONNECTOR/DISTRIBUTOR/FASTBREAK/FLEX/HYBRID/SWITCH/SWITCHBOARD) — 83 total. Fallback renders, never user-facing names. - 9 classifier-legacy marks added to `glyphDefs()` + `assets/glyphs/`: BRUSH #E0B84A, WHIFF #E86A6A, CONNECTOR #9AB0C4, DISTRIBUTOR #7AB8D8, FASTBREAK #4AA0E8, FLEX #A08AC8, HYBRID #C88AB0, SWITCH #C0B08A, SWITCHBOARD #90A0E8. Same 24-grid duotone language; colors deduped off signal-green/amber/red.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,274 @@
<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<script src="./support.js"></script>
</head>
<body>
<x-dc>
<helmet>
<meta name="viewport" content="width=device-width, initial-scale=1" />
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700;800;900&family=JetBrains+Mono:wght@400;500;600;700;800&display=swap" rel="stylesheet" />
<style>
*{box-sizing:border-box;margin:0;padding:0}
html,body{background:#06060B;color:#F0F0F0;font-family:'Inter',system-ui,sans-serif;-webkit-font-smoothing:antialiased}
::selection{background:rgba(0,212,160,.28);color:#F0F0F0}
a{color:#00D4A0;text-decoration:none}
a:hover{color:#5cf0cf}
.mono{font-family:'JetBrains Mono',monospace;font-variant-numeric:tabular-nums}
@keyframes vy-livedot{0%,100%{opacity:1;transform:scale(1)}50%{opacity:.35;transform:scale(.82)}}
@keyframes vy-syncdot{0%,100%{opacity:1}50%{opacity:.2}}
@keyframes vy-rxup{0%{background:rgba(0,212,160,.28)}100%{background:rgba(0,212,160,0)}}
@keyframes vy-rxdown{0%{background:rgba(255,71,87,.26)}100%{background:rgba(255,71,87,0)}}
@keyframes vy-rowin{0%{opacity:0;transform:translateY(9px)}100%{opacity:var(--o,1);transform:translateY(0)}}
@keyframes vy-glitch{0%,90%,100%{transform:translate(0,0);clip-path:inset(0 0 0 0);text-shadow:none}91%{transform:translate(-1.5px,0);clip-path:inset(12% 0 42% 0);text-shadow:1.6px 0 #FF4757,-1.6px 0 #00D4A0}93%{transform:translate(1.5px,0);clip-path:inset(58% 0 8% 0);text-shadow:-1.6px 0 #FF4757,1.6px 0 #00D4A0}95%{transform:translate(-1px,0);clip-path:inset(32% 0 30% 0);text-shadow:1px 0 #00D4A0}96%{clip-path:inset(70% 0 4% 0)}}
@keyframes vy-wirein{0%{opacity:0;transform:translateY(4px)}8%,86%{opacity:1;transform:none}100%{opacity:0;transform:translateY(-3px)}}
.rx-up{animation:vy-rxup .75s ease-out}
.rx-down{animation:vy-rxdown .75s ease-out}
</style>
</helmet>
<div ref="{{ rootRef }}" style="min-height:100vh;background:radial-gradient(1100px 640px at 74% -8%,rgba(0,212,160,.07),transparent 60%),#06060B">
<!-- top bar -->
<div style="display:flex;align-items:center;justify-content:space-between;padding:16px 48px;border-bottom:1px solid #14141E">
<div style="display:flex;align-items:center;gap:12px">
<svg width="28" height="28" viewBox="0 0 32 32" fill="none"><rect x="1" y="1" width="30" height="30" rx="7" fill="#0E0E14" stroke="#2A2A38" stroke-width="1"></rect><path d="M8 9 L16 24 L24 9" fill="none" stroke="#F0F0F0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><path d="M16 24 L24 9" fill="none" stroke="#00D4A0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><circle cx="24" cy="9" r="2" fill="#00D4A0"></circle></svg>
<span class="mono" style="font-weight:800;font-size:17px;letter-spacing:.3em;animation:vy-glitch 6.5s steps(1) infinite">VYND<span style="color:#00D4A0">R</span></span>
</div>
<div style="display:flex;align-items:center;gap:26px">
<a href="#proof" class="mono" style="font-size:10px;letter-spacing:.18em;color:#B8BCC8">THE RECORD</a>
<a href="#pricing" class="mono" style="font-size:10px;letter-spacing:.18em;color:#B8BCC8">PRICING</a>
<div style="display:flex;align-items:center;gap:7px">
<span style="width:6px;height:6px;border-radius:50%;background:#00D4A0;animation:vy-syncdot 1.4s ease-in-out infinite"></span>
<span class="mono" data-sync-clock style="font-size:11px;color:#B8BCC8;letter-spacing:.06em">--:--:--</span>
</div>
<button style="cursor:pointer;height:36px;padding:0 18px;border-radius:9px;background:#00D4A0;border:none;color:#06060B;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:10.5px;letter-spacing:.1em">OPEN THE TERMINAL</button>
</div>
</div>
<!-- hero -->
<div style="max-width:1240px;margin:0 auto;padding:76px 48px 0;display:grid;grid-template-columns:1.1fr 1fr;gap:56px;align-items:center">
<div>
<div class="mono" style="font-size:10.5px;letter-spacing:.3em;color:#00D4A0;margin-bottom:20px">SPORTS INTELLIGENCE TERMINAL</div>
<h1 style="font-size:56px;font-weight:800;letter-spacing:-.025em;line-height:1.02;text-wrap:pretty">Every line, graded before you bet it.</h1>
<p style="margin-top:18px;max-width:460px;color:#B8BCC8;font-size:16px;line-height:1.55">312 props graded tonight across 47 games — ranked by edge, verified against closing lines, settled in public. Misses included.</p>
<div style="display:flex;align-items:center;gap:12px;margin-top:30px">
<button style="cursor:pointer;height:48px;padding:0 26px;border-radius:11px;background:#00D4A0;border:none;color:#06060B;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:12px;letter-spacing:.1em">OPEN THE TERMINAL</button>
<button style="cursor:pointer;height:48px;padding:0 22px;border-radius:11px;background:transparent;border:1px solid #2A2A38;color:#B8BCC8;font-family:'JetBrains Mono',monospace;font-weight:700;font-size:11px;letter-spacing:.1em">SEE THE PUBLIC RECORD</button>
</div>
<div style="display:flex;align-items:center;gap:20px;margin-top:26px">
<span class="mono" style="display:inline-flex;align-items:center;gap:6px;font-size:9.5px;letter-spacing:.12em;color:#00D4A0"><span style="width:5px;height:5px;border-radius:50%;background:#00D4A0"></span>CLV-VERIFIED</span>
<span class="mono" style="font-size:9.5px;letter-spacing:.12em;color:#707080">RECORD PUBLIC</span>
<span class="mono" style="font-size:9.5px;letter-spacing:.12em;color:#707080">MISSES INCLUDED</span>
</div>
</div>
<!-- live board proof -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;overflow:hidden;box-shadow:0 40px 80px -40px rgba(0,0,0,.9),inset 0 1px 0 rgba(255,255,255,.04)">
<div style="display:flex;align-items:center;justify-content:space-between;padding:13px 16px;border-bottom:1px solid #14141E">
<span class="mono" style="font-size:10px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">EDGE BOARD</span>
<span style="display:flex;align-items:center;gap:6px"><span style="width:6px;height:6px;border-radius:50%;background:#00D4A0;animation:vy-livedot 1.5s ease-in-out infinite"></span><span class="mono" style="font-size:9px;letter-spacing:.16em;color:#00D4A0">LIVE</span></span>
</div>
<div style="--o:1;opacity:1;display:flex;align-items:center;gap:11px;padding:12px 16px;box-shadow:inset 2px 0 0 rgba(0,212,160,.55);animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.15s">
<span class="mono" style="font-size:10.5px;color:#00D4A0;font-weight:700;width:18px">01</span>
<div style="flex:1;min-width:0"><div style="font-size:13.5px;font-weight:700">Jokić <span style="color:#B8BCC8;font-weight:500">o27.5 PTS</span></div><div class="mono" style="font-size:8.5px;color:#707080;margin-top:2px">NBA · <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#0E4DA4,#4C8BF5);vertical-align:-1.5px;margin-right:3px"></span>DEN @ <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#0C2340,#236192);vertical-align:-1.5px;margin-right:3px"></span>MIN</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:rgba(0,212,160,.14);color:#00D4A0;font-weight:800;font-size:10.5px;border:1px solid rgba(0,212,160,.3)">A+</span>
<span class="mono" data-react data-fmt="edge" style="font-size:12.5px;font-weight:700;color:#00D4A0;width:52px;text-align:right;padding:1px 3px;border-radius:4px">+8.4%</span>
</div>
<div style="--o:.9;opacity:.9;display:flex;align-items:center;gap:11px;padding:12px 16px;border-top:1px solid #101018;animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.22s">
<span class="mono" style="font-size:10.5px;color:#707080;width:18px">02</span>
<div style="flex:1;min-width:0"><div style="font-size:13.5px;font-weight:700">Skenes <span style="color:#B8BCC8;font-weight:500">o6.5 K</span></div><div class="mono" style="font-size:8.5px;color:#707080;margin-top:2px">MLB · <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#2b2b2b,#c9a227);vertical-align:-1.5px;margin-right:3px"></span>PIT @ <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#0E3386,#CC3433);vertical-align:-1.5px;margin-right:3px"></span>CHC</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:rgba(0,212,160,.1);color:#00D4A0;font-weight:800;font-size:10.5px;border:1px solid rgba(0,212,160,.24)">A</span>
<span class="mono" data-react data-fmt="edge" style="font-size:12.5px;font-weight:700;color:#00D4A0;width:52px;text-align:right;padding:1px 3px;border-radius:4px">+6.1%</span>
</div>
<div style="--o:.9;opacity:.9;display:flex;align-items:center;gap:11px;padding:12px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.4);animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.29s">
<span class="mono" style="font-size:10.5px;color:#707080;width:18px">03</span>
<div style="flex:1;min-width:0"><div style="display:flex;align-items:center;gap:7px"><span style="width:6px;height:6px;border-radius:50%;background:#00D4A0;animation:vy-livedot 1.5s ease-in-out infinite"></span><span style="font-size:13.5px;font-weight:700">Edwards <span style="color:#B8BCC8;font-weight:500">o5.5 AST</span></span></div><div class="mono" style="font-size:8.5px;color:#00D4A0;margin-top:2px">LIVE · Q3 4:12</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:rgba(0,212,160,.1);color:#00D4A0;font-weight:800;font-size:10.5px;border:1px solid rgba(0,212,160,.24)">A</span>
<span class="mono" data-react data-fmt="edge" style="font-size:12.5px;font-weight:700;color:#00D4A0;width:52px;text-align:right;padding:1px 3px;border-radius:4px">+5.2%</span>
</div>
<div style="--o:.65;opacity:.65;display:flex;align-items:center;gap:11px;padding:12px 16px;border-top:1px solid #101018;animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.36s">
<span class="mono" style="font-size:10.5px;color:#707080;width:18px">04</span>
<div style="flex:1;min-width:0"><div style="font-size:13.5px;font-weight:600">Judge <span style="color:#B8BCC8;font-weight:500">o1.5 TB</span></div><div class="mono" style="font-size:8.5px;color:#707080;margin-top:2px">MLB · <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#1b3a6b,#0C2340);vertical-align:-1.5px;margin-right:3px"></span>NYY @ <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#BD3039,#6b1c22);vertical-align:-1.5px;margin-right:3px"></span>BOS</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">B</span>
<span class="mono" data-react data-fmt="edge" style="font-size:12.5px;font-weight:700;color:#00D4A0;width:52px;text-align:right;padding:1px 3px;border-radius:4px">+3.4%</span>
</div>
<div style="--o:.45;opacity:.45;display:flex;align-items:center;gap:11px;padding:12px 16px;border-top:1px solid #101018;animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.43s">
<span class="mono" style="font-size:10.5px;color:#707080;width:18px">05</span>
<div style="flex:1;min-width:0"><div style="font-size:13.5px;font-weight:600">LaVine <span style="color:#B8BCC8;font-weight:500">o22.5 PTS</span></div><div class="mono" style="font-size:8.5px;color:#707080;margin-top:2px">NBA · <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#CE1141,#6d0a22);vertical-align:-1.5px;margin-right:3px"></span>CHI @ <span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#00471B,#0a7a3a);vertical-align:-1.5px;margin-right:3px"></span>MIL</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:rgba(255,71,87,.1);color:#FF4757;font-weight:800;font-size:10.5px;border:1px solid rgba(255,71,87,.28)">D</span>
<span class="mono" style="font-size:12.5px;font-weight:700;color:#FF4757;width:52px;text-align:right">-1.2%</span>
</div>
<div style="display:flex;align-items:center;justify-content:space-between;padding:10px 16px;border-top:1px solid #101018;background:#08080D">
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58">+ 42 GRADED READS · PRO</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#707080">RANKED BY EDGE</span>
</div>
</div>
</div>
<!-- proof: tier record -->
<div id="proof" style="max-width:1240px;margin:0 auto;padding:96px 48px 0">
<div style="display:grid;grid-template-columns:1fr 1.3fr;gap:56px;align-items:center">
<div>
<div class="mono" style="font-size:10px;letter-spacing:.28em;color:#00D4A0;margin-bottom:14px">THE PROOF</div>
<h2 style="font-size:34px;font-weight:800;letter-spacing:-.02em;line-height:1.1;text-wrap:pretty">The grades calibrate. That's the product.</h2>
<p style="margin-top:14px;max-width:420px;color:#B8BCC8;font-size:14.5px;line-height:1.6">A+ reads hit 80%. D reads lose — so we tell you to fade them. Every grade settles against the closing line in public, misses included.</p>
</div>
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;overflow:hidden">
<div style="display:flex;align-items:center;justify-content:space-between;padding:13px 20px;border-bottom:1px solid #14141E">
<span class="mono" style="font-size:10px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">TIER RECORD</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">SEASON · ALL SPORTS</span>
</div>
<div style="display:grid;grid-template-columns:56px 1fr 90px 70px;gap:12px;align-items:center;padding:10px 20px 4px">
<span class="mono" style="font-size:8px;letter-spacing:.16em;color:#4a4a58">TIER</span><span class="mono" style="font-size:8px;letter-spacing:.16em;color:#4a4a58">HIT RATE</span><span class="mono" style="font-size:8px;letter-spacing:.16em;color:#4a4a58;text-align:right">RECORD</span><span class="mono" style="font-size:8px;letter-spacing:.16em;color:#4a4a58;text-align:right">ROI</span>
</div>
<div style="display:grid;grid-template-columns:56px 1fr 90px 70px;gap:12px;align-items:center;padding:9px 20px">
<span class="mono" style="font-size:15px;font-weight:800;color:#00D4A0;text-shadow:0 0 12px rgba(0,212,160,.4)">A+</span>
<div style="height:6px;border-radius:3px;background:#14141E;overflow:hidden"><div style="width:80%;height:100%;background:#00D4A0"></div></div>
<span class="mono" style="font-size:12px;color:#F0F0F0;text-align:right">246</span><span class="mono" style="font-size:12px;color:#00D4A0;text-align:right;font-weight:700">+31%</span>
</div>
<div style="display:grid;grid-template-columns:56px 1fr 90px 70px;gap:12px;align-items:center;padding:9px 20px">
<span class="mono" style="font-size:15px;font-weight:800;color:#00D4A0">A</span>
<div style="height:6px;border-radius:3px;background:#14141E;overflow:hidden"><div style="width:68%;height:100%;background:rgba(0,212,160,.75)"></div></div>
<span class="mono" style="font-size:12px;color:#F0F0F0;text-align:right">4119</span><span class="mono" style="font-size:12px;color:#00D4A0;text-align:right;font-weight:700">+18%</span>
</div>
<div style="display:grid;grid-template-columns:56px 1fr 90px 70px;gap:12px;align-items:center;padding:9px 20px">
<span class="mono" style="font-size:15px;font-weight:800;color:#F0F0F0">B</span>
<div style="height:6px;border-radius:3px;background:#14141E;overflow:hidden"><div style="width:58%;height:100%;background:#B8BCC8"></div></div>
<span class="mono" style="font-size:12px;color:#B8BCC8;text-align:right">5842</span><span class="mono" style="font-size:12px;color:#B8BCC8;text-align:right">+6%</span>
</div>
<div style="display:grid;grid-template-columns:56px 1fr 90px 70px;gap:12px;align-items:center;padding:9px 20px">
<span class="mono" style="font-size:15px;font-weight:800;color:#B8BCC8">C</span>
<div style="height:6px;border-radius:3px;background:#14141E;overflow:hidden"><div style="width:52%;height:100%;background:#707080"></div></div>
<span class="mono" style="font-size:12px;color:#707080;text-align:right">3028</span><span class="mono" style="font-size:12px;color:#707080;text-align:right">+1%</span>
</div>
<div style="display:grid;grid-template-columns:56px 1fr 90px 70px;gap:12px;align-items:center;padding:9px 20px 14px">
<span class="mono" style="font-size:15px;font-weight:800;color:#FF4757">D</span>
<div style="height:6px;border-radius:3px;background:#14141E;overflow:hidden"><div style="width:38%;height:100%;background:#FF4757"></div></div>
<span class="mono" style="font-size:12px;color:#707080;text-align:right">1220</span><span class="mono" style="font-size:12px;color:#FF4757;text-align:right;font-weight:700">FADE ✓</span>
</div>
</div>
</div>
</div>
<!-- feature triptych -->
<div style="max-width:1240px;margin:0 auto;padding:96px 48px 0">
<div style="display:grid;grid-template-columns:repeat(3,1fr);gap:16px">
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:24px">
<div class="mono" style="font-size:9.5px;letter-spacing:.24em;color:#00D4A0;margin-bottom:12px">RANKED, NOT NOISY</div>
<div style="font-size:17px;font-weight:700;letter-spacing:-.01em">One board. Every sport. Sorted by edge.</div>
<p style="margin-top:9px;color:#B8BCC8;font-size:13px;line-height:1.55">The A-tier read floats to the top at full color; everything else dims. You see the one thing that matters in a glance.</p>
</div>
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:24px">
<div class="mono" style="font-size:9.5px;letter-spacing:.24em;color:#00D4A0;margin-bottom:12px">ALIVE DURING THE SLATE</div>
<div style="font-size:17px;font-weight:700;letter-spacing:-.01em">Grades that react while the game runs.</div>
<p style="margin-top:9px;color:#B8BCC8;font-size:13px;line-height:1.55">A number pulses the instant a line moves. Live grades firm and ease with pace and usage — on the board, and on your lock screen.</p>
</div>
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:24px">
<div class="mono" style="font-size:9.5px;letter-spacing:.24em;color:#FFB347;margin-bottom:12px">MATH THAT SAYS NO</div>
<div style="font-size:17px;font-weight:700;letter-spacing:-.01em">Correlation warnings that trim your stake.</div>
<p style="margin-top:9px;color:#B8BCC8;font-size:13px;line-height:1.55">Same-game legs move together — books pay as if they don't. VYNDR prices the true parlay and sizes the stake to the real risk.</p>
</div>
</div>
</div>
<!-- pricing -->
<div id="pricing" style="max-width:880px;margin:0 auto;padding:96px 48px 90px">
<div style="text-align:center;margin-bottom:30px">
<div class="mono" style="font-size:10px;letter-spacing:.28em;color:#00D4A0;margin-bottom:12px">ONE DECISION</div>
<h2 style="font-size:32px;font-weight:800;letter-spacing:-.02em">The free tier is the trial.</h2>
</div>
<div style="display:grid;grid-template-columns:1fr 1.25fr;gap:16px;align-items:stretch">
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:24px;display:flex;flex-direction:column">
<div style="display:flex;align-items:center;justify-content:space-between"><span class="mono" style="font-size:10px;letter-spacing:.24em;color:#B8BCC8;font-weight:700">FREE</span><span class="mono" style="font-size:15px;color:#707080">$0</span></div>
<div style="display:flex;flex-direction:column;gap:9px;margin-top:16px">
<span style="display:flex;align-items:center;gap:9px;font-size:13px;color:#B8BCC8"><span style="color:#707080"></span> Live scores + the wire</span>
<span style="display:flex;align-items:center;gap:9px;font-size:13px;color:#B8BCC8"><span style="color:#707080"></span> Tonight's top 3 grades</span>
<span style="display:flex;align-items:center;gap:9px;font-size:13px;color:#707080"><span></span> Board ranks 447 stay sealed</span>
</div>
<div style="flex:1"></div>
<button style="cursor:pointer;width:100%;margin-top:20px;height:44px;border-radius:11px;background:transparent;border:1px solid #2A2A38;color:#B8BCC8;font-family:'JetBrains Mono',monospace;font-weight:700;font-size:11px;letter-spacing:.1em">START FREE</button>
</div>
<div style="border-radius:16px;background:linear-gradient(180deg,#0F1512,#0A0A10);border:1px solid rgba(0,212,160,.4);box-shadow:0 30px 60px -34px rgba(0,212,160,.28),inset 0 1px 0 rgba(0,212,160,.12);padding:24px">
<div style="display:flex;align-items:center;justify-content:space-between">
<span class="mono" style="font-size:10px;letter-spacing:.24em;color:#00D4A0;font-weight:800">THE DESK</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#707080;border:1px solid #2A2A38;border-radius:5px;padding:2px 7px">ALL SPORTS</span>
</div>
<div style="display:flex;align-items:flex-end;gap:8px;margin-top:12px"><span class="mono" style="font-size:44px;font-weight:800;line-height:.85">$44.99</span><span class="mono" style="font-size:13px;color:#707080;padding-bottom:4px">/mo</span></div>
<div style="display:flex;flex-direction:column;gap:9px;margin-top:16px">
<span style="display:flex;align-items:center;gap:9px;font-size:13px"><span style="color:#00D4A0;font-weight:700"></span> The full edge board — every graded read, nightly</span>
<span style="display:flex;align-items:center;gap:9px;font-size:13px"><span style="color:#00D4A0;font-weight:700"></span> Live grade-shift + steam alerts on lock screen</span>
<span style="display:flex;align-items:center;gap:9px;font-size:13px"><span style="color:#00D4A0;font-weight:700"></span> CLV-verified public record + share cards</span>
<span style="display:flex;align-items:center;gap:9px;font-size:13px"><span style="color:#00D4A0;font-weight:700"></span> Correlation math on every slip</span>
</div>
<button style="cursor:pointer;width:100%;margin-top:20px;height:46px;border-radius:11px;background:#00D4A0;border:none;color:#06060B;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:12px;letter-spacing:.1em">UPGRADE TO PRO</button>
<div class="mono" style="font-size:9px;letter-spacing:.06em;color:#707080;text-align:center;margin-top:10px">cancel anytime · no trial — the free tier is the trial</div>
</div>
</div>
</div>
<!-- footer -->
<div style="border-top:1px solid #14141E">
<div style="max-width:1240px;margin:0 auto;padding:22px 48px;display:flex;flex-wrap:wrap;gap:10px 24px;align-items:center;justify-content:space-between">
<div style="display:flex;align-items:center;gap:10px">
<span class="mono" style="font-weight:800;font-size:12px;letter-spacing:.26em">VYND<span style="color:#00D4A0">R</span></span>
<span class="mono" data-wire style="font-size:9.5px;color:#707080;letter-spacing:.04em">Session open — 312 props graded.</span>
</div>
<div style="display:flex;align-items:center;gap:18px">
<a href="Vyndr System.dc.html" class="mono" style="font-size:9px;letter-spacing:.16em">DESKTOP SYSTEM</a>
<a href="Vyndr Mobile.dc.html" class="mono" style="font-size:9px;letter-spacing:.16em">MOBILE SPEC</a>
</div>
</div>
</div>
</div>
</x-dc>
<script type="text/x-dc" data-dc-script data-props="{&quot;$preview&quot;:{&quot;width&quot;:1440,&quot;height&quot;:1000}}">
class Component extends DCLogic {
renderVals(){ return { rootRef:(n)=>{ this.rootNode=n; } }; }
componentDidMount(){
const tick=()=>{
const el=this.rootNode && this.rootNode.querySelector('[data-sync-clock]');
if(el){const d=new Date();const p=n=>String(n).padStart(2,'0');el.textContent=p(d.getHours())+':'+p(d.getMinutes())+':'+p(d.getSeconds());}
};
tick(); this._clock=setInterval(tick,1000);
this._booted=false; this._bt=setTimeout(()=>{this._booted=true;},1400);
this._rx=setInterval(()=>{
if(!this._booted || !this.rootNode) return;
const cells=[...this.rootNode.querySelectorAll('[data-react]')];
if(!cells.length) return;
const c=cells[Math.floor(Math.random()*cells.length)];
const cur=parseFloat(c.textContent.replace(/[^0-9.\-]/g,''));
if(isNaN(cur)) return;
const up=Math.random()<0.6;
const n=cur+(up?0.2:-0.2);
c.textContent=(n>=0?'+':'')+n.toFixed(1)+'%';
c.style.color=n>=0?'#00D4A0':'#FF4757';
c.classList.remove('rx-up','rx-down');void c.offsetWidth;
c.classList.add(up?'rx-up':'rx-down');
setTimeout(()=>c.classList.remove('rx-up','rx-down'),760);
},2800);
this._we=[
['STEAM','Jokić 25.5 → 27.5 since open. Grade holds A+.','#00D4A0'],
['LIVE','Edwards o5.5 AST pacing ▲ — usage +6%.','#00D4A0'],
['FLAG','Same-game legs correlated — stake trimmed.','#FFB347'],
['CLV','Tonight: +2.8% closing-line value across 14 edges.','#B8BCC8']
];
this._wi=0;
this._wire=setInterval(()=>{
const w=this.rootNode && this.rootNode.querySelector('[data-wire]');
if(!w) return;
const [tag,msg,col]=this._we[this._wi%this._we.length];this._wi++;
w.innerHTML='<span style="color:'+col+';font-weight:700">'+tag+'</span><span style="color:#2A2A38"> · </span>'+msg;
w.style.animation='none';void w.offsetWidth;
w.style.animation='vy-wirein 6.4s ease-out';
},6500);
}
componentWillUnmount(){ clearInterval(this._clock); clearInterval(this._rx); clearInterval(this._wire); clearTimeout(this._bt); }
}
</script>
</body>
</html>
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,961 @@
<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<script src="./support.js"></script>
</head>
<body>
<x-dc>
<helmet>
<meta name="viewport" content="width=device-width, initial-scale=1" />
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700;800;900&family=JetBrains+Mono:wght@400;500;600;700;800&display=swap" rel="stylesheet" />
<style>
*{box-sizing:border-box;margin:0;padding:0}
html,body{background:#06060B;color:#F0F0F0;font-family:'Inter',system-ui,sans-serif;-webkit-font-smoothing:antialiased}
::selection{background:rgba(0,212,160,.28);color:#F0F0F0}
a{color:#00D4A0;text-decoration:none}
a:hover{color:#5cf0cf}
.mono{font-family:'JetBrains Mono',monospace;font-variant-numeric:tabular-nums}
@keyframes vy-livedot{0%,100%{opacity:1;transform:scale(1)}50%{opacity:.35;transform:scale(.82)}}
@keyframes vy-syncdot{0%,100%{opacity:1}50%{opacity:.2}}
@keyframes vy-rxup{0%{background:rgba(0,212,160,.28)}100%{background:rgba(0,212,160,0)}}
@keyframes vy-rxdown{0%{background:rgba(255,71,87,.26)}100%{background:rgba(255,71,87,0)}}
@keyframes vy-rowin{0%{opacity:0;transform:translateY(9px)}100%{opacity:var(--o,1);transform:translateY(0)}}
@keyframes vy-bootline{0%{transform:scaleX(0)}100%{transform:scaleX(1)}}
@keyframes vy-arrive{0%{opacity:0;transform:scale(.88) translateY(10px);filter:blur(6px)}60%{opacity:1;transform:scale(1.015) translateY(0);filter:blur(0)}100%{opacity:1;transform:scale(1) translateY(0)}}
@keyframes vy-glitch{0%,90%,100%{transform:translate(0,0);clip-path:inset(0 0 0 0);text-shadow:none}91%{transform:translate(-1.5px,0);clip-path:inset(12% 0 42% 0);text-shadow:1.6px 0 #FF4757,-1.6px 0 #00D4A0}93%{transform:translate(1.5px,0);clip-path:inset(58% 0 8% 0);text-shadow:-1.6px 0 #FF4757,1.6px 0 #00D4A0}95%{transform:translate(-1px,0);clip-path:inset(32% 0 30% 0);text-shadow:1px 0 #00D4A0}96%{transform:translate(1px,0);clip-path:inset(70% 0 4% 0);text-shadow:none}}
@keyframes vy-cursor{0%,49%{opacity:1}50%,100%{opacity:0}}
@keyframes vy-wirein{0%{opacity:0;transform:translateY(4px)}8%,86%{opacity:1;transform:none}100%{opacity:0;transform:translateY(-3px)}}
.rx-up{animation:vy-rxup .75s ease-out}
.rx-down{animation:vy-rxdown .75s ease-out}
.vy-grain{position:fixed;inset:0;z-index:5;pointer-events:none;mix-blend-mode:overlay;opacity:.05;background-image:url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='140' height='140'%3E%3Cfilter id='n'%3E%3CfeTurbulence type='fractalNoise' baseFrequency='0.82' numOctaves='2' stitchTiles='stitch'/%3E%3C/filter%3E%3Crect width='100%25' height='100%25' filter='url(%23n)'/%3E%3C/svg%3E");background-size:140px}
.vy-vignette{position:fixed;inset:0;z-index:4;pointer-events:none;box-shadow:inset 0 0 200px 40px rgba(0,0,0,.55),inset 0 90px 120px -60px rgba(0,212,160,.05)}
[data-reveal]{opacity:0;transform:translateY(16px);transition:opacity .6s cubic-bezier(.2,.8,.2,1),transform .6s cubic-bezier(.2,.8,.2,1)}
[data-reveal].vy-in{opacity:1;transform:none}
.vy-still *{animation-play-state:paused !important}
html{scroll-behavior:smooth}
</style>
<script src="image-slot.js"></script>
<style>image-slot::part(empty){opacity:0}</style>
</helmet>
<div ref="{{ rootRef }}" class="{{ stillClass }}" style="min-height:100vh;background:radial-gradient(1200px 700px at 78% -8%,rgba(0,212,160,.06),transparent 60%),#06060B;padding:0 0 120px">
<div class="vy-vignette" aria-hidden="true"></div>
<div class="vy-grain" aria-hidden="true"></div>
<!-- ============ THE WIRE ============ -->
<div style="position:fixed;left:0;right:0;bottom:0;z-index:45;display:flex;align-items:center;gap:14px;height:36px;padding:0 40px;background:rgba(6,6,11,.88);backdrop-filter:blur(14px);border-top:1px solid #14141E">
<span class="mono" style="font-size:9px;letter-spacing:.26em;color:#00D4A0;font-weight:800;flex:none">THE WIRE</span>
<span style="width:1px;height:14px;background:#1E1E2A;flex:none"></span>
<span class="mono" data-wire-time style="font-size:10px;color:#707080;letter-spacing:.06em;flex:none">--:--:--</span>
<span class="mono" data-wire style="font-size:11px;color:#B8BCC8;letter-spacing:.03em;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1">Offseason desk open — 214 NFL season lines repriced on news.</span>
<span class="mono" style="font-size:9px;letter-spacing:.16em;color:#4a4a58;flex:none">LOG ▸</span>
</div>
<div style="max-width:1240px;margin:0 auto;padding:0 40px">
<!-- ============ MASTHEAD ============ -->
<div style="padding:52px 0 30px">
<div class="mono" style="font-size:11px;letter-spacing:.34em;color:#00D4A0;margin-bottom:16px">SESSION 2 · SURFACE 1 · THE OFFSEASON HUB</div>
<h1 style="font-size:38px;font-weight:800;letter-spacing:-.02em;line-height:1.05">Every sport is always in season.<span class="mono" style="color:#00D4A0;animation:vy-cursor 1.1s steps(1) infinite"></span></h1>
<p style="margin-top:12px;max-width:640px;color:#B8BCC8;font-size:14px;line-height:1.55">A sport is in one of two states: LIVE SLATE or OFFSEASON HUB. The hub lives inside the sport — tapping NFL gets NFL, whatever NFL is today. Season-long is a different read: no "tonight," longer horizons, lines that move on news. The story is the movement — open, current, model.</p>
<div class="mono" style="margin-top:14px;font-size:10px;letter-spacing:.16em;color:#707080"><a href="Vyndr System.dc.html">◂ DESKTOP SYSTEM</a> · <a href="Vyndr Mobile.dc.html">MOBILE</a> · <a href="Vyndr Intelligence.dc.html">INTELLIGENCE (S2-5) ▸</a></div>
</div>
<!-- ============ ACT 01 · HUB HOME ============ -->
<div data-reveal style="margin-top:14px;display:flex;align-items:baseline;gap:18px;border-bottom:1px solid #14141E;padding-bottom:18px">
<span class="mono" style="font-size:64px;font-weight:800;line-height:.8;color:#14141E;-webkit-text-stroke:1px #2A2A38">01</span>
<div><div class="mono" style="font-size:13px;letter-spacing:.3em;color:#F0F0F0;font-weight:800">HUB HOME</div><div style="font-size:12.5px;color:#707080;margin-top:5px">Per sport. NFL primary · NBA variant. Countdown is ambient context, never the hero — the hero is what changed today.</div></div>
</div>
<div id="hub" data-screen-label="Hub home · NFL desktop" style="margin-top:44px">
<div style="display:flex;align-items:center;gap:12px;margin-bottom:14px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">NFL HUB · DESKTOP</span>
<span class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#707080">SPORT TAB STATE LIVES IN THE TAB · NEVER A SEPARATE "OFFSEASON" TAB</span>
</div>
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;overflow:hidden;box-shadow:0 30px 60px -34px rgba(0,0,0,.9),inset 0 1px 0 rgba(255,255,255,.03)">
<!-- terminal top bar w/ sport states -->
<div style="display:flex;align-items:center;justify-content:space-between;gap:18px;padding:13px 20px;border-bottom:1px solid #14141E;background:rgba(6,6,11,.6)">
<div style="display:flex;align-items:center;gap:10px;flex:none">
<svg width="22" height="22" viewBox="0 0 32 32" fill="none"><rect x="1" y="1" width="30" height="30" rx="7" fill="#0E0E14" stroke="#2A2A38" stroke-width="1"></rect><path d="M8 9 L16 24 L24 9" fill="none" stroke="#F0F0F0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><path d="M16 24 L24 9" fill="none" stroke="#00D4A0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><circle cx="24" cy="9" r="2" fill="#00D4A0"></circle></svg>
<span class="mono" style="font-weight:800;font-size:13px;letter-spacing:.26em;animation:vy-glitch 7s steps(1) infinite">VYND<span style="color:#00D4A0">R</span></span>
</div>
<div style="display:flex;align-items:center;gap:6px;flex:1;justify-content:center">
<span class="mono" style="display:inline-flex;align-items:center;gap:6px;height:24px;padding:0 10px;border-radius:7px;font-size:9.5px;letter-spacing:.12em;color:#B8BCC8"><span style="width:5px;height:5px;border-radius:50%;background:#00D4A0;animation:vy-livedot 1.5s ease-in-out infinite"></span>MLB · 9 LIVE</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:6px;height:24px;padding:0 10px;border-radius:7px;background:#14141E;border:1px solid #2A2A38;font-size:9.5px;letter-spacing:.12em;color:#F0F0F0;font-weight:700">NFL · CAMP 5D</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:6px;height:24px;padding:0 10px;border-radius:7px;font-size:9.5px;letter-spacing:.12em;color:#707080">NBA · SL DAY 8</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:6px;height:24px;padding:0 10px;border-radius:7px;font-size:9.5px;letter-spacing:.12em;color:#707080">NHL · FA</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:6px;height:24px;padding:0 10px;border-radius:7px;font-size:9.5px;letter-spacing:.12em;color:#707080">UFC · SAT</span>
</div>
<div style="display:flex;align-items:center;gap:8px;flex:none">
<span style="width:6px;height:6px;border-radius:50%;background:#00D4A0;animation:vy-syncdot 1.4s ease-in-out infinite"></span>
<span class="mono" data-sync-clock style="font-size:10px;color:#B8BCC8;letter-spacing:.06em">--:--:--</span>
</div>
</div>
<!-- hub header: countdown = ambient -->
<div style="display:flex;align-items:baseline;justify-content:space-between;gap:18px;padding:22px 24px 0;flex-wrap:wrap">
<div>
<div style="display:flex;align-items:baseline;gap:12px">
<span style="font-size:26px;font-weight:800;letter-spacing:-.01em">NFL</span>
<span class="mono" style="font-size:10px;letter-spacing:.26em;color:#707080">THE OFFSEASON DESK</span>
</div>
<div class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#4a4a58;margin-top:6px">OUTLOOKS REPRICE ON NEWS · NOT GAME ODDS · AS OF JUL 17</div>
</div>
<div class="mono" style="font-size:11px;letter-spacing:.2em;color:#707080"><span data-kickoff style="color:#B8BCC8;font-weight:700">{{ kickoffDays }}</span> DAYS TO KICKOFF · THU SEP 10</div>
</div>
<!-- key dates strip -->
<div style="padding:18px 24px 4px">
<div style="position:relative;height:34px">
<div style="position:absolute;left:0;right:0;top:8px;height:1px;background:#1E1E2A"></div>
<div style="position:absolute;left:0;top:8px;height:1px;width:22%;background:#2A2A38"></div>
<div style="position:absolute;left:0;right:0;top:0;display:grid;grid-template-columns:repeat(6,1fr)">
<div><span style="display:block;width:7px;height:7px;border-radius:50%;background:#2A2A38;margin-top:5px"></span><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58;margin-top:6px">TAG DDL · JUL 15</div></div>
<div><span style="display:block;width:9px;height:9px;border-radius:50%;background:#F0F0F0;margin-top:4px;box-shadow:0 0 0 3px rgba(240,240,240,.12)"></span><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#F0F0F0;margin-top:5px;font-weight:700">CAMPS · JUL 22</div></div>
<div><span style="display:block;width:7px;height:7px;border-radius:50%;background:#1E1E2A;border:1px solid #2A2A38;margin-top:5px"></span><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#707080;margin-top:6px">HOF GAME · JUL 30</div></div>
<div><span style="display:block;width:7px;height:7px;border-radius:50%;background:#1E1E2A;border:1px solid #2A2A38;margin-top:5px"></span><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#707080;margin-top:6px">PRESEASON · AUG 6</div></div>
<div><span style="display:block;width:7px;height:7px;border-radius:50%;background:#1E1E2A;border:1px solid #2A2A38;margin-top:5px"></span><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#707080;margin-top:6px">FINAL CUTS · SEP 1</div></div>
<div style="justify-self:end;text-align:right"><span style="display:inline-block;width:9px;height:9px;border-radius:50%;background:#00D4A0;margin-top:4px;box-shadow:0 0 10px rgba(0,212,160,.5)"></span><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#00D4A0;margin-top:5px;font-weight:700">KICKOFF · SEP 10</div></div>
</div>
</div>
</div>
<div style="display:grid;grid-template-columns:1.25fr 1fr;gap:16px;padding:16px 24px 22px">
<!-- WHAT CHANGED TODAY -->
<div style="border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;overflow:hidden">
<div style="display:flex;align-items:center;justify-content:space-between;padding:14px 16px 12px">
<span class="mono" style="font-size:10.5px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">WHAT CHANGED TODAY</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">4 EVENTS · JUL 17</span>
</div>
<div style="height:1px;background:#14141E;transform-origin:left;animation:vy-bootline .6s ease-out both"></div>
<div style="--o:1;opacity:1;padding:12px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.55);animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.14s">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080;flex:none">11:42 AM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:8px;letter-spacing:.14em;color:#B8BCC8;flex:none">CLEARED</span></div>
<div style="font-size:13.5px;font-weight:700;margin-top:6px">Nabers cleared for camp — no restrictions</div>
<div class="mono" style="font-size:9px;letter-spacing:.08em;color:#707080;margin-top:5px"><span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#0B2265,#5a7fd4);vertical-align:-1.5px;margin-right:3px"></span>NYG · REC YDS <span style="color:#4a4a58">1,050.5</span><span data-react data-fmt="line" style="color:#F0F0F0;font-weight:700;padding:1px 3px;border-radius:4px">1,120.5</span> · <span style="color:#00D4A0">VYNDR 1,188</span> · <a href="#reveal" class="mono" style="font-size:9px;letter-spacing:.12em">READ ▸</a></div>
</div>
<div style="--o:.86;opacity:.86;padding:12px 16px;border-top:1px solid #101018;animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.22s">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080;flex:none">10:15 AM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:8px;letter-spacing:.14em;color:#B8BCC8;flex:none">CAMP REPORT</span></div>
<div style="font-size:13.5px;font-weight:600;margin-top:6px">Worthy running clear WR1 reps in KC installs</div>
<div class="mono" style="font-size:9px;letter-spacing:.08em;color:#707080;margin-top:5px"><span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#E31837,#FFB81C);vertical-align:-1.5px;margin-right:3px"></span>KC · REC YDS <span style="color:#4a4a58">990.5</span><span data-react data-fmt="line" style="color:#F0F0F0;font-weight:700;padding:1px 3px;border-radius:4px">1,024.5</span> · <span style="color:#00D4A0">VYNDR 1,061</span> · <span class="mono" style="color:#4a4a58">READ ▸</span></div>
</div>
<div style="--o:.86;opacity:.86;padding:12px 16px;border-top:1px solid #101018;animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.3s">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080;flex:none">9:30 AM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:rgba(255,179,71,.08);border:1px solid rgba(255,179,71,.25);font-size:8px;letter-spacing:.14em;color:#FFB347;flex:none">CONTRACT</span></div>
<div style="font-size:13.5px;font-weight:600;margin-top:6px">Higgins restructure done — reports Tuesday</div>
<div class="mono" style="font-size:9px;letter-spacing:.08em;color:#707080;margin-top:5px"><span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#FB4F14,#1a1a1a);vertical-align:-1.5px;margin-right:3px"></span>CIN · BURROW PASS TD <span style="color:#4a4a58">27.5</span><span data-react data-fmt="half" style="color:#F0F0F0;font-weight:700;padding:1px 3px;border-radius:4px">28.5</span> · <span style="color:#00D4A0">VYNDR 29.8</span> · <span class="mono" style="color:#4a4a58">READ ▸</span></div>
</div>
<div style="--o:.64;opacity:.64;padding:12px 16px;border-top:1px solid #101018;animation:vy-rowin .5s cubic-bezier(.2,.8,.2,1) both;animation-delay:.38s">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080;flex:none">8:05 AM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:8px;letter-spacing:.14em;color:#B8BCC8;flex:none">DEPTH</span></div>
<div style="font-size:13.5px;font-weight:600;margin-top:6px">Bills name Coleman the starting X receiver</div>
<div class="mono" style="font-size:9px;letter-spacing:.08em;color:#707080;margin-top:5px"><span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#00338D,#C60C30);vertical-align:-1.5px;margin-right:3px"></span>BUF · REC YDS <span style="color:#4a4a58">815.5</span><span style="color:#F0F0F0;font-weight:700">815.5</span> · <span style="color:#707080">FLAT · MODEL HOLDS 862</span></div>
</div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;justify-content:space-between">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#4a4a58">SOURCE-LINKED · EVERY IMPACT SHOWS THE REPRICE</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">FULL FEED ▸</span>
</div>
</div>
<!-- SEASON BOARD PREVIEW + report signup -->
<div style="display:flex;flex-direction:column;gap:16px">
<div style="border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;overflow:hidden">
<div style="display:flex;align-items:center;justify-content:space-between;padding:14px 16px 12px">
<span class="mono" style="font-size:10.5px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">SEASON BOARD</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">BIGGEST MODELMARKET GAPS</span>
</div>
<div style="height:1px;background:#14141E"></div>
<div style="--o:1;opacity:1;display:flex;align-items:center;gap:10px;padding:10px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.55)">
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:700">Nabers <span style="color:#B8BCC8;font-weight:500">o1,120.5 Rec Yds</span></div><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#707080;margin-top:2px">O 1,050.5 · NOW 1,120.5 · <span style="color:#00D4A0">M 1,188</span></div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:rgba(0,212,160,.14);color:#00D4A0;font-weight:800;font-size:10px;border:1px solid rgba(0,212,160,.3)">A+</span>
<span class="mono" data-react data-fmt="edge" style="font-size:11.5px;font-weight:700;color:#00D4A0;width:46px;text-align:right;padding:1px 3px;border-radius:4px">+6.0%</span>
</div>
<div style="--o:.86;opacity:.86;display:flex;align-items:center;gap:10px;padding:10px 16px;border-top:1px solid #101018">
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:600">Allen <span style="color:#B8BCC8;font-weight:500">o37.5 Pass+Rush TD</span></div><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#707080;margin-top:2px">O 36.5 · NOW 37.5 · <span style="color:#00D4A0">M 40.2</span></div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:rgba(0,212,160,.1);color:#00D4A0;font-weight:800;font-size:10px;border:1px solid rgba(0,212,160,.24)">A</span>
<span class="mono" data-react data-fmt="edge" style="font-size:11.5px;font-weight:700;color:#00D4A0;width:46px;text-align:right;padding:1px 3px;border-radius:4px">+4.8%</span>
</div>
<div style="--o:.64;opacity:.64;display:flex;align-items:center;gap:10px;padding:10px 16px;border-top:1px solid #101018">
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:600">Chiefs <span style="color:#B8BCC8;font-weight:500">u11.5 Wins</span></div><div class="mono" style="font-size:8px;letter-spacing:.1em;color:#707080;margin-top:2px">O 11.5 · NOW 11.5 · <span style="color:#00D4A0">M 10.8</span></div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10px;border:1px solid #1E1E2A">B</span>
<span class="mono" style="font-size:11.5px;font-weight:700;color:#00D4A0;width:46px;text-align:right">+3.1%</span>
</div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;justify-content:space-between">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#4a4a58">214 SEASON LINES</span>
<a href="#board" class="mono" style="font-size:9px;letter-spacing:.14em">FULL BOARD ▸</a>
</div>
</div>
<!-- report signup slim (spec'd in Intelligence S5) -->
<div style="border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;padding:14px 16px;display:flex;align-items:center;gap:12px">
<div style="flex:1;min-width:0">
<div class="mono" style="font-size:9.5px;letter-spacing:.22em;color:#F0F0F0;font-weight:700">THE VYNDR REPORT</div>
<div style="font-size:11px;color:#707080;margin-top:3px">The offseason, read daily. One email, no noise.</div>
</div>
<span class="mono" style="display:inline-flex;align-items:center;height:28px;padding:0 12px;border-radius:7px;background:#00D4A0;color:#06060B;font-weight:800;font-size:9.5px;letter-spacing:.14em;flex:none">SUBSCRIBE</span>
</div>
</div>
</div>
</div>
</div>
<!-- hub mobile: NFL + NBA variant -->
<div data-screen-label="Hub home · mobile NFL + NBA" style="margin-top:44px">
<div style="display:flex;align-items:center;gap:12px;margin-bottom:14px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">HUB HOME · 390</span>
<span class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#707080">NEW TAB BAR — SLATE · EXPLORE · READ FAB · LEDGER · MORE</span>
</div>
<div style="display:flex;flex-wrap:wrap;gap:34px;justify-content:center;align-items:flex-start">
<div>
<div class="mono" style="font-size:10px;letter-spacing:.2em;color:#707080;margin-bottom:12px;text-align:center">S1-A · NFL HUB</div>
<x-import component-from-global-scope="IOSDevice" from="./ios-frame.jsx" dark="{{ true }}" hint-size="390px,844px">
<div style="display:flex;flex-direction:column;height:100%;background:#06060B;font-family:'Inter',system-ui,sans-serif;color:#F0F0F0;padding-top:54px">
<div style="display:flex;align-items:center;justify-content:space-between;padding:10px 16px;border-bottom:1px solid #14141E">
<div style="display:flex;align-items:center;gap:9px">
<svg width="24" height="24" viewBox="0 0 32 32" fill="none"><rect x="1" y="1" width="30" height="30" rx="7" fill="#0E0E14" stroke="#2A2A38" stroke-width="1"></rect><path d="M8 9 L16 24 L24 9" fill="none" stroke="#F0F0F0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><path d="M16 24 L24 9" fill="none" stroke="#00D4A0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><circle cx="24" cy="9" r="2" fill="#00D4A0"></circle></svg>
<span class="mono" style="font-weight:800;font-size:14px;letter-spacing:.26em">VYND<span style="color:#00D4A0">R</span></span>
</div>
<div style="display:flex;align-items:center;gap:7px">
<span style="width:6px;height:6px;border-radius:50%;background:#00D4A0;animation:vy-syncdot 1.4s ease-in-out infinite"></span>
<span class="mono" data-sync-clock style="font-size:10px;color:#B8BCC8;letter-spacing:.06em">--:--:--</span>
</div>
</div>
<div style="display:flex;align-items:center;gap:6px;padding:9px 16px;border-bottom:1px solid #14141E;overflow:hidden">
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:22px;padding:0 8px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#B8BCC8;flex:none"><span style="width:4px;height:4px;border-radius:50%;background:#00D4A0"></span>MLB 9</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 8px;border-radius:6px;background:#14141E;border:1px solid #2A2A38;font-size:8.5px;letter-spacing:.1em;color:#F0F0F0;font-weight:700;flex:none">NFL · CAMP 5D</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 8px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">NBA · SL</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 8px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">NHL</span>
</div>
<div style="display:flex;align-items:baseline;justify-content:space-between;padding:14px 16px 10px">
<div style="display:flex;align-items:baseline;gap:8px"><span style="font-size:19px;font-weight:800">NFL</span><span class="mono" style="font-size:8px;letter-spacing:.22em;color:#707080">OFFSEASON DESK</span></div>
<span class="mono" style="font-size:9px;letter-spacing:.12em;color:#707080"><span data-kickoff style="color:#B8BCC8;font-weight:700">{{ kickoffDays }}</span>D TO KICKOFF</span>
</div>
<div style="flex:1;overflow:hidden">
<div style="display:flex;align-items:center;justify-content:space-between;padding:4px 16px 6px">
<span class="mono" style="font-size:9.5px;letter-spacing:.22em;color:#F0F0F0;font-weight:700">WHAT CHANGED TODAY</span>
<span class="mono" style="font-size:8px;letter-spacing:.12em;color:#707080">JUL 17</span>
</div>
<div style="padding:10px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.55)">
<div style="font-size:12.5px;font-weight:700;line-height:1.3">Nabers cleared for camp — no restrictions</div>
<div class="mono" style="font-size:8.5px;letter-spacing:.06em;color:#707080;margin-top:4px"><span style="display:inline-block;width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#0B2265,#5a7fd4);vertical-align:-1px;margin-right:3px"></span>NYG · REC YDS <span style="color:#4a4a58">1,050.5</span><span style="color:#F0F0F0;font-weight:700">1,120.5</span> · <span style="color:#00D4A0">M 1,188</span></div>
</div>
<div style="padding:10px 16px;border-top:1px solid #101018;opacity:.85">
<div style="font-size:12.5px;font-weight:600;line-height:1.3">Worthy running clear WR1 reps in KC installs</div>
<div class="mono" style="font-size:8.5px;letter-spacing:.06em;color:#707080;margin-top:4px"><span style="display:inline-block;width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#E31837,#FFB81C);vertical-align:-1px;margin-right:3px"></span>KC · REC YDS <span style="color:#4a4a58">990.5</span><span style="color:#F0F0F0;font-weight:700">1,024.5</span> · <span style="color:#00D4A0">M 1,061</span></div>
</div>
<div style="padding:10px 16px;border-top:1px solid #101018;opacity:.7">
<div style="font-size:12.5px;font-weight:600;line-height:1.3">Higgins restructure done — reports Tuesday</div>
<div class="mono" style="font-size:8.5px;letter-spacing:.06em;color:#707080;margin-top:4px"><span style="display:inline-block;width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#FB4F14,#1a1a1a);vertical-align:-1px;margin-right:3px"></span>CIN · PASS TD <span style="color:#4a4a58">27.5</span><span style="color:#F0F0F0;font-weight:700">28.5</span> · <span style="color:#00D4A0">M 29.8</span></div>
</div>
<div style="display:flex;align-items:center;justify-content:space-between;padding:12px 16px 6px;border-top:1px solid #14141E;margin-top:2px">
<span class="mono" style="font-size:9.5px;letter-spacing:.22em;color:#F0F0F0;font-weight:700">SEASON BOARD</span>
<span class="mono" style="font-size:8px;letter-spacing:.12em;color:#00D4A0">FULL BOARD ▸</span>
</div>
<div style="display:flex;align-items:center;gap:10px;padding:10px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.55)">
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:700">Nabers <span style="color:#B8BCC8;font-weight:500">o1,120.5 Rec Yds</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">O 1,050.5 → 1,120.5 · <span style="color:#00D4A0">M 1,188</span></div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:rgba(0,212,160,.14);color:#00D4A0;font-weight:800;font-size:10px;border:1px solid rgba(0,212,160,.3)">A+</span>
</div>
<div style="display:flex;align-items:center;gap:10px;padding:10px 16px;border-top:1px solid #101018;opacity:.85">
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:600">Allen <span style="color:#B8BCC8;font-weight:500">o37.5 P+R TD</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">O 36.5 → 37.5 · <span style="color:#00D4A0">M 40.2</span></div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:rgba(0,212,160,.1);color:#00D4A0;font-weight:800;font-size:10px;border:1px solid rgba(0,212,160,.24)">A</span>
</div>
<!-- key dates compact -->
<div style="display:flex;align-items:center;gap:0;padding:12px 16px 0">
<div style="flex:1;text-align:left"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#F0F0F0;font-weight:700">CAMPS</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">JUL 22</div></div>
<div style="flex:1;text-align:center"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#707080">PRESEASON</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">AUG 6</div></div>
<div style="flex:1;text-align:center"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#707080">CUTS</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">SEP 1</div></div>
<div style="flex:1;text-align:right"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#00D4A0;font-weight:700">KICKOFF</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">SEP 10</div></div>
</div>
</div>
<!-- wire -->
<div style="display:flex;align-items:center;gap:9px;padding:9px 16px;border-top:1px solid #14141E;background:#08080D">
<span class="mono" style="font-size:8px;letter-spacing:.22em;color:#00D4A0;font-weight:800;flex:none">WIRE</span>
<span class="mono" data-wire style="font-size:9.5px;color:#B8BCC8;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1">Nabers 1,050.5 → 1,120.5 on camp clearance.</span>
</div>
<!-- tab bar w/ READ fab -->
<div style="position:relative;display:flex;align-items:center;justify-content:space-around;padding:10px 8px 6px;border-top:1px solid #14141E;background:rgba(6,6,11,.9)">
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><path stroke="#00D4A0" stroke-width="2" stroke-linecap="round" d="M4 6h16M4 12h16M4 18h9"></path></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#00D4A0;font-weight:700">SLATE</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><circle cx="11" cy="11" r="6.5" stroke="#707080" stroke-width="2"></circle><path stroke="#707080" stroke-width="2" stroke-linecap="round" d="M19.5 19.5 16 16"></path></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">EXPLORE</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:3px;min-width:60px;padding:0;transform:translateY(-14px)">
<span style="width:50px;height:50px;border-radius:50%;background:#00D4A0;display:flex;align-items:center;justify-content:center;box-shadow:0 8px 24px -6px rgba(0,212,160,.55),0 0 0 6px #06060B"><svg width="24" height="24" viewBox="0 0 32 32" fill="none"><path d="M8 9 L16 24 L24 9" fill="none" stroke="#06060B" stroke-width="3.2" stroke-linecap="round" stroke-linejoin="round"></path><circle cx="24" cy="9" r="2.2" fill="#06060B"></circle></svg></span>
<span class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#00D4A0;font-weight:800">READ</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><path stroke="#707080" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" d="M5 4h14v16H5zM8 9h8M8 13h6"></path></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">LEDGER</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><circle cx="5" cy="12" r="1.8" fill="#707080"></circle><circle cx="12" cy="12" r="1.8" fill="#707080"></circle><circle cx="19" cy="12" r="1.8" fill="#707080"></circle></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">MORE</span>
</div>
</div>
</div>
</x-import>
</div>
<div>
<div class="mono" style="font-size:10px;letter-spacing:.2em;color:#707080;margin-bottom:12px;text-align:center">S1-B · NBA VARIANT · SUMMER LEAGUE HONESTY</div>
<x-import component-from-global-scope="IOSDevice" from="./ios-frame.jsx" dark="{{ true }}" hint-size="390px,844px">
<div style="display:flex;flex-direction:column;height:100%;background:#06060B;font-family:'Inter',system-ui,sans-serif;color:#F0F0F0;padding-top:54px">
<div style="display:flex;align-items:center;justify-content:space-between;padding:10px 16px;border-bottom:1px solid #14141E">
<div style="display:flex;align-items:center;gap:9px">
<svg width="24" height="24" viewBox="0 0 32 32" fill="none"><rect x="1" y="1" width="30" height="30" rx="7" fill="#0E0E14" stroke="#2A2A38" stroke-width="1"></rect><path d="M8 9 L16 24 L24 9" fill="none" stroke="#F0F0F0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><path d="M16 24 L24 9" fill="none" stroke="#00D4A0" stroke-width="2.6" stroke-linecap="round" stroke-linejoin="round"></path><circle cx="24" cy="9" r="2" fill="#00D4A0"></circle></svg>
<span class="mono" style="font-weight:800;font-size:14px;letter-spacing:.26em">VYND<span style="color:#00D4A0">R</span></span>
</div>
<div style="display:flex;align-items:center;gap:7px">
<span style="width:6px;height:6px;border-radius:50%;background:#00D4A0;animation:vy-syncdot 1.4s ease-in-out infinite"></span>
<span class="mono" data-sync-clock style="font-size:10px;color:#B8BCC8;letter-spacing:.06em">--:--:--</span>
</div>
</div>
<div style="display:flex;align-items:center;gap:6px;padding:9px 16px;border-bottom:1px solid #14141E;overflow:hidden">
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:22px;padding:0 8px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#B8BCC8;flex:none"><span style="width:4px;height:4px;border-radius:50%;background:#00D4A0"></span>MLB 9</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 8px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">NFL · CAMP</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 8px;border-radius:6px;background:#14141E;border:1px solid #2A2A38;font-size:8.5px;letter-spacing:.1em;color:#F0F0F0;font-weight:700;flex:none">NBA · SL DAY 8</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 8px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">NHL</span>
</div>
<div style="display:flex;align-items:baseline;justify-content:space-between;padding:14px 16px 10px">
<div style="display:flex;align-items:baseline;gap:8px"><span style="font-size:19px;font-weight:800">NBA</span><span class="mono" style="font-size:8px;letter-spacing:.22em;color:#707080">SUMMER DESK</span></div>
<span class="mono" style="font-size:9px;letter-spacing:.12em;color:#707080">95D TO OPENING NIGHT</span>
</div>
<div style="flex:1;overflow:hidden">
<!-- summer league: outlook only, engine refuses -->
<div style="margin:2px 16px 0;border-radius:11px;background:#0E0E14;border:1px solid #1E1E2A;overflow:hidden">
<div style="display:flex;align-items:center;justify-content:space-between;padding:11px 13px 9px">
<span class="mono" style="font-size:9.5px;letter-spacing:.2em;color:#F0F0F0;font-weight:700">SUMMER LEAGUE · LAS VEGAS</span>
<span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:rgba(255,179,71,.08);border:1px solid rgba(255,179,71,.25);font-size:7.5px;letter-spacing:.12em;color:#FFB347">OUTLOOK ONLY</span>
</div>
<div style="padding:0 13px 10px;font-size:10.5px;line-height:1.45;color:#707080">We don't grade Summer League props. Ten-day samples against fringe rosters aren't signal — they're noise wearing a jersey. We read minutes, roles, and rotation intent instead.</div>
<div style="padding:10px 13px;border-top:1px solid #101018">
<div style="display:flex;align-items:center;gap:9px"><span style="width:26px;height:26px;border-radius:7px;background:linear-gradient(135deg,#1E1E2A,#2A2A38);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:9px;color:#B8BCC8;flex:none">AD</span><div style="flex:1;min-width:0"><div style="font-size:12px;font-weight:700">Dybantsa <span class="mono" style="font-size:8px;color:#707080;letter-spacing:.08em">F · R</span></div><div class="mono" style="font-size:8px;letter-spacing:.06em;color:#707080;margin-top:1px">31 MIN/G · USAGE-HEAVY · STARTER-TRACK</div></div><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:7.5px;letter-spacing:.12em;color:#707080;flex:none">NOT GRADED</span></div>
</div>
<div style="padding:10px 13px;border-top:1px solid #101018;opacity:.85">
<div style="display:flex;align-items:center;gap:9px"><span style="width:26px;height:26px;border-radius:7px;background:linear-gradient(135deg,#1E1E2A,#2A2A38);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:9px;color:#B8BCC8;flex:none">CB</span><div style="flex:1;min-width:0"><div style="font-size:12px;font-weight:600">Boozer <span class="mono" style="font-size:8px;color:#707080;letter-spacing:.08em">F · R</span></div><div class="mono" style="font-size:8px;letter-spacing:.06em;color:#707080;margin-top:1px">28 MIN/G · CLOSING LINEUPS · ROTATION LOCK</div></div><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:7.5px;letter-spacing:.12em;color:#707080;flex:none">NOT GRADED</span></div>
</div>
<div style="padding:10px 13px;border-top:1px solid #101018;opacity:.7">
<div style="display:flex;align-items:center;gap:9px"><span style="width:26px;height:26px;border-radius:7px;background:linear-gradient(135deg,#1E1E2A,#2A2A38);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:9px;color:#B8BCC8;flex:none">DP</span><div style="flex:1;min-width:0"><div style="font-size:12px;font-weight:600">Peterson <span class="mono" style="font-size:8px;color:#707080;letter-spacing:.08em">G · R</span></div><div class="mono" style="font-size:8px;letter-spacing:.06em;color:#707080;margin-top:1px">24 MIN/G · SECOND UNIT · WATCHLIST</div></div><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:7.5px;letter-spacing:.12em;color:#707080;flex:none">NOT GRADED</span></div>
</div>
</div>
<!-- futures preview -->
<div style="display:flex;align-items:center;justify-content:space-between;padding:14px 16px 6px">
<span class="mono" style="font-size:9.5px;letter-spacing:.22em;color:#F0F0F0;font-weight:700">SEASON BOARD</span>
<span class="mono" style="font-size:8px;letter-spacing:.12em;color:#00D4A0">FULL BOARD ▸</span>
</div>
<div style="display:flex;align-items:center;gap:10px;padding:10px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.55)">
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:700">Wembanyama <span style="color:#B8BCC8;font-weight:500">MVP</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">O +420 → +330 · <span style="color:#00D4A0">FAIR +270</span></div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:rgba(0,212,160,.1);color:#00D4A0;font-weight:800;font-size:10px;border:1px solid rgba(0,212,160,.24)">A</span>
</div>
<div style="display:flex;align-items:center;gap:10px;padding:10px 16px;border-top:1px solid #101018;opacity:.85">
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:600">Thunder <span style="color:#B8BCC8;font-weight:500">o61.5 Wins</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">O 60.5 → 61.5 · <span style="color:#00D4A0">M 60.9</span> · THIN</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:#14141E;color:#B8BCC8;font-weight:800;font-size:10px;border:1px solid #1E1E2A">C</span>
</div>
<!-- key dates -->
<div style="display:flex;align-items:center;gap:0;padding:12px 16px 0">
<div style="flex:1;text-align:left"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#F0F0F0;font-weight:700">SL ENDS</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">JUL 19</div></div>
<div style="flex:1;text-align:center"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#707080">MEDIA DAY</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">SEP 28</div></div>
<div style="flex:1;text-align:center"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#707080">CAMPS</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">SEP 29</div></div>
<div style="flex:1;text-align:right"><div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#00D4A0;font-weight:700">OPENER</div><div class="mono" style="font-size:7.5px;color:#4a4a58;margin-top:1px">OCT 20</div></div>
</div>
</div>
<div style="display:flex;align-items:center;gap:9px;padding:9px 16px;border-top:1px solid #14141E;background:#08080D">
<span class="mono" style="font-size:8px;letter-spacing:.22em;color:#00D4A0;font-weight:800;flex:none">WIRE</span>
<span class="mono" style="font-size:9.5px;color:#B8BCC8;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1">Wemby MVP steams +420 → +330 since June.</span>
</div>
<div style="position:relative;display:flex;align-items:center;justify-content:space-around;padding:10px 8px 6px;border-top:1px solid #14141E;background:rgba(6,6,11,.9)">
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><path stroke="#00D4A0" stroke-width="2" stroke-linecap="round" d="M4 6h16M4 12h16M4 18h9"></path></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#00D4A0;font-weight:700">SLATE</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><circle cx="11" cy="11" r="6.5" stroke="#707080" stroke-width="2"></circle><path stroke="#707080" stroke-width="2" stroke-linecap="round" d="M19.5 19.5 16 16"></path></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">EXPLORE</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:3px;min-width:60px;padding:0;transform:translateY(-14px)">
<span style="width:50px;height:50px;border-radius:50%;background:#00D4A0;display:flex;align-items:center;justify-content:center;box-shadow:0 8px 24px -6px rgba(0,212,160,.55),0 0 0 6px #06060B"><svg width="24" height="24" viewBox="0 0 32 32" fill="none"><path d="M8 9 L16 24 L24 9" fill="none" stroke="#06060B" stroke-width="3.2" stroke-linecap="round" stroke-linejoin="round"></path><circle cx="24" cy="9" r="2.2" fill="#06060B"></circle></svg></span>
<span class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#00D4A0;font-weight:800">READ</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><path stroke="#707080" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" d="M5 4h14v16H5zM8 9h8M8 13h6"></path></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">LEDGER</span>
</div>
<div style="display:flex;flex-direction:column;align-items:center;gap:4px;min-width:52px;padding:2px 0">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none"><circle cx="5" cy="12" r="1.8" fill="#707080"></circle><circle cx="12" cy="12" r="1.8" fill="#707080"></circle><circle cx="19" cy="12" r="1.8" fill="#707080"></circle></svg>
<span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">MORE</span>
</div>
</div>
</div>
</x-import>
</div>
</div>
</div>
<!-- ============ ACT 02 · SEASON-LONG BOARD ============ -->
<div data-reveal style="margin-top:110px;display:flex;align-items:baseline;gap:18px;border-bottom:1px solid #14141E;padding-bottom:18px">
<span class="mono" style="font-size:64px;font-weight:800;line-height:.8;color:#14141E;-webkit-text-stroke:1px #2A2A38">02</span>
<div><div class="mono" style="font-size:13px;letter-spacing:.3em;color:#F0F0F0;font-weight:800">SEASON-LONG BOARD</div><div style="font-size:12.5px;color:#707080;margin-top:5px">The open → current → model triplet is the signature. Ladder pattern: one row per player, lines nested.</div></div>
</div>
<div id="board" data-screen-label="Season board · desktop" style="margin-top:44px">
<div style="display:flex;align-items:center;gap:12px;margin-bottom:14px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">NFL SEASON BOARD · DESKTOP</span>
<span class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#707080">OPEN DIM · NOW WHITE · MODEL GREEN — ONE READING ORDER, EVERY ROW</span>
</div>
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;overflow:hidden;box-shadow:0 30px 60px -34px rgba(0,0,0,.9),inset 0 1px 0 rgba(255,255,255,.03)">
<!-- filters -->
<div style="display:flex;align-items:center;gap:6px;padding:14px 16px;border-bottom:1px solid #14141E;flex-wrap:wrap">
<span class="mono" style="display:inline-flex;align-items:center;height:24px;padding:0 11px;border-radius:7px;background:#14141E;border:1px solid #2A2A38;font-size:9px;letter-spacing:.14em;color:#F0F0F0;font-weight:700">ALL MARKETS</span>
<span class="mono" style="display:inline-flex;align-items:center;height:24px;padding:0 11px;border-radius:7px;font-size:9px;letter-spacing:.14em;color:#707080">PASSING</span>
<span class="mono" style="display:inline-flex;align-items:center;height:24px;padding:0 11px;border-radius:7px;font-size:9px;letter-spacing:.14em;color:#707080">RUSH / REC</span>
<span class="mono" style="display:inline-flex;align-items:center;height:24px;padding:0 11px;border-radius:7px;font-size:9px;letter-spacing:.14em;color:#707080">TDS</span>
<span class="mono" style="display:inline-flex;align-items:center;height:24px;padding:0 11px;border-radius:7px;font-size:9px;letter-spacing:.14em;color:#707080">WIN TOTALS</span>
<span class="mono" style="display:inline-flex;align-items:center;height:24px;padding:0 11px;border-radius:7px;font-size:9px;letter-spacing:.14em;color:#707080">DIVISIONS</span>
<span class="mono" style="display:inline-flex;align-items:center;height:24px;padding:0 11px;border-radius:7px;font-size:9px;letter-spacing:.14em;color:#707080">AWARDS</span>
<span class="mono" style="margin-left:auto;font-size:9px;letter-spacing:.16em;color:#707080">SORTED · MODEL GAP ▼</span>
</div>
<!-- column heads -->
<div style="display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:9px 16px 7px">
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58">PLAYER / MARKET</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58;justify-self:end">OPEN</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58;justify-self:end">NOW</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#00D4A0;justify-self:end">VYNDR</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58;justify-self:end">EDGE</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58;justify-self:end">GRD</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58;justify-self:end">MOVE · SINCE OPEN</span>
</div>
<div style="height:1px;background:#14141E;transform-origin:left;animation:vy-bootline .6s ease-out both"></div>
<!-- ladder: NABERS -->
<div style="display:flex;align-items:center;gap:10px;padding:11px 16px 7px;border-top:1px solid #101018">
<span style="position:relative;width:26px;height:26px;border-radius:7px;background:linear-gradient(135deg,#0B2265,#5a7fd4);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:9px;color:#cdd9ff;flex:none">MN<image-slot id="hs-nabers" shape="rounded" radius="7" style="position:absolute;inset:0" placeholder="Nabers headshot"></image-slot></span>
<span style="font-size:13.5px;font-weight:700">Malik Nabers</span>
<span class="mono" style="font-size:9px;color:#707080;letter-spacing:.08em">WR · NYG</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:18px;padding:0 8px;border-radius:5px;background:rgba(255,122,74,.09);border:1px solid rgba(255,122,74,.28);font-size:8.5px;letter-spacing:.14em;color:#FF7A4A;font-weight:700"><svg width="11" height="11" viewBox="0 0 24 24" fill="none" style="color:#FF7A4A"><path fill="none" stroke="currentColor" stroke-width="2.4" stroke-linecap="round" stroke-linejoin="round" d="M3 19 19 5M14 5h5v5"></path><path stroke="currentColor" stroke-width="2" stroke-linecap="round" d="M4 14l2 2M8 10l2 2"></path></svg>BURNER</span>
<span class="mono" style="margin-left:auto;font-size:8.5px;letter-spacing:.12em;color:#4a4a58">2 LINES</span>
</div>
<div style="--o:1;opacity:1;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 8px 52px;box-shadow:inset 2px 0 0 rgba(0,212,160,.55)">
<span style="font-size:12.5px;font-weight:600;color:#B8BCC8">Receiving Yards <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">O/U</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">1,050.5</span>
<span class="mono" data-react data-fmt="line" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0;padding:1px 4px;border-radius:5px">1,120.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">1,188</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+6.0%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:rgba(0,212,160,.14);color:#00D4A0;font-weight:800;font-size:10.5px;border:1px solid rgba(0,212,160,.3)">A+</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,16 18,16 18,12 40,12 40,9 62,9 62,4 86,4" fill="none" stroke="#00D4A0" stroke-width="1.5" stroke-linejoin="round" opacity=".85"></polyline><circle cx="86" cy="4" r="2.2" fill="#00D4A0"></circle></svg><span class="mono" style="font-size:9px;color:#00D4A0;font-weight:700">▲ +70.0</span></span>
</div>
<div style="--o:.86;opacity:.86;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 10px 52px;border-bottom:1px solid #101018">
<span style="font-size:12.5px;font-weight:500;color:#B8BCC8">Receptions <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">O/U</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">88.5</span>
<span class="mono" data-react data-fmt="half" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0;padding:1px 4px;border-radius:5px">92.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">97.0</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+3.2%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">B</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,14 26,14 26,11 56,11 56,7 86,7" fill="none" stroke="#B8BCC8" stroke-width="1.5" stroke-linejoin="round" opacity=".6"></polyline><circle cx="86" cy="7" r="2" fill="#B8BCC8"></circle></svg><span class="mono" style="font-size:9px;color:#B8BCC8;font-weight:700">▲ +4.0</span></span>
</div>
<!-- ladder: ALLEN -->
<div style="display:flex;align-items:center;gap:10px;padding:11px 16px 7px;border-top:1px solid #101018">
<span style="width:26px;height:26px;border-radius:7px;background:linear-gradient(135deg,#00338D,#C60C30);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:9px;color:#bcd0ff;flex:none">JA</span>
<span style="font-size:13.5px;font-weight:700">Josh Allen</span>
<span class="mono" style="font-size:9px;color:#707080;letter-spacing:.08em">QB · BUF</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:18px;padding:0 8px;border-radius:5px;background:rgba(217,164,65,.09);border:1px solid rgba(217,164,65,.28);font-size:8.5px;letter-spacing:.14em;color:#D9A441;font-weight:700"><svg width="11" height="11" viewBox="0 0 24 24" fill="none" style="color:#D9A441"><path fill="none" stroke="currentColor" stroke-width="2.2" stroke-linecap="round" d="M5 19 16 6"></path><circle cx="17" cy="5" r="1.9" fill="currentColor"></circle><path fill="none" stroke="currentColor" stroke-width="1.8" stroke-linecap="round" d="M5 9c2-2 5-2 7 0M4 13c2.5-2.5 6-2.5 9 0"></path></svg>MAESTRO</span>
<span class="mono" style="margin-left:auto;font-size:8.5px;letter-spacing:.12em;color:#4a4a58">3 LINES</span>
</div>
<div style="--o:.86;opacity:.86;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 8px 52px">
<span style="font-size:12.5px;font-weight:500;color:#B8BCC8">Pass + Rush TD <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">O/U</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">36.5</span>
<span class="mono" data-react data-fmt="half" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0;padding:1px 4px;border-radius:5px">37.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">40.2</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+4.8%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:rgba(0,212,160,.1);color:#00D4A0;font-weight:800;font-size:10.5px;border:1px solid rgba(0,212,160,.24)">A</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,13 34,13 34,9 86,9" fill="none" stroke="#B8BCC8" stroke-width="1.5" stroke-linejoin="round" opacity=".6"></polyline><circle cx="86" cy="9" r="2" fill="#B8BCC8"></circle></svg><span class="mono" style="font-size:9px;color:#B8BCC8;font-weight:700">▲ +1.0</span></span>
</div>
<div style="--o:.86;opacity:.86;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 8px 52px">
<span style="font-size:12.5px;font-weight:500;color:#B8BCC8">Passing Yards <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">O/U</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">3,895.5</span>
<span class="mono" data-react data-fmt="line" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0;padding:1px 4px;border-radius:5px">3,940.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">4,080</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+3.6%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">B</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,13 40,13 40,10 86,10" fill="none" stroke="#B8BCC8" stroke-width="1.5" stroke-linejoin="round" opacity=".6"></polyline><circle cx="86" cy="10" r="2" fill="#B8BCC8"></circle></svg><span class="mono" style="font-size:9px;color:#B8BCC8;font-weight:700">▲ +45.0</span></span>
</div>
<div style="--o:.64;opacity:.64;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 10px 52px;border-bottom:1px solid #101018">
<span style="font-size:12.5px;font-weight:500;color:#B8BCC8">MVP <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">AWARDS</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">+650</span>
<span class="mono" data-react data-fmt="odds" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0;padding:1px 4px;border-radius:5px">+600</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">+490</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+2.6%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">B</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,12 50,12 50,9 86,9" fill="none" stroke="#B8BCC8" stroke-width="1.5" stroke-linejoin="round" opacity=".6"></polyline><circle cx="86" cy="9" r="2" fill="#B8BCC8"></circle></svg><span class="mono" style="font-size:9px;color:#B8BCC8;font-weight:700">▲ 650→600</span></span>
</div>
<!-- ladder: BIJAN -->
<div style="display:flex;align-items:center;gap:10px;padding:11px 16px 7px;border-top:1px solid #101018">
<span style="width:26px;height:26px;border-radius:7px;background:linear-gradient(135deg,#A71930,#1a1a1a);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:9px;color:#ffb8c4;flex:none">BR</span>
<span style="font-size:13.5px;font-weight:700">Bijan Robinson</span>
<span class="mono" style="font-size:9px;color:#707080;letter-spacing:.08em">RB · ATL</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:18px;padding:0 8px;border-radius:5px;background:rgba(181,137,90,.09);border:1px solid rgba(181,137,90,.28);font-size:8.5px;letter-spacing:.14em;color:#B5895A;font-weight:700"><svg width="11" height="11" viewBox="0 0 24 24" fill="none" style="color:#B5895A"><path fill="currentColor" fill-opacity=".16" d="M6 17c0-6 2-9 6-9s6 3 6 9z"></path><path fill="none" stroke="currentColor" stroke-width="2" stroke-linejoin="round" d="M6 17c0-6 2-9 6-9s6 3 6 9zM4 17h16M12 8V5"></path><circle cx="12" cy="20" r="1.6" fill="currentColor"></circle></svg>BELL COW</span>
<span class="mono" style="margin-left:auto;font-size:8.5px;letter-spacing:.12em;color:#4a4a58">2 LINES</span>
</div>
<div style="--o:.64;opacity:.64;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 8px 52px">
<span style="font-size:12.5px;font-weight:500;color:#B8BCC8">Rushing Yards <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">O/U</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">1,275.5</span>
<span class="mono" data-react data-fmt="line" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0;padding:1px 4px;border-radius:5px">1,290.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">1,342</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+2.2%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#B8BCC8;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">C</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,12 60,12 60,10 86,10" fill="none" stroke="#B8BCC8" stroke-width="1.5" stroke-linejoin="round" opacity=".5"></polyline><circle cx="86" cy="10" r="2" fill="#B8BCC8"></circle></svg><span class="mono" style="font-size:9px;color:#B8BCC8;font-weight:700">▲ +15.0</span></span>
</div>
<div style="--o:.48;opacity:.48;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 10px 52px;border-bottom:1px solid #101018">
<span style="font-size:12.5px;font-weight:500;color:#B8BCC8">Rushing TD <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">O/U</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">11.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0">11.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">12.4</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+1.4%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#B8BCC8;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">C</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,11 86,11" fill="none" stroke="#4a4a58" stroke-width="1.5" opacity=".7"></polyline></svg><span class="mono" style="font-size:9px;color:#4a4a58;font-weight:700">FLAT · 41D</span></span>
</div>
<!-- ladder: WIN TOTALS -->
<div style="display:flex;align-items:center;gap:10px;padding:11px 16px 7px;border-top:1px solid #101018">
<span class="mono" style="font-size:10px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">TEAM WIN TOTALS</span>
<span class="mono" style="margin-left:auto;font-size:8.5px;letter-spacing:.12em;color:#4a4a58">32 TEAMS</span>
</div>
<div style="--o:.86;opacity:.86;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 8px 52px">
<span style="font-size:12.5px;font-weight:600;color:#B8BCC8"><span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#0076B6,#B0B7BC);vertical-align:-1.5px;margin-right:6px"></span>DET Lions <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">WINS O/U</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">10.5</span>
<span class="mono" data-react data-fmt="half" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0;padding:1px 4px;border-radius:5px">11.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">11.9</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+1.9%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#B8BCC8;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">C</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,14 44,14 44,8 86,8" fill="none" stroke="#B8BCC8" stroke-width="1.5" stroke-linejoin="round" opacity=".6"></polyline><circle cx="86" cy="8" r="2" fill="#B8BCC8"></circle></svg><span class="mono" style="font-size:9px;color:#B8BCC8;font-weight:700">▲ +1.0</span></span>
</div>
<div style="--o:.86;opacity:.86;display:grid;grid-template-columns:1fr 88px 88px 88px 64px 44px 150px;align-items:center;gap:12px;padding:8px 16px 10px 52px;border-bottom:1px solid #101018">
<span style="font-size:12.5px;font-weight:600;color:#B8BCC8"><span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,#E31837,#FFB81C);vertical-align:-1.5px;margin-right:6px"></span>KC Chiefs <span class="mono" style="font-size:8.5px;color:#4a4a58;letter-spacing:.1em">WINS O/U · MODEL LEANS UNDER</span></span>
<span class="mono" style="justify-self:end;font-size:11.5px;color:#707080">11.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:700;color:#F0F0F0">11.5</span>
<span class="mono" style="justify-self:end;font-size:12.5px;font-weight:800;color:#00D4A0">10.8</span>
<span class="mono" style="justify-self:end;font-size:11.5px;font-weight:700;color:#00D4A0">+3.1%</span>
<span class="mono" style="justify-self:end;display:inline-flex;align-items:center;height:18px;padding:0 7px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10.5px;border:1px solid #1E1E2A">B</span>
<span style="justify-self:end;display:flex;align-items:center;gap:8px"><svg width="86" height="20" viewBox="0 0 86 20" preserveAspectRatio="none" style="display:block"><polyline points="0,10 86,10" fill="none" stroke="#4a4a58" stroke-width="1.5" opacity=".7"></polyline></svg><span class="mono" style="font-size:9px;color:#4a4a58;font-weight:700">FLAT · 62D</span></span>
</div>
<div style="padding:12px 16px;border-top:1px solid #101018;display:flex;align-items:center;justify-content:space-between;flex-wrap:wrap;gap:8px">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#4a4a58">+ 207 SEASON LINES · REPRICED ON NEWS · LAST REPRICE JUL 17 · 11:42 AM ET</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">EDGE = MODEL VS CURRENT · ODDS MARKETS GRADED ON FAIR PRICE</span>
</div>
</div>
</div>
<!-- board mobile -->
<div data-screen-label="Season board · mobile" style="margin-top:44px">
<div style="display:flex;flex-wrap:wrap;gap:34px;justify-content:center;align-items:flex-start">
<div>
<div class="mono" style="font-size:10px;letter-spacing:.2em;color:#707080;margin-bottom:12px;text-align:center">S1-C · SEASON BOARD · 390</div>
<x-import component-from-global-scope="IOSDevice" from="./ios-frame.jsx" dark="{{ true }}" hint-size="390px,844px">
<div style="display:flex;flex-direction:column;height:100%;background:#06060B;font-family:'Inter',system-ui,sans-serif;color:#F0F0F0;padding-top:54px">
<div style="display:flex;align-items:center;justify-content:space-between;padding:10px 16px;border-bottom:1px solid #14141E">
<div style="display:flex;align-items:center;gap:9px">
<span class="mono" style="font-size:11px;color:#707080"></span>
<span class="mono" style="font-weight:800;font-size:11px;letter-spacing:.2em">NFL <span style="color:#707080;font-weight:600">· SEASON BOARD</span></span>
</div>
<span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080">214 LINES</span>
</div>
<div style="display:flex;align-items:center;gap:6px;padding:9px 16px;border-bottom:1px solid #14141E">
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 9px;border-radius:6px;background:#14141E;border:1px solid #2A2A38;font-size:8.5px;letter-spacing:.1em;color:#F0F0F0;font-weight:700;flex:none">ALL</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 9px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">PASS</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 9px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">RUSH/REC</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 9px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">WINS</span>
<span class="mono" style="display:inline-flex;align-items:center;height:22px;padding:0 9px;border-radius:6px;font-size:8.5px;letter-spacing:.1em;color:#707080;flex:none">AWARDS</span>
</div>
<div style="flex:1;overflow:hidden">
<!-- player group -->
<div style="display:flex;align-items:center;gap:8px;padding:12px 16px 6px">
<span style="width:22px;height:22px;border-radius:6px;background:linear-gradient(135deg,#0B2265,#5a7fd4);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:8px;color:#cdd9ff;flex:none">MN</span>
<span style="font-size:13px;font-weight:700">Nabers</span>
<span class="mono" style="font-size:8px;color:#707080;letter-spacing:.08em">WR · NYG</span>
<span class="mono" style="margin-left:auto;display:inline-flex;align-items:center;gap:4px;height:16px;padding:0 6px;border-radius:4px;background:rgba(255,122,74,.09);border:1px solid rgba(255,122,74,.28);font-size:7.5px;letter-spacing:.12em;color:#FF7A4A;font-weight:700">BURNER</span>
</div>
<div style="padding:9px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.55);display:flex;align-items:center;gap:10px">
<div style="flex:1;min-width:0">
<div style="font-size:12.5px;font-weight:600;color:#B8BCC8">Rec Yds</div>
<div class="mono" style="font-size:9px;letter-spacing:.04em;color:#707080;margin-top:3px"><span style="color:#4a4a58">1,050.5</span><span style="color:#F0F0F0;font-weight:700;font-size:11px">1,120.5</span> · <span style="color:#00D4A0;font-weight:700">V 1,188</span></div>
</div>
<span class="mono" style="font-size:10px;font-weight:700;color:#00D4A0;flex:none">▲ +70</span>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:rgba(0,212,160,.14);color:#00D4A0;font-weight:800;font-size:10px;border:1px solid rgba(0,212,160,.3);flex:none">A+</span>
</div>
<div style="padding:9px 16px;border-top:1px solid #101018;opacity:.85;display:flex;align-items:center;gap:10px">
<div style="flex:1;min-width:0">
<div style="font-size:12.5px;font-weight:500;color:#B8BCC8">Receptions</div>
<div class="mono" style="font-size:9px;letter-spacing:.04em;color:#707080;margin-top:3px"><span style="color:#4a4a58">88.5</span><span style="color:#F0F0F0;font-weight:700;font-size:11px">92.5</span> · <span style="color:#00D4A0;font-weight:700">V 97.0</span></div>
</div>
<span class="mono" style="font-size:10px;font-weight:700;color:#B8BCC8;flex:none">▲ +4</span>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10px;border:1px solid #1E1E2A;flex:none">B</span>
</div>
<div style="display:flex;align-items:center;gap:8px;padding:12px 16px 6px;border-top:1px solid #101018">
<span style="width:22px;height:22px;border-radius:6px;background:linear-gradient(135deg,#00338D,#C60C30);display:flex;align-items:center;justify-content:center;font-family:'JetBrains Mono',monospace;font-weight:800;font-size:8px;color:#bcd0ff;flex:none">JA</span>
<span style="font-size:13px;font-weight:700">Allen</span>
<span class="mono" style="font-size:8px;color:#707080;letter-spacing:.08em">QB · BUF</span>
<span class="mono" style="margin-left:auto;display:inline-flex;align-items:center;gap:4px;height:16px;padding:0 6px;border-radius:4px;background:rgba(217,164,65,.09);border:1px solid rgba(217,164,65,.28);font-size:7.5px;letter-spacing:.12em;color:#D9A441;font-weight:700">MAESTRO</span>
</div>
<div style="padding:9px 16px;border-top:1px solid #101018;opacity:.85;display:flex;align-items:center;gap:10px">
<div style="flex:1;min-width:0">
<div style="font-size:12.5px;font-weight:500;color:#B8BCC8">Pass + Rush TD</div>
<div class="mono" style="font-size:9px;letter-spacing:.04em;color:#707080;margin-top:3px"><span style="color:#4a4a58">36.5</span><span style="color:#F0F0F0;font-weight:700;font-size:11px">37.5</span> · <span style="color:#00D4A0;font-weight:700">V 40.2</span></div>
</div>
<span class="mono" style="font-size:10px;font-weight:700;color:#B8BCC8;flex:none">▲ +1</span>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:rgba(0,212,160,.1);color:#00D4A0;font-weight:800;font-size:10px;border:1px solid rgba(0,212,160,.24);flex:none">A</span>
</div>
<div style="padding:9px 16px;border-top:1px solid #101018;opacity:.7;display:flex;align-items:center;gap:10px">
<div style="flex:1;min-width:0">
<div style="font-size:12.5px;font-weight:500;color:#B8BCC8">MVP <span class="mono" style="font-size:7.5px;color:#4a4a58;letter-spacing:.1em">FAIR PRICE</span></div>
<div class="mono" style="font-size:9px;letter-spacing:.04em;color:#707080;margin-top:3px"><span style="color:#4a4a58">+650</span><span style="color:#F0F0F0;font-weight:700;font-size:11px">+600</span> · <span style="color:#00D4A0;font-weight:700">V +490</span></div>
</div>
<span class="mono" style="font-size:10px;font-weight:700;color:#B8BCC8;flex:none"></span>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:#14141E;color:#F0F0F0;font-weight:800;font-size:10px;border:1px solid #1E1E2A;flex:none">B</span>
</div>
<div style="display:flex;align-items:center;gap:8px;padding:12px 16px 6px;border-top:1px solid #101018">
<span class="mono" style="font-size:9px;letter-spacing:.2em;color:#F0F0F0;font-weight:700">TEAM WIN TOTALS</span>
</div>
<div style="padding:9px 16px;border-top:1px solid #101018;opacity:.85;display:flex;align-items:center;gap:10px">
<div style="flex:1;min-width:0">
<div style="font-size:12.5px;font-weight:500;color:#B8BCC8"><span style="display:inline-block;width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#0076B6,#B0B7BC);vertical-align:-1px;margin-right:4px"></span>DET Wins</div>
<div class="mono" style="font-size:9px;letter-spacing:.04em;color:#707080;margin-top:3px"><span style="color:#4a4a58">10.5</span><span style="color:#F0F0F0;font-weight:700;font-size:11px">11.5</span> · <span style="color:#00D4A0;font-weight:700">V 11.9</span></div>
</div>
<span class="mono" style="font-size:10px;font-weight:700;color:#B8BCC8;flex:none">▲ +1</span>
<span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 6px;border-radius:5px;background:#14141E;color:#B8BCC8;font-weight:800;font-size:10px;border:1px solid #1E1E2A;flex:none">C</span>
</div>
</div>
<div style="display:flex;align-items:center;gap:9px;padding:9px 16px;border-top:1px solid #14141E;background:#08080D">
<span class="mono" style="font-size:8px;letter-spacing:.22em;color:#00D4A0;font-weight:800;flex:none">WIRE</span>
<span class="mono" style="font-size:9.5px;color:#B8BCC8;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1">Last reprice 11:42 AM — Nabers rec yds.</span>
</div>
</div>
</x-import>
</div>
<!-- reveal mobile -->
<div>
<div class="mono" style="font-size:10px;letter-spacing:.2em;color:#707080;margin-bottom:12px;text-align:center">S1-D · SEASON READ REVEAL · 390</div>
<x-import component-from-global-scope="IOSDevice" from="./ios-frame.jsx" dark="{{ true }}" hint-size="390px,844px">
<div style="display:flex;flex-direction:column;height:100%;background:#06060B;font-family:'Inter',system-ui,sans-serif;color:#F0F0F0;padding-top:54px">
<div style="display:flex;align-items:center;justify-content:space-between;padding:10px 16px;border-bottom:1px solid #14141E">
<span class="mono" style="font-size:11px;color:#707080">◂ BOARD</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.16em;color:#707080">SEASON READ</span>
<span class="mono" style="font-size:11px;color:#707080"></span>
</div>
<div style="flex:1;overflow:hidden;padding:18px 16px 0">
<div style="display:flex;align-items:center;gap:11px">
<span style="position:relative;width:44px;height:44px;border-radius:11px;background:linear-gradient(135deg,#0B2265,#5a7fd4);display:flex;align-items:center;justify-content:center;font-weight:800;font-size:14px;color:#cdd9ff;flex:none">MN<image-slot id="hs-nabers-m" shape="rounded" radius="11" style="position:absolute;inset:0" placeholder="Nabers headshot"></image-slot></span>
<div style="flex:1;min-width:0">
<div style="font-weight:700;font-size:16px">Malik Nabers</div>
<div class="mono" style="font-size:9px;color:#707080;letter-spacing:.08em;margin-top:2px">WR · NYG · <span style="color:#FF7A4A">BURNER</span></div>
</div>
<span class="mono" style="font-size:34px;font-weight:800;color:#00D4A0;text-shadow:0 0 22px rgba(0,212,160,.5);flex:none">A+</span>
</div>
<div style="font-size:14.5px;font-weight:700;margin-top:14px">Over 1,120.5 Receiving Yards <span style="color:#707080;font-weight:500">· 2026 season</span></div>
<!-- triplet band -->
<div style="display:grid;grid-template-columns:repeat(3,1fr);gap:1px;margin-top:12px;background:#1E1E2A;border:1px solid #1E1E2A;border-radius:10px;overflow:hidden">
<div style="background:#0E0E14;padding:10px 12px"><div class="mono" style="font-size:8px;letter-spacing:.16em;color:#707080;margin-bottom:4px">OPEN · MAR 4</div><div class="mono" style="font-size:15px;font-weight:700;color:#707080">1,050.5</div></div>
<div style="background:#0E0E14;padding:10px 12px"><div class="mono" style="font-size:8px;letter-spacing:.16em;color:#707080;margin-bottom:4px">NOW</div><div class="mono" style="font-size:15px;font-weight:800;color:#F0F0F0">1,120.5</div></div>
<div style="background:#0E0E14;padding:10px 12px"><div class="mono" style="font-size:8px;letter-spacing:.16em;color:#00D4A0;margin-bottom:4px">VYNDR</div><div class="mono" style="font-size:15px;font-weight:800;color:#00D4A0">1,188</div></div>
</div>
<!-- movement with news markers -->
<div style="margin-top:14px">
<svg width="100%" height="46" viewBox="0 0 358 46" preserveAspectRatio="none" style="display:block">
<polyline points="0,36 40,36 40,30 110,30 110,24 200,24 200,16 290,16 290,8 358,8" fill="none" stroke="#00D4A0" stroke-width="1.6" stroke-linejoin="round" opacity=".9"></polyline>
<polyline points="0,36 40,36 40,30 110,30 110,24 200,24 200,16 290,16 290,8 358,8 358,46 0,46" fill="rgba(0,212,160,.06)" stroke="none"></polyline>
<circle cx="40" cy="30" r="2.4" fill="#B8BCC8"></circle>
<circle cx="200" cy="16" r="2.4" fill="#B8BCC8"></circle>
<circle cx="290" cy="8" r="2.8" fill="#00D4A0" style="animation:vy-livedot 1.5s ease-in-out infinite"></circle>
</svg>
<div style="display:flex;align-items:center;justify-content:space-between;margin-top:4px">
<span class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#4a4a58">MAR 12 · FA</span>
<span class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#4a4a58">APR 25 · DRAFT</span>
<span class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#00D4A0">JUL 16 · CLEARED</span>
</div>
</div>
<!-- season context -->
<div style="margin-top:14px;background:#0A0A10;border:1px solid #14141E;border-radius:10px;padding:12px 14px">
<div style="display:flex;align-items:center;gap:8px;margin-bottom:7px">
<span style="width:5px;height:5px;border-radius:50%;background:#00D4A0"></span>
<span class="mono" style="font-size:8.5px;letter-spacing:.22em;color:#B8BCC8">VYNDR INTELLIGENCE</span>
</div>
<p style="font-size:11.5px;line-height:1.5;color:#B8BCC8">Pre-injury target share <span class="mono" style="color:#F0F0F0">28%</span> with no added competition. Market repriced the clearance <span class="mono" style="color:#00D4A0">+70</span> but still carries the injury discount — on current form the number is <span class="mono" style="color:#00D4A0">68</span> yards light.</p>
</div>
<!-- what would change this read -->
<div style="margin-top:12px">
<div class="mono" style="font-size:8.5px;letter-spacing:.2em;color:#707080;margin-bottom:7px">WHAT WOULD CHANGE THIS READ</div>
<div style="display:flex;align-items:center;gap:8px;padding:7px 0;border-top:1px solid #101018"><span class="mono" style="font-size:10px;color:#FF4757;flex:none"></span><span style="font-size:11px;color:#B8BCC8;flex:1">Soft-tissue setback in camp</span><span class="mono" style="font-size:8px;color:#4a4a58;letter-spacing:.1em">RE-GRADE ≤ B</span></div>
<div style="display:flex;align-items:center;gap:8px;padding:7px 0;border-top:1px solid #101018"><span class="mono" style="font-size:10px;color:#FF4757;flex:none"></span><span style="font-size:11px;color:#B8BCC8;flex:1">QB room change before Week 1</span><span class="mono" style="font-size:8px;color:#4a4a58;letter-spacing:.1em">RE-RUN MODEL</span></div>
<div style="display:flex;align-items:center;gap:8px;padding:7px 0;border-top:1px solid #101018"><span class="mono" style="font-size:10px;color:#B8BCC8;flex:none"></span><span style="font-size:11px;color:#B8BCC8;flex:1">Slot usage in preseason installs</span><span class="mono" style="font-size:8px;color:#4a4a58;letter-spacing:.1em">MODEL ↑</span></div>
</div>
</div>
<div style="padding:10px 16px 14px">
<div class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#4a4a58;text-align:center;margin-bottom:8px">MODEL RUN JUL 17 · 6:00 AM ET · SETTLES WK 18 · CLV TRACKED FROM LOCK</div>
<div style="display:flex;gap:10px">
<span class="mono" style="flex:1;display:inline-flex;align-items:center;justify-content:center;height:44px;border-radius:11px;background:#00D4A0;color:#06060B;font-weight:800;font-size:11px;letter-spacing:.14em">LOCK THE READ</span>
<span class="mono" style="display:inline-flex;align-items:center;justify-content:center;height:44px;width:88px;border-radius:11px;background:#14141E;border:1px solid #2A2A38;color:#B8BCC8;font-weight:700;font-size:10px;letter-spacing:.12em">SHOP ▸</span>
</div>
</div>
</div>
</x-import>
</div>
</div>
</div>
<!-- ============ ACT 03 · SEASON READ REVEAL · DESKTOP ============ -->
<div data-reveal style="margin-top:110px;display:flex;align-items:baseline;gap:18px;border-bottom:1px solid #14141E;padding-bottom:18px">
<span class="mono" style="font-size:64px;font-weight:800;line-height:.8;color:#14141E;-webkit-text-stroke:1px #2A2A38">03</span>
<div><div class="mono" style="font-size:13px;letter-spacing:.3em;color:#F0F0F0;font-weight:800">SEASON READ REVEAL</div><div style="font-size:12.5px;color:#707080;margin-top:5px">The grade reveal at a season horizon — no "tonight." News-driven factors, movement annotated with events, and what would change the read.</div></div>
</div>
<div id="reveal" data-screen-label="Season reveal · desktop" style="margin-top:44px">
<div style="display:grid;grid-template-columns:1.4fr 1fr;gap:16px;align-items:start">
<div style="position:relative;border-radius:16px;background:linear-gradient(180deg,#12121C,#0C0C12);border:1px solid #2A2A38;overflow:hidden;box-shadow:0 30px 60px -34px rgba(0,0,0,.9),inset 0 1px 0 rgba(255,255,255,.05);animation:vy-arrive .7s cubic-bezier(.2,.8,.2,1) both">
<div style="position:relative;padding:22px 24px">
<div style="display:flex;align-items:center;justify-content:space-between;margin-bottom:20px">
<span class="mono" style="font-size:10.5px;letter-spacing:.28em;color:#707080">SEASON READ · 2026 · NFL</span>
<span class="mono" style="font-size:10px;letter-spacing:.18em;color:#00D4A0">CLV TRACKED FROM LOCK</span>
</div>
<div style="display:flex;align-items:center;gap:20px">
<div style="display:flex;flex-direction:column;align-items:center;gap:6px">
<div class="mono" style="font-size:82px;line-height:.82;font-weight:800;color:#00D4A0;text-shadow:0 0 34px rgba(0,212,160,.5)">A+</div>
<div class="mono" style="font-size:9.5px;letter-spacing:.24em;color:#707080">TIER</div>
</div>
<div style="flex:1;min-width:0">
<div style="display:flex;align-items:center;gap:10px;margin-bottom:8px">
<span style="position:relative;width:38px;height:38px;border-radius:9px;background:linear-gradient(135deg,#0B2265,#5a7fd4);display:flex;align-items:center;justify-content:center;font-weight:800;font-size:13px;color:#cdd9ff">MN<image-slot id="hs-nabers-d" shape="rounded" radius="9" style="position:absolute;inset:0" placeholder="Nabers headshot"></image-slot></span>
<div>
<div style="font-weight:700;font-size:16px;line-height:1.1">Malik Nabers</div>
<div class="mono" style="font-size:10.5px;color:#707080;letter-spacing:.08em">WR · NYG · <span style="color:#FF7A4A">BURNER</span></div>
</div>
</div>
<div style="font-size:15px;font-weight:600;line-height:1.35">Over 1,120.5 Receiving Yards <span style="color:#707080;font-weight:500">· season</span></div>
<div style="margin-top:12px">
<svg width="100%" height="40" viewBox="0 0 420 40" preserveAspectRatio="none" style="display:block">
<polyline points="0,32 48,32 48,27 130,27 130,21 236,21 236,14 340,14 340,7 420,7" fill="none" stroke="#00D4A0" stroke-width="1.6" stroke-linejoin="round" opacity=".9"></polyline>
<polyline points="0,32 48,32 48,27 130,27 130,21 236,21 236,14 340,14 340,7 420,7 420,40 0,40" fill="rgba(0,212,160,.07)" stroke="none"></polyline>
<circle cx="48" cy="27" r="2.4" fill="#B8BCC8"></circle>
<circle cx="236" cy="14" r="2.4" fill="#B8BCC8"></circle>
<circle cx="340" cy="7" r="2.8" fill="#00D4A0" style="animation:vy-livedot 1.5s ease-in-out infinite"></circle>
</svg>
<div style="display:flex;align-items:center;justify-content:space-between;margin-top:4px">
<span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#4a4a58">OPEN 1,050.5 · MAR 4</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#4a4a58">MAR 12 FA · APR 25 DRAFT</span>
<span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#00D4A0">JUL 16 CLEARED · NOW 1,120.5</span>
</div>
</div>
</div>
</div>
<div style="display:grid;grid-template-columns:repeat(3,1fr);gap:1px;margin-top:20px;background:#1E1E2A;border:1px solid #1E1E2A;border-radius:10px;overflow:hidden">
<div style="background:#0E0E14;padding:12px 14px"><div class="mono" style="font-size:9.5px;letter-spacing:.18em;color:#707080;margin-bottom:6px">OPEN</div><div class="mono" style="font-size:20px;font-weight:700;color:#707080">1,050.5</div></div>
<div style="background:#0E0E14;padding:12px 14px"><div class="mono" style="font-size:9.5px;letter-spacing:.18em;color:#707080;margin-bottom:6px">NOW</div><div class="mono" data-react data-fmt="line" style="font-size:20px;font-weight:800;color:#F0F0F0;padding:1px 4px;border-radius:5px">1,120.5</div></div>
<div style="background:#0E0E14;padding:12px 14px"><div class="mono" style="font-size:9.5px;letter-spacing:.18em;color:#00D4A0;margin-bottom:6px">VYNDR</div><div class="mono" style="font-size:20px;font-weight:800;color:#00D4A0">1,188</div></div>
</div>
<div style="margin-top:16px;background:#0A0A10;border:1px solid #14141E;border-radius:10px;padding:14px 16px">
<div style="display:flex;align-items:center;gap:8px;margin-bottom:8px">
<span style="width:5px;height:5px;border-radius:50%;background:#00D4A0"></span>
<span class="mono" style="font-size:9.5px;letter-spacing:.24em;color:#B8BCC8">VYNDR INTELLIGENCE</span>
</div>
<p style="font-size:12.5px;line-height:1.5;color:#B8BCC8">On current form: pre-injury target share <span class="mono" style="color:#F0F0F0">28%</span>, no added target competition, QB continuity intact. The market repriced the clearance <span class="mono" style="color:#00D4A0">+70</span> yards but still carries last season's injury discount. Our season sim puts the number <span class="mono" style="color:#00D4A0">68</span> yards light — the news, not the noise, moved this line.</p>
</div>
<div class="mono" style="margin-top:14px;font-size:8.5px;letter-spacing:.12em;color:#4a4a58">MODEL RUN JUL 17 · 6:00 AM ET · SETTLES WK 18 · JAN 3 2027 · NO "TONIGHT" ON A SEASON READ</div>
</div>
</div>
<!-- factors + change-the-read -->
<div style="display:flex;flex-direction:column;gap:16px">
<div style="border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;overflow:hidden">
<div style="padding:14px 16px 12px"><span class="mono" style="font-size:10.5px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">NEWS-DRIVEN FACTORS</span></div>
<div style="height:1px;background:#14141E"></div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;gap:10px"><span class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58;width:44px;flex:none">JUL 16</span><span style="font-size:11.5px;color:#B8BCC8;flex:1">Cleared for camp — no restrictions</span><span class="mono" style="font-size:9.5px;color:#00D4A0;font-weight:700">MODEL ▲</span></div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;gap:10px;opacity:.86"><span class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58;width:44px;flex:none">APR 25</span><span style="font-size:11.5px;color:#B8BCC8;flex:1">No WR drafted in first four rounds</span><span class="mono" style="font-size:9.5px;color:#00D4A0;font-weight:700">MODEL ▲</span></div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;gap:10px;opacity:.64"><span class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58;width:44px;flex:none">MAR 12</span><span style="font-size:11.5px;color:#B8BCC8;flex:1">O-line additions in free agency</span><span class="mono" style="font-size:9.5px;color:#B8BCC8;font-weight:700">NEUTRAL</span></div>
</div>
<div style="border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;overflow:hidden">
<div style="padding:14px 16px 12px"><span class="mono" style="font-size:10.5px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">WHAT WOULD CHANGE THIS READ</span></div>
<div style="height:1px;background:#14141E"></div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;gap:10px"><span class="mono" style="font-size:11px;color:#FF4757;flex:none"></span><span style="font-size:11.5px;color:#B8BCC8;flex:1">Soft-tissue setback in camp</span><span class="mono" style="font-size:8px;color:#4a4a58;letter-spacing:.1em">RE-GRADE ≤ B</span></div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;gap:10px"><span class="mono" style="font-size:11px;color:#FF4757;flex:none"></span><span style="font-size:11.5px;color:#B8BCC8;flex:1">QB room change before Week 1</span><span class="mono" style="font-size:8px;color:#4a4a58;letter-spacing:.1em">RE-RUN MODEL</span></div>
<div style="padding:11px 16px;border-top:1px solid #101018;display:flex;align-items:center;gap:10px"><span class="mono" style="font-size:11px;color:#B8BCC8;flex:none"></span><span style="font-size:11.5px;color:#B8BCC8;flex:1">Heavy slot usage in preseason installs</span><span class="mono" style="font-size:8px;color:#4a4a58;letter-spacing:.1em">MODEL ↑</span></div>
<div style="padding:11px 16px;border-top:1px solid #101018"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#4a4a58">EVERY SEASON READ NAMES ITS KILL CONDITIONS — HONESTY IS THE PRODUCT</span></div>
</div>
</div>
</div>
</div>
<!-- ============ ACT 04 · NEWS / OUTLOOK FEED ============ -->
<div data-reveal style="margin-top:110px;display:flex;align-items:baseline;gap:18px;border-bottom:1px solid #14141E;padding-bottom:18px">
<span class="mono" style="font-size:64px;font-weight:800;line-height:.8;color:#14141E;-webkit-text-stroke:1px #2A2A38">04</span>
<div><div class="mono" style="font-size:13px;letter-spacing:.3em;color:#F0F0F0;font-weight:800">NEWS / OUTLOOK FEED</div><div style="font-size:12.5px;color:#707080;margin-top:5px">Event → affected entities → line impact → article. One anatomy, reused in the hub, the reveal, and /blog.</div></div>
</div>
<div data-screen-label="News feed + anatomy" style="margin-top:44px;display:grid;grid-template-columns:1.4fr 1fr;gap:16px;align-items:start">
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;overflow:hidden">
<div style="display:flex;align-items:center;justify-content:space-between;padding:15px 16px 13px">
<span class="mono" style="font-size:10.5px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">THE FEED · NFL</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">EVENTS THAT MOVED OUTLOOKS ONLY</span>
</div>
<div style="height:1px;background:#14141E"></div>
<div style="padding:9px 16px 5px"><span class="mono" style="font-size:8.5px;letter-spacing:.2em;color:#4a4a58">TODAY · JUL 17</span></div>
<div style="--o:1;opacity:1;padding:12px 16px;border-top:1px solid #101018;box-shadow:inset 2px 0 0 rgba(0,212,160,.55)">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080">11:42 AM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:8px;letter-spacing:.14em;color:#B8BCC8">CLEARED</span><span class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58">SOURCE · TEAM REPORT</span></div>
<div style="font-size:14px;font-weight:700;margin-top:6px;line-height:1.35">Nabers cleared for camp with no restrictions — the injury discount is now the market's mistake</div>
<div style="display:flex;align-items:center;gap:10px;margin-top:8px;flex-wrap:wrap">
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:20px;padding:0 8px;border-radius:6px;background:#14141E;border:1px solid #1E1E2A;font-size:9px;letter-spacing:.08em;color:#B8BCC8"><span style="width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#0B2265,#5a7fd4)"></span>NYG</span>
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:20px;padding:0 8px;border-radius:6px;background:#14141E;border:1px solid #1E1E2A;font-size:9px;letter-spacing:.08em;color:#B8BCC8">M. NABERS</span>
<span class="mono" style="font-size:9.5px;color:#707080">REC YDS <span style="color:#4a4a58">1,050.5</span><span style="color:#F0F0F0;font-weight:700">1,120.5</span> · <span style="color:#00D4A0">V 1,188</span></span>
<a href="Vyndr Intelligence.dc.html#articles" class="mono" style="margin-left:auto;font-size:9px;letter-spacing:.14em">FULL ARTICLE ▸</a>
</div>
</div>
<div style="--o:.86;opacity:.86;padding:12px 16px;border-top:1px solid #101018">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080">10:15 AM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:#14141E;border:1px solid #1E1E2A;font-size:8px;letter-spacing:.14em;color:#B8BCC8">CAMP REPORT</span><span class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58">SOURCE · BEAT POOL</span></div>
<div style="font-size:14px;font-weight:600;margin-top:6px;line-height:1.35">Worthy running clear WR1 reps in KC installs</div>
<div style="display:flex;align-items:center;gap:10px;margin-top:8px;flex-wrap:wrap">
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:20px;padding:0 8px;border-radius:6px;background:#14141E;border:1px solid #1E1E2A;font-size:9px;letter-spacing:.08em;color:#B8BCC8"><span style="width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#E31837,#FFB81C)"></span>KC</span>
<span class="mono" style="font-size:9.5px;color:#707080">REC YDS <span style="color:#4a4a58">990.5</span><span style="color:#F0F0F0;font-weight:700">1,024.5</span> · <span style="color:#00D4A0">V 1,061</span></span>
<span class="mono" style="margin-left:auto;font-size:9px;letter-spacing:.14em;color:#4a4a58">FULL ARTICLE ▸</span>
</div>
</div>
<div style="--o:.7;opacity:.7;padding:12px 16px;border-top:1px solid #101018">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080">9:30 AM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:rgba(255,179,71,.08);border:1px solid rgba(255,179,71,.25);font-size:8px;letter-spacing:.14em;color:#FFB347">CONTRACT</span><span class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58">SOURCE · LEAGUE FILING</span></div>
<div style="font-size:14px;font-weight:600;margin-top:6px;line-height:1.35">Higgins restructure keeps the WR2 in stripes — Burrow's ceiling holds</div>
<div style="display:flex;align-items:center;gap:10px;margin-top:8px;flex-wrap:wrap">
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:20px;padding:0 8px;border-radius:6px;background:#14141E;border:1px solid #1E1E2A;font-size:9px;letter-spacing:.08em;color:#B8BCC8"><span style="width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#FB4F14,#1a1a1a)"></span>CIN</span>
<span class="mono" style="font-size:9.5px;color:#707080">PASS TD <span style="color:#4a4a58">27.5</span><span style="color:#F0F0F0;font-weight:700">28.5</span> · <span style="color:#00D4A0">V 29.8</span></span>
<span class="mono" style="margin-left:auto;font-size:9px;letter-spacing:.14em;color:#4a4a58">FULL ARTICLE ▸</span>
</div>
</div>
<div style="padding:9px 16px 5px;border-top:1px solid #101018"><span class="mono" style="font-size:8.5px;letter-spacing:.2em;color:#4a4a58">WED · JUL 16</span></div>
<div style="--o:.55;opacity:.55;padding:12px 16px;border-top:1px solid #101018">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080">4:40 PM</span><span class="mono" style="display:inline-flex;align-items:center;height:15px;padding:0 6px;border-radius:4px;background:rgba(255,71,87,.08);border:1px solid rgba(255,71,87,.25);font-size:8px;letter-spacing:.14em;color:#FF4757">INJURY</span></div>
<div style="font-size:14px;font-weight:600;margin-top:6px;line-height:1.35">Kincaid tweaks hamstring in conditioning — monitoring, not moving</div>
<div style="display:flex;align-items:center;gap:10px;margin-top:8px;flex-wrap:wrap">
<span class="mono" style="display:inline-flex;align-items:center;gap:5px;height:20px;padding:0 8px;border-radius:6px;background:#14141E;border:1px solid #1E1E2A;font-size:9px;letter-spacing:.08em;color:#B8BCC8"><span style="width:9px;height:9px;border-radius:3px;background:linear-gradient(135deg,#00338D,#C60C30)"></span>BUF</span>
<span class="mono" style="font-size:9.5px;color:#707080">REC YDS <span style="color:#F0F0F0;font-weight:700">625.5</span> · <span style="color:#707080">HELD · WATCHING</span></span>
</div>
</div>
<!-- honest empty -->
<div style="margin:14px 16px 16px;border:1px dashed #1E1E2A;border-radius:10px;padding:16px;text-align:center">
<div class="mono" style="font-size:9px;letter-spacing:.22em;color:#707080">QUIET WIRE</div>
<div style="font-size:12px;color:#707080;margin-top:6px">No outlook-moving news since 11:42 AM. We don't manufacture movement.</div>
<div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:6px">EMPTY STATE · RENDERS WHEN A DAY HAS NO EVENTS</div>
</div>
</div>
<!-- anatomy spec -->
<div style="border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;padding:16px">
<div class="mono" style="font-size:10.5px;letter-spacing:.24em;color:#F0F0F0;font-weight:700;margin-bottom:12px">ROW ANATOMY</div>
<div style="display:flex;flex-direction:column;gap:10px">
<div style="display:flex;gap:10px;align-items:baseline"><span class="mono" style="font-size:9px;color:#00D4A0;width:18px;flex:none">1</span><div><div class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#B8BCC8">TIME + EVENT TAG + SOURCE</div><div style="font-size:11px;color:#707080;margin-top:2px">Tag color = meaning: neutral grey, CONTRACT amber, INJURY red. Never green — green is edge only.</div></div></div>
<div style="display:flex;gap:10px;align-items:baseline"><span class="mono" style="font-size:9px;color:#00D4A0;width:18px;flex:none">2</span><div><div class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#B8BCC8">HEADLINE</div><div style="font-size:11px;color:#707080;margin-top:2px">Editorial voice — states what happened and why the desk cares. Weight 700 only when the event repriced a line.</div></div></div>
<div style="display:flex;gap:10px;align-items:baseline"><span class="mono" style="font-size:9px;color:#00D4A0;width:18px;flex:none">3</span><div><div class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#B8BCC8">AFFECTED CHIPS</div><div style="font-size:11px;color:#707080;margin-top:2px">Existing TeamChip + player chips. Tap target ≥ 44px on mobile.</div></div></div>
<div style="display:flex;gap:10px;align-items:baseline"><span class="mono" style="font-size:9px;color:#00D4A0;width:18px;flex:none">4</span><div><div class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#B8BCC8">LINE IMPACT</div><div style="font-size:11px;color:#707080;margin-top:2px">old → new · VYNDR model. FLAT + HELD are honest values — a non-move is information.</div></div></div>
<div style="display:flex;gap:10px;align-items:baseline"><span class="mono" style="font-size:9px;color:#00D4A0;width:18px;flex:none">5</span><div><div class="mono" style="font-size:9.5px;letter-spacing:.14em;color:#B8BCC8">FULL ARTICLE ▸</div><div style="font-size:11px;color:#707080;margin-top:2px">Present only when an article exists. No dead links, no "coming soon."</div></div></div>
</div>
<div style="margin-top:14px;padding-top:12px;border-top:1px solid #14141E">
<div class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#4a4a58;line-height:1.7">REUSED IN · HUB "WHAT CHANGED TODAY" · REVEAL FACTORS · /BLOG FEED · THE REPORT EMAIL</div>
</div>
</div>
</div>
<div data-screen-label="footer" style="margin-top:64px">
<div style="padding-top:22px;border-top:1px solid #14141E;display:flex;flex-wrap:wrap;align-items:center;justify-content:space-between;gap:10px 24px">
<span class="mono" style="font-size:10px;letter-spacing:.16em;color:#707080">HUB HOME · SEASON BOARD · SEASON REVEAL · THE FEED · NFL + NBA · DESKTOP + 390</span>
<span class="mono" style="font-size:10px;letter-spacing:.16em;color:#00D4A0">VYNDR · SESSION 2 · SURFACE 1</span>
</div>
</div>
</div>
</div>
</x-dc>
<script type="text/x-dc" data-dc-script data-props="{&quot;motion&quot;:{&quot;editor&quot;:&quot;boolean&quot;,&quot;default&quot;:true,&quot;tsType&quot;:&quot;boolean&quot;,&quot;section&quot;:&quot;Behavior&quot;},&quot;reactions&quot;:{&quot;editor&quot;:&quot;boolean&quot;,&quot;default&quot;:true,&quot;tsType&quot;:&quot;boolean&quot;,&quot;section&quot;:&quot;Behavior&quot;}}">
class Component extends DCLogic {
renderVals() {
const kick = Math.max(0, Math.ceil((new Date('2026-09-10T13:00:00-04:00') - Date.now()) / 864e5));
return {
rootRef: (n) => { this.rootNode = n; },
kickoffDays: String(kick),
stillClass: (this.props.motion ?? true) ? '' : 'vy-still'
};
}
componentDidMount() {
this._t = [];
const tick = () => {
const now = new Date();
const s = now.toLocaleTimeString('en-US', { hour12: false });
document.querySelectorAll('[data-sync-clock]').forEach(el => el.textContent = s);
const wt = document.querySelector('[data-wire-time]');
if (wt) wt.textContent = s;
};
tick();
this._t.push(setInterval(tick, 1000));
const entries = [
'Nabers 1,050.5 → 1,120.5 on camp clearance. Grade holds A+.',
'Worthy rec yds repriced +34.0 on WR1 camp reps.',
'KC win total: market 11.5, model 10.8 — under lean logged.',
'NBA · Wemby MVP steams +420 → +330 since June.',
'Summer League stays ungraded — low-signal market refused.',
'214 NFL season lines live · last reprice 11:42 AM ET.'
];
let wi = 0;
this._t.push(setInterval(() => {
const el = document.querySelector('[data-wire]');
if (!el) return;
wi = (wi + 1) % entries.length;
el.textContent = entries[wi];
el.style.animation = 'none';
void el.offsetWidth;
el.style.animation = 'vy-wirein 6s ease both';
}, 6000));
// reactions: gated post-boot, fire on "data events" only
const fmt = {
line: v => v.toLocaleString('en-US', { minimumFractionDigits: 1, maximumFractionDigits: 1 }),
half: v => v.toFixed(1),
edge: v => (v >= 0 ? '+' : '') + v.toFixed(1) + '%',
odds: v => (v >= 0 ? '+' : '') + Math.round(v),
int: v => String(Math.round(v))
};
const nudge = () => {
if (!(this.props.reactions ?? true)) return;
const cells = [...document.querySelectorAll('[data-react]')];
if (!cells.length) return;
const el = cells[Math.floor(Math.random() * cells.length)];
const f = el.getAttribute('data-fmt') || 'half';
let v = parseFloat(el.textContent.replace(/[,%+]/g, ''));
if (isNaN(v)) return;
let up;
if (f === 'edge') { const d = (Math.random() - .45) * .3; up = d >= 0; v += d; }
else if (f === 'odds') { const d = Math.random() < .5 ? -25 : 25; up = d < 0; v += d; }
else { const step = f === 'line' && v > 500 ? (Math.random() < .5 ? -1 : 1) * (Math.random() < .7 ? 1 : 2) : (Math.random() < .5 ? -.5 : .5); up = step > 0; v += step; }
el.textContent = fmt[f] ? fmt[f](v) : v;
el.classList.remove('rx-up', 'rx-down');
void el.offsetWidth;
el.classList.add(up ? 'rx-up' : 'rx-down');
};
this._t.push(setTimeout(() => { this._t.push(setInterval(nudge, 5200)); }, 1800));
// scroll reveal
const secs = [...(this.rootNode || document).querySelectorAll('[data-reveal],[data-screen-label]')];
secs.forEach(s => { if (!s.hasAttribute('data-reveal')) s.setAttribute('data-reveal', ''); });
this._io = new IntersectionObserver(ents => {
ents.forEach(e => { if (e.isIntersecting) { e.target.classList.add('vy-in'); this._io.unobserve(e.target); } });
}, { rootMargin: '0px 0px -8% 0px' });
secs.forEach(s => this._io.observe(s));
}
componentWillUnmount() {
(this._t || []).forEach(t => { clearInterval(t); clearTimeout(t); });
if (this._io) this._io.disconnect();
}
}
</script>
</body>
</html>
@@ -0,0 +1,426 @@
<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<script src="./support.js"></script>
</head>
<body>
<x-dc>
<helmet>
<meta name="viewport" content="width=device-width, initial-scale=1" />
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700;800;900&family=JetBrains+Mono:wght@400;500;600;700;800&display=swap" rel="stylesheet" />
<style>
*{box-sizing:border-box;margin:0;padding:0}
html,body{background:#06060B;color:#F0F0F0;font-family:'Inter',system-ui,sans-serif;-webkit-font-smoothing:antialiased}
::selection{background:rgba(0,212,160,.28);color:#F0F0F0}
a{color:#00D4A0;text-decoration:none}
a:hover{color:#5cf0cf}
.mono{font-family:'JetBrains Mono',monospace;font-variant-numeric:tabular-nums}
@keyframes vy-livedot{0%,100%{opacity:1;transform:scale(1)}50%{opacity:.35;transform:scale(.82)}}
@keyframes vy-syncdot{0%,100%{opacity:1}50%{opacity:.2}}
@keyframes vy-rxup{0%{background:rgba(0,212,160,.28)}100%{background:rgba(0,212,160,0)}}
@keyframes vy-rxdown{0%{background:rgba(255,71,87,.26)}100%{background:rgba(255,71,87,0)}}
@keyframes vy-rowin{0%{opacity:0;transform:translateY(9px)}100%{opacity:var(--o,1);transform:translateY(0)}}
@keyframes vy-bootline{0%{transform:scaleX(0)}100%{transform:scaleX(1)}}
@keyframes vy-arrive{0%{opacity:0;transform:scale(.88) translateY(10px);filter:blur(6px)}60%{opacity:1;transform:scale(1.015) translateY(0);filter:blur(0)}100%{opacity:1;transform:scale(1) translateY(0)}}
@keyframes vy-glitch{0%,90%,100%{transform:translate(0,0);clip-path:inset(0 0 0 0);text-shadow:none}91%{transform:translate(-1.5px,0);clip-path:inset(12% 0 42% 0);text-shadow:1.6px 0 #FF4757,-1.6px 0 #00D4A0}93%{transform:translate(1.5px,0);clip-path:inset(58% 0 8% 0);text-shadow:-1.6px 0 #FF4757,1.6px 0 #00D4A0}95%{transform:translate(-1px,0);clip-path:inset(32% 0 30% 0);text-shadow:1px 0 #00D4A0}96%{transform:translate(1px,0);clip-path:inset(70% 0 4% 0);text-shadow:none}}
@keyframes vy-cursor{0%,49%{opacity:1}50%,100%{opacity:0}}
@keyframes vy-wirein{0%{opacity:0;transform:translateY(4px)}8%,86%{opacity:1;transform:none}100%{opacity:0;transform:translateY(-3px)}}
.rx-up{animation:vy-rxup .75s ease-out}
.rx-down{animation:vy-rxdown .75s ease-out}
.vy-grain{position:fixed;inset:0;z-index:5;pointer-events:none;mix-blend-mode:overlay;opacity:.05;background-image:url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='140' height='140'%3E%3Cfilter id='n'%3E%3CfeTurbulence type='fractalNoise' baseFrequency='0.82' numOctaves='2' stitchTiles='stitch'/%3E%3C/filter%3E%3Crect width='100%25' height='100%25' filter='url(%23n)'/%3E%3C/svg%3E");background-size:140px}
.vy-vignette{position:fixed;inset:0;z-index:4;pointer-events:none;box-shadow:inset 0 0 200px 40px rgba(0,0,0,.55),inset 0 90px 120px -60px rgba(0,212,160,.05)}
[data-reveal]{opacity:0;transform:translateY(16px);transition:opacity .6s cubic-bezier(.2,.8,.2,1),transform .6s cubic-bezier(.2,.8,.2,1)}
[data-reveal].vy-in{opacity:1;transform:none}
.vy-still *{animation-play-state:paused !important}
html{scroll-behavior:smooth}
</style>
<script src="image-slot.js"></script>
<style>image-slot::part(empty){opacity:0}</style>
</helmet>
<div ref="{{ rootRef }}" class="{{ stillClass }}" style="min-height:100vh;background:radial-gradient(1200px 700px at 78% -8%,rgba(0,212,160,.06),transparent 60%),#06060B;padding:0 0 120px">
<div class="vy-vignette" aria-hidden="true"></div>
<div class="vy-grain" aria-hidden="true"></div>
<div style="position:fixed;left:0;right:0;bottom:0;z-index:45;display:flex;align-items:center;gap:14px;height:36px;padding:0 40px;background:rgba(6,6,11,.88);backdrop-filter:blur(14px);border-top:1px solid #14141E">
<span class="mono" style="font-size:9px;letter-spacing:.26em;color:#00D4A0;font-weight:800;flex:none">THE WIRE</span>
<span style="width:1px;height:14px;background:#1E1E2A;flex:none"></span>
<span class="mono" data-wire-time style="font-size:10px;color:#707080;letter-spacing:.06em;flex:none">--:--:--</span>
<span class="mono" data-wire style="font-size:11px;color:#B8BCC8;letter-spacing:.03em;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1">Intelligence desk open — 7 books held per prop, CLV tracked from lock.</span>
<span class="mono" style="font-size:9px;letter-spacing:.16em;color:#4a4a58;flex:none">LOG ▸</span>
</div>
<div style="max-width:1240px;margin:0 auto;padding:0 40px">
<div style="padding:52px 0 30px">
<div class="mono" style="font-size:11px;letter-spacing:.34em;color:#00D4A0;margin-bottom:16px">VYNDR · ACT 01 · THE PRICE TRIPLET</div>
<h1 style="font-size:38px;font-weight:800;letter-spacing:-.02em;line-height:1.05">The honest layer, rendered.<span class="mono" style="color:#00D4A0;animation:vy-cursor 1.1s steps(1) infinite"></span></h1>
<p style="margin-top:12px;max-width:660px;color:#B8BCC8;font-size:14px;line-height:1.55">The price triplet leads — <span class="mono" style="color:#F0F0F0">book · fair · model</span>, the de-vigged number nobody else shows consumers. Then line shopping, article media, CLV &amp; calibration, and The Report — extensions sharing primitives. The movement strip is defined once here and reused everywhere a line has a past.</p>
<div class="mono" style="margin-top:14px;font-size:10px;letter-spacing:.16em;color:#707080">BOOK · FAIR · MODEL · FIVE HONESTY STATES · STAT PROJECTION + PRICE LAYER</div>
</div>
<!-- ============ ACT 01 · THE PRICE TRIPLET (HERO FEATURE) ============ -->
<div data-reveal style="margin-top:14px;display:flex;align-items:baseline;gap:18px;border-bottom:1px solid #14141E;padding-bottom:18px">
<span class="mono" style="font-size:64px;font-weight:800;line-height:.8;color:#14141E;-webkit-text-stroke:1px #2A2A38">01</span>
<div><div class="mono" style="font-size:13px;letter-spacing:.3em;color:#F0F0F0;font-weight:800">THE PRICE TRIPLET <span style="color:#FFB347">· HERO</span></div><div style="font-size:12.5px;color:#707080;margin-top:5px">Book · fair · model, per prop — the price-honesty layer beside the stat projection. Fair is the amber hero, honest on every row: value, edge-but-priced-out, or no edge. Renders on the read card, board and hero; the close lives in the ledger.</div></div>
</div>
<div id="triplet" data-screen-label="Price triplet · anatomy + four honesty states" style="margin-top:44px">
<!-- anatomy · inherits the parlay TRUE FAIR VALUE language -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:22px 24px">
<div style="display:flex;align-items:center;gap:12px;flex-wrap:wrap;margin-bottom:16px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">ANATOMY</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">EXTENDS THE PARLAY "TRUE FAIR VALUE" — 2-UP GRID → 3-LEG PER-PROP</span>
</div>
<div style="max-width:560px">
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:12px;overflow:hidden;background:#0E0E14">
<div style="flex:1;min-width:0;text-align:center;padding:16px 10px">
<div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">BOOK</div>
<div class="mono" style="font-size:22px;font-weight:800;color:#F0F0F0;margin-top:8px;letter-spacing:-.01em">+125</div>
<div class="mono" style="font-size:7px;letter-spacing:.12em;color:#4a4a58;margin-top:6px">OFFERED</div>
</div>
<div style="flex:1.16;min-width:0;text-align:center;padding:16px 10px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A">
<div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div>
<div class="mono" style="font-size:31px;font-weight:800;color:#FFB347;margin-top:5px;letter-spacing:-.02em;text-shadow:0 0 22px rgba(255,179,71,.24)">+110</div>
<div class="mono" style="font-size:7px;letter-spacing:.12em;color:#8a7a55;margin-top:5px">DE-VIGGED · HERO</div>
</div>
<div style="flex:1;min-width:0;text-align:center;padding:16px 10px">
<div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">MODEL</div>
<div class="mono" style="font-size:22px;font-weight:800;color:#00D4A0;margin-top:8px;letter-spacing:-.01em">+98</div>
<div class="mono" style="font-size:7px;letter-spacing:.12em;color:#4a4a58;margin-top:6px">OUR PRICE</div>
</div>
</div>
</div>
<div style="display:grid;grid-template-columns:repeat(3,1fr);gap:14px;margin-top:16px;max-width:820px">
<div><div class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#F0F0F0;font-weight:700">BOOK</div><div style="font-size:10.5px;color:#707080;line-height:1.55;margin-top:5px">The market's offered line. Honest on every row.</div></div>
<div><div class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#FFB347;font-weight:700">FAIR · THE HERO</div><div style="font-size:10.5px;color:#B8BCC8;line-height:1.55;margin-top:5px">The de-vigged fair number. The whole product. Never hidden — not even when the model leg is poisoned.</div></div>
<div><div class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#00D4A0;font-weight:700">MODEL</div><div style="font-size:10.5px;color:#707080;line-height:1.55;margin-top:5px">VYNDR's own price. Green only for takeable edge — never a juiced one. Suppressible on quarantined rows.</div></div>
</div>
<div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:16px;line-height:1.7">SAME TOKENS AS THE PARLAY · LABEL 7.5PX / .14EM / #707080 · VALUE MONO 800 · FAIR AMBER #FFB347 VS BOOK #F0F0F0 · GREEN #00D4A0 = TAKEABLE EDGE ONLY · BLUE = EDGE PRICED OUT</div>
</div>
<!-- four honesty states · desktop -->
<div style="display:flex;align-items:center;gap:12px;margin:30px 0 14px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">FIVE HONESTY STATES · DESKTOP</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">THREE VERDICTS — VALUE · EDGE-NOT-TAKEABLE · NO EDGE — PLUS QUARANTINE &amp; REFUSAL</span>
</div>
<div style="display:grid;grid-template-columns:1fr 1fr;gap:16px">
<!-- STATE 1 · POSITIVE VALUE -->
<div style="border-radius:14px;background:#0E0E14;border:1px solid #1E1E2A;padding:18px;display:flex;flex-direction:column;gap:13px">
<div style="display:flex;align-items:flex-start;justify-content:space-between;gap:10px">
<div style="min-width:0"><div style="font-size:13.5px;font-weight:700">Nabers <span style="color:#B8BCC8;font-weight:500">o 1,115.5 Rec Yds</span></div><div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:3px">NYG · WK 1</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:rgba(0,212,160,.1);border:1px solid rgba(0,212,160,.28);font-size:8px;letter-spacing:.12em;color:#00D4A0;font-weight:700;flex:none">1 · VALUE</span>
</div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:11px;overflow:hidden;background:#0A0A10">
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">BOOK</div><div class="mono" style="font-size:20px;font-weight:800;color:#F0F0F0;margin-top:7px">+125</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">OFFERED</div></div>
<div style="flex:1.16;min-width:0;text-align:center;padding:14px 8px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:29px;font-weight:800;color:#FFB347;margin-top:4px;letter-spacing:-.02em;text-shadow:0 0 20px rgba(255,179,71,.22)">+110</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#8a7a55;margin-top:4px">DE-VIGGED</div></div>
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">MODEL</div><div class="mono" style="font-size:20px;font-weight:800;color:#00D4A0;margin-top:7px;text-shadow:0 0 16px rgba(0,212,160,.35)">+98</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">OUR PRICE</div></div>
</div>
<div style="padding:11px 13px;border-radius:10px;background:rgba(0,212,160,.08);border:1px solid rgba(0,212,160,.28)">
<div style="display:flex;align-items:center;gap:8px"><span style="width:7px;height:7px;border-radius:50%;background:#00D4A0;flex:none;box-shadow:0 0 8px rgba(0,212,160,.7)"></span><span class="mono" style="font-size:10.5px;font-weight:800;letter-spacing:.05em;color:#00D4A0">VALUE · +2.9% VS FAIR</span></div>
<div style="font-size:10.5px;color:#B8BCC8;line-height:1.5;margin-top:6px">Model prices it shorter than the honest number — the book is paying longer than the leg is worth. Real edge.</div>
</div>
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58">THE PRICE STORY · READ CARD</span><a href="#clv" class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">CLOSES TRACKED IN YOUR LEDGER ▸</a></div>
</div>
<!-- STATE 2 · EDGE BUT NOT TAKEABLE (priced out) -->
<div style="border-radius:14px;background:#0E0E14;border:1px solid #1E1E2A;padding:18px;display:flex;flex-direction:column;gap:13px">
<div style="display:flex;align-items:flex-start;justify-content:space-between;gap:10px">
<div style="min-width:0"><div style="font-size:13.5px;font-weight:700">Chase <span style="color:#B8BCC8;font-weight:500">o 6.5 Receptions</span></div><div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:3px">CIN · WK 1</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:rgba(106,147,200,.12);border:1px solid rgba(106,147,200,.34);font-size:8px;letter-spacing:.1em;color:#8fb2de;font-weight:700;flex:none">2 · PRICED OUT</span>
</div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:11px;overflow:hidden;background:#0A0A10">
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">BOOK</div><div class="mono" style="font-size:20px;font-weight:800;color:#F0F0F0;margin-top:7px">210</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">JUICED</div></div>
<div style="flex:1.16;min-width:0;text-align:center;padding:14px 8px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:29px;font-weight:800;color:#FFB347;margin-top:4px;letter-spacing:-.02em;text-shadow:0 0 20px rgba(255,179,71,.22)">175</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#8a7a55;margin-top:4px">DE-VIGGED</div></div>
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">MODEL</div><div class="mono" style="font-size:20px;font-weight:800;color:#8fb2de;margin-top:7px">240</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">OUR PRICE</div></div>
</div>
<div style="padding:11px 13px;border-radius:10px;background:rgba(106,147,200,.08);border:1px solid rgba(106,147,200,.3)">
<div style="display:flex;align-items:center;justify-content:space-between;gap:10px"><div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:12px;color:#8fb2de;flex:none;line-height:1"></span><span class="mono" style="font-size:10.5px;font-weight:800;letter-spacing:.05em;color:#8fb2de">EDGE · NOT TAKEABLE</span></div><span class="mono" style="font-size:9px;letter-spacing:.06em;color:#7a93b0">+11.7% EV AT 210</span></div>
<div style="font-size:10.5px;color:#B8BCC8;line-height:1.5;margin-top:6px">The math shows edge — but 210 is too juiced to bet. We won't call a juiced price a value, and we won't pretend the edge wasn't there.</div>
</div>
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58">TAKEABLE BAND 160 TO +200</span><a href="#clv" class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">CLOSES TRACKED IN YOUR LEDGER ▸</a></div>
</div>
<!-- STATE 3 · NO / NEGATIVE VALUE -->
<div style="border-radius:14px;background:#0E0E14;border:1px solid #1E1E2A;padding:18px;display:flex;flex-direction:column;gap:13px">
<div style="display:flex;align-items:flex-start;justify-content:space-between;gap:10px">
<div style="min-width:0"><div style="font-size:13.5px;font-weight:700">Achane <span style="color:#B8BCC8;font-weight:500">o 71.5 Rush Yds</span></div><div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:3px">MIA · WK 1</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:#14141E;border:1px solid #2A2A38;font-size:8px;letter-spacing:.12em;color:#B8BCC8;font-weight:700;flex:none">3 · NO EDGE</span>
</div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:11px;overflow:hidden;background:#0A0A10">
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">BOOK</div><div class="mono" style="font-size:20px;font-weight:800;color:#F0F0F0;margin-top:7px">+118</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">OFFERED</div></div>
<div style="flex:1.16;min-width:0;text-align:center;padding:14px 8px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:29px;font-weight:800;color:#FFB347;margin-top:4px;letter-spacing:-.02em;text-shadow:0 0 20px rgba(255,179,71,.22)">+104</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#8a7a55;margin-top:4px">DE-VIGGED</div></div>
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">MODEL</div><div class="mono" style="font-size:20px;font-weight:800;color:#B8BCC8;margin-top:7px">+112</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">OUR PRICE</div></div>
</div>
<div style="padding:11px 13px;border-radius:10px;background:#101018;border:1px solid #1E1E2A">
<div style="display:flex;align-items:center;gap:8px"><span style="width:7px;height:7px;border-radius:50%;border:1.5px solid #707080;flex:none"></span><span class="mono" style="font-size:10.5px;font-weight:800;letter-spacing:.05em;color:#B8BCC8">NO EDGE HERE · 1.8% VS FAIR</span></div>
<div style="font-size:10.5px;color:#707080;line-height:1.5;margin-top:6px">Model lands short of fair — the honest number is already priced. We're not calling this a play. That "no" is the instrument working.</div>
</div>
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58">THE PRICE STORY · READ CARD</span><a href="#clv" class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">CLOSES TRACKED IN YOUR LEDGER ▸</a></div>
</div>
<!-- STATE 4 · MODEL-LEG SUPPRESSED (quarantined) · PRIMARY -->
<div style="border-radius:14px;background:#0E0E14;border:1px solid #2A2A38;padding:18px;display:flex;flex-direction:column;gap:13px;box-shadow:inset 3px 0 0 rgba(255,179,71,.4)">
<div style="display:flex;align-items:flex-start;justify-content:space-between;gap:10px">
<div style="min-width:0"><div style="font-size:13.5px;font-weight:700">Bijan <span style="color:#B8BCC8;font-weight:500">o 84.5 Rush Yds</span></div><div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:3px">ATL · WK 1</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:rgba(255,179,71,.1);border:1px solid rgba(255,179,71,.3);font-size:8px;letter-spacing:.1em;color:#FFB347;font-weight:700;flex:none">4 · QUARANTINE · 38%</span>
</div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:11px;overflow:hidden;background:#0A0A10">
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">BOOK</div><div class="mono" style="font-size:20px;font-weight:800;color:#F0F0F0;margin-top:7px">+130</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">OFFERED</div></div>
<div style="flex:1.16;min-width:0;text-align:center;padding:14px 8px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:29px;font-weight:800;color:#FFB347;margin-top:4px;letter-spacing:-.02em;text-shadow:0 0 20px rgba(255,179,71,.22)">+116</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#8a7a55;margin-top:4px">DE-VIGGED</div></div>
<div style="flex:1;min-width:0;text-align:center;padding:14px 8px;display:flex;flex-direction:column;align-items:center;justify-content:center;background:repeating-linear-gradient(45deg,transparent,transparent 5px,rgba(255,255,255,.014) 5px,rgba(255,255,255,.014) 10px)"><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">MODEL</div><div class="mono" style="font-size:19px;font-weight:800;color:#4a4a58;margin-top:8px;letter-spacing:.1em">— · —</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#5a5a68;margin-top:6px">WITHHELD</div></div>
</div>
<div style="padding:11px 13px;border-radius:10px;background:#101018;border:1px solid #2A2A38">
<div style="display:flex;align-items:center;gap:8px"><span style="width:7px;height:7px;border-radius:2px;background:#FFB347;flex:none;opacity:.85"></span><span class="mono" style="font-size:10.5px;font-weight:800;letter-spacing:.05em;color:#B8BCC8">MODEL READ WITHHELD</span></div>
<div style="font-size:10.5px;color:#707080;line-height:1.5;margin-top:6px">Graded against an unverified matchup — so we suppressed our price. Book &amp; fair stand; we never hide the honest fair number because a different leg is poisoned.</div>
</div>
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58">PRIMARY STATE · 200 / 520 ROWS</span><a href="#clv" class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">CLOSES TRACKED IN YOUR LEDGER ▸</a></div>
</div>
<!-- STATE 5 · REFUSAL / CAN'T-GRADE · spans full width -->
<div style="grid-column:1 / -1;border-radius:14px;background:#0E0E14;border:1px solid #1E1E2A;padding:18px;display:flex;flex-direction:column;gap:13px">
<div style="display:flex;align-items:flex-start;justify-content:space-between;gap:10px">
<div style="min-width:0"><div style="font-size:13.5px;font-weight:700">Egbuka <span style="color:#B8BCC8;font-weight:500">o 44.5 Rec Yds</span></div><div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:3px">TB · NFL DEBUT · NO PRIORS</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:rgba(255,71,87,.08);border:1px solid rgba(255,71,87,.28);font-size:8px;letter-spacing:.1em;color:#FF6B78;font-weight:700;flex:none">5 · REFUSAL</span>
</div>
<div style="flex:1;display:flex;align-items:center;justify-content:center;border:1px dashed #2A2A38;border-radius:11px;padding:20px 16px;text-align:center;background:#0A0A10">
<div>
<div class="mono" style="font-size:9px;letter-spacing:.22em;color:#FF6B78">CAN'T GRADE THIS ONE</div>
<div style="font-size:11.5px;color:#B8BCC8;margin-top:9px;line-height:1.55;max-width:340px">No fair price we'd defend yet — the inputs are too thin. When they clear, the triplet returns. We'd rather show nothing than a number we can't stand behind.</div>
<div class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58;margin-top:11px">NO FABRICATED PRICE · NO EMPTY GAUGE · TRUTH LAW</div>
</div>
</div>
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58">NOTHING TO TRACK YET</span><span class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58">LEDGER OPENS ON FIRST HONEST LOCK</span></div>
</div>
</div>
<!-- in context · where it renders -->
<div style="display:flex;align-items:center;gap:12px;margin:30px 0 14px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">IN CONTEXT · WHERE IT RENDERS</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">READ-CARD HERO · BOARD ROW · FREE-TIER GATE</span>
</div>
<!-- A · read-card hero -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;overflow:hidden">
<div style="display:flex;align-items:center;gap:8px;flex-wrap:wrap;padding:16px 22px 0">
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:#14141E;border:1px solid #1E1E2A;font-size:8.5px;letter-spacing:.14em;color:#B8BCC8">READ CARD · HERO</span>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:#14141E;border:1px solid #1E1E2A;font-size:8.5px;letter-spacing:.14em;color:#B8BCC8">NFL</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#4a4a58;margin-left:auto">JUL 17, 2026 · THE DESK</span>
</div>
<div style="padding:14px 22px 20px;display:grid;grid-template-columns:minmax(0,1fr) minmax(0,1.08fr);gap:24px;align-items:start">
<div style="min-width:0">
<div style="display:flex;align-items:center;gap:12px">
<span class="mono" style="display:inline-flex;align-items:center;justify-content:center;width:52px;height:52px;border-radius:12px;background:rgba(0,212,160,.12);border:1px solid rgba(0,212,160,.32);font-size:22px;font-weight:800;color:#00D4A0;flex:none;text-shadow:0 0 18px rgba(0,212,160,.4)">A+</span>
<div style="min-width:0"><div style="font-size:18px;font-weight:800;letter-spacing:-.01em;line-height:1.15">Nabers over 1,115.5 receiving yards</div><div class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080;margin-top:4px">NYG · WK 1 · OPENED 1,050.5</div></div>
</div>
<p style="font-size:12.5px;line-height:1.65;color:#B8BCC8;margin-top:14px;max-width:420px">The market moved 70 on the clearance and stopped — our sim says the discount was worth twice that. Read the projection first (does it clear), then the price (is it fair). Two questions, one card.</p>
<div style="display:flex;align-items:center;gap:10px;margin-top:14px"><span class="mono" style="display:inline-flex;align-items:center;justify-content:center;height:42px;padding:0 18px;border-radius:11px;background:#00D4A0;color:#06060B;font-weight:800;font-size:10.5px;letter-spacing:.08em">TAKE +125 AT FANDUEL ▸</span><a href="#clv" class="mono" style="font-size:8px;letter-spacing:.12em;color:#707080;line-height:1.5">CLOSES TRACKED<br>IN YOUR LEDGER ▸</a></div>
</div>
<div style="min-width:0;display:flex;flex-direction:column;gap:12px">
<div style="border:1px solid #1E1E2A;border-radius:13px;background:#0E0E14;padding:13px 16px">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px;margin-bottom:11px"><span class="mono" style="font-size:8px;letter-spacing:.16em;color:#707080">① PROJECTION · WILL IT CLEAR</span><span class="mono" style="font-size:8px;letter-spacing:.1em;color:#00D4A0;font-weight:700">CLEARS BY 72.5</span></div>
<div style="display:flex;align-items:flex-end;gap:16px">
<div><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">MODEL PROJ</div><div class="mono" style="font-size:28px;font-weight:800;color:#00D4A0;line-height:1;margin-top:5px;text-shadow:0 0 18px rgba(0,212,160,.28)">1,188</div></div>
<div class="mono" style="padding-bottom:5px;color:#4a4a58;font-size:11px">vs</div>
<div><div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080">LINE</div><div class="mono" style="font-size:28px;font-weight:800;color:#F0F0F0;line-height:1;margin-top:5px">1,115.5</div></div>
<div style="flex:1;min-width:0;padding-bottom:6px"><div style="position:relative;height:8px;border-radius:4px;background:#14141E;overflow:hidden"><div style="position:absolute;left:0;top:0;bottom:0;width:64%;background:#00D4A0;opacity:.5"></div></div><div class="mono" style="font-size:6.5px;letter-spacing:.1em;color:#4a4a58;margin-top:5px">PROJECTED ABOVE THE LINE</div></div>
</div>
</div>
<div>
<div class="mono" style="font-size:8px;letter-spacing:.16em;color:#707080;margin-bottom:9px">② PRICE · IS IT FAIR</div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:13px;overflow:hidden;background:#0E0E14">
<div style="flex:1;min-width:0;text-align:center;padding:16px 8px"><div class="mono" style="font-size:8px;letter-spacing:.14em;color:#707080">BOOK</div><div class="mono" style="font-size:24px;font-weight:800;color:#F0F0F0;margin-top:8px">+125</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:6px">OFFERED</div></div>
<div style="flex:1.2;min-width:0;text-align:center;padding:16px 8px;background:rgba(255,179,71,.06);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:8px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:36px;font-weight:800;color:#FFB347;margin-top:4px;letter-spacing:-.02em;text-shadow:0 0 24px rgba(255,179,71,.26)">+110</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#8a7a55;margin-top:4px">DE-VIGGED · HERO</div></div>
<div style="flex:1;min-width:0;text-align:center;padding:16px 8px"><div class="mono" style="font-size:8px;letter-spacing:.14em;color:#707080">MODEL</div><div class="mono" style="font-size:24px;font-weight:800;color:#00D4A0;margin-top:8px;text-shadow:0 0 18px rgba(0,212,160,.35)">+98</div><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#4a4a58;margin-top:6px">OUR PRICE</div></div>
</div>
<div style="margin-top:10px;padding:12px 14px;border-radius:11px;background:rgba(0,212,160,.08);border:1px solid rgba(0,212,160,.28)">
<div style="display:flex;align-items:center;justify-content:space-between;gap:10px"><div style="display:flex;align-items:center;gap:8px"><span style="width:7px;height:7px;border-radius:50%;background:#00D4A0;flex:none;box-shadow:0 0 8px rgba(0,212,160,.7)"></span><span class="mono" style="font-size:11px;font-weight:800;letter-spacing:.05em;color:#00D4A0">VALUE · +2.9% VS FAIR</span></div><span class="mono" style="font-size:9px;letter-spacing:.06em;color:#8fb8ac">+6.1% EV AT +125</span></div>
<div style="font-size:10.5px;color:#B8BCC8;line-height:1.5;margin-top:6px">Model beats the honest number and the price is takeable — the edge is real at what you can actually get.</div>
</div>
</div>
</div>
</div>
</div>
<div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:8px;line-height:1.7">HIERARCHY · STAT PROJECTION READS FIRST (WILL IT CLEAR) → PRICE TRIPLET QUALIFIES IT (IS IT FAIR). TWO QUESTIONS, TWO STACKED BLOCKS — NEITHER CROWDS THE OTHER. FAIR STAYS THE AMBER HERO OF THE PRICE LAYER.</div>
<!-- B board row · C free-tier gate -->
<div style="display:grid;grid-template-columns:minmax(0,1.55fr) minmax(0,1fr);gap:16px;margin-top:16px;align-items:start">
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;overflow:hidden;min-width:0">
<div style="display:flex;align-items:center;justify-content:space-between;padding:15px 18px 13px"><span class="mono" style="font-size:10.5px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">TONIGHT'S BOARD · NBA</span><span class="mono" style="font-size:8.5px;letter-spacing:.12em;color:#707080">RANKED BY EDGE · FAIR IS THE NUMBER</span></div>
<div style="display:flex;align-items:center;gap:12px;padding:11px 18px;border-top:1px solid #101018">
<span class="mono" style="display:inline-flex;align-items:center;justify-content:center;width:30px;height:20px;border-radius:5px;background:rgba(0,212,160,.14);border:1px solid rgba(0,212,160,.3);font-size:9px;font-weight:800;color:#00D4A0;flex:none">A+</span>
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:700">Jokić <span style="color:#B8BCC8;font-weight:500">o 27.5 Pts</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">PROJ <span style="color:#00D4A0;font-weight:700">28.9</span> ▸ CLEARS · BOOK +125 · MDL <span style="color:#00D4A0;font-weight:700">+98</span></div></div>
<div style="text-align:right;flex:none"><div class="mono" style="font-size:7px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:16px;font-weight:800;color:#FFB347;line-height:1">+110</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:20px;padding:0 8px;border-radius:5px;background:rgba(0,212,160,.1);border:1px solid rgba(0,212,160,.28);font-size:9px;font-weight:800;color:#00D4A0;flex:none;width:52px;justify-content:center">+2.9%</span>
</div>
<div style="display:flex;align-items:center;gap:12px;padding:11px 18px;border-top:1px solid #101018;opacity:.9">
<span class="mono" style="display:inline-flex;align-items:center;justify-content:center;width:30px;height:20px;border-radius:5px;background:rgba(0,212,160,.1);border:1px solid rgba(0,212,160,.24);font-size:9px;font-weight:800;color:#00D4A0;flex:none">A</span>
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:700">Edwards <span style="color:#B8BCC8;font-weight:500">o 5.5 Ast</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">PROJ <span style="color:#00D4A0;font-weight:700">6.4</span> ▸ CLEARS · BOOK +112 · MDL <span style="color:#00D4A0;font-weight:700">+96</span></div></div>
<div style="text-align:right;flex:none"><div class="mono" style="font-size:7px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:16px;font-weight:800;color:#FFB347;line-height:1">+103</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:20px;padding:0 8px;border-radius:5px;background:rgba(0,212,160,.1);border:1px solid rgba(0,212,160,.28);font-size:9px;font-weight:800;color:#00D4A0;flex:none;width:52px;justify-content:center">+1.6%</span>
</div>
<div style="display:flex;align-items:center;gap:12px;padding:11px 18px;border-top:1px solid #101018;box-shadow:inset 3px 0 0 rgba(255,179,71,.4)">
<span class="mono" style="display:inline-flex;align-items:center;justify-content:center;width:30px;height:20px;border-radius:5px;background:#14141E;border:1px solid #2A2A38;font-size:9px;font-weight:800;color:#B8BCC8;flex:none">B</span>
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:700">Gobert <span style="color:#B8BCC8;font-weight:500">o 11.5 Reb</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">PROJ <span style="color:#B8BCC8;font-weight:700">12.1</span> ▸ CLEARS · BOOK +108 · MDL <span style="color:#FFB347;font-weight:700">withheld</span></div></div>
<div style="text-align:right;flex:none"><div class="mono" style="font-size:7px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:16px;font-weight:800;color:#FFB347;line-height:1">+99</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:20px;padding:0 8px;border-radius:5px;background:#14141E;border:1px solid #2A2A38;font-size:8px;font-weight:700;color:#707080;flex:none;width:52px;justify-content:center">QUAR</span>
</div>
<div style="display:flex;align-items:center;gap:12px;padding:11px 18px;border-top:1px solid #101018;opacity:.72">
<span class="mono" style="display:inline-flex;align-items:center;justify-content:center;width:30px;height:20px;border-radius:5px;background:#14141E;border:1px solid #1E1E2A;font-size:9px;font-weight:800;color:#B8BCC8;flex:none">C</span>
<div style="flex:1;min-width:0"><div style="font-size:12.5px;font-weight:700">Murray <span style="color:#B8BCC8;font-weight:500">o 22.5 Pts</span></div><div class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-top:2px">PROJ <span style="color:#B8BCC8;font-weight:700">22.9</span> ▸ THIN · BOOK +118 · MDL <span style="color:#B8BCC8;font-weight:700">+124</span></div></div>
<div style="text-align:right;flex:none"><div class="mono" style="font-size:7px;letter-spacing:.14em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:16px;font-weight:800;color:#FFB347;line-height:1">+109</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:20px;padding:0 8px;border-radius:5px;background:#101018;border:1px solid #1E1E2A;font-size:8px;font-weight:700;color:#707080;flex:none;width:52px;justify-content:center">NO EDGE</span>
</div>
<div style="padding:10px 18px;border-top:1px solid #101018"><span class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58">PROJECTION LEADS THE ROW (WILL IT CLEAR) · FAIR IS THE PRICE · EDGE CHIP IS THE VERDICT</span></div>
</div>
<!-- C · free-tier gate -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:20px;min-width:0">
<div class="mono" style="font-size:9px;letter-spacing:.2em;color:#707080;margin-bottom:14px">FREE-TIER GATE · MODEL LEG LOCKED</div>
<div style="border-radius:13px;background:#0E0E14;border:1px solid #1E1E2A;padding:16px">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px;margin-bottom:12px"><div style="font-size:12.5px;font-weight:700;min-width:0;overflow:hidden;text-overflow:ellipsis;white-space:nowrap">Nabers <span style="color:#B8BCC8;font-weight:500">o 1,115.5</span></div><span class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#707080;font-weight:700;flex:none">FREE</span></div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:11px;overflow:hidden;background:#0A0A10">
<div style="flex:1;text-align:center;padding:13px 6px"><div class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">BOOK</div><div class="mono" style="font-size:19px;font-weight:800;color:#F0F0F0;margin-top:6px">+125</div></div>
<div style="flex:1.16;text-align:center;padding:13px 6px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:26px;font-weight:800;color:#FFB347;margin-top:3px;letter-spacing:-.02em;text-shadow:0 0 18px rgba(255,179,71,.22)">+110</div></div>
<div style="flex:1;text-align:center;padding:13px 6px;display:flex;flex-direction:column;align-items:center;justify-content:center;position:relative;background:#0C0C12"><div class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#707080">MODEL</div><svg width="15" height="15" viewBox="0 0 24 24" fill="none" style="margin-top:8px;color:#5a5a68"><rect x="4" y="10" width="16" height="10" rx="2" stroke="currentColor" stroke-width="2"></rect><path d="M8 10V7a4 4 0 0 1 8 0v3" stroke="currentColor" stroke-width="2"></path></svg><div class="mono" style="font-size:11px;font-weight:800;color:#5a5a68;letter-spacing:.14em;margin-top:5px">+ █ █</div></div>
</div>
<div style="display:flex;align-items:center;gap:8px;margin-top:11px;padding:10px 12px;border-radius:10px;background:#101018;border:1px solid #1E1E2A"><span style="width:7px;height:7px;border-radius:50%;border:1.5px solid #707080;flex:none"></span><span class="mono" style="font-size:9.5px;font-weight:800;letter-spacing:.04em;color:#B8BCC8;flex:1">MODEL PRICE ON ANALYST</span><span class="mono" style="font-size:8.5px;letter-spacing:.08em;color:#00D4A0;font-weight:700;flex:none">UNLOCK ▸</span></div>
</div>
<div style="font-size:10.5px;color:#707080;line-height:1.55;margin-top:14px">The honest de-vigged number stays free — that's the hook. Only VYNDR's own price gates. The fair leg is never the paywall.</div>
</div>
</div>
<!-- mobile 390 + tier -->
<div style="display:grid;grid-template-columns:minmax(0,1fr) minmax(0,1.25fr);gap:16px;margin-top:16px;align-items:start">
<!-- mobile 390 -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:20px;min-width:0">
<div class="mono" style="font-size:9px;letter-spacing:.2em;color:#707080;margin-bottom:14px">MOBILE · 390 · THREE LEGS HOLD, FAIR STAYS HERO</div>
<div style="width:340px;max-width:100%;margin:0 auto;display:flex;flex-direction:column;gap:12px">
<!-- positive -->
<div style="border-radius:13px;background:#0E0E14;border:1px solid #1E1E2A;padding:14px">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px;margin-bottom:11px"><div style="font-size:12.5px;font-weight:700;min-width:0;overflow:hidden;text-overflow:ellipsis;white-space:nowrap">Nabers <span style="color:#B8BCC8;font-weight:500">o 1,115.5 Rec</span></div><span class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#00D4A0;font-weight:700;flex:none">VALUE</span></div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:10px;overflow:hidden;background:#0A0A10">
<div style="flex:1;text-align:center;padding:11px 4px"><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#707080">BOOK</div><div class="mono" style="font-size:16px;font-weight:800;color:#F0F0F0;margin-top:5px">+125</div></div>
<div style="flex:1.14;text-align:center;padding:11px 4px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:23px;font-weight:800;color:#FFB347;margin-top:3px;letter-spacing:-.02em">+110</div></div>
<div style="flex:1;text-align:center;padding:11px 4px"><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#707080">MODEL</div><div class="mono" style="font-size:16px;font-weight:800;color:#00D4A0;margin-top:5px">+98</div></div>
</div>
<div style="display:flex;align-items:center;gap:7px;margin-top:10px;padding:8px 10px;border-radius:9px;background:rgba(0,212,160,.08);border:1px solid rgba(0,212,160,.26)"><span style="width:6px;height:6px;border-radius:50%;background:#00D4A0;flex:none"></span><span class="mono" style="font-size:9px;font-weight:800;letter-spacing:.04em;color:#00D4A0">+2.9% VS FAIR</span><span style="font-size:9px;color:#707080;margin-left:auto">real edge</span></div>
</div>
<!-- quarantine -->
<div style="border-radius:13px;background:#0E0E14;border:1px solid #2A2A38;padding:14px;box-shadow:inset 3px 0 0 rgba(255,179,71,.4)">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px;margin-bottom:11px"><div style="font-size:12.5px;font-weight:700;min-width:0;overflow:hidden;text-overflow:ellipsis;white-space:nowrap">Bijan <span style="color:#B8BCC8;font-weight:500">o 84.5 Rush</span></div><span class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#FFB347;font-weight:700;flex:none">QUARANTINE</span></div>
<div style="display:flex;align-items:stretch;border:1px solid #1E1E2A;border-radius:10px;overflow:hidden;background:#0A0A10">
<div style="flex:1;text-align:center;padding:11px 4px"><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#707080">BOOK</div><div class="mono" style="font-size:16px;font-weight:800;color:#F0F0F0;margin-top:5px">+130</div></div>
<div style="flex:1.14;text-align:center;padding:11px 4px;background:rgba(255,179,71,.055);border-left:1px solid #1E1E2A;border-right:1px solid #1E1E2A"><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#FFB347">◆ FAIR</div><div class="mono" style="font-size:23px;font-weight:800;color:#FFB347;margin-top:3px;letter-spacing:-.02em">+116</div></div>
<div style="flex:1;text-align:center;padding:11px 4px;background:repeating-linear-gradient(45deg,transparent,transparent 4px,rgba(255,255,255,.014) 4px,rgba(255,255,255,.014) 8px)"><div class="mono" style="font-size:7px;letter-spacing:.1em;color:#707080">MODEL</div><div class="mono" style="font-size:15px;font-weight:800;color:#4a4a58;margin-top:6px"></div></div>
</div>
<div style="display:flex;align-items:center;gap:7px;margin-top:10px;padding:8px 10px;border-radius:9px;background:#101018;border:1px solid #2A2A38"><span style="width:6px;height:6px;border-radius:2px;background:#FFB347;opacity:.85;flex:none"></span><span class="mono" style="font-size:9px;font-weight:800;letter-spacing:.04em;color:#B8BCC8">MODEL WITHHELD</span><span style="font-size:9px;color:#707080;margin-left:auto">book &amp; fair stand</span></div>
</div>
</div>
<div class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58;margin-top:14px;line-height:1.6;text-align:center">44PX MIN TAP ON THE LEDGER LINK · NO HORIZONTAL SCROLL AT 390 · FAIR NEVER DROPS BELOW BOOK/MODEL IN SIZE</div>
</div>
<!-- tier placement -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:20px 22px;min-width:0">
<div style="display:flex;align-items:center;gap:10px;flex-wrap:wrap;margin-bottom:6px"><span class="mono" style="font-size:11px;letter-spacing:.24em;color:#F0F0F0;font-weight:700">TIER PLACEMENT</span><span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 7px;border-radius:5px;background:rgba(255,179,71,.1);border:1px solid rgba(255,179,71,.28);font-size:7.5px;letter-spacing:.1em;color:#FFB347;font-weight:700">RECOMMENDED · CONFIRM</span></div>
<div style="font-size:11.5px;color:#707080;line-height:1.55;margin-bottom:16px">The de-vigged fair number is the free-tier hook per prior product intent — worth surfacing on Free. Flagged for confirmation before ship.</div>
<div style="display:flex;flex-direction:column;gap:10px">
<div style="display:flex;align-items:flex-start;gap:12px;padding:13px 14px;border-radius:11px;background:#0E0E14;border:1px solid #1E1E2A">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#FFB347;font-weight:800;width:62px;flex:none;padding-top:1px">FREE</span>
<div style="min-width:0"><div style="font-size:11.5px;color:#F0F0F0;font-weight:600">Book · fair — model leg locked.</div><div style="font-size:10.5px;color:#707080;line-height:1.5;margin-top:3px">The hook: the honest de-vigged number, free. Model shows as a locked teaser that upsells the read.</div></div>
</div>
<div style="display:flex;align-items:flex-start;gap:12px;padding:13px 14px;border-radius:11px;background:#0E0E14;border:1px solid #1E1E2A">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#00D4A0;font-weight:800;width:62px;flex:none;padding-top:1px">ANALYST</span>
<div style="min-width:0"><div style="font-size:11.5px;color:#F0F0F0;font-weight:600">Full triplet + VALUE marker.</div><div style="font-size:10.5px;color:#707080;line-height:1.5;margin-top:3px">All three legs, the verdict fires both ways — value and no-edge. Quarantine renders as the primary two-leg state.</div></div>
</div>
<div style="display:flex;align-items:flex-start;gap:12px;padding:13px 14px;border-radius:11px;background:#0E0E14;border:1px solid #1E1E2A">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#B8BCC8;font-weight:800;width:62px;flex:none;padding-top:1px">DESK</span>
<div style="min-width:0"><div style="font-size:11.5px;color:#F0F0F0;font-weight:600">Triplet + ledger cross-ref + calibration.</div><div style="font-size:10.5px;color:#707080;line-height:1.5;margin-top:3px">The price story wired to its resolved close and the confidence-bucket record behind the model leg.</div></div>
</div>
</div>
</div>
</div>
<div style="margin-top:16px;padding:13px 18px;border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;display:flex;align-items:center;justify-content:space-between;flex-wrap:wrap;gap:8px">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#FFB347;font-weight:700">THE TRIPLET NEVER PRINTS A NUMBER IT CAN'T STAND BEHIND</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">NO FABRICATED MODEL PRICE · NO HIDDEN "NO EDGE" · FAIR IS HONEST ON EVERY ROW</span>
</div>
</div>
</div>
</div>
</x-dc>
<script type="text/x-dc" data-dc-script data-props="{&quot;motion&quot;:{&quot;editor&quot;:&quot;boolean&quot;,&quot;default&quot;:true,&quot;tsType&quot;:&quot;boolean&quot;,&quot;section&quot;:&quot;Behavior&quot;},&quot;reactions&quot;:{&quot;editor&quot;:&quot;boolean&quot;,&quot;default&quot;:true,&quot;tsType&quot;:&quot;boolean&quot;,&quot;section&quot;:&quot;Behavior&quot;}}">
class Component extends DCLogic {
renderVals() {
return {
rootRef: (n) => { this.rootNode = n; },
stillClass: (this.props.motion ?? true) ? '' : 'vy-still'
};
}
componentDidMount() {
this._t = [];
const tick = () => {
const s = new Date().toLocaleTimeString('en-US', { hour12: false });
document.querySelectorAll('[data-sync-clock]').forEach(el => el.textContent = s);
const wt = document.querySelector('[data-wire-time]');
if (wt) wt.textContent = s;
};
tick();
this._t.push(setInterval(tick, 1000));
const entries = [
'FD holds best number on Nabers o rec yds — 1,115.5 112.',
'Market spread 10.0 across 7 books. Split logged as signal.',
'CLV update: beat close on 61% of settled reads (N 305).',
'Calibration holds — stated 70% bucket hitting 68.2%.',
'The Report Nº 128 out — 12,408 readers.'
];
let wi = 0;
this._t.push(setInterval(() => {
const el = document.querySelector('[data-wire]');
if (!el) return;
wi = (wi + 1) % entries.length;
el.textContent = entries[wi];
el.style.animation = 'none';
void el.offsetWidth;
el.style.animation = 'vy-wirein 6s ease both';
}, 6000));
const nudge = () => {
if (!(this.props.reactions ?? true)) return;
const cells = [...document.querySelectorAll('[data-react]')];
if (!cells.length) return;
const el = cells[Math.floor(Math.random() * cells.length)];
const f = el.getAttribute('data-fmt') || 'odds';
const raw = el.textContent;
const pct = raw.includes('%');
let v = parseFloat(raw.replace(/[^0-9.\-+]/g, ''));
if (isNaN(v)) return;
let up;
if (f === 'edge') { const d = (Math.random() - .45) * .2; up = d >= 0; v += d; el.textContent = (raw.startsWith('AVG') ? 'AVG CLV ' : '') + (v >= 0 ? '+' : '') + v.toFixed(1) + (pct ? '%' : ''); }
else { const d = Math.random() < .5 ? -2 : 2; up = d > 0; v += d; el.textContent = (v > 0 ? '+' : '') + Math.round(v); }
el.classList.remove('rx-up', 'rx-down');
void el.offsetWidth;
el.classList.add(up ? 'rx-up' : 'rx-down');
};
this._t.push(setTimeout(() => { this._t.push(setInterval(nudge, 6400)); }, 1800));
const secs = [...(this.rootNode || document).querySelectorAll('[data-reveal],[data-screen-label]')];
secs.forEach(s => { if (!s.hasAttribute('data-reveal')) s.setAttribute('data-reveal', ''); });
this._io = new IntersectionObserver(ents => {
ents.forEach(e => { if (e.isIntersecting) { e.target.classList.add('vy-in'); this._io.unobserve(e.target); } });
}, { rootMargin: '0px 0px -8% 0px' });
secs.forEach(s => this._io.observe(s));
}
componentWillUnmount() {
(this._t || []).forEach(t => { clearInterval(t); clearTimeout(t); });
if (this._io) this._io.disconnect();
}
}
</script>
</body>
</html>
@@ -0,0 +1,305 @@
<!DOCTYPE html>
<html>
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<script src="./support.js"></script>
</head>
<body>
<x-dc>
<helmet>
<meta name="viewport" content="width=device-width, initial-scale=1" />
<link rel="preconnect" href="https://fonts.googleapis.com" />
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin />
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600;700;800;900&family=JetBrains+Mono:wght@400;500;600;700;800&display=swap" rel="stylesheet" />
<style>
*{box-sizing:border-box;margin:0;padding:0}
html,body{background:#06060B;color:#F0F0F0;font-family:'Inter',system-ui,sans-serif;-webkit-font-smoothing:antialiased}
::selection{background:rgba(0,212,160,.28);color:#F0F0F0}
a{color:#00D4A0;text-decoration:none}
a:hover{color:#5cf0cf}
.mono{font-family:'JetBrains Mono',monospace;font-variant-numeric:tabular-nums}
@keyframes vy-cursor{0%,49%{opacity:1}50%,100%{opacity:0}}
@keyframes vy-syncdot{0%,100%{opacity:1}50%{opacity:.2}}
@keyframes vy-wirein{0%{opacity:0;transform:translateY(4px)}8%,86%{opacity:1;transform:none}100%{opacity:0;transform:translateY(-3px)}}
.vy-grain{position:fixed;inset:0;z-index:5;pointer-events:none;mix-blend-mode:overlay;opacity:.05;background-image:url("data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' width='140' height='140'%3E%3Cfilter id='n'%3E%3CfeTurbulence type='fractalNoise' baseFrequency='0.82' numOctaves='2' stitchTiles='stitch'/%3E%3C/filter%3E%3Crect width='100%25' height='100%25' filter='url(%23n)'/%3E%3C/svg%3E");background-size:140px}
.vy-vignette{position:fixed;inset:0;z-index:4;pointer-events:none;box-shadow:inset 0 0 200px 40px rgba(0,0,0,.55),inset 0 90px 120px -60px rgba(0,212,160,.05)}
[data-reveal]{opacity:0;transform:translateY(16px);transition:opacity .6s cubic-bezier(.2,.8,.2,1),transform .6s cubic-bezier(.2,.8,.2,1)}
[data-reveal].vy-in{opacity:1;transform:none}
.vy-still *{animation-play-state:paused !important}
html{scroll-behavior:smooth}
</style>
</helmet>
<div ref="{{ rootRef }}" class="{{ stillClass }}" style="min-height:100vh;background:radial-gradient(1200px 700px at 78% -8%,rgba(106,147,200,.06),transparent 60%),#06060B;padding:0 0 120px">
<div class="vy-vignette" aria-hidden="true"></div>
<div class="vy-grain" aria-hidden="true"></div>
<div style="position:fixed;left:0;right:0;bottom:0;z-index:45;display:flex;align-items:center;gap:14px;height:36px;padding:0 40px;background:rgba(6,6,11,.88);backdrop-filter:blur(14px);border-top:1px solid #14141E">
<span class="mono" style="font-size:9px;letter-spacing:.26em;color:#00D4A0;font-weight:800;flex:none">THE WIRE</span>
<span style="width:1px;height:14px;background:#1E1E2A;flex:none"></span>
<span class="mono" data-wire-time style="font-size:10px;color:#707080;letter-spacing:.06em;flex:none">--:--:--</span>
<span class="mono" data-wire style="font-size:11px;color:#B8BCC8;letter-spacing:.03em;white-space:nowrap;overflow:hidden;text-overflow:ellipsis;flex:1">Scanner spec open — honest absence is a designed state, not a bug.</span>
<span class="mono" style="font-size:9px;letter-spacing:.16em;color:#4a4a58;flex:none">LOG ▸</span>
</div>
<div style="max-width:1240px;margin:0 auto;padding:0 40px">
<div style="padding:52px 0 30px">
<div class="mono" style="font-size:11px;letter-spacing:.34em;color:#00D4A0;margin-bottom:16px">VYNDR · SCANNER SURFACES · S6 + S7</div>
<h1 style="font-size:38px;font-weight:800;letter-spacing:-.02em;line-height:1.05">Honest absence, designed.<span class="mono" style="color:#00D4A0;animation:vy-cursor 1.1s steps(1) infinite"></span></h1>
<p style="margin-top:12px;max-width:680px;color:#B8BCC8;font-size:14px;line-height:1.55">Two scanner states flagged missing from the bundle. <span class="mono" style="color:#F0F0F0">S6</span> — the marketless scan: what a read card renders when the board never priced that line. <span class="mono" style="color:#F0F0F0">S7</span> — the priced-line display: the scanner surfaces the board's real lines so nobody free-types a line that doesn't exist. Both extend the boundary channel: <span class="mono" style="color:#8fb2de">blue</span> is the system being honest about what it can't hand you.</p>
<div class="mono" style="margin-top:14px;font-size:10px;letter-spacing:.16em;color:#707080">NO FABRICATED MARKETS · ABSENCE IS THE DEFAULT CASE · EVERY DEAD END POINTS SOMEWHERE REAL</div>
</div>
<!-- ============ S6 · SCAN EMPTY STATE ============ -->
<div data-reveal style="margin-top:14px;display:flex;align-items:baseline;gap:18px;border-bottom:1px solid #14141E;padding-bottom:18px">
<span class="mono" style="font-size:64px;font-weight:800;line-height:.8;color:#14141E;-webkit-text-stroke:1px #2A2A38">S6</span>
<div><div class="mono" style="font-size:13px;letter-spacing:.3em;color:#F0F0F0;font-weight:800">SCAN EMPTY STATE <span style="color:#8fb2de">· MARKETLESS</span></div><div style="font-size:12.5px;color:#707080;margin-top:5px">The user scanned a prop the board never priced. No market numbers render — there is no market. The card says so plainly, then hands the user tonight's real board. Help, never a wall.</div></div>
</div>
<div id="s6" data-screen-label="S6 · Scan empty state" style="margin-top:44px">
<div style="display:grid;grid-template-columns:minmax(0,1.35fr) minmax(0,1fr);gap:16px;align-items:start">
<!-- anatomy · the marketless read card -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:22px 24px;min-width:0">
<div style="display:flex;align-items:center;gap:12px;flex-wrap:wrap;margin-bottom:16px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">ANATOMY</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">READ-CARD FRAME · ZERO MARKET NUMBERS · PATH FORWARD BUILT IN</span>
</div>
<div style="max-width:520px;margin:0 auto;border-radius:14px;background:#0E0E14;border:1px solid #1E1E2A;padding:18px;display:flex;flex-direction:column;gap:13px">
<div style="display:flex;align-items:flex-start;justify-content:space-between;gap:10px">
<div style="min-width:0"><div style="font-size:13.5px;font-weight:700">Herbert <span style="color:#B8BCC8;font-weight:500">o 287.5 Pass Yds</span></div><div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:3px">LAC · SCANNED · JUL 22, 2026</div></div>
<span class="mono" style="display:inline-flex;align-items:center;height:18px;padding:0 8px;border-radius:5px;background:rgba(106,147,200,.12);border:1px solid rgba(106,147,200,.34);font-size:8px;letter-spacing:.1em;color:#8fb2de;font-weight:700;flex:none">NO MARKET</span>
</div>
<div style="display:flex;align-items:center;justify-content:center;border:1px dashed rgba(106,147,200,.35);border-radius:11px;padding:22px 16px;text-align:center;background:#0A0A10">
<div>
<div class="mono" style="font-size:9px;letter-spacing:.22em;color:#8fb2de">NO LIVE MARKET FOR THIS LINE</div>
<div style="font-size:11.5px;color:#B8BCC8;margin-top:9px;line-height:1.55;max-width:330px">No book prices this exact line right now — so there's no vig to strip and no fair number to show. We won't invent a market that isn't there.</div>
<div class="mono" style="font-size:7.5px;letter-spacing:.12em;color:#4a4a58;margin-top:11px">NO BOOK · NO FAIR · NO MODEL · NOTHING FABRICATED</div>
</div>
</div>
<div style="padding:12px 14px;border-radius:10px;background:rgba(0,212,160,.06);border:1px solid rgba(0,212,160,.22)">
<div style="display:flex;align-items:center;justify-content:space-between;gap:10px"><span class="mono" style="font-size:8px;letter-spacing:.16em;color:#707080">WHAT WE PRICE NOW</span><span class="mono" style="font-size:8px;letter-spacing:.08em;color:#00D4A0;font-weight:700">142 PROPS LIVE</span></div>
<div style="display:flex;align-items:center;gap:10px;margin-top:10px">
<span class="mono" style="display:inline-flex;align-items:center;justify-content:center;height:38px;padding:0 16px;border-radius:10px;background:#00D4A0;color:#06060B;font-weight:800;font-size:10px;letter-spacing:.08em;flex:none">TONIGHT'S BOARD ▸</span>
<span style="font-size:10.5px;color:#B8BCC8;line-height:1.5;min-width:0">Herbert has <span class="mono" style="color:#F0F0F0;font-weight:700">3 priced props</span> on tonight's board — full triplets on each.</span>
</div>
</div>
</div>
<div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:16px;line-height:1.7">CHIP · BLUE BOUNDARY CHANNEL rgba(106,147,200,.12) / BORDER .34 / TEXT #8FB2DE — SAME TOKENS AS PRICED-OUT · DASHED VOID BOX BORROWS THE REFUSAL FRAME BUT BLUE, NOT RED: NOTHING FAILED, THE MARKET JUST ISN'T THERE · GREEN RESERVED FOR THE ONE REAL PATH FORWARD (PRIMARY CTA LAW)</div>
</div>
<!-- laws + distinctions -->
<div style="display:flex;flex-direction:column;gap:16px;min-width:0">
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:20px 22px">
<div class="mono" style="font-size:11px;letter-spacing:.24em;color:#F0F0F0;font-weight:700;margin-bottom:14px">STATE LAWS</div>
<div style="display:flex;flex-direction:column;gap:10px">
<div style="display:flex;align-items:flex-start;gap:11px"><span class="mono" style="color:#8fb2de;font-size:11px;flex:none;padding-top:1px">01</span><div style="font-size:11px;color:#B8BCC8;line-height:1.55;min-width:0"><span style="color:#F0F0F0;font-weight:600">Zero market numbers.</span> No book, no fair, no model, no gauge at 0. An empty triplet frame is a lie — the frame doesn't render.</div></div>
<div style="display:flex;align-items:flex-start;gap:11px"><span class="mono" style="color:#8fb2de;font-size:11px;flex:none;padding-top:1px">02</span><div style="font-size:11px;color:#B8BCC8;line-height:1.55;min-width:0"><span style="color:#F0F0F0;font-weight:600">True copy only.</span> "No live market for this line" — never "coming soon", never "check back". We don't promise markets we don't control.</div></div>
<div style="display:flex;align-items:flex-start;gap:11px"><span class="mono" style="color:#8fb2de;font-size:11px;flex:none;padding-top:1px">03</span><div style="font-size:11px;color:#B8BCC8;line-height:1.55;min-width:0"><span style="color:#F0F0F0;font-weight:600">Always a path forward.</span> The card ends at tonight's board — with the live priced-prop count and, when the player is on the slate, their real priced props. A scan never dead-ends.</div></div>
<div style="display:flex;align-items:flex-start;gap:11px"><span class="mono" style="color:#8fb2de;font-size:11px;flex:none;padding-top:1px">04</span><div style="font-size:11px;color:#B8BCC8;line-height:1.55;min-width:0"><span style="color:#F0F0F0;font-weight:600">Help framing.</span> "What we price now" — the product showing its inventory, never apologizing for what a book didn't post.</div></div>
</div>
</div>
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:20px 22px">
<div class="mono" style="font-size:11px;letter-spacing:.24em;color:#F0F0F0;font-weight:700;margin-bottom:12px">NOT THE SAME STATE AS</div>
<div style="display:flex;flex-direction:column;gap:9px">
<div style="display:flex;align-items:center;gap:10px;padding:10px 12px;border-radius:10px;background:#0E0E14;border:1px solid #1E1E2A"><span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 7px;border-radius:5px;background:rgba(255,71,87,.08);border:1px solid rgba(255,71,87,.28);font-size:7.5px;letter-spacing:.1em;color:#FF6B78;font-weight:700;flex:none">REFUSAL</span><span style="font-size:10.5px;color:#707080;line-height:1.45;min-width:0">Market exists, our inputs are too thin to grade. Red — a model decision.</span></div>
<div style="display:flex;align-items:center;gap:10px;padding:10px 12px;border-radius:10px;background:#0E0E14;border:1px solid #1E1E2A"><span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 7px;border-radius:5px;background:rgba(255,179,71,.1);border:1px solid rgba(255,179,71,.3);font-size:7.5px;letter-spacing:.1em;color:#FFB347;font-weight:700;flex:none">QUARANTINE</span><span style="font-size:10.5px;color:#707080;line-height:1.45;min-width:0">Market exists, model leg suppressed. Amber — book &amp; fair still stand.</span></div>
<div style="display:flex;align-items:center;gap:10px;padding:10px 12px;border-radius:10px;background:#0E0E14;border:1px solid rgba(106,147,200,.3)"><span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 7px;border-radius:5px;background:rgba(106,147,200,.12);border:1px solid rgba(106,147,200,.34);font-size:7.5px;letter-spacing:.1em;color:#8fb2de;font-weight:700;flex:none">NO MARKET</span><span style="font-size:10.5px;color:#B8BCC8;line-height:1.45;min-width:0">No market exists at all. Blue boundary — nothing to price, nothing failed.</span></div>
</div>
</div>
</div>
</div>
<!-- S6 mobile 390 -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:20px;margin-top:16px">
<div class="mono" style="font-size:9px;letter-spacing:.2em;color:#707080;margin-bottom:14px">MOBILE · 390 · SAME ANATOMY, 44PX PATHS</div>
<div style="width:340px;max-width:100%;margin:0 auto;border-radius:13px;background:#0E0E14;border:1px solid #1E1E2A;padding:14px;display:flex;flex-direction:column;gap:11px">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><div style="font-size:12.5px;font-weight:700;min-width:0;overflow:hidden;text-overflow:ellipsis;white-space:nowrap">Herbert <span style="color:#B8BCC8;font-weight:500">o 287.5 Pass Yds</span></div><span class="mono" style="font-size:7.5px;letter-spacing:.1em;color:#8fb2de;font-weight:700;flex:none">NO MARKET</span></div>
<div style="border:1px dashed rgba(106,147,200,.35);border-radius:10px;padding:16px 12px;text-align:center;background:#0A0A10">
<div class="mono" style="font-size:8px;letter-spacing:.18em;color:#8fb2de">NO LIVE MARKET FOR THIS LINE</div>
<div style="font-size:10.5px;color:#B8BCC8;margin-top:7px;line-height:1.5">No book prices this line right now. We won't invent one.</div>
</div>
<span class="mono" style="display:flex;align-items:center;justify-content:center;height:44px;border-radius:11px;background:#00D4A0;color:#06060B;font-weight:800;font-size:10.5px;letter-spacing:.08em">HERBERT · 3 PRICED PROPS ▸</span>
<span class="mono" style="display:flex;align-items:center;justify-content:center;height:44px;border-radius:11px;background:#14141E;border:1px solid #2A2A38;color:#B8BCC8;font-weight:700;font-size:10px;letter-spacing:.08em">TONIGHT'S BOARD · 142 LIVE ▸</span>
</div>
<div class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58;margin-top:14px;line-height:1.6;text-align:center">BOTH PATHS 44PX · PLAYER-SPECIFIC PATH LEADS WHEN THE PLAYER IS ON THE SLATE, BOARD PATH OTHERWISE PROMOTES TO GREEN</div>
</div>
</div>
<!-- ============ S7 · PRICED-LINE DISPLAY ============ -->
<div data-reveal style="margin-top:56px;display:flex;align-items:baseline;gap:18px;border-bottom:1px solid #14141E;padding-bottom:18px">
<span class="mono" style="font-size:64px;font-weight:800;line-height:.8;color:#14141E;-webkit-text-stroke:1px #2A2A38">S7</span>
<div><div class="mono" style="font-size:13px;letter-spacing:.3em;color:#F0F0F0;font-weight:800">SCANNER PRICED-LINE DISPLAY <span style="color:#8fb2de">· THE NUDGE</span></div><div style="font-size:12.5px;color:#707080;margin-top:5px">As the user picks player + stat, the scanner shows the board's real priced lines — one tap to a real triplet instead of free-typing a line that isn't priced. Only genuinely-priced lines render. "None priced" is the common case and the default design, not an edge.</div></div>
</div>
<div id="s7" data-screen-label="S7 · Scanner priced-line display" style="margin-top:44px">
<!-- the strip anatomy, three cases -->
<div style="display:flex;align-items:center;gap:12px;margin-bottom:14px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">THE STRIP · THREE CASES</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">CASE A IS THE DEFAULT — MOST PLAYER+STAT PICKS HAVE NO PRICED LINE</span>
</div>
<div style="display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:16px;align-items:start">
<!-- CASE A · none priced (DEFAULT) -->
<div style="border-radius:14px;background:#0E0E14;border:1px solid rgba(106,147,200,.3);padding:18px;display:flex;flex-direction:column;gap:12px">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#8fb2de;font-weight:700">A · NONE PRICED</span><span class="mono" style="display:inline-flex;align-items:center;height:17px;padding:0 7px;border-radius:5px;background:rgba(106,147,200,.12);border:1px solid rgba(106,147,200,.34);font-size:7px;letter-spacing:.1em;color:#8fb2de;font-weight:700">DEFAULT · COMMON CASE</span></div>
<div style="border-radius:11px;background:#0A0A10;border:1px solid #1E1E2A;padding:13px">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:11px;color:#F0F0F0;font-weight:700">Herbert</span><span class="mono" style="font-size:9px;color:#707080">·</span><span class="mono" style="display:inline-flex;align-items:center;height:19px;padding:0 8px;border-radius:5px;background:#14141E;border:1px solid #2A2A38;font-size:8.5px;color:#F0F0F0;font-weight:700">COMPLETIONS</span></div>
<div style="margin-top:11px;padding:12px;border-radius:9px;border:1px dashed rgba(106,147,200,.3);text-align:center">
<div class="mono" style="font-size:8px;letter-spacing:.14em;color:#8fb2de">NO PRICED LINE FOR THIS STAT</div>
<div style="font-size:10px;color:#707080;margin-top:6px;line-height:1.5">The board doesn't price Herbert completions tonight.</div>
</div>
<div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080;margin:11px 0 7px">WE PRICE THESE FOR HERBERT</div>
<div style="display:flex;gap:6px;flex-wrap:wrap">
<span class="mono" style="font-size:9px;color:#F0F0F0;background:#14141E;border:1px solid #2A2A38;border-radius:7px;padding:7px 10px;font-weight:700">PASS YDS <span style="color:#707080;font-weight:400">o 275.5</span></span>
<span class="mono" style="font-size:9px;color:#F0F0F0;background:#14141E;border:1px solid #2A2A38;border-radius:7px;padding:7px 10px;font-weight:700">PASS TD <span style="color:#707080;font-weight:400">o 1.5</span></span>
<span class="mono" style="font-size:9px;color:#F0F0F0;background:#14141E;border:1px solid #2A2A38;border-radius:7px;padding:7px 10px;font-weight:700">INT <span style="color:#707080;font-weight:400">o 0.5</span></span>
</div>
</div>
<div style="font-size:10.5px;color:#707080;line-height:1.55">Quiet, structured, never alarmed — the strip keeps its shape and offers the player's real priced stats as the way through. This state must look intentional at scale: it's most of the traffic.</div>
</div>
<!-- CASE B · priced lines exist -->
<div style="border-radius:14px;background:#0E0E14;border:1px solid #1E1E2A;padding:18px;display:flex;flex-direction:column;gap:12px">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#00D4A0;font-weight:700">B · PRICED — ONE TAP TO THE TRIPLET</span></div>
<div style="border-radius:11px;background:#0A0A10;border:1px solid #1E1E2A;padding:13px">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:11px;color:#F0F0F0;font-weight:700">Jokić</span><span class="mono" style="font-size:9px;color:#707080">·</span><span class="mono" style="display:inline-flex;align-items:center;height:19px;padding:0 8px;border-radius:5px;background:#14141E;border:1px solid #2A2A38;font-size:8.5px;color:#F0F0F0;font-weight:700">POINTS</span></div>
<div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080;margin:11px 0 7px">WE PRICE THESE LINES</div>
<div style="display:flex;align-items:center;gap:7px;height:44px;padding:0 10px;border-radius:10px;background:#0E0E14;border:1px solid rgba(0,212,160,.28)">
<span class="mono" style="font-size:10.5px;font-weight:800;color:#F0F0F0;flex:none">o 27.5</span>
<span class="mono" style="font-size:7.5px;color:#707080;flex:none">BK 114</span>
<span class="mono" style="font-size:7.5px;color:#FFB347;font-weight:700;flex:none">105</span>
<span class="mono" style="font-size:10px;color:#00D4A0;font-weight:800;margin-left:auto;flex:none"></span>
</div>
<div style="display:flex;align-items:center;gap:7px;height:44px;padding:0 10px;border-radius:10px;background:#0E0E14;border:1px solid #1E1E2A;margin-top:7px">
<span class="mono" style="font-size:10.5px;font-weight:800;color:#F0F0F0;flex:none">u 27.5</span>
<span class="mono" style="font-size:7.5px;color:#707080;flex:none">BK 106</span>
<span class="mono" style="font-size:7.5px;color:#FFB347;font-weight:700;flex:none">115</span>
<span class="mono" style="font-size:10px;color:#707080;font-weight:800;margin-left:auto;flex:none"></span>
</div>
</div>
<div style="font-size:10.5px;color:#707080;line-height:1.55">Only lines the board genuinely prices — per exact player + stat. Each row is a 44px tap straight into the real triplet. Fair previews amber; green edges only when the read behind it is takeable.</div>
</div>
<!-- CASE C · off-slate -->
<div style="border-radius:14px;background:#0E0E14;border:1px solid #1E1E2A;padding:18px;display:flex;flex-direction:column;gap:12px">
<div style="display:flex;align-items:center;justify-content:space-between;gap:8px"><span class="mono" style="font-size:8.5px;letter-spacing:.14em;color:#B8BCC8;font-weight:700">C · OFF-SLATE PLAYER</span></div>
<div style="border-radius:11px;background:#0A0A10;border:1px solid #1E1E2A;padding:13px">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:11px;color:#F0F0F0;font-weight:700">Skenes</span><span class="mono" style="font-size:9px;color:#707080">·</span><span class="mono" style="display:inline-flex;align-items:center;height:19px;padding:0 8px;border-radius:5px;background:#14141E;border:1px solid #2A2A38;font-size:8.5px;color:#F0F0F0;font-weight:700">STRIKEOUTS</span></div>
<div style="margin-top:11px;padding:14px 12px;text-align:center">
<div class="mono" style="font-size:8px;letter-spacing:.14em;color:#707080">NOT ON TONIGHT'S SLATE</div>
<div style="font-size:10px;color:#4a4a58;margin-top:6px;line-height:1.5">No game, no markets, no prices to show.</div>
</div>
</div>
<div style="font-size:10.5px;color:#707080;line-height:1.55">Nothing renders — no chips, no suggestions, no dashed frame. Off-slate is quieter than case A: the strip states the fact in text tokens only and stops. No blue; there's no boundary to mark, just no game.</div>
</div>
</div>
<!-- typed-line mismatch, in scanner context -->
<div style="display:flex;align-items:center;gap:12px;margin:30px 0 14px">
<span class="mono" style="font-size:12px;letter-spacing:.28em;color:#F0F0F0;font-weight:700">IN CONTEXT · TYPED LINE ≠ PRICED LINE</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">THE NUDGE — REDIRECT TO THE REAL LINE, NEVER PRETEND THE TYPED ONE EXISTS</span>
</div>
<div style="display:grid;grid-template-columns:minmax(0,1.35fr) minmax(0,1fr);gap:16px;align-items:start">
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:22px 24px;min-width:0">
<div style="max-width:520px;margin:0 auto">
<div style="display:flex;align-items:center;gap:10px;height:44px;padding:0 14px;border-radius:12px;background:#0E0E14;border:1px solid rgba(0,212,160,.4)">
<svg width="15" height="15" viewBox="0 0 24 24" fill="none"><circle cx="11" cy="11" r="7" stroke="#707080" stroke-width="2"></circle><path stroke="#707080" stroke-width="2" stroke-linecap="round" d="M20 20l-4-4"></path></svg>
<span class="mono" style="font-size:13px;color:#F0F0F0">jokic o 30.5 pts<span style="display:inline-block;width:2px;height:15px;background:#00D4A0;vertical-align:-2px;animation:vy-syncdot 1.1s steps(1) infinite"></span></span>
</div>
<div style="margin-top:12px;border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;padding:13px">
<div style="display:flex;align-items:center;gap:8px"><span class="mono" style="font-size:9px;color:#8fb2de;flex:none"></span><span class="mono" style="font-size:8.5px;letter-spacing:.1em;color:#8fb2de;font-weight:700">o 30.5 ISN'T PRICED</span><span class="mono" style="font-size:8px;letter-spacing:.08em;color:#707080;margin-left:auto">NEAREST PRICED LINE ↓</span></div>
<div style="display:flex;align-items:center;gap:10px;height:44px;padding:0 12px;border-radius:10px;background:#0A0A10;border:1px solid rgba(0,212,160,.28);margin-top:9px">
<span class="mono" style="font-size:11px;font-weight:800;color:#F0F0F0;flex:none">o 27.5 PTS</span>
<span class="mono" style="font-size:8.5px;color:#707080;flex:none">BOOK 114</span>
<span class="mono" style="font-size:8.5px;color:#FFB347;font-weight:700;flex:none">◆ FAIR 105</span>
<span class="mono" style="font-size:8.5px;color:#00D4A0;font-weight:700;margin-left:auto;flex:none">OPEN READ ▸</span>
</div>
<div style="font-size:10px;color:#707080;line-height:1.5;margin-top:9px">The typed line stays visible in the query — the strip never rewrites the user's intent, it shows what's actually priced next to it.</div>
</div>
</div>
<div class="mono" style="font-size:8px;letter-spacing:.12em;color:#4a4a58;margin-top:16px;line-height:1.7">"ISN'T PRICED" IS A FACT LINE, NOT AN ERROR — BLUE TEXT, NO RED, NO SHAKE · NEAREST LINE = SAME PLAYER + STAT ONLY, NEVER A DIFFERENT STAT PASSED OFF AS CLOSE</div>
</div>
<!-- S7 mobile 390 -->
<div style="border-radius:16px;background:#0A0A10;border:1px solid #1E1E2A;padding:20px;min-width:0">
<div class="mono" style="font-size:9px;letter-spacing:.2em;color:#707080;margin-bottom:14px">MOBILE · 390 · STRIP UNDER THE SCAN FIELD</div>
<div style="width:300px;max-width:100%;margin:0 auto;display:flex;flex-direction:column;gap:9px">
<div style="display:flex;align-items:center;gap:9px;height:44px;padding:0 13px;border-radius:12px;background:#0E0E14;border:1px solid rgba(0,212,160,.4)">
<svg width="14" height="14" viewBox="0 0 24 24" fill="none"><circle cx="11" cy="11" r="7" stroke="#707080" stroke-width="2"></circle><path stroke="#707080" stroke-width="2" stroke-linecap="round" d="M20 20l-4-4"></path></svg>
<span class="mono" style="font-size:12px;color:#F0F0F0">herbert comp<span style="display:inline-block;width:2px;height:14px;background:#00D4A0;vertical-align:-2px;animation:vy-syncdot 1.1s steps(1) infinite"></span></span>
</div>
<div style="border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;padding:12px">
<div style="padding:11px 10px;border-radius:9px;border:1px dashed rgba(106,147,200,.3);text-align:center">
<div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#8fb2de">NO PRICED LINE FOR THIS STAT</div>
</div>
<div class="mono" style="font-size:7.5px;letter-spacing:.14em;color:#707080;margin:10px 0 7px">WE PRICE THESE FOR HERBERT</div>
<div style="display:flex;flex-direction:column;gap:6px">
<span class="mono" style="display:flex;align-items:center;height:44px;padding:0 12px;border-radius:9px;background:#14141E;border:1px solid #2A2A38;font-size:9.5px;color:#F0F0F0;font-weight:700;gap:8px">PASS YDS <span style="color:#707080;font-weight:400">o 275.5</span><span style="color:#FFB347;font-weight:700;margin-left:auto">108</span></span>
<span class="mono" style="display:flex;align-items:center;height:44px;padding:0 12px;border-radius:9px;background:#14141E;border:1px solid #2A2A38;font-size:9.5px;color:#F0F0F0;font-weight:700;gap:8px">PASS TD <span style="color:#707080;font-weight:400">o 1.5</span><span style="color:#FFB347;font-weight:700;margin-left:auto">◆ +102</span></span>
</div>
</div>
</div>
<div class="mono" style="font-size:8px;letter-spacing:.1em;color:#4a4a58;margin-top:14px;line-height:1.6;text-align:center">ALL SUGGESTION ROWS 44PX · FAIR PREVIEW RIDES EVERY SUGGESTED STAT — THE FREE HOOK WORKS EVEN IN THE EMPTY STATE</div>
</div>
</div>
<div style="margin-top:16px;padding:13px 18px;border-radius:12px;background:#0E0E14;border:1px solid #1E1E2A;display:flex;align-items:center;justify-content:space-between;flex-wrap:wrap;gap:8px">
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#8fb2de;font-weight:700">THE SCANNER NEVER IMPLIES A LINE EXISTS WHERE IT DOESN'T</span>
<span class="mono" style="font-size:9px;letter-spacing:.14em;color:#707080">ABSENCE IS THE DEFAULT DESIGN · EVERY EMPTY POINTS TO A REAL PRICED PROP · BLUE = THE HONEST BOUNDARY</span>
</div>
</div>
</div>
</div>
</x-dc>
<script type="text/x-dc" data-dc-script data-props="{&quot;motion&quot;:{&quot;editor&quot;:&quot;boolean&quot;,&quot;default&quot;:true,&quot;tsType&quot;:&quot;boolean&quot;,&quot;section&quot;:&quot;Behavior&quot;}}">
class Component extends DCLogic {
renderVals() {
return {
rootRef: (n) => { this.rootNode = n; },
stillClass: (this.props.motion ?? true) ? '' : 'vy-still'
};
}
componentDidMount() {
this._t = [];
const tick = () => {
const wt = document.querySelector('[data-wire-time]');
if (wt) wt.textContent = new Date().toLocaleTimeString('en-US', { hour12: false });
};
tick();
this._t.push(setInterval(tick, 1000));
const entries = [
'Marketless scans tonight: 61% of scanner traffic — S6 is the front door.',
'Board carries 142 priced props · scanner nudge routes dead-ends there.',
'Blue boundary channel: priced-out edge + unpriced lines share one meaning.',
'No fabricated markets. No "coming soon". Truth law holds.'
];
let wi = 0;
this._t.push(setInterval(() => {
const el = document.querySelector('[data-wire]');
if (!el) return;
wi = (wi + 1) % entries.length;
el.textContent = entries[wi];
el.style.animation = 'none';
void el.offsetWidth;
el.style.animation = 'vy-wirein 6s ease both';
}, 6000));
const secs = [...(this.rootNode || document).querySelectorAll('[data-reveal],[data-screen-label]')];
secs.forEach(s => { if (!s.hasAttribute('data-reveal')) s.setAttribute('data-reveal', ''); });
this._io = new IntersectionObserver(ents => {
ents.forEach(e => { if (e.isIntersecting) { e.target.classList.add('vy-in'); this._io.unobserve(e.target); } });
}, { rootMargin: '0px 0px -8% 0px' });
secs.forEach(s => this._io.observe(s));
}
componentWillUnmount() {
(this._t || []).forEach(t => { clearInterval(t); clearTimeout(t); });
if (this._io) this._io.disconnect();
}
}
</script>
</body>
</html>
File diff suppressed because it is too large Load Diff
@@ -86,4 +86,8 @@
| SWITCHBOARD | `#90A0E8` | `assets/glyphs/switchboard.svg` | | SWITCHBOARD | `#90A0E8` | `assets/glyphs/switchboard.svg` |
| RETURNER | `#6C9CB0` | `assets/glyphs/returner.svg` | | RETURNER | `#6C9CB0` | `assets/glyphs/returner.svg` |
Marks use `currentColor`. 74 display names + 9 classifier-legacy aliases (never user-facing). Marks use `currentColor`. 74 display names + 9 classifier-legacy aliases (BRUSH, WHIFF, CONNECTOR, DISTRIBUTOR, FASTBREAK, FLEX, HYBRID, SWITCH, SWITCHBOARD) — legacy names are classifier-side only, never displayed; map them to renamed display archetypes (BRUSH→SLASH etc.) or render their own mark where the backend emits the legacy key.
## SESSION 2 NOTE (Jul 17 2026)
No new marks. Surfaces 2-5 reuse the 83 existing marks only (hero-scale usage in article media: render the SVG at up to 72px, duotone unchanged). New Session 2 assets are chart/viz templates and book wordmark chips, defined in `Vyndr Intelligence.dc.html` — not glyph assets.
@@ -1 +1 @@
<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24" width="24" height="24" fill="none"><path fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" d="M4 20 15 9M15 9h-5M15 9v5"/><path stroke="currentColor" stroke-width="2" stroke-linecap="round" d="M17 16l3 3"/><circle cx="19" cy="6" r="1.8" fill="currentColor"/></svg> <svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" viewBox="0 0 24 24" fill="none" color="#FF6B4A"><path fill="none" stroke="currentColor" stroke-width="2.5" stroke-linecap="round" stroke-linejoin="round" d="M4 20 15 9M15 9h-5M15 9v5"></path><path stroke="currentColor" stroke-width="2" stroke-linecap="round" d="M17 16l3 3"></path><circle cx="19" cy="6" r="1.8" fill="currentColor"></circle></svg>

Before

Width:  |  Height:  |  Size: 375 B

After

Width:  |  Height:  |  Size: 410 B

Some files were not shown because too many files have changed in this diff Show More